Learning about writing: A consideration of the recently revised asTTle: Writing
Judy Parr and Gavin Brown
http://dx.doi.org/10.18296/cm.0008
Abstract
The recent revision of a national writing assessment tool, e-asTTle: Writing (2012), is viewed from theoretical, design and implementation, and practice perspectives, and considered in relation to the original assessment concept and design. Aspects of the revision are questioned. These include a reduction in the scope and complexity of writing, with fewer functions or communicative purposes for writing included, and narrowing of the dimensions of writing. Arguably, these changes limit the diagnostic information available for teacher and student learning and, importantly, opportunities for valuable content learning about writing for teachers. Regarding design and implementation, questions relate to the necessity for statistical manipulation of teacher judgement. The increased emphasis in scoring on the more technical aspects of writing risks sending inappropriate messages to teachers and students about what is valued in quality writing and curriculum implementation.
Writing has an important, dual position within The New Zealand Curriculum (Ministry of Education, 2007) (NZC), the broad, guiding document intended to “set the direction for student learning” (p. 6). The aim that students learn “how to communicate knowledge and ideas in appropriate ways” (p. 16) nominally places writing within all learning areas, although the knowledge and skills for writing largely are described within the guidelines of the English learning area. The English curriculum focuses on meaning making within two groupings, broadly receptive and productive—the latter includes writing. For each grouping, processes and strategies are identified and, using these, outcomes are described for students to achieve. A broad aim is specified for each of four areas: purposes and audiences; ideas; language features, and structure and for each aim a small number of indicators for different levels of the curriculum are specified. The more detailed desired levels of performance in writing at various curriculum levels (Years 1–10) are described in the New Zealand Curriculum Reading and Writing Standards (Ministry of Education, 2009). The standards are not standards for writing as a subject, but are based on an analysis of curriculum documents in the learning areas that make up the national curriculum and describe the level of writing required to meet curriculum demands in each area. The emphasis is on writing in the service of learning.
Importantly, in line with the policy of adapting the national curriculum to local needs, NZC positions teaching as involving a process of inquiry and teachers as inquiring practitioners (p. 35). Teachers are seen as responsible for their own learning, for inquiring into student learning and making appropriate adjustments to practice to meet the needs of students. This is also consistent with policy which states that the decision about the level of achievement of each student in relation to standards will be based on an overall teacher judgement (OTJ). Arguably, with a non-prescriptive approach to how the curriculum might be implemented and with an emphasis on teaching as inquiry, it is necessary to provide teachers with opportunities for ongoing professional learning and with supports or scaffolds for inquiry, in the form of quality tools and resources that instantiate a sound theory of the task of writing. The curriculum and the assessment tools, resources, and standards interact; they are mutually influential.
We examine the recent revision of an assessment tool, e-asTTle: Writing (New Zealand Council for Educational Research (NZCER), 2012), arguing that the revision represents a significant departure from the previous versions of asTTle (versions 1–4) and e-asTTle. Both authors were involved in the design of asTTle. The first author led the small team of writing research and curriculum experts who designed asTTle:Writing, specifically the conceptual framing, scoring rubrics, and prompts, and who were part of the teacher feedback and reliability testing processes. The second author was the senior project manager, under John Hattie the originator of asTTle and e-asTTle, for the asTTle versions 1 to 4 team from 2000 to 2005, and was involved in design of the asTTle reporting system, norm data collection, statistical data analysis, standard setting, and report writing for the research and development processes.
We discuss both the original asTTle: Writing (Hattie et al., 2004) and its versions, and the recent revision from three perspectives: theoretical or conceptual; design and implementation; and practice, while acknowledging that these perspectives are not mutually exclusive. And, we evaluate the likely attendant outcomes of the revision, particularly in terms of backwash (Alderson & Wall, 1993), the notion of the effects, both positive and negative of assessment on curriculum implementation and teaching and learning. Our purpose here in drawing contrasts is not only to stimulate debate about the concept of writing that assessment may represent but also about the broader functions of assessment tools such as asTTle within the curriculum.
A theoretical perspective
A key component of a valid assessment system is: “a model (a theory) of student learning and cognition in the domain” (National Research Council, 2001, p. 44). Test content has to be relevant to the construct(s) it is to measure. In writing, obtaining a definitive model of learning and cognition in the domain is challenging. What develops and what it develops towards (Marshall, 2006) is problematic; a detailed model that specifies development of both product and processes in writing does not exist (Almargot & Fahol, 2009). Theoretical views do exist; they basically parallel three movements reflecting the textual, the cognitive, and the social (Witte, 1992). Theoretically, views of writing have encompassed an emphasis on the product, a focus on the cognitive processes involved in writing, and a turn to the social, recognising the role of context where writing is essentially a social construction designed to meet communicative purposes.
A significant point of difference between the recent 2012 revision and earlier versions of asTTle is that the theoretical conception of writing which underpinned the original design of the assessment tool reflected the view that writing is a highly complex socio-cognitive act, far more complex than reading. The incorporation of at least some of that complexity in the original design was purposeful to provide a tool to inform teacher practice bothfor learning about their students and for learning about writing. The theoretical stance that was adopted in designing the original Assessment Tools for Teaching and Learning (asTTle):Writing(Glasswell, Parr, & Aikman, 2001; Hattie et al., 2004) is that writing is a social practice; genre is the processes involved in getting things done through language (Kress, 1993). The content and structure of a genre emerge from the social process of communicating for a functional purpose that is meaningful within the context of use. The work of functional genre theorists who identified common patterned responses and common linguistic features associated with certain genres (e.g. Martin, Christie & Rothery, 1987) informed the design of asTTle writing. A written text can be seen from two perspectives: “a thing in itself that can be recorded, analysed and discussed, and also a process that is the outcome of a socially produced occasion” (Knapp & Watkins, 2005, p. 13). The notion of social context adds a further layer of complexity. Students are expected to navigate a complex semiotic world whereby different discourses (Gee, 1996) operate in different contexts, like the characteristics of writing that are seen in particular disciplines within schooling or academic writing more generally. A piece of writing that describes, organises, and classifies information in order to report could, arguably, look different in science compared with social studies; a piece that aims to evaluate a piece of literature in English class may be quite different to an evaluative review written in a history class. These ideas are significant when teachers consider engaging students in writing through communicative purposes that are meaningful and authentic.
To take account of these ideas while making the task of designing rubrics manageable, writing was viewed through a curriculum lens as serving six (for Levels 1–4 of the curriculum) and seven (for Levels 5–8) major communicative purposes that encapsulate what the text is doing (Knapp & Watkins, 1994; Knapp & Watkins, 2005). For each of the major purposes, an analytic rubric was developed. Descriptions of features and text structures commonly associated with a generic communicative purpose were used to inform the rubric criteria for each communicative purpose for the dimensions of audience awareness and purpose, content and ideas, and structure and language resources, while the criteria for grammar, spelling, and punctuation were the same across the different communicative purposes. The criteria were differentiated by levels of the national curriculum, thus showing a progression. Development may be patterned differently across students at the same overall level and individual students may exhibit strengths and weaknesses in the different dimensions of writing as well as across purposes for writing.
The recent revision (NZCER, 2012) is a departure with a move to a one-size rubric that fits all communicative purposes. This suggests a view of writing as a generic set of cognitive and linguistic skills; that the features of writing contained in the scoring rubric criteria apply equally, for example, to writing that aims to argue or persuade and to writing that has as its primary purpose to classify, organise, and describe in order to report. And, when criteria describing features are mastered in one context, the implication could be that the learning is readily transferable. In reducing complexity and scope in a number of ways, as will be noted following, the revision does not recognise and make explicit through criteria the particular ways in which language commonly works to meet each different social communicative purpose. It thus does not acknowledge that the content, structure and features of a text arise from the socially contextualised process of communicating for a particular purpose.
A design-implementation perspective
The revised asTTle: Writing (NZCER, 2012) appears to differ from the original, as discussed above, regarding the underpinning theoretical premises. But it also differs in the approach to development and implementation. We consider whose judgement and how this is obtained; the scope of judgements; how judgements are processed; and the implications of these design decisions.
asTTle:Writing (versions 1–4)
Underpinning the design of the original asTTle was a strong belief in the professional judgement of teachers and their ability to develop understanding of constructs and make judgements collegially. The scoring mechanism was a best-fit judgement process in which the marker evaluated the writing using a progress map or rubric aligned to Levels 2 to 6 of the curriculum. The rubric defined the characteristics of the Proficient range within each level. Markers were to first ascertain which level fit best and then decide whether to adjust their curriculum level decision to Basic, if the work was close to, but weaker than Proficient, or to Advanced, if the work was close to but stronger than Proficient. Each piece of work earned seven potentially different scores; the approach was based on the possibility that students might not have a flat, consistent profile across all attributes. The intention of this approach to diagnostic and analytic marking was to provide information to improve the pedagogical responses.
Considerable efforts were made to ensure the writing would be evaluated in a way that maximised consistent scoring across teachers and schools. The asTTle v4 manual reminded teachers that “Scoring is against criteria and the piece of writing actually written; not against the teacher’s prior knowledge of the student or what the student might have done in different circumstances” (Hattie et al., 2004, p. 59). The marking process was supplemented with annotated exemplars which helped illustrate characteristics of each level (Hattie et al., 2004, Appendix). There was strong evidence in research reports (Brown, Irving, & Sussex, 2004; Glasswell & Brown, 2003) and in a published peer-reviewed article (Brown, Glasswell, & Harland, 2004) of the consistency with which New Zealand teachers, trained or untrained, could use the rubrics to reach similar ratings of common scripts.
A professional resource was developed for in-school use on how to conduct moderation (Carlisle, Absolum, Brown, Irving, & Hattie, 2005), based on the methods used in the normative marking of asTTle. It emphasised the importance of regular feedback and discussion within schools as the basis of reaching consensus prior to the dissemination of scores to students and families. These routines were intended to ensure a coherent vision and a plausible basis for confidence that school-based practices were being conducted in a consistent manner. The challenge of ensuring appropriate operational use of the asTTle writing rubrics became an integral part of the regional Assess to Learn school advisory teams and other professional development projects such as the Literacy Professional Development Project (LPDP) (Parr et al., 2007).
The scoring of asTTle: Writing involved a guided, best-fit judgement of the student’s level of performance in relation to the National Curriculum Framework. This is unlike the asTTle and e-asTTle mathematics and reading tests which used item-response theory to estimate a score based on the relative difficulty of items. The asTTle: Writing tasks were cross-sectional samples and no child did more than one task. This means that the norm sample scoring could not estimate repeated performances by the same child (i.e., test–retest reliability). Only the consistency of markers to each other when they scored matching items could be used to establish the consistency of scoring. The evidence of the moderated norms writing panels was that the marking of the norms scripts had been done in a consistent and dependable fashion. That is, each marker was cross-checked each hour, and each day began with refresher training with exemplar scripts for which agreed scores existed.
However, it could be seen from the norms scoring that it was easier to reach a higher curriculum level depending on the specific task or genre (or both) being written. The degree of difference between tasks was not great, consistent with other research which indicates that the actual task or prompt has only a moderate impact on the quality of scoring (Brennan, 1996). While some might consider that this differential difficulty invalidates the scoring, the asTTle v4 team position was that the curriculum-based rubric, combined with moderation processes, led to accurate assignment of scores to student performances. Hence, it was concluded that the level assigned was accurate and that, as in the scoring of reading and mathematics items, some tasks were easier than others. This is not a case of unreliability, but rather a confirmation that some factors in the task design or the random assignment of tasks to students led to greater or lesser performance, and that this had been accurately detected by the markers.
The asTTle development team rejected the idea that the assignment of a level by a trained and moderated marker should be modified by a statistical approach to account for the relative ease or difficulty of the task being scored. Such adjustment was seen as incompatible within a framework of teachers making judgements against a scoring rubric with clear descriptors. This is not a case of ignoring reliability in favour of validity; rather it was a commitment to the professionalism of teachers who had demonstrated ability to score in a consistent and dependable fashion with sufficient training, or moderation, or both. Quality assurance was achieved by judgement moderation by the teaching community rather than by statistical manipulation. The original asTTle: Writing relied on teacher judgement and professional dialogue directly around assigning a curriculum level, and the design, process, and output aimed to provide teachers with optimal information.
The revised asTTle: Writing
With the most recent revision, a significant number of changes have been made to the design of the scoring rubrics, to the marking process, and to the statistical processing of scores (NZCER, 2012). Instead of seven purposes of writing (i.e., Persuade, Instruct, Narrate, Describe, Explain, Recount, and Analyse) only five purposes (i.e., Describe, Narrate, Recount, Explain, and Persuade) are addressed. There is no clear explanation given for why Instruct and Analyse have been removed, leaving the teacher to wonder if these are no longer important curriculum goals. Additionally, there are just four prompts for each purpose, instead of the five to nine tasks for each purpose in asTTle v4. This means a much smaller range of writing opportunities in the bank of available prompts, raising questions about whether the number is sufficient to capture generalisable information about student ability to write for a purpose. As Brennan (1996) has made clear, five or more samples of writing are needed to make strong claims about the quality of performance.
The revision involves using a rubric that is no longer customised to the different purposes of written communication. Instead, a single common rubric applies. This approach treats all functions and purposes of writing in the bank as fundamentally equivalent, a matter of some considerable simplification relative to the purpose-based scoring rubrics in asTTle: Writing v4. Although the revision still targets ostensibly the same seven dimensions (now elements) of writing that the original design focused on, the scope and interpretations of the dimensions have narrowed. The notion of “addressing the audience, awareness of the rhetorical context as a dimension” appears not to be foregrounded; it has become structure and language. Grammar has become sentence structure, a much narrower focus than asTTle: Writing v4 understanding that grammar “refers to accepted patterns in language use … refer[s] to aspects of grammar such as subject-verb agreement, the use of complete verbs/verb groups, and the appropriate and consistent use of tense-choices for verbs” (Hattie et al., 2004, Appendix, p. 54). Language resources has become vocabulary, which is a much narrower interpretation. Language resources reflects:
three main considerations (1) What are we writing about? (content influences vocabulary, idioms or phrases). (2) What is our purpose? (language choices and grammatical structures that are associated with a desire to argue, to entertain, to instruct, etc.) (3) Who are we writing for? (language choice and grammatical choices that acknowledge different ways of addressing our parents, our friends, the teacher, the principal, etc. (Hattie et al., 2004, Appendix, pp. 3–4)
By contrast, vocabulary is glossed as “The range, precision and effectiveness of word choices appropriate to the topic” (NZCER, 2012, p. 10). These revisions to the range of purposes and to the dimensions-cum-elements weaken the ability of teachers to investigate student writing performance and the opportunities for students to demonstrate their abilities.
Other alterations to design features arguably detract from the tool’s ability to provide stable and robust estimates of the population against which students are benchmarked. While empirical evidence is a good way to test the design of a scoring rubric, it is a much weaker way to develop a new method since it is dependent on the sample of work available. In this case a much smaller and less representative sample was used than in asTTle V4. The new single rubric was developed from 4,755 pieces of student work, collected as part of the norming process for the 20 new writing prompts. This sample is just 23 percent of the norming sample of the 20,838 pieces used to develop the asTTle V4 norms and rubrics. Since it was the objective of the government in funding the development of asTTle that robust norms for many possible combinations were available, large scale representative sampling was undertaken and documented in Chapter 4 of the asTTle V4 manual (Hattie et al., 2004). In writing, 20838 students were surveyed which means that, on average, 347 students completed each writing task, compared to the average 237 for the current revised asTTle: Writing sample size per item. Standard errors of means are smaller with larger samples and thus greater accuracy in discriminating between candidates is supported.
Samples also need to be sufficiently representative of the various sub-populations for which norms need to be reported. Teachers and school leaders are able to use the asTTle reporting engine to query their own students’ performance against norms that should allow for the interaction of two demographic variables (e.g., between sex and ethnicity). Assuming the revised asTTle norms accurately reflect the population distributions, the current norming sample of less than 250 candidates per item would have a margin of error of approximately 6 percent. The asTTle v4 sample size, without any oversampling, would have been at least 887 candidates in the same cell generating a margin of error of only 3 percent. Significantly, in the revised version, interaction norms for the Pasifika and other ethnic groups by sex would be even smaller with much larger margins of error. Hence, the revised e-asTTle norms have larger margins of error than the previous asTTle writing norms and may not actually be as representative of student characteristics, as shown in Table 4.4 of the asTTle v4 manual (Hattie et al., 2004). The sample cells for gender by year and by ethnicity were so small that a linear regression model was needed to create norms (NZCER, 2012); as opposed to the large sampling of asTTle V4 which ensured statistical adjustments were not needed to impute norms for all the cells that are possible in e-asTTle reporting. So, while it might be argued that the much smaller sampling of the revised asTTle writing is sufficient to establish reliability, it is questionable whether it is large enough to establish comprehensive norms for all 20 prompts. The implication is that the benchmarks against which the schools compare their own students’ performance are not as stable and robust estimates of the population performance as the original asTTle v4 norms.
Instead of assigning curriculum levels directly, the new writing rubric assigns R (rubric) scores ranging from R1 to R6 or R7 depending on the element being marked. Once the teacher has assigned an R score for each element of writing the “online e-asTTle application is able to convert the rubric scores to scores on an e-asTTle writing scale and subsequently to curriculum levels, and then to produce a range of reporting at the individual and group level” (NZCER, 2012, p. 5). The updated linking results study makes this separation of the curriculum from the scoring of writing by teachers explicit, although no theoretical or practical reason seems to be given, other than the implication that within the previous system teachers were not using observable characteristics of writing to reach a curriculum level decision.
Rather than first placing a student’s writing within a broad curriculum level then considering finer distinctions within this (i.e., the three-level Basic, Proficient, Advanced—BPA—system of the previous version of e-asTTle), the current rubric allows for a consideration of observable skill levels within each element. This consideration is independent of the curriculum. It is the student’s overall performance that is compared with the curriculum (e-asTTle, 2012, p. 7). This process of first assigning an R score and allowing the computer to impute a curriculum level seems to remove the teacher even more from his or her ability as a professional to judge student work relative to curriculum levels, the very thing teachers are meant to be teaching toward, marking against, and reporting on.
The 2012 manual correctly indicates that the difference between R1 and R2 is not necessarily identical between any other pair of scores in any of the seven scales. The manual also points out that some prompts are easier than others, meaning that there is a greater probability of a higher score for some prompts. Figure 4 in the manual (NZCER, 2012, p. 29) shows that student scores ranged from 684 (39 students) to 1782 (39 students), prompts ranged on average from 1416 (six prompts) to 1538 (six prompts), markers ranged on average from 1416 (one marker) to 1538 (three markers), and elements ranged from 1355 (Spelling) to 1520 (Organisation, Sentence Structure, Punctuation, and Structure and Language). The most important inference from these data is that score distinctions attributable to prompt or element are not small relative to the standard error, which is 40 points at the mid-point of the range and closer to 100 points at the lowest and highest ends of the scale (Ministry of Education, 2013). Since two standard errors represent a 95 percent confidence interval, differences greater than that are considered to be beyond chance. Hence, the difference attributable to these various facets at the middle of the range is greater than would normally occur by chance; although, at the tails of the distribution (i.e., the weakest and strongest writers) these facets probably make no difference to a student’s score.
To account for this variability in the probability of assigning an R score, the current, revised e-asTTle: Writing uses an item-response theory transformation (i.e., Multifacet Rasch Modeling, MFRM) to calibrate scores on a unidimensional latent-trait scale which captures simultaneously the ability of the student, the difficulty of the elements, the difficulty of the prompts, and the harshness of the marker. The Rasch approach transforms monotonically “the ‘dependent’ variable” (the probability that a certain item response will occur) as “an additive function of two ‘independent’ variables, namely person ability and item difficulty” (Borsboom & Zand Scholten, 2008, p. 112). While some advocates of Rasch procedures insist that only this dual-calibration (i.e., ability with difficulty) provides true measurement, other item-response theory researchers consider that this approach is unnecessarily narrow in its presumption that all discrimination indices must be equal (de Gruijter & van der Kamp, 2003) or that latent-trait theories may provide a better description and explanation of item responses (Borsboom, 2005).
While the MFRM approach can determine the contributing variance of scores due to task, purpose, element, and marker in the norming sample, this is not the only statistical approach to this problem. Generalisability theory (g theory) can also apportion variance to score, rater, error, and interaction components; and thus provides a robust indicator of degree of agreement attributable to the similarity of raters’ scores (Shavelson & Webb, 1991). Within g theory, the Brennan and Kane dependability index (φ) was calculated for the asTTle v4 panel rating scores (Brown, Glasswell & Harland, 2004); values ranged from .67 to .95 with an average of .77, indicating that the scores assigned to the norm scripts marked this way were dependably scored and could be used as benchmark scripts for subsequent curriculum-level-based marking. Hence, we conclude that the use of MFRM is not necessarily the best approach to handling teacher ratings, nor are we convinced that the introduction of R scores as opposed to direct ratings of curriculum levels adds any superior accuracy or validity to teacher rating of student writing. This whole mechanism was unnecessary when teachers directly assigned curriculum levels, guided by rubrics, exemplars, and moderation.
A further weakness of the Ministry’s R score system can be seen in the Ministry of Education’s (2013) own document that guides teachers who want to use the new e-asTTle writing rubrics without the e-asTTle software. Teachers are advised to find the sum of all seven attributes (ranging from 7 to a maximum of 44) to infer an e-asTTle: Writing score and curriculum level. This approach treats all seven elements of writing, and all prompts, as if they were equally difficult and all score thresholds as if they were equally probable—with no adjustments for difficulty. This is at odds with the approaches advocated by previous asTTle developers. It makes no sense to take a multidimensional rating system that is designed to account for multiple elements of writing and reduce it to a single score. We consider that the path to this practice may have been inadvertently laid in the revised writing approach in creating a numeric rating scale rather than persisting with direct judgement against curriculum-level descriptors. Conceivably, teachers could take the path of least resistance and eschew the computer to manually calculate e-asTTle: Writing scores and levels. The authors of the revision claim that an “e-asTTle result is not sufficient to determine an overall teacher judgement” (NZCER, 2012, p. 22), even though e-asTTle results were drawn from teacher judgements. A more defensible claim would be that evidence from a single piece of writing is insufficient to make an overall teacher judgement about where a student’s performance lies relative to the expectations of a grade; a point eloquently and frequently made in research into performance assessment (Brennan, 1996).
A practice perspective
Teachers, as professionals, have a major role in curriculum implementation to meet the needs of the students in front of them. Unfortunately, curriculum documents at the time of the original development of asTTle: Writing gave a murky portrayal of writing, particularly regarding the relationship between form and function. In this context, the original asTTle: Writing was designed to help build teacher content knowledge about writing. Internationally, it is suggested that teachers’ knowledge of language is important but is lacking (Jones & Chen, 2012; Wong-Fillimore & Snow, 2002). In general, teachers have lacked a shared language for talking about writing. A recent survey of New Zealand primary teachers (Parr & Jesson, 2015) suggests there are ongoing teacher concerns—largely related to pre-service preparation—with content and pedagogical knowledge in writing, although in-service support in relation to teaching writing is more favourably perceived.
A major consideration in designing the original scoring rubrics was to provide a tool to aid teacher learning about writing. The theoretical notion of a tool derives from both Vygotskian and cognitive theory. Tools are considered to be externalised representations of ideas that people use in their practice (Norman, 1988; Spillane, Reiser, & Reimer, 2002). When the ideas represented in the tool are valid and represented in a quality way, then the tool serves to incorporate sound theory about how to achieve the purpose of the task in question (often referred to as a smart tool). In the case of writing, the design of scoring rubrics should incorporate sound theory about achieving the writing purpose. The original version of asTTle: Writing was designed to alert teachers to the communicative purposes encountered within curricula in compulsory schooling (and beyond) and illustrate the ways in which language commonly works to achieve each of these communicative purposes.
A tool is significant in promoting learning partly as a function of how it is integrated into the routines of practice (Wenger, 1998). A characteristic of practice as a source of community coherence is the development of shared repertoires; the community’s set of shared resources created and used to negotiate meaning. In writing, an example of a routine might be the processes followed by a group of teachers to moderate writing samples, considering how they meet the criteria for quality performance at any given level, as described in a sound rubric. Informed collegial discussion helps teachers to negotiate meaning regarding constructs of quality writing (Sadler, 1989).
In New Zealand, there are several key tools that could be regarded as resources to help negotiate meaning in writing that have become part of the repertoire of teachers. These include the Assessment Tools for Teaching and Learning (asTTle) (Hattie et al., 2004), as a detailed diagnostic tool for assessing writing. More recently, the English Literacy Learning Progressions or LLP (Ministry of Education, 2010) have been a source of considerable learning for teachers (see Parr, 2011).
The original asTTle: Writing was designed primarily to provide information to empower teachers with knowledge of their students’ patterns of performance in writing and to allow them to compare, with confidence, their students against similar others. Our discussion questions the extent to which changes in the asTTle: Writing revision continue to support diagnostic purposes. The original tool provided a detailed, diagnostic profile of student writing in addition to a summary performance score or curriculum level, and we question whether the narrowing of purposes and of dimensions, coupled with a one-size-fits-all rubric, provides such detail. In writing there is considerable intra-individual variability in performance related to features of the context, including content knowledge of topic and the different purposes for writing. Students need multiple and varied opportunities to demonstrate what they know. A theoretically sound and detailed series of profiles showing the strengths and gaps in performance in different contexts is necessary for assessment for learning purposes.
The act of interpreting the data obtained from an assessment is only the first part of their use in a complex teaching act. While building knowledge to interpret data from assessment tools is relatively straightforward, working to apply to practice is much more problematic (Parr & Timperley, 2008). Detailed diagnostic data of student profiles in writing better inform the process of making decisions about what actions to take to meet learning needs in writing, but considerable knowledge of writing is also needed. The latter involves being explicit about the way language works to make meaning and how that meaning is socially constructed, together with pedagogical knowledge that includes a disposition to seek out research evidence of what has been shown to work, in similar situations. Such expertise would, arguably, characterise the inquiring teacher referred to in The New Zealand Curriculum (Ministry of Education, 2007). There is evidence from a major national professional learning project, the Literacy Professional Learning Project 2004–2010 (see for example Parr, Timperley, Reddish, Jesson, & Adams, 2007; Timperley, Parr, & Meissel, 2010) of considerably accelerated progress in writing where asTTle: Writing was adopted, initially, in the first cohort as a new assessment tool for participants. Significantly, the greatest progress in writing for each of the three 2-year cohorts of schools was made in the first year of the professional development. This was the period when teachers developed understanding, as part of an analysis of learning needs (theirs and their students), about concepts of writing as they used the asTTle: Writing rubrics and engaged in facilitated moderation procedures to obtain reliable writing data. From this, discussion was facilitated by experts regarding what the data showed and what it meant for honing their teaching and optimising student learning. The writing assessment tool was a key resource in supporting facilitated teacher learning and enhanced practice.
While the aims of assessing student writing involve using information to better tailor support to move students forward, learning from assessment also involves the writers themselves. For students to drive their own learning, they need an understanding of the qualities of the desired performance, where they are currently in relation to that performance, and what they have to do to achieve it (Hattie & Timperley, 2007). To assist these understandings involves providing quality feedback to student writers (Parr & Timperley, 2010) and also addressing interpretive challenges in the use of feedback. Such challenges arise because students need to understand the criteria; they also need sufficient working knowledge of some fundamental concepts, the type of working knowledge that knowledgeable, expert teachers draw on when they compose feedback (Sadler, 2010). This learning helps build their own internal self-regulatory mechanisms and enhance their evaluative insights as well as their writing. There is evidence that teachers were able to use the original rubrics to produce student-friendly versions to aid peer and self assessment (Parr & Jesson, 2011).
However, teacher moves and the artefacts in the environment have to provide consistent messages to students. Teachers, through various means, some explicit, some implicit, give students messages about understandings of the purposes of writing; of what is valued in writing; indications of our notions of quality, and glimpses into teacher tacit understandings. There is potential for confusion where there is a lack of alignment between what is valued in writing, and various pedagogical moves, particularly in regard to assessment. Assessments provide powerful messages to students of the theory held of the nature of writing, of what is valued in writing, and what are considered the hallmarks of a competent writer. This is not a new idea. In The Testing Trap, George Hillocks (2002) pointed out that assessments influence what happens in classrooms in terms of rhetorical stance, instructional mode, and writing process; they privilege curricula content that appears related to the assessment—such as particular “forms” of writing—and may narrow the definition of writing held by students. The nature of assessments, like the revision of e-asTTle, could downplay the social- communicative intent of writing, the writer–reader interaction, and potentially send mixed or even undesirable messages about what is valued and constitutes quality writing.
Summary
Writing is a multidisciplinary, and also a much contested, field of study. In the scholarship on the evaluation of writing, the making of largely evaluative, end-point decisions have been shown to be a shaping influence in the field and issues in the assessment of writing have, perhaps as a result, often been framed technically (Huot & Neal, 2006). Most significantly, research has seldom considered how assessment might link to teaching and learning in writing. The nature of writing and how proficiency develops in different contexts are crucially part of diagnostic, formative assessment. We assert that a rubric intended to encompass all social-communicative purposes for writing, with narrower definitions of dimensions like audience and language use and with greater weight to technical aspects of writing, is a less rich source of diagnostic information for teachers, their students and other stakeholders. The new version potentially aligns to a more product and skills-oriented view of writing; the design and implementation represents a swing back to what Huot (2006) called a technicist view of writing assessment and, with a one-size-fits all rubric, suggests a view that writing is a generic skill able to be transferred across different functions and contexts.
Implications of the revised version also relate to messages about what learning is valued. Ideas about writing that assessments validate are, we argue, as important as the assessment formats, procedures, and technical properties. The new asTTle: Writing assessment, where a model of writing is not consciously articulated and where the notion that language works in different ways to meet varied communicative purposes seems to have all but disappeared, and where there is greater emphasis on technical elements of writing, potentially contributes to a shift in ideas about what is important in teaching and learning to write.
We contend the revised assessment offers less to teachers in relation to professional learning opportunities around developing knowledge about writing and how language works to meet varied communicative purposes; about the nature of quality performance at different curriculum levels, and about the development of their student writers. A narrowing of the nature of writing (through purposes and elements) means teachers have a less-nuanced diagnostic profile for individuals and groups from which to devise targeted teaching. The issue of a less-detailed diagnosis, and perhaps therefore a greater tendency to use the total score as a summative measure, is in contrast to the original rationale for the asTTle suite.
As asserted earlier, research views writing as a highly complex socio-cognitive act and our curriculum assigns it a key and dual role in learning. Teachers need the best possible resources to enable them to prepare students for this complexity. Our discussion has been intended to question whether the concept of writing represented in the revised asTTle: Writing (2012) assessment is sound and imparts a view consistent with current thinking and practice; to question aspects of implementation, including the distancing of professional judgement, are warranted and also to question whether the revision optimally supports the broader teaching and learning functions of assessment tools such as asTTle within the curriculum, including whether it remains a useful tool for teacher content learning about writing. We suggest the current revision falls short and issue a challenge for informed, systematic research to show otherwise.
References
Almargot, D., & Fahol, M. (2009). Modeling the development of written composition. In R. Beard, D. Myhill, M. Nystrand & J. Riley (Eds.) Handbook of writing development (pp. 23–47). Thousand Oaks, CA: Sage. http://dx.doi.org/10.4135/9780857021069.n3
Alderson, J. C., & Wall, D. (1993). Does washback exist? Applied Linguistics, 14, 115–129. http://dx.doi.org/10.1093/applin/14.2.115
Borsboom, D. (2005). Measuring the mind: Conceptual issues in contemporary psychometrics. Cambridge, UK: Cambridge University Press. http://dx.doi.org/10.1017/CBO9780511490026
Borsboom, D., & Zand Scholten, A. (2008). The Rasch model and conjoint measurement theory from the perspective of psychometrics. Theory & Psychology, 18, 111–117. http://dx.doi.org/10.1177/0959354307086925
Brennan, R. L. (1996). Generalizability of performance assessments. In G. W. Phillips (Ed.), Technical issues in large-scale performance assessment (NCES 96-802) (pp. 19–58). Washington, DC: National Center for Education Statistics.
Brown, G. T. L., Glasswell, K., & Harland, D. (2004). Accuracy in the scoring of writing: Studies of reliability and validity using a New Zealand writing assessment system. Assessing Writing, 9(2), 105–121. http://dx.doi.org/10.1016/j.asw.2004.07.001
Brown, G.T.L., Irving, S.E. & Sussex, K. (2004, August). Scoring of a nationally representative sample of student writing at Level 5–6 of the New Zealand English curriculum: Use and refinement of the asTTle progress indicators & tasks (asTTle Tech. Rep. #48). Auckland: University of Auckland/Ministry of Education.
Carlisle, A., Absolum, M., Brown, G. T. L., Irving, S. E., & Hattie, J. A. C. (2005). Moderating asTTle writing (asTTle Tutorial). Retrieved from http://www.breezeserver.co.nz/p11095800/
de Gruijter, D. N. M., & van der Kamp, L. J. T. (2003). Statistical test theory for education and psychology. Leiden, The Netherlands: Graduate School of Education, Universiteit Leiden.
e-asTTle. (2012). e-asTTle linking study. Wellington: Ministry of Education.
Gee, J. P. (1996). Social linguistics and literacies: Ideology in discourses (2nd ed.). London, England: Taylor-Francis.
Glasswell, K., & Brown, G. T. L. (2003, August). Accuracy in the scoring of writing: Study in large-scale scoring of asTTle writing assessments (asTTle Tech. Rep. #26). Auckland: University of Auckland/Ministry of Education.
Glasswell, K., Parr, J., & Aikman, M. (2001). Development of the asTTle writing assessment rubrics for scoring extended writing tasks (asTTle Tech. Rep. #6). University of Auckland/Ministry of Education.
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. http://dx.doi.org/10.3102/003465430298487
Hattie, J. A. C., Brown, G. T. L., Keegan, P. J., MacKay, A. J., Irving, S. E., Cutforth, S., … Yu, J. (2004, December). Assessment tools for teaching and learning (asTTle) manual (Version 4, 2005). Wellington: University of Auckland/ Ministry of Education.
Hillocks, G. (2002). The testing trap: How state writing assessments control learning. New York, NY: Teachers College Press.
Huot, B., & Neal, M (2006). Writing assessment: A techno-history. In C.A. MacArthur, S. Graham & J. Fitzgerald (Eds.). Handbook of writing research (pp. 417–432). New York, NY: Guilford Press.
Jones, P., & Chen, H. (2012). Teachers’ knowledge about language: Issues of pedagogy and expertise. Australian Journal of Language and Literacy, 35 (1), 147–168.
Knapp, P., & Watkins, M. (1994). Context, text, grammar: Teaching the genres and grammar of school writing in infant and primary classrooms. Sydney, Australia: Text Productions.
Knapp, P., & Watkins, M. (2005). Genre, text, grammar: Technologies for teaching and assessing writing. Sydney, Australia: University of New South Wales Press.
Kress, G. (1993). Genre as social process. In B. Cope & M. Kalantzis (Eds.), The powers of literacy: A genre approach to teaching writing (pp. 22–37). Pittsburg, PA: University of Pittsburg Press.
Marshall, B. (2004). Goals or horizons—The conundrum of progression in English: Or a possible way of understanding formative assessment in English. The Curriculum Journal, 15, 101–113. http://dx.doi.org/10.1080/0958517042000226784
Martin, J. R., Christie, F., & Rothery, J. (1987). Social processes in education. In: I. Reid (Ed.), The place of genre in learning: Current debates. Geelong, Australia: Deakin University Press.
Ministry of Education (2007). The New Zealand curriculum. Wellington: Learning Media.
Ministry of Education. (2009). The New Zealand curriculum reading and writing standards for Years 1–8. Wellington: Learning Media.
Ministry of Education. (2010). The literacy learning progressions: Meeting the reading and writing demands of the curriculum. Wellington: Learning Media.
Ministry of Education and New Zealand Council for Educational Research. (2012). e-asTTle: Writing (revised). Retrieved from http://e-asttle.tki.org.nz/Teacher-resources/PLD-resources-for-e-asTTle-writing
Ministry of Education. (2013, April). The e-asTTle writing score conversion table. PLD resources for e-asTTle writing (revised). Retrieved from http://bit.ly/1F86Zf0
National Research Council. (2001). Knowing what students know: The science and design of educational assessment. Washington, DC: National Academy Press.
New Zealand Council for Educational Research (NZCER). (2012). e-asTTle writing (revised) manual. Wellington: Ministry of Education.
Norman, D. (1988). The psychology of everyday things. New York, NY: Basic Books.
Parr, J. M. (2011). Repertoires to scaffold teacher learning and practice in assessment of writing. Assessing Writing, 16, 32–48. http://dx.doi.org/10.1016/j.asw.2010.11.002
Parr, J. M., & Jesson, R. (2011). If students are not succeeding as writers, teach them to self assess using a rubric. In D. Lapp & B. Moss (Eds.) Exemplary instruction in the middle grades (pp. 174–188). New York, NY: Guilford Press.
Parr, J. M., & Jesson, R. (2015). Mapping the landscape of writing instruction in New Zealand primary school classrooms. Reading and Writing. http://dx.doi.org/10.1007/s11145-015-9589-5
Parr, J. M., & Timperley, H. (2008). Teachers, schools and using evidence: Considerations of preparedness. Assessment in Education: Principles, Policy and Practice, 15, 57–71.
Parr, J. M., & Timperley, H. (2010). Feedback to writing, assessment for teaching and learning and student progress. Assessing Writing, 15, 68–85.
Parr, J. M., Timperley, H., Reddish, P., Jesson, R., & Adams, R. (2007). Literacy professional development project: Identifying effective teaching and professional development practices for enhanced student learning (851). Wellington: Learning Media.
Sadler, R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18, 119–144. http://dx.doi.org/10.1007/BF00117714
Sadler, R. (2010). Beyond feedback: Developing student capability in complex appraisal. Assessment and Evaluation, 35, 535–550. http://dx.doi.org/10.1080/02602930903541015
Shavelson, R. J., & Webb, N. M. (1991). Generalizability theory: A primer. Newbury Park, CA: Sage.
Spillane, J. P., Reiser, B. J., & Reimer, T. (2002). Policy implementation and cognition: Reframing and refocusing implementation research. Review of Educational Research, 72, 387–431. http://dx.doi.org/10.3102/00346543072003387
Timperley, H., Parr, J. M., & Meissel, K. (2010). Making a difference to student achievement in literacy: Final research report on the literacy professional development project. Report to Learning Media and the Ministry of Education. Auckland: UniServices, University of Auckland.
Wenger, E. (1998). Communities of practice: Learning, meaning and identity. New York, NY: Cambridge University Press. http://dx.doi.org/10.1017/CBO9780511803932
Witte, S. (1992). Context, text, intertext: Towards a constructivist semiotic of writing. Written Communication, 9, 237–308. http://dx.doi.org/10.1177/0741088392009002003
Wong-Fillimore, L, & Snow, C. (2002). What teachers need to know about language. In C.T. Adger, C. E. Snow & D. Christian (Eds.). What teachers need to know about language (pp. 7–54). McHenry, ILL: Delta Systems.
The authors
Judy Parr is a professor in the Faculty of Education and Social Work, the University of Auckland. Judy’s research focuses on enhancing teacher practice and raising student achievement in literacy. She has particular expertise in writing, encompassing how writing develops and is assessed and considerations of instructional issues like teacher knowledge and practice.
Email: jm.parr@auckland.ac.nz
Gavin Brown is an associate professor in the Faculty of Education and Social Work, the University of Auckland. His expertise is in educational assessment and measurement as well as the psychological study of how assessment affects teachers and students.