The untold story of assessment
Jim Neyland
Abstract
If one were to name one year as the year that assessment was invented, that year would be 1980. This proposition is not fully proven here, but a plausible argument is presented in its defence.
Introduction
What is the difference between very good teaching and assessment? This is a defining question. The answering of it separates people into two groups. For some, the answer is obvious. For others—and I include myself in this group—it is impossible to answer. Let me illustrate why. The best swimmers have what is referred to as “a feel for the water”. What is the difference between a good swimmer and someone who has a feel for the water? The two cannot be separated; they are part of each other. A top netballer or rugby player can “read a game”. What is the difference between a top player and someone who can read a game? Impossible to answer. Good teaching is similar. Good teaching does not include assessment as a subcategory. Nor is good teaching complemented by assessment. Good teaching and assessment cannot be distinguished or prised apart. The moment one begins to think that assessment is a category of its own, one no longer has good teaching. Throughout the history of education no one seriously thought of tearing apart teaching and assessment; no one, that is, until 1980.
* * *
Most readers will be able to remember the time before the word “cellphone” was used. The arrival of this word accompanied the arrival of a particular device, and named it. The expression “rap music” accompanied the emergence of a new style of music, and before this time it was unused. It is not hard to think of other examples where new words became current. But it is harder to bring to mind other types of new word, especially when the new word is an abstract concept and not identifiable with a new device or event. There are, however, notable examples of the latter.
For example, it is hard for us to conceive that the word “risk” is a relatively new one. In fact it can truly be called a modern word. The word “risk”, according to the sociologist Niklas Luhmann, appeared in English for the first time around the late 17th century during the early-modern enlightenment period. It replaced what was formerly thought of as fate. Having confidence in one’s fate “refers to a more or less taken-for-granted attitude that familiar things will remain stable” (Giddens, 1990, pp. 30–31). Risk, the new idea, refers to a different attitude, born of a growing ability to quantify and control future events using new techniques in mathematical probability and in science. Risk, unlike fate, involves an awareness of the probable contingencies that attend certain anticipated events. We say today, “What are the risks involved in doing such and such?” When we say this we do not mean, “What will our fate be if we do such and such?” Unlike fate, there is an atmosphere of calculability and controllability associated with the notion of risk. It is no coincidence that mathematical probability and fate were invented around the same time.
What we have, then, is not just the birth of a new word, but the birth of a new concept; a new way of thinking, unknown—or should I say, unthought—before that time. The word “sincerity” is another example of a modern word. The philosopher Isaiah Berlin pointed out that this word and the notion associated with it was unknown/unthought before around the turn of the 18th century (Berlin, 1998). Prior to this, people had no notion of sincerity; the idea, for instance, that one might hold an untrue belief, but, despite this, be given some credit for holding it “sincerely”, was unthinkable. In earlier times such an idea would have been thought absurd: there was only trueness; sincerity had no meaning. It would have been viewed in much the same way we would view someone today saying: “She believes that two plus two equals five, but this is not so bad, because she believes it sincerely.” There are many other examples of words and concepts, taken for granted today, that were unthought prior to the dawn of modernity.
The untold story of assessment
Can you remember the first time you heard the word “assessment” as it is commonly used today? I can, probably because it was accompanied by some embarrassment on my part. It was during the mid-1980s. I was expected to write something meaningful on assessment in education. I had a vague idea what it meant, based on what I could infer from its more general meaning—a mountaineer might make an assessment of the viability of a new route up a mountain, for example. But despite the fact that I was reasonably well read, I could not remember coming across anything on assessment in education. Perhaps I was reading the wrong books!
When I did try to make sense of this idea, I struggled. I could not see the point. I had been taught at school by some superb teachers, who seemed not to need this concept. My father had been a successful teacher, and, nearing retirement at the time, was unsure what assessment meant beyond the obvious extrapolation from the more general meaning. And I had taught at the school level for a decade—with some accomplishment, I thought—without needing this idea. I thought in terms of being a good teacher; not in terms of being a teacher and an assessor. Even the best of what is now called assessment does not seem to me to be as good as what I mean now, and meant then, by “good teacher”.
When I write of assessment I am not referring to the notions of testing or objective measurement of achievement. These last two have been evident since at least 1845. In that year, written examinations were used in Boston. The Reverend G. Fischer, in 1864, used “objective” measures of achievement and compared student work to “standard” specimens. During the psychometric period in the early 1900s, tests of general intelligence and achievement tests were first used (Webb, 1992). Assessment does include something of the notions of testing and objective measurement, much in the way that risk has something of fate and sincerity something of trueness associated with them. But, like the distinctions between fate and risk, and trueness and sincerity, there is a difference between objective testing and assessment. In fact, there is something of the earlier-mentioned change in attitude from fate/trueness to risk/sincerity involved in the shift from objective testing to assessment.
Objective testing was of the student and was periodic. Associated with the earlier idea of objective testing is the conviction that what was being measured was an inherent and stable capacity of the individual, and that this capacity, because it was stable, could be known with a high degree of certainty. Similarly, the knowledge tested was taken as having an unquestionably real value as worthwhile knowledge. The atmosphere here was not too dissimilar to that associated with fate and trueness. Assessment is different. Students, teachers, the school, and the education system are all objects of assessment, and assessment is a continual presence. With assessment there is a more humanistic orientation. Capacities are seen as changeable; in fact, education is aimed precisely at changing capacities. This is much more like managing risk than accepting fate. As with sincerity, more allowance is given for personal orientations. There is a shift from the real value of knowledge to its exchange value via credentials and tokens awarded for behaving in certain prescribed ways.
All this is background. How does the concept of assessment function in society? In seeking an answer it seemed sensible to check the indexes and content listings of a random selection of books on education from the university library, and to look for entries based on the root “assess”; we also checked “evaluate”. I have to confess that the random selection was not entirely random, but where it was not, the motive was laudable. I deliberately added into the otherwise random selection books on mathematics education. Why? I have long thought that mathematics education is an indicator subject. Frogs are considered by environmental scientists to be an indicator species. If the climate changes, frogs are often affected first. If frogs begin to die in a region of Brazil, scientists know that it is likely that a delicate balance has been tipped, and they ought to look closely at pollution levels there. How is mathematics education similar? Of all the mainstream subjects, the teaching and learning of mathematics is widely thought to be the easiest to study. It is thought to be the most suited for experimental trials. Accordingly, new ideas are often tested first in mathematics education, and if one wants to find the origin of a new idea, one can do worse than consult mathematics education books.
What did we find? I’ll arrange the titles in publication order for dramatic effect. The following books, published between 1960 and 1970, were selected at random and examined:
•&;&;&;Bruner (1960), The Process of Education
•&;&;&;Musgrave (1965), The Sociology of Education
•&;&;&;Mitchell (Ed.) (1968), New Zealand Education Today
•&;&;&;MacGinitie and Ball (1968), Readings in Psychological Foundations of Education
•&;&;&;Hirst and Peters (1970), The Logic of Education.
Not one of these contains any mention of assessment or evaluation in the contents listings or indexes; not even the book on the psychology of education. However, we were bound to find it in the next book, a well-known text on the psychology of mathematics education:
•&;&;Skemp (1971), The Psychology of Learning Mathematics.
It, too, contains nothing on assessment or evaluation in the contents listing or index.
The following books were selected from the period 1971 to 1980:
•&;&;&;Hooper (Ed.) (1971), The Curriculum: Context, design and development
•&;&;&;Dakin (1973), Education in New Zealand
•&;&;&;Lawton (1973), Social Change, Educational Theory and Curriculum Planning
•&;&;&;Peters (Ed.) (1973), Philosophy of Education
•&;&;&;Tanner and Tanner (1975), Curriculum Development
•&;&;&;Navarick (1979), Principles of Learning.
We did not expect to see anything in the book on the philosophy of education, but there were three that dealt with curriculum, and another with learning. Surely these would contain something on the subject? Apart from Hooper, and Tanner and Tanner, none of these makes mention of assessment or evaluation in the contents listings or indexes. Hooper contains nothing in the contents listing. It has one minor mention in the index for p. 189, where it states: “It will be more usual … for the teacher … to find it necessary to devise a special instrument for the assessment of a particular aspect of the new curriculum. Few teachers have the necessary skills.”
Tanner and Tanner, a comprehensive 743-page treatment of curriculum and related aspects of education, contains only a very minor reference to assessment in the index. It does have a concluding chapter on “Evaluation for curriculum improvement”, which addresses larger questions about whether education is working. There is a subheading “Assessing the effects of education”. Under it are two short paragraphs, which discuss whether or not education is combating social inequality. There is a sub-subsection, “Assessment and accountability” under the subheading “The deficiency of efficiency”. The sub-subsection discusses a new proposal to conduct a “national assessment of education” using tests “constructed by testing agencies” (p. 696). The one mention of assessment in the index says: “see the National Assessment of Educational Progress (NAEP)”. Looking up these references reveals a common theme: a report that proposes that “The National Assessment of Educational Progress should become a bulwark of educational accountability” (p. 516). We will see shortly that these are early references to what I call the birth of assessment; more on this soon. I was going to write that Tanner and Tanner is a small exception that proves the rule. It is not. It is an exception that indicates that something was up; that restructuring forces were moving into position.
During 1980 the evident pattern—that books on education tend to neglect assessment—changed suddenly. There was a quantum shift during that year and it is best illustrated by two books on mathematics education commissioned by the highly influential National Council of Teachers of Mathematics, and published that year. The first, Research in Mathematics Education (Shumway, 1980), was a summary of all important research in mathematics education prior to 1980. It contains nothing on assessment in the contents listing or index. The second book, An Agenda for Action, heralds a change. A significant proportion of this document, which aimed to be a visionary programme for mathematics education, is devoted to assessment. So, we have two books—one retrospective, the other prospective—published by the same prestigious educational institution; the first, a summary of all that was important prior to 1980, the second, indicating what the future ought to be. The first does not mention assessment. The second gives it considerable prominence.
This change is sustained in subsequent publications in both mathematics education and education more generally, and signs of interest can be detected outside of the United States. For instance, the British report, Mathematics Counts: The Cockcroft Report (1982), contains both a chapter on assessment and many references in the index. The Australian book on, yes, the philosophy of education, Kleinig’s (1982) Philosophical Issues in Education, also contains a chapter on assessment and many references in the index. Since 1982 the literature on the subject has continued to expand. It is not being claimed that something equivalent to what we now call assessment did not occur—except in nascent form—prior to 1980, or that words such as “testing”, “examination”, “assessment”, and “evaluation” were not used. It is being argued that during the 1970s a class of practices was cordoned off for unprecedented attention, and named in order to distinguish it as a new class of activities. It is noteworthy that in the next edition of the National Council of Teachers of Mathematics’ summary of research, Handbook of Research on Mathematics Teaching and Learning (Grouws, 1992)—a large-dimensioned, 771-page book with a tiny font—there is a 23-page chapter on assessment, and a 13.5 cm-long listing on assessment in the index.
It is hard to escape the conclusion that 1980 is the beginning of a dramatic new interest in assessment. One could say that 1980 seems to mark the birth of the assessment enterprise. This raises an obvious question. Why? What happened just before 1980 to cause this change? I will argue that around 1980 there was a carving off of assessment from teaching; a bisecting of what was formerly called good teaching; an artificial separation of teaching into instruction and assessment. In another paper (Neyland, 2007a) I discuss the origins of the notion of “diversity”. I argue that before the enlightenment there was diversity, but it was unremarkable and unremarked. No one saw the need to name it as such, or to draw particular attention to it. It was only when diversity came to be seen as the product of educational activity, and therefore could and should be purposefully shaped—this, in enlightenment times, boiled down to attempting to systematically eradicate diversity—that a particular focus was placed on this notion. The same is true of assessment. Before 1980, one could say, assessment existed but it was unremarkable and unremarked. During 1980 it became a focus of attention and subject to legislation.
Some readers may quibble with me saying 1980 marks the birth of the assessment enterprise. There is, after all, evidence of its presence before then. This, of course, is true. Nonetheless, it can be argued that 1980 is the birth date. Consider, for instance, the invention of the moveable-type printing press. Many historians name 1456 as the birth of the press. This was the date that Gutenberg’s Bible was printed.
But the press was up and running by 1450 and a number of small pamphlets were printed between 1450 and 1456. In addition, the two or three decades leading up to 1450 also featured a number of similar moveable-type machines of various kinds. Does all this make it wrong to name 1456 as the birth of the press? if one were to mark the date of its launching, or of its unequivocal passing from concept to reality, 1456 would seem to be that date. Of course a case could be put for 1450, too, but that only strengthens the point I am trying to make. In a similar way I am arguing that if one were to choose one year that best represents the birth of the assessment enterprise, the launching of a new invention, that year is 1980.
A new approach to curriculum
So, what happened during 1980 or just prior to it that led to the turn to assessment? The birth of these notions can be traced to an event in the United States during 1969. During that year the Federal Government instigated a project, the already-mentioned National Assessment of Educational Progress (NAEP), which aimed to: (i) examine achievement in 10 learning areas; (ii) spot changes in levels of achievement over the years; and (iii) apply the implications of these to changes in national education policy. This project resulted in an unprecedented incursion of legislative authority into the heart of education. The legislation in question was based on the principles of Scientific Management Theory (SMT). This legislative turn in education began with the Education Improvement Act, passed in California in the same year. Other legislation, also based on the principles of SMT, was passed in other states in subsequent years (Wise, 1979).
The impact of the American legislation of 1969, and subsequently, was eventually felt beyond those shores. In 1976 the British educationist Malcolm Skilbeck, for example, observed in his Curriculum Design and Development that in addition to the three ideological models outlined in his monograph—classical humanism, progressivism, and reconstructionism—there appeared to be an emerging new model, which he called the “technocratic-bureaucratic ideology” (Skilbeck, 1976). In addition, the influence of the legislative turn in education can be seen in the growing globalisation of an educational adaptation of the language of SMT. Terms such as “competency-based education”, “performance-based education”, and “assessment systems” became ubiquitous.
SMT was created by Fredrick Taylor, a man obsessed by control and driven by a compulsive need to subject almost every aspect of his life to accurate measurement (Morgan, 1986). Wise (1979, ch. 1) argues that SMT has resulted in the following notions becoming commonplace:
•&;&;&;accountability
•&;&;&;planning
•&;&;&;programming
•&;&;&;budget systems (PPBS)
•&;&;&;management-by-objectives (MBO)
•&;&;&;operational analysis
•&;&;&;systems analysis
•&;&;&;programme evaluation and review techniques
•&;&;&;management information systems (MIS)
•&;&;&;management science
•&;&;&;planning models
•&;&;&;cost-benefit analysis
•&;&;&;cost-effectiveness analysis
•&;&;&;economic analysis
•&;&;&;systems engineering
•&;&;&;zero-based budgeting.
Some of these have been adapted for education purposes, and most of the following are now familiar:
•&;&;&;competency-based education (CBE)
•&;&;&;performance-based education (PBE)
•&;&;&;competency-based teacher education (CBTE)
•&;&;&;assessment systems
•&;&;&;programme evaluation
•&;&;&;learner verification (textbooks have to be proven to increase student achievement before they can be sold)
•&;&;&;behavioural objectives
•&;&;&;mastery learning
•&;&;&;criterion-referenced testing
•&;&;&;outcomes-based education (OBE)
•&;&;&;standards-based education (SBE)
•&;&;&;benchmarking
•&;&;&;educational indicators
•&;&;&;performance contracting (payment on results).
The fact that the notions and terms of SMT are central to American state legislation on education between 1969 and 1976 is evident from an examination of the actual legislation. The following is an illustration of how these ideas shifted from being merely a set of notions circulating in educational theory and implemented sporadically, to becoming the subject of mandatory laws and close supervision.
•&;&;&;&;1969—California: the Education Improvement Act mandated that project grants were to focus on basic skills. These were to be evaluated, and the evaluation was to assess student improvement and cost effectiveness.
•&;&;&;&;1971—Colorado: the Educational Accountability Act introduced the term “accountability” into policy discourse as a generic term and referred to the legislative tools of scientific management in education.
•&;&;&;&;1971—Florida: the Educational Accountability Act directed the Commissioner of Education and the State Board of Education to establish grade-by-grade standards in the basic skills and to use tests based on specific theories of testing.
•&;&;&;&;1971—California: an Educational Management and Evaluation Commission was appointed. Nine public members were to be appointed— three to represent the field of economics, three to represent the learning sciences, and three to represent the management sciences.
•&;&;&;&;1971—Colorado: there was a call to determine whether decisions affecting the educational process are improving or retarding achievement.
•&;&;&;&;1971—California: the Stull Act mandated that the evaluation of teachers be based on the performance of their students. School districts were required to adopt evaluation and assessment guidelines to do this.
•&;&;&;&;1971—Virginia enacted a law based on the principles of management by objectives. There are clear implications that schools are not trying hard enough or are not succeeding, and there is an additional implication that the legislation will change matters.
•&;&;&;&;1971—California: the Guaranteed Learning Achievement Act provided for “performance contracting” between private contractors and school authorities for the teaching of reading and mathematics. Payment of the contractor was contingent on the students attaining specific learning.
•&;&;&;&;1972—Ohio: educational efficiency was sought through the use of computer-based management information systems.
•&;&;&;&;1973—Texas: the Legislature enacted a bill which directed the Legislative Budget Board to establish a planning, programming, and budgeting system. The emphasis was on performance and the intent was to develop measurable output standards. Subsequent resolutions directed “program budgeting” and a study of “zero-based budgeting” and “cost-benefit analysis”.
•&;&;&;&;1973—Rhode Island adopted a law that called for the approval of a master plan to encompass not only elementary and secondary schools but also colleges and universities.
•&;&;&;&;1973—Oklahoma passed an act which imposed systems analysis, originally developed to manage defence expenditures, on the school districts, and also introduced the term “needs assessment”.
•&;&;&;&;1974—Georgia enacted legislation calling for performance-based criteria for operating the instructional programme of each school.
•&;&;&;&;1974—Florida enacted a law seeking to guarantee in advance that textbooks and other instructional material will be effective. The term used is “learner verification”.
•&;&;&;&;1976—Florida: the Educational Accountability Act mandated the specification of a system that guarantees each student will attain a specified standard, and specified the contingencies should a student fail to attain the minimum performance standards.
•&;&;&;&;1976—Virginia mandated state-wide testing of the basic learning skills (Wise, 1979, pp. 12–26.)
In order to fully understand the significance of all this one needs to examine in more detail the consequences of this legislative turn in education. There are four requirements that go hand in hand with the decision to legislate within education along the lines of SMT. The legislator needs to provide a clear statement of what is expected to occur as a result of the legislation. There needs to be a reason to be confident that those who are required to achieve these outcomes will in fact do so. In addition, both a mechanism for auditing the achievement of the outcomes, and the provision for precise and dependable information to aid rationalistic decision making, are necessary. That is, each of the following is required: (i) a statement of unambiguous outcomes; (ii) a theory (probably implicit) of control; (iii) a system of auditing; and (iv) the provision for instrumentally-oriented research. Without a clear statement of outcomes the legislation will lack focus and intent, and without a theory of compliance it will lack teeth. The need for monitoring and the demand for precise information arise from the fact that the legislation is founded on SMT.
The outcomes-based curriculum is intended to provide the required clear statement of outcomes. But what is the probable theory of control? In order to answer this question we need to examine the social theories implicit in the legislative turn: functionalism and individualism. Functionalism provides the global rationale for the necessity to achieve control over individuals. Individualism provides the mechanism by which this is achieved through the contemporary variants of the early-modern idea of the social contract. These are: agency theory; public choice theory; new public management theory; and transaction-cost analysis. Contemporary public policy has been substantially based on these theories. Their impact has been considerable, and they have changed the way we speak about education (Boston, 1991). Once the processes of auditing and information gathering are added to the specification of outcomes and the notion that teachers are contractors who deliver these outcomes, we have the conditions for the élite-ruled, artificially designed, and privatised curriculum (Neyland, 2007). So, in summary, what did happen just prior to 1980? In an unprecedented development, federal and state governments mobilised to scientifically manage education, backed by the power of legislation. The 1970s saw the groundwork completed. By 1980 the project could be said to be fully launched.
A further dimension of the untold story
Any reader who is a committed educator and enthusiastic about good assessment will either have given up reading this paper by now or will be quietly fuming at what she will feel is my misrepresentation of assessment. In exasperation she will want to say to me: What about assessment for better learning? What about those inadequate teachers who had no idea whether or not their students were learning anything? My answer is this. First, good teaching is sufficient; if you have good teaching, nothing more is needed. Second, good assessment can never fully escape the unpleasant odour of assessment for accountability. Even the very best assessment, once it is artificially separated from good teaching, is already tainted.
However, there is another dimension to the untold story of assessment that I have not included here. It is the story of a counter-movement against the “assessment for accountability” enterprise. This counter-movement could be called the “assessment for better learning” movement. Those who recognised the harm the new approach to curriculum would cause began to reshape the meaning of the word “assessment”.1 In effect this boiled down to an attempt—using a series of devices well known to discourse analysts—to subvert the legislative-based meaning associated with assessment and re-colonise it for educational purposes. This was all laudable, and certainly did a great deal to blunt the weapons of the accountability movement. The problem is this: subversion never escapes the assumptions that lie at the heart of that which is subverted. The counter-movement never escapes the orbit of the mass of that which is opposed. The details of why this is the case require more space than is available here and so will need to wait for a subsequent paper. An example must suffice. It can be argued that idealism and realism are not completely different orientations, but two sides of the same coin. Both depend on the assumption that the mind and body are separated. Subverting realism by appropriating idealism does not show that the coin is a valueless currency. It strengthens the coin’s value; it draws the attention further away from the transformational idea that the mind and body may in fact be connected.
References
Berlin, I. (1998). The apotheosis of the romantic will. In I. Berlin, H. Hardy, & R. Hausheer (Eds.), The proper study of mankind (pp. 553–580). London: Random House.
Boston, J. (1991). The theoretical underpinnings of public sector restructuring in New Zealand. In J. Boston, J. Martin, J. Pallot, & P. Walsh (Eds.), Reshaping the state: New Zealand’s bureaucratic revolution (pp. 1–26). London: Oxford University Press.
Bruner, J. (1960). The process of education. Cambridge: Harvard University Press.
Dakin, J. (1973). Education in New Zealand. Auckland: Leonard Fullerton.
Giddens, A. (1990). The consequences of modernity. Stanford: Stanford University Press.
Grouws, D. (Ed.). (1992). Handbook of research on mathematics teaching and learning. New York: Macmillan.
Her Majesty’s Stationery Office. Mathematics counts: The Cockcroft Report. (1982). London: Author.
Hirst, P., & Peters, R. (1970). The logic of education. London: Routledge & Kegan Paul.
Hooper, R. (Ed.). (1971). The curriculum: Context, design and development. Edinburgh: Open University Press.
Kleinig, J. (1982). Philosophical issues in education. London: Crown Helm.
Lawton, D. (1973). Social change, educational theory and curriculum planning. London: University of London Press.
MacGinitie, W., & Ball, S. (1968). Readings in psychological foundations of education. New York: McGraw-Hill.
Morgan, G. (1986). Images of organisation. London: Sage Publications.
Musgrave, P. (1965). The sociology of education. London: Methuen.
National Council of Teachers of Mathematics. An agenda for action. (1980). Reston, Virginia: Author.
Neyland, J. (2007a). Globalisation, ethics and mathematics education. In B. Atweh, A. Barton, M. Borba, N. Gough, C. Keitel, C. Vistro-Yu, et al. (Eds.), Internationalisation and globalisation in mathematics and science education (pp. 113–128). Dordrecht: Springer.
Neyland, J. (2007b). The hidden injuries of the literary curriculum (unpublished).
Shumway, R. (Ed.). (1980). Research in mathematics education. Reston, Virginia: National Council of Teachers of Mathematics.
Skilbeck, M. (1976). Ideologies and values: Unit 3 of course E203 Curriculum design and development. Milton Keynes: Open University.
Webb, N. (1992). Assessment of students’ knowledge of mathematics: Steps toward a theory. In D. Grouws (Ed.), Handbook of research on mathematics teaching and learning (pp. 661–683). New York: Macmillan.
Wise, A. (1979). Legislating learning: The bureaucratization of the American classroom. Los Angeles: University of California Press.
Note
1&;&;&;Some of this harm is to teachers and students themselves (Neyland, unpublished).
The author
Jim Neyland has been a classroom teacher and a preservice and ongoing teacher educator. He has worked in curriculum development at the national level. He is currently a senior lecturer in the School of Education Studies, Faculty of Education, Victoria University, Wellington.
Email: jim.neyland@vuw.ac.nz