NZQA Automated Text Scoring (ATS) review—Review of model marking performance, quality assurance, and retraining approach for AE2 2025 and AE1 2026

Image
NZQA Automated Text Scoring (ATS) review—Review of model marking performance quality assurance.png

NZQA commissioned this review to provide an independent evaluation of its Automated Text Scoring (ATS) model and to support decisions about its continued use in marking assessments for the corequisite Literacy Writing standard 32405. The review examined model marking performance across Assessment Event 2, 2025 (AE2 2025) and Assessment Event 1, 2026 (AE1 2026) alongside the quality assurance, human marking, and retraining processes, and considered the potential transferability of the ATS approach to other assessment standards.

Overall, the ATS model is considered fit for continued operational use in marking the Literacy Writing
assessment. In AE1 2026, when ATS was used as the production marking model, it demonstrated strong operational performance. At the rubric element level, exact match between ATS and human scores was 79.6%, exact-or-adjacent agreement reached 99.8%, and the mean Quadratic Weighted Kappa (QWK) was 0.763. At the total score level, 84.5% of ATS scores were within one point of the corresponding human score, and the QWK was 0.872. Performance in AE2 2025 was similarly strong. 

On average, ATS scored slightly lower than human markers. The mean ATS–human score difference was −0.07 with a small effect size in AE1 2026. The ATS model also appears to differentiate somewhat less strongly than human markers among the “Content”, “Language”, and “Structure” rubric elements, which is an area worth continuing to monitor. The subgroup analysis did not identify substantial differences in ATS–human agreement across demographic groups, although ongoing monitoring is recommended. 

The quality assurance and retraining approaches are generally sound. Recommended refinements include strengthening validity checks, automating the identification of empty responses, broadening the selection of high-scoring scripts, and considering a random or stratified-random sample to provide a more representative estimate of overall performance. Retraining data should be sampled separately by week and question, with continued attention to less common scores and weaker performing rubric elements. Given the high ATS–human agreement for scripts scored near the cut scores, the size and human marking approach for this cohort could also be reviewed. 

While the broader ATS approach may be transferable to other assessment standards, this would need to be established separately for each standard through model redesign, appropriate training data, rubric alignment, validation, fairness and robustness testing, and ongoing human oversight. 

The improvements observed from AE2 2025 to AE1 2026 are consistent with a beneficial retraining effect, although differences in questions, student populations, human marking, and dataset selection may also have contributed. The datasets used in this review were purposively selected quality assurance samples focused on more challenging marking cases, so the reported statistics should not be treated as estimates of performance across all responses. 

Publication type
Book
Publication year
2025