The purpose of this report is to present an independent analysis of the New Zealand Qualifications Authority’s (NZQA’s) Automated Text Scoring (ATS) model.1 It evaluates the model’s performance, bias/fairness, and readiness for marking literacy writing assessments associated with US 32405, as well as its potential transferability to other assessment standards.
The commercial marking model was used for live scoring in both the 2025 assessment events (AE1 and AE2). To date, the NZQA ATS model has been used to support the quality assurance process of marked responses.
This review includes statistical analyses and visualisations of the ATS model’s scores against human marking from the first 2025 live assessment events (AE1). Data from the 2024 Pilot were reviewed as a supplement to the 2025 AE1 analysis, where it helped identify any significant differences or emerging patterns.