AI & Assessment

Last updated

When AI Supports Assessment: Transparency and Educator Oversight

Automated scoring and feedback can scale support, but educators need to understand how systems work and when to override or supplement them. This article discusses principles for transparent AI-assisted assessment and the role of human judgment in interpretation and action.

Transparency Is a Working Practice

AI can help an assessment team classify responses, suggest feedback, identify patterns and direct human attention. None of those uses transfers responsibility to the system. A score still affects a learner, a feedback message still shapes what they practise next, and a flag may still influence an educator’s decision. Transparency therefore cannot mean publishing a general statement that “AI is used”. It means giving each person enough information, at the right point in the workflow, to understand the system’s role and challenge an outcome.

This is consistent with the British Council’s 2025 human-centred position: technology should remain supportive rather than become an autonomous decision-maker. It also reflects the longer-standing automated-scoring principle that usefulness depends on evidence for the particular assessment purpose, not simply agreement with a set of historical human scores.

Start With the Decision, Not the Model

Before selecting a model, an organisation should state what decision is being supported. Is the output low-stakes practice feedback, a teacher-facing diagnostic signal, a provisional mark, or evidence used in certification? The higher the consequence, the stronger the requirements for validation, security, human review and appeal. A tool that is useful for formative feedback is not automatically suitable for awarding a grade.

The NIST AI Risk Management Framework offers a practical discipline: Govern, Map, Measure and Manage. For assessment, “Map” is especially important. Teams need to record the intended population, construct, context, foreseeable misuse and consequences of error before evaluating technical performance. That prevents a good-looking accuracy figure from answering the wrong educational question.

Explain the Role of AI at Two Levels

Learners need a plain-language explanation: what material is processed, whether AI contributes to a result, what the output means, and how to request review. Educators and administrators need deeper operational information: which rubric or construct the system targets, what evidence supports its use, which responses are routed to people, and what limitations have been observed. Giving everyone the same technical document serves neither audience well.

Cambridge’s 2025 ethical AI principles connect transparency with explainability, oversight and test integrity. That connection matters. An explanation is useful only when it helps somebody act: interpret a result correctly, notice an unsupported inference, protect personal data, or escalate a questionable case.

Keep Feedback Close to the Evidence

When AI produces feedback, teams should test more than whether its comments sound plausible. Feedback should point to observable features of the learner’s response, stay within the intended criteria, and avoid diagnosing a learner from one performance. “This response gives one supporting reason” is an observation; “you cannot develop arguments” is a much broader inference. Showing the relevant response evidence helps an educator check the claim and helps the learner understand it. Suggested next steps should also be proportionate: an option for practice or discussion, not an automated prescription that closes off professional judgment.

Design Human Control Into the Workflow

“Human in the loop” is too vague unless the human has time, evidence and authority. A credible workflow defines which cases require review, who can change an outcome, how disagreement is resolved, and how a learner can appeal. Review thresholds might include low system confidence, unusual response patterns, missing or poor-quality input, a large discrepancy between evidence sources, or a decision close to an important boundary.

Ofqual’s 2026 working paper on AI in high-stakes marking is a useful reality check. It emphasises validity, fairness, transparency and accountability, and notes that using AI as the sole mechanism for determining a student’s mark does not comply with Ofqual’s regulations. Even outside regulated qualifications, that distinction is sound: automation may support a professional decision without becoming an unaccountable substitute for one.

Show Evidence and Uncertainty Honestly

A single correlation with human marks is not a complete validity argument. Teams should ask whether the system attends to the intended language abilities, performs consistently across relevant tasks and learner groups, and fails in educationally significant ways. Williamson, Xi and Breyer’s framework for evaluating automated scoring remains relevant because it treats evaluation and use as an evidence-based argument rather than a one-off model benchmark.

Interfaces should not turn uncertain evidence into false precision. Where appropriate, they can show a review status, the evidence considered, a confidence category, or a warning that the output is advisory. Confidence should never be presented as a guarantee of correctness. More importantly, low confidence must trigger a defined action rather than merely decorate a dashboard.

Treat Data, Fairness and Change as Continuing Duties

Assessment data can contain identifiable speech, writing, demographic information and records of performance. Transparent practice explains what is collected, why it is needed, how long it is retained, who can access it, and whether it is used to improve a model. Consent language should not hide consequential reuse inside broad terms. Data minimisation is preferable to collecting information simply because it may become useful later.

Fairness also requires monitoring after launch. Performance can change when tasks, cohorts, recording conditions or language varieties differ from development data. Model and rubric updates should therefore be versioned; material changes should trigger re-evaluation; and audit records should connect each result to the system version, evidence and human actions involved. Without that chain, an organisation may know what score was issued but not why it was defensible at the time.

An Operational Standard for Responsible Use

For EduZMS, the practical standard is straightforward: state the purpose; validate against the intended construct and population; disclose the AI role; route exceptions to qualified people; preserve a usable audit trail; provide correction and appeal paths; and monitor performance, access and fairness over time. If any link is missing, the system should be described as experimental or advisory and kept away from decisions it cannot yet support.

The aim is not to make every learner understand a model’s mathematics. It is to make institutional responsibility visible. People should know when automation is involved, what reliance is reasonable, where human judgment enters, and who remains accountable. In assessment, trust should follow evidence and review—not a polished interface or the presence of AI itself.

References

  1. Felice, M., Spiby, R., O’Sullivan, B., & Edmett, A. (2025). Human-centred AI: lessons for English learning and assessment. British Council.
  2. Pastorino-Campos, C., & Saville, N. (2025). Ethical AI for Language Learning and Assessment. Cambridge University Press & Assessment.
  3. Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology.
  4. Ofqual. (2026). Principles of AI use in marking. Working paper Ofqual/26/7292/1.
  5. Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A Framework for Evaluation and Use of Automated Scoring. Educational Measurement: Issues and Practice, 31(1), 2–13.

← Back to Insights