Assessment in the age of ChatGPT has a blunt new constraint: any task that can be completed by pasting the question into a chatbot no longer measures what it used to measure. That does not mean assessment is broken — it means a particular style of assessment, the unsupervised written product, has lost its evidentiary value. This article works through what still functions, what needs redesigning, and how to assess in a way that treats AI as a fact of the environment rather than an emergency.
Why detection is the wrong hill to die on
The instinctive response to AI-written work is to detect and punish it. The problem is that detection tools are unreliable in both directions: they miss AI text that has been lightly edited, and they flag human text — particularly from non-native English speakers and from writers with very tidy styles — as machine-made. Building academic-integrity cases on unreliable evidence is unfair to students and indefensible when challenged.
There is also an arms-race problem. Every improvement in detection is met by paraphrasing tools and newer models, and the cycle consumes staff energy without ever producing certainty. The institutions handling this well have largely stopped asking "how do we catch AI writing?" and started asking "how do we design tasks where AI use is either irrelevant, impossible, or explicitly part of the work?" That reframing is the foundation everything below builds on.
Assess the process, not just the product
A finished essay or program is now weak evidence of capability; the process that produced it remains strong evidence. Requiring visible process — outline, draft, revision, final — turns one unverifiable artefact into a sequence that is genuinely hard to fake coherently. Version history in shared documents, short reflective notes on what changed between drafts, and checkpoints where students discuss work-in-progress all thicken the evidential record while simultaneously teaching better working habits.
The oral layer is the strongest tool available. A five-minute conversation — "walk me through your argument", "why this structure?", "what would break if you removed this section?" — distinguishes authors from assemblers with remarkable reliability, and it scales better than feared: brief structured vivas for a class often cost less time than the marking disputes they prevent. Students who did the work generally enjoy the chance to show it; students who did not reveal it within a question or two.
Move the measurable work to where you can see it
Some capabilities still need to be certified as the student's own unaided work — foundational writing, core problem-solving, basic fluency in a discipline's methods. The honest response is to assess those in supervised conditions: in-class writing, invigilated practicals, whiteboard problem-solving, timed exercises in controlled environments. This is not nostalgia; it is matching the assessment condition to the claim being made. If the claim is "this student can do X unaided", the evidence must come from a setting where aid was absent.
The corollary is to stop pretending unsupervised homework certifies anything it cannot. Homework still has enormous value — as practice, as preparation, as formative feedback — but its role shifts from proving capability to building it. Grading weight moves toward supervised checkpoints and process evidence; homework becomes the training ground where using AI assistance may be perfectly acceptable, because the certification happens elsewhere.
Design tasks where AI use is the skill
The workplaces students are heading into will not forbid AI; they will expect fluent, critical use of it. Assessment can get ahead of that by making AI part of the task. Ask students to generate an AI draft and submit it alongside their critique and improved version, with the marks attached to the critique. Ask for the prompt they used, and grade the prompt — writing a precise, well-scoped prompt is itself a demonstration of understanding the problem. Ask them to verify an AI-produced analysis against sources and document what survived scrutiny.
These designs have a pleasant property: they are hard to shortcut with the very tool they incorporate, because the assessed skill sits one level up from generation. Evaluating output requires the domain knowledge the course teaches. Platforms are emerging that grade this kind of work directly — Square 1 AI, for example, uses its tutor Nova to grade both code and the prompts students write, against explicit per-exercise criteria — and the underlying pattern generalises: define the criteria precisely, and AI-era work becomes assessable again.
Rebuild your rubrics around judgement
Old rubrics rewarded qualities AI produces cheaply: fluency, organisation, coverage, correct formatting. Leaving those weightings intact inflates the value of machine-assisted work. Rebalanced rubrics push weight toward what remains scarce: specificity to the local context, quality of reasoning under questioning, appropriate handling of uncertainty, choices defended with evidence, and genuine engagement with counter-arguments rather than token acknowledgement.
It helps to state AI policy per task, in writing, using plain categories — not allowed (and assessed under supervision), allowed with disclosure, or required. Ambiguity is where both cheating and unfair accusations breed. Students are markedly more cooperative with rules that come with reasons, and "this task builds a muscle you need, so it is AI-free; that task mirrors real work, so AI is expected" is a reason they recognise as honest.
Frequently asked questions
Can AI detectors be trusted for misconduct cases?
No — not as sole evidence. They produce both false negatives and false positives, and the false positives fall unevenly across student groups, which makes them dangerous foundations for high-stakes accusations. If detection output is used at all, it should only ever be a prompt for a conversation, with conclusions resting on process evidence and the student's ability to discuss their own work.
Does assessing process take much more staff time?
Less than expected, once weightings shift. Checkpoints and brief orals add contact time, but they displace time currently spent marking long final products in forensic detail and adjudicating misconduct disputes. Marking a five-minute defence plus a shorter supervised piece is often faster than marking a long take-home essay you cannot trust anyway.
Should students be allowed AI for homework?
Increasingly, the workable answer is yes with disclosure — provided certification of unaided capability happens in supervised settings. Forbidding AI in unsupervised conditions creates a rule you cannot enforce, which teaches students that rules are theatre. Permitting it, requiring disclosure, and assessing what they did with it keeps the honesty intact and the learning visible.
Where to go from here
If you are redesigning assessment and want hands-on fluency with these tools first, the AI for Teachers course walks educators through practical AI use with graded exercises. Curious what criteria-based AI grading feels like from the student's side? The free 3-minute skill check will show you in three minutes.
