Perspective
AI in education: what we let it mark, and what we never let it decide
Assessly uses AI in three places. Here is each one, what it is trusted with, what it is not, and the feature we removed because it could not be trusted at all.
Every education product now says it uses AI, and the claims have run ahead of the candour. Assessly is one of those products, so this is an attempt at precision. We use AI in three places. For each one: what it does, what it is not allowed to do, and what happens when it gets something wrong.
1. The code mark: the tests decide, not the model
The easiest place to trust AI too much is the code itself, so we do not ask it. When a problem has test cases, the code part of the mark is the share of hidden tests the submission passes. The student's program runs in a sandbox against inputs they never saw, and either prints the right answer or does not. A language model's opinion of whether the code looks correct is not part of that number.
The model still reads the code, to write feedback: what the approach is, where it is likely to break, what to try next. If the model is slow or unavailable, the student still gets their mark from the tests, and the result says plainly that the written feedback is missing. Nobody's grade waits on an AI provider.
2. The explanation: where AI does mark
Faculty can ask students to explain their solution in their own words, and set how much of the mark the explanation carries. The default is 70% code and 30% explanation. This part is marked by AI, against a rubric, because nothing else can read four hundred explanations the same way.
Consistency is the real benefit. A person marking the fiftieth explanation in a stack gives it less attention than the first, anchors on whatever came before, and reads fatigue as strictness or generosity depending on the hour. That is not a criticism of faculty. It is what attention does under load. The model applies the same rubric to every script at every position in the pile, and writes down its reasons, so nobody is asked to trust a bare number.
It also has a real weakness. A student who solves a problem in an unusual but sound way can be marked too conservatively, because the approach sits far from what the model sees most often. So any faculty member can change any score. The change needs a written reason, and it is logged with who made it and when. The AI does the volume; people keep the last word on every mark, not only the edge cases.
The useful question is not whether the AI is ever wrong. It is whether anyone notices, and who has the power to fix it.
3. The tutor: hints, not answers
Ace, the tutor inside Learn, answers students while they work on a problem. It is built to ask the question that gets them unstuck (where does the left pointer go when the second a arrives?) rather than to hand over the solution. A tutor that writes the code teaches the student to ask the tutor.
Each college sets how many messages a student can send in a day, so the tutor supplements the course without replacing the work. For staff, Ace drafts: a notice to the students below a drive's bar, a practice set, a form. It sends nothing on its own. A person reads the draft and approves it.
What we took out
The fourth place we used AI no longer exists. Until August 2026, Assessly put a percentage on how likely a student's written explanation was to be AI-generated. When we audited it, it did not do what its label said. It read the explanation but never the code, the model underneath was built to catch prompt-injection attacks rather than to tell who wrote a paragraph, and across more than three thousand graded submissions it produced a number for fewer than forty. Every one of those numbers was 0 or 100.
We withdrew it from every screen on 26 August 2026 and have not replaced it. Nobody can currently tell reliably whether a paragraph came from a person or a model, and a tool that pretends it can will accuse real students. The full account is in What our proctoring records, and what it deliberately doesn't.
What it cannot do
AI marking cannot mentor a struggling student, notice that the real problem is at home rather than in the code, or make someone curious about recursion. It evaluates work; it does not teach a class. The reason to automate the marking is to hand hours back to the people who do the parts no model can.
Five questions to ask about any AI claim
If you are choosing a platform, these separate a careful use of AI from a label on the box:
- What does the AI decide on its own? Ask which marks come from a model and which from something checkable, like test cases
- What happens when the model is down? A grade that waits on an AI provider will fail on exam day
- Who can change an AI mark, and is the change recorded? An override nobody can audit is its own problem
- Does it produce any number nobody can explain? A probability with no reasons behind it should not reach a student's record
- What have you switched off? A vendor who has never removed an AI feature has probably never checked one
The honest case for AI in education is narrow and strong: it makes the most mechanical part of teaching consistent, fast and reviewable. Everything else is still teaching. Book a demo to see where the line sits on a real paper.
Was this page helpful?