Human vs. Machine: The AI Grading Dilemma in Technical Classrooms
Sunday Reflection — 2026-09-20
First time an AI grader told me a student’s Python was “correct” because the function returned the right value, I nearly spat my tea. No comment about the hard-coded input, the missing docstring, or why the loop broke on the third case. The machine ticked its box. I failed the code, called the student in, and spent half an hour on how to break things properly. I doubt any algorithm would have found the real teaching moment there.
That’s the problem with AI assessments coming in hot this year. I understand why management loves them—scaling, ticking boxes, defending consistency when the EU AI Act inspectors knock. But in the classroom, context matters. I’ve watched a student copy-paste the perfect answer straight off Reddit, get an auto-pass, and still have no idea what a decorator does. Equally, I’ve seen a transfer student struggle through a networking practical, miss points on the bot’s rubric, but approach the core problem with more insight than half the group. The AI flagged her as “incomplete.” My gut flagged her as “thinking.”
mermaid flowchart TD Q[Student submits assignment] --> GR{Grading method?} GR -- Manual by Teacher --> M[Human reviews and gives feedback] GR -- AI Assessment --> AI[Machine checks rubric, returns score] AI --> H[Optional human spot-check?] H -- Yes --> M H -- No --> F[Final automated grade] M --> F
We use a blend at Masterschool. For repeatable, basic stuff—like syntax checks—I’ll let the AI clear the weeds. For project code, we sample grades: one part human, one part AI, sometimes both disagree, and we argue it out in the staff room. Occasionally, peer review surfaces mistakes all of us missed. I wouldn’t want to go full auto. The army taught me mechanisation works for logistics, not for judgments that shape people. There’s something about catching a student’s half-right answer and nudging them, quietly, in the right direction. The software doesn’t spot those moments. I’m not sure it ever will.
One day the machines might get better at picking up nuance. For now, the best results come from deciding what’s safe to automate—and what needs a human eye. That line isn’t fixed. But if education means anything, it’s spotting what people are truly thinking, not just whether they followed instructions.