Published calibration evidence
A mark you can test.
Students test essay markers the obvious way: submit the same essay twice and see if the number holds. We think that's exactly right — so we run that test continuously against a frozen benchmark, and publish the results. To our knowledge, no other GAMSAT marker publishes any calibration evidence at all.
±1
Re-mark consistency
The same essay re-marked lands within 1 point (median 1) on our 40-essay frozen benchmark. Measured 19 July 2026.
4 pts
Cross-examiner agreement
Average gap between our two independent AI examiners on the same essay, across the benchmark. Every Pro mark is this cross-check; a real disagreement calls in a third examiner. Measured 19 July 2026.
39/40
Benchmark accuracy
Essays landing inside their expected score range on the frozen benchmark, which anchors to reviewed exemplars. A run outside budget blocks release. Measured 19 July 2026.
collecting
Tracking against real ACER results
Students report official S2 scores after each sitting (0 so far). We publish the aggregate our-mark-vs-ACER figure once at least 10 reports are in — not before, because a small sample would mislead in either direction.
The method, plainly
ACER publishes exactly two assessment criteria for Section II — the quality of the thinking, and the control of language in expressing it — and no marking rubric. Our instrument is six sub-criteria built from those two published criteria, calibrated against reviewed exemplar essays, and locked with regression tests: a frozen benchmark of 40 essays with expected score ranges that every engine change must pass before it ships. When we deliberately change marking behaviour, the benchmark is re-frozen in the same commit — score movement always comes with a recorded reason.
Every Pro essay is marked by two independent AI examiners and cross-checked; ACER itself states that Section II responses are “scored independently three times” — the same principle. When our two examiners genuinely disagree, a third examiner from a different model family is called in and each criterion is decided by the middle mark, disclosed on the grade.
Honest limits: these are calibrated practice estimates, not official ACER marks — only ACER issues those, on exam day. Consistency is measured on our frozen benchmark, not on your individual essay. GAMSAT® is a registered trademark of ACER, which is not affiliated with and does not endorse Aptavia.