Meet Felix, the AI examiner
Write your answer to a real exam task and Felix marks it sentence by sentence on the four Cambridge scales — quoting your exact words, computing your band from the official rubric, and telling you plainly whether it would pass. An examiner, not a yes-man.
Writing tasks and Felix’s marking come with the passes that cover Writing — one-time timed access, not a subscription. Compare the passes.
Felix does not return “good job”. It scores each of the four Cambridge scales from 0 to 5 and plots them against the line a typical pass sits on. In the example opposite, the writing communicates well but its Language sits below the pass line — precisely the kind of gap a real examiner would dock, and a chatbot would gloss over.
Example B2 assessment — one answer against the typical pass line (3/5).
Felix is not a chat window. It is a marking pipeline that forces the model to behave like an examiner and to prove everything it claims.
Every part is marked on the four Cambridge scales — Content, Communicative Achievement, Organisation and Language — each from 0 to 5, exactly as a real examiner does.
B2 and C1 are marked to different standards. At C1 it demands advanced vocabulary, complex structures and sophisticated cohesion — so a B2-level answer will not be handed a C1 pass.
Every note quotes the exact words you wrote: mistakes to fix, weak-but-correct phrasing to upgrade, and the phrases you used well. It has to show its working, character for character.
The examiner is required to find real errors and at least two concrete upgrades on every part — never a wall of compliments. Flattery is designed out of it.
The model reports the sub-scores; your band and totals are then calculated in our own code from the official scale. The number cannot be quietly inflated, and it stays consistent between attempts.
Both parts are marked together, so it catches the mistakes you repeat across your writing and gives one coherent verdict — plus a model answer showing what a strong response looks like.
Did you cover every part of the task? Miss a prompt point and the mark drops.
Is the register and tone right for the genre, and is the reader fully informed?
Is the answer logically ordered, paragraphed and linked together?
How wide and how accurate is your grammar and vocabulary?
Felix is not a prompt someone typed once. The examiner brief it marks against is written out in full — the published Cambridge band descriptors for each scale, the task rules an examiner applies before reading a single sentence, and the stance it has to hold — and every change to that brief is a reviewable edit, not a tweak buried in code. Here is what is actually in it.
The published descriptors
Bands 5, 3 and 1 of every Cambridge scale, as the B2 First and C1 Advanced handbooks word them, are in the brief — so a band means the same thing here as on the day. The C1 scale starts where the B2 scale peaks, and Felix is told so.
Task rules first
Off-topic or not in English is 0 across the board. A missing content point caps Content. Under-length, memorised filler and the wrong register each have a fixed ceiling. These apply before the language is judged, however fluent the text feels.
Mark, not coach
Felix is briefed to give the band a real examiner would give: no cushioning phrases, no praise the text has not earned, and every remark tied to words you actually wrote. A strong answer still gets its 5 — accuracy cuts both ways.
Rubric
Encode the four Cambridge scales into a machine-readable marking schema.
Calibration
Tune the examiner brief level by level, against real B2 and C1 answers.
Evidence
Force verbatim quoting and a strict, validated output structure.
Scoring
Move all band maths out of the model and into deterministic code.
Live
Ship it — then keep tightening how strictly it holds the band.
Gemini is a brilliant general assistant — built to help with anything. Felix is built to do one thing: mark Cambridge writing to the band. Those are different jobs, and it shows across every capability a marker needs.
How each tool is designed to behave by default — illustrative of approach, not a lab benchmark.
| Capability | Felix | Gemini | Chatbot |
|---|---|---|---|
| Marks on the four Cambridge scales (0–5) | |||
| Calibrated to your CEFR level (B2 vs C1) | |||
| Quotes your exact words as evidence | |||
| Score computed from a locked rubric | |||
| Required to flag real errors (won't only praise) | |||
| Hard task rules applied before the language is judged | |||
| Whole-paper marking + repeated-mistake detection |
Felix runs on a frontier AI model, the same kind of intelligence behind the best assistants. But raw intelligence, dropped into a chat box, makes a flatterer — not an examiner. The difference is the harness we built around it over that long development.
Before it ever sees your writing, the model is locked to a level-specific examiner brief and the Cambridge rubric. It is forced to return a strict, validated structure — sub-scores, quoted annotations, strengths, fixes and repeated mistakes — rather than a free-flowing chat. Then our own code, not the model, turns those sub-scores into your band. The intelligence is the same; the discipline is what you cannot get from a chat window.
Pick a task, write your answer, and Felix shows you exactly where you stand — and exactly what to fix before it counts.
Felix gives an AI estimate to guide your practice — a genuinely rigorous one, but not an official Cambridge result. Only a real examiner can award your final grade.