The only TCF Canada examiner that publishes its error rate
Here are the numbers behind that promise: how often the AI examiner's score lands on the right level, measured on productions of known level, with the method and its limits. Numbers come before claims.
Last updated: · Changelog
Small samples (24 writing, 36 speaking) and known reference levels, not official TCF results. Read the ranges, not only the percentages.
Our numbers
Writing N = 24 samples
Speaking N = 36 samples
The figures on the cards are held out: each sample is scored through a curve fitted without it. Before and after the curve:
| Writing | No curve | Curve, held out | Curve, own data |
|---|---|---|---|
| Exact level | 42% | 54% | 50% |
| Within one level | 88% | 100% | 100% |
| Weighted kappa | 0.76 | 0.90 | 0.88 |
| Rank correlation | 0.95 | 0.94 | 0.95 |
| False “ready” | 0/14 | 0/14 | 0/14 |
| Speaking | No curve | Curve, held out | Curve, own data |
|---|---|---|---|
| Exact level | 31% | 44% | 44% |
| Within one level | 81% | 94% | 94% |
| Weighted kappa | 0.63 | 0.80 | 0.80 |
| Rank correlation | 0.92 | 0.89 | 0.92 |
| False “ready” | 0/16 | 0/16 | 0/16 |
How to read these numbers
For a sense of scale: in a published study of CEFR speaking ratings, two trained raters agreed exactly about 39 to 44% of the time and within one sub-band about 88 to 93% of the time (Huang, 2018, Language Testing in Asia). Figures from the published excerpt, to be confirmed: the full text is paywalled and we have not read it.
This is neither a bar to clear nor a like-for-like comparison, for three reasons:
- the scales differ: that study counts sub-bands, finer than our six CEFR levels. Our “within one level” is a looser tolerance than “within one sub-band”, and exact agreement is easier on a coarser grid;
- the populations differ: our sample spans every level from A1 to C2, which makes agreement and correlation easier;
- the target differs: in time we want to match official TCF Canada results, a harder reference than a second rater scoring the same recording. We do not measure that yet (see below).
Another industry reference: for TOEFL Junior, ETS reports a machine-human correlation of .81 against .89 between two humans in speaking, and .83 against .90 in writing (ETS research report RR-15). We publish two measures in the same spirit, weighted kappa and rank correlation, but on a different test and different data: do not compare the numbers one to one.
We do not claim the examiner is as good as a human rater, let alone better. We publish where it is wrong so you can judge for yourself.
Method
- Known-level samples, not official results. 24 texts (19 written productions and 5 oral productions given as transcripts) and 36 speaking clips: public productions whose source states the CEFR level (calibrated samples from CIEP and France Éducation international, expert-rated recordings, tutor-graded practice exams). None is an official TCF Canada result: none are published. A few levels come from prep sites (weaker labels).
- Levels. Scores out of 20 are read on France Éducation international's grid (A1 1, A2 2–5, B1 6–9, B2 10–13, C1 14–17, C2 18–20) and then as NCLC (NCLC 7 starts at 10 out of 20). “Exact”: the examiner's level is the known level (a known range such as C1/C2 counts as exact anywhere inside it). “Within one level”: at most one CEFR level apart, for example B1 instead of B2. One level apart can cross the NCLC 7 line, which is why the false “ready” rate is measured separately.
- What the curve does. The model ranks candidates well but compresses scores above B1: a B2 often gets 6 to 10 out of 20. The curve is a table that lifts those scores (a raw 8 becomes 10, 9 becomes 11, 10 becomes 12). It is fitted by isotonic regression on the 60 samples, limited to 40% of the full correction so it does not manufacture false “ready” calls, and leaves scores 0 to 5 alone.
- Cross-validation. The numbers on the cards above are held-out: each sample is scored through a curve fitted on the other 59. The table above also gives scores without the curve and the curve on its own data (optimistic) so you can see the gap.
- Weighted kappa measures agreement while penalising large misses more (1 is perfect, 0 is chance). Rank correlation (Spearman) says whether the examiner puts candidates in the right order.
Which way the examiner errs
Errors by known level (read at the middle of the known level's band), after the curve and held out:
| Writing | N | Too low | Right level | Too high |
|---|---|---|---|---|
| NCLC 6 and below | 14 | 1 | 10 | 3 |
| NCLC 7–8 | 4 | 2 | 2 | 0 |
| NCLC 9 and above | 6 | 5 | 1 | 0 |
| Speaking | N | Too low | Right level | Too high |
|---|---|---|---|---|
| NCLC 6 and below | 16 | 5 | 10 | 1 |
| NCLC 7–8 | 6 | 3 | 3 | 0 |
| NCLC 9 and above | 14 | 11 | 3 | 0 |
At higher levels the error almost always runs low: the examiner under-estimates rather than over-estimates. The risk for you is therefore more often thinking you are below the bar when you have reached it. C-level samples are few and partly labelled by prep sites: take this trend as a signal, not a precise measurement.
Same answer, scored three times
An examiner that changes its mind from one try to the next is worthless, right on average or not. So we sent the same 6 written answers to the real model three times, unchanged, and compared the raw scores out of 20 (before the curve).
On the displayed score (after the curve): 83% within one point, largest spread 2 points. The curve lifts scores in steps from 6/20 (a raw 6 has shown 7 since 2026-10-11, instead of 8, which cut the largest displayed spread from 3 to 2 points), so a one-point gap before it can still widen to two after it.
At higher levels the error almost always runs low: the examiner under-estimates rather than over-estimates. See “Which way the examiner errs” above.
Only 6 public answers of known level: read this as an order of magnitude. Model claude-opus-5-5.
The false “ready”: thinking you are NCLC 7 when you are not
This is the costliest error: booking a paid exam believing you are ready. Among productions whose known level is below B2 (under NCLC 7), the examiner scored 0 of 14 in writing and 0 of 16 in speaking at B2 or above.
With so few cases, zero does not rule out a true rate of up to 22% (95% range). Only 5 writing and 9 speaking cases are at B1, just under the bar; the rest are easier to tell apart. The opposite error exists too: 2 of 10 (writing) and 4 of 20 (speaking) productions at B2 or above were given a lower level.
Scores near the NCLC 6/7 line
When an estimate falls within a point of a level boundary, we do not show a single number but a range, for example “9–10/20, borderline NCLC 6/7”, because the official result can fall either side. The NCLC 7 bar is 10 out of 20, so a displayed 10 does not promise NCLC 7.
Official results received
Official results received: 0 — 0 of them with a prior estimate
0 in writing, 0 in speaking. Entered by candidates who sat the real test, never shown individually.
These results will replace the public samples as the reference as they arrive. The numbers on this page do not use them yet. They arrive roughly 15 business days to 5 weeks after the sitting, so the counter grows slowly.
Our promise: this page is updated on every model or curve change. Current figures computed on 10 October 2026.
Listening and reading: indicative
For listening and reading no AI scores anything: the estimate for a ten-question set is a fixed rule. The estimated level is the highest level L such that at least two thirds of the questions up to level L, and at least half of those at level L, are right. That level is converted to an NCLC range with IRCC's table.
This rule has not yet been calibrated against official results, so we publish no accuracy rate for these two skills, and their estimate stays indicative, drawn from only ten questions.
Changelog
- 2026-10-11 · Smoother curve
Raw 6 now maps to 7 (was 8): removes a 3-point swing seen in the test-retest study (same answer shown 8, 5, 5). Level accuracy unchanged: 7 and 8 are in the same NCLC 6 and CEFR B1 band. - 2026-10-10 · Curve 2026-10-10-iso-s0.4
Isotonic curve at 40% strength activated. Writing: exact 42% → 54%, within one level 88% → 100%. Speaking: exact 31% → 44%, within one level 81% → 94%. No false “ready” in the cross-checks. - 2026-10-10 · Before the curve
Raw scores from the claude-opus-5-5 model. Writing: exact 42%, within one level 88%, kappa 0.76. Speaking: exact 31%, within one level 81%, kappa 0.63. The model compressed scores above B1.
What these numbers do not tell you
- Small samples: 24 writing and 36 speaking. The 95% intervals (ci95) are wide and the headline can move several points with each new sample.
- Labels are known levels from the sample's source, not official TCF Canada results. 5 of the 24 writing samples and 5 speaking clips are labelled by prep sites or tutors (weaker labels); the CIEP calibrated and official-calibration ones are stronger.
- The curve was designed on these same 60 samples. The cross-validated figures hold each sample out of the fit, but the shrink setting was chosen by looking at the cross-checks, so some optimism remains.
- Speaking is scored from an automatic transcript: pronunciation is not assessed, and pair interactions and excerpts are scored as one transcript.
- Most samples sit at A1-B2; there are few C1/C2 examples, so the direction of error at NCLC 9+ rests on very few samples.
- The model that scored these samples is the default scoring model at the time of the run; the report does not prove the same figures for a later model.
Question
Are the tasks I am scored on original?
Yes. Every task is written for this product and we do not reproduce exam questions, so your score is what you would get on material you have never seen. About the AI
Try it yourself
Your first speaking task and your first writing task are free, with full feedback. No card needed.
Scores are an AI estimate, not an official result.
NCLC 7 is an independent practice tool. It is not affiliated with, endorsed by or connected to France Éducation international or the Government of Canada. "TCF" and "TCF Canada" belong to their owners and are used only to say which test this practice is for.
About the AI · How the TCF is assessed (France Éducation international)