NCLC 7

The only TCF Canada examiner that publishes its error rate

Here are the numbers behind that promise: how often the AI examiner's score lands on the right level, measured on productions of known level, with the method and its limits. Numbers come before claims.

Last updated: · Changelog

Small samples (24 writing, 36 speaking) and known reference levels, not official TCF results. Read the ranges, not only the percentages.

Our numbers

Writing N = 24 samples

54%Exact level13/24 · 95% range 35–72%
100%Within one level24/24 · 95% range 86–100%
0.90Weighted kappaquadratic, 1 = perfect
0.94Rank correlationSpearman, 1 = same order
Samples
24 of known level, public
Collected and scored
found 9 October 2026, scored 10 October 2026
Model
claude-opus-5-5
Score curve
2026-10-11-iso-s0.4-smooth
Figures computed
10 October 2026

Speaking N = 36 samples

44%Exact level16/36 · 95% range 30–60%
94%Within one level34/36 · 95% range 82–99%
0.80Weighted kappaquadratic, 1 = perfect
0.89Rank correlationSpearman, 1 = same order
Samples
36 of known level, public
Collected and scored
found 9 October 2026, scored 10 October 2026
Model
claude-opus-5-5
Score curve
2026-10-11-iso-s0.4-smooth
Figures computed
10 October 2026

The figures on the cards are held out: each sample is scored through a curve fitted without it. Before and after the curve:

Writing
WritingNo curveCurve, held outCurve, own data
Exact level42%54%50%
Within one level88%100%100%
Weighted kappa0.760.900.88
Rank correlation0.950.940.95
False “ready”0/140/140/14
Speaking
SpeakingNo curveCurve, held outCurve, own data
Exact level31%44%44%
Within one level81%94%94%
Weighted kappa0.630.800.80
Rank correlation0.920.890.92
False “ready”0/160/160/16

How to read these numbers

For a sense of scale: in a published study of CEFR speaking ratings, two trained raters agreed exactly about 39 to 44% of the time and within one sub-band about 88 to 93% of the time (Huang, 2018, Language Testing in Asia). Figures from the published excerpt, to be confirmed: the full text is paywalled and we have not read it.

This is neither a bar to clear nor a like-for-like comparison, for three reasons:

Another industry reference: for TOEFL Junior, ETS reports a machine-human correlation of .81 against .89 between two humans in speaking, and .83 against .90 in writing (ETS research report RR-15). We publish two measures in the same spirit, weighted kappa and rank correlation, but on a different test and different data: do not compare the numbers one to one.

We do not claim the examiner is as good as a human rater, let alone better. We publish where it is wrong so you can judge for yourself.

Method

Which way the examiner errs

Errors by known level (read at the middle of the known level's band), after the curve and held out:

Writing
WritingNToo lowRight levelToo high
NCLC 6 and below141103
NCLC 7–84220
NCLC 9 and above6510
Speaking
SpeakingNToo lowRight levelToo high
NCLC 6 and below165101
NCLC 7–86330
NCLC 9 and above141130

At higher levels the error almost always runs low: the examiner under-estimates rather than over-estimates. The risk for you is therefore more often thinking you are below the bar when you have reached it. C-level samples are few and partly labelled by prep sites: take this trend as a signal, not a precise measurement.

Same answer, scored three times

An examiner that changes its mind from one try to the next is worthless, right on average or not. So we sent the same 6 written answers to the real model three times, unchanged, and compared the raw scores out of 20 (before the curve).

100%Within one point6/6 answers · all three scores 1 point or less apart
1 ptLargest spreadbiggest difference between two of the three scores of one answer
6Answers3 scorings each · 10 October 2026

On the displayed score (after the curve): 83% within one point, largest spread 2 points. The curve lifts scores in steps from 6/20 (a raw 6 has shown 7 since 2026-10-11, instead of 8, which cut the largest displayed spread from 3 to 2 points), so a one-point gap before it can still widen to two after it.

At higher levels the error almost always runs low: the examiner under-estimates rather than over-estimates. See “Which way the examiner errs” above.

Only 6 public answers of known level: read this as an order of magnitude. Model claude-opus-5-5.

The false “ready”: thinking you are NCLC 7 when you are not

This is the costliest error: booking a paid exam believing you are ready. Among productions whose known level is below B2 (under NCLC 7), the examiner scored 0 of 14 in writing and 0 of 16 in speaking at B2 or above.

With so few cases, zero does not rule out a true rate of up to 22% (95% range). Only 5 writing and 9 speaking cases are at B1, just under the bar; the rest are easier to tell apart. The opposite error exists too: 2 of 10 (writing) and 4 of 20 (speaking) productions at B2 or above were given a lower level.

Scores near the NCLC 6/7 line

When an estimate falls within a point of a level boundary, we do not show a single number but a range, for example “9–10/20, borderline NCLC 6/7”, because the official result can fall either side. The NCLC 7 bar is 10 out of 20, so a displayed 10 does not promise NCLC 7.

Official results received

Official results received: 0 — 0 of them with a prior estimate

0 in writing, 0 in speaking. Entered by candidates who sat the real test, never shown individually.

These results will replace the public samples as the reference as they arrive. The numbers on this page do not use them yet. They arrive roughly 15 business days to 5 weeks after the sitting, so the counter grows slowly.

Our promise: this page is updated on every model or curve change. Current figures computed on 10 October 2026.

Listening and reading: indicative

For listening and reading no AI scores anything: the estimate for a ten-question set is a fixed rule. The estimated level is the highest level L such that at least two thirds of the questions up to level L, and at least half of those at level L, are right. That level is converted to an NCLC range with IRCC's table.

This rule has not yet been calibrated against official results, so we publish no accuracy rate for these two skills, and their estimate stays indicative, drawn from only ten questions.

Changelog

  1. 2026-10-11 · Smoother curve
    Raw 6 now maps to 7 (was 8): removes a 3-point swing seen in the test-retest study (same answer shown 8, 5, 5). Level accuracy unchanged: 7 and 8 are in the same NCLC 6 and CEFR B1 band.
  2. 2026-10-10 · Curve 2026-10-10-iso-s0.4
    Isotonic curve at 40% strength activated. Writing: exact 42% → 54%, within one level 88% → 100%. Speaking: exact 31% → 44%, within one level 81% → 94%. No false “ready” in the cross-checks.
  3. 2026-10-10 · Before the curve
    Raw scores from the claude-opus-5-5 model. Writing: exact 42%, within one level 88%, kappa 0.76. Speaking: exact 31%, within one level 81%, kappa 0.63. The model compressed scores above B1.

What these numbers do not tell you

Question

Are the tasks I am scored on original?

Yes. Every task is written for this product and we do not reproduce exam questions, so your score is what you would get on material you have never seen. About the AI

Try it yourself

Your first speaking task and your first writing task are free, with full feedback. No card needed.

Scores are an AI estimate, not an official result.

NCLC 7 is an independent practice tool. It is not affiliated with, endorsed by or connected to France Éducation international or the Government of Canada. "TCF" and "TCF Canada" belong to their owners and are used only to say which test this practice is for.

About the AI · How the TCF is assessed (France Éducation international)