Trust panel
Every commitment on the landing page maps to a metric here. Values come from the evaluation pipeline output, never typed by hand. Last run 2026-09-25.
Practice difficulty vs. published 2024 levels
Not yet measured
Target deviation within ±0.25 level · Published when the measurement exists
Median estimated level on our practice tests compared with the published mean for each component.
Scoring consistency (the estimate you see, same response scored repeatedly)
0.13 CLB
Target < 0.5 CLB · n = 30 · measured 2026-09-24
3 scorings per response · 30 non-anchor responses of content/eval · each scoring as in production: 3 passes (claude-opus-5-5), +2 on claude-opus-5-5 when they disagree by ≥ 1 level (1/90 scorings), per-dimension median · agent-session, one session per pass, no tools (ADR 0011, 0017)
The same response is scored from scratch several times, each time the way the product scores it (several passes, a second model when they disagree, the median); this is the spread between the highest and lowest of those estimates. The scoring count is published with the measurement.
Raw pass spread (single scoring passes, same response scored repeatedly)
0.35 CLB
Target < 0.5 CLB · n = 30 · measured 2026-09-24
9 passes per response · 30 non-anchor responses of content/eval · the 9 base passes of the estimate run (claude-opus-5-5, effort high), escalation passes excluded · agent-session, one session per pass, no tools (ADR 0011, 0017)
Spread between the highest and lowest of single scoring passes on the same response, before any passes are combined. Stricter than what you see: the product never shows a single pass. The pass count is published with the measurement.
Answer-key accuracy
100%
Target ≈ 100% (independent solver agrees with the key) · n = 1014 · measured 2026-09-25
f-001 + f-002 + f-003 + f-004 + f-005 + f-006 + f-007 + f-008 + f-009 + f-010 + f-011 + f-012 + f-013 (the thirteen bank forms): 1014 MCQ items across 130 parts · solver: agent-session · run in a separate agent session per part, not per item (ADR 0011)
Share of items where an independent solver reproduces the answer key. Items that fail never enter the bank.
Reported and corrected items
Not yet measured
Target published monthly · Published when the measurement exists
Items flagged by users, verified, and corrected or quarantined — with the time it took.
Estimate vs. reported result (beta cohort)
Not yet measured
Target MAE ≤ 1 CLB · Published when the measurement exists
Mean absolute error between the estimate we gave before the test and the level test takers report afterwards, with their consent. Published from 30 test takers, only as a total — never an individual result.
The reference: published 2024 levels
Our difficulty target is the published distribution, restated in our own words. Source: CELPIP Data Report of 2024 Test Takers (Paragon Research Reports, PDF).
| Component | Mean level | Reached 9+ | Reached 7+ |
|---|---|---|---|
| Listening | 8.22 | 52.0% | 74.5% |
| Reading | 7.21 | 33.3% | 58.0% |
| Writing | 7.91 | 29.8% | 80.4% |
| Speaking | 7.76 | 32.6% | 72.6% |
How each metric is computed: methodology.