Skip to content
celpip9

Trust panel

Every commitment on the landing page maps to a metric here. Values come from the evaluation pipeline output, never typed by hand. Last run 2026-09-25.

Practice difficulty vs. published 2024 levels

Not yet measured

Target deviation within ±0.25 level · Published when the measurement exists

Median estimated level on our practice tests compared with the published mean for each component.

Source: celpip9 evaluation pipeline output (packages/eval)

Scoring consistency (the estimate you see, same response scored repeatedly)

0.13 CLB

Target < 0.5 CLB · n = 30 · measured 2026-09-24

3 scorings per response · 30 non-anchor responses of content/eval · each scoring as in production: 3 passes (claude-opus-5-5), +2 on claude-opus-5-5 when they disagree by ≥ 1 level (1/90 scorings), per-dimension median · agent-session, one session per pass, no tools (ADR 0011, 0017)

The same response is scored from scratch several times, each time the way the product scores it (several passes, a second model when they disagree, the median); this is the spread between the highest and lowest of those estimates. The scoring count is published with the measurement.

Source: celpip9 evaluation pipeline output (packages/eval)

Raw pass spread (single scoring passes, same response scored repeatedly)

0.35 CLB

Target < 0.5 CLB · n = 30 · measured 2026-09-24

9 passes per response · 30 non-anchor responses of content/eval · the 9 base passes of the estimate run (claude-opus-5-5, effort high), escalation passes excluded · agent-session, one session per pass, no tools (ADR 0011, 0017)

Spread between the highest and lowest of single scoring passes on the same response, before any passes are combined. Stricter than what you see: the product never shows a single pass. The pass count is published with the measurement.

Source: celpip9 evaluation pipeline output (packages/eval)

Answer-key accuracy

100%

Target ≈ 100% (independent solver agrees with the key) · n = 1014 · measured 2026-09-25

f-001 + f-002 + f-003 + f-004 + f-005 + f-006 + f-007 + f-008 + f-009 + f-010 + f-011 + f-012 + f-013 (the thirteen bank forms): 1014 MCQ items across 130 parts · solver: agent-session · run in a separate agent session per part, not per item (ADR 0011)

Share of items where an independent solver reproduces the answer key. Items that fail never enter the bank.

Source: celpip9 evaluation pipeline output (packages/eval)

Reported and corrected items

Not yet measured

Target published monthly · Published when the measurement exists

Items flagged by users, verified, and corrected or quarantined — with the time it took.

Source: celpip9 evaluation pipeline output (packages/eval)

Estimate vs. reported result (beta cohort)

Not yet measured

Target MAE ≤ 1 CLB · Published when the measurement exists

Mean absolute error between the estimate we gave before the test and the level test takers report afterwards, with their consent. Published from 30 test takers, only as a total — never an individual result.

Source: celpip9 evaluation pipeline output (packages/eval)

The reference: published 2024 levels

Our difficulty target is the published distribution, restated in our own words. Source: CELPIP Data Report of 2024 Test Takers (Paragon Research Reports, PDF).

ComponentMean levelReached 9+Reached 7+
Listening8.2252.0%74.5%
Reading7.2133.3%58.0%
Writing7.9129.8%80.4%
Speaking7.7632.6%72.6%

How each metric is computed: methodology.