Skip to content
celpip9

Methodology

What we build, how we check it, and what we do not claim.

Content: generated, then gated

Every item is written by an AI model from the published format — the component, part or task type, question counts and time limits — and never from the provider's own materials. No test-provider items, passages, audio or sample answers are copied, paraphrased or used as training data, including free samples.

Before an item can be used it passes independent gates: an independent solver must reproduce the answer key; every wrong option must have a defensible reason to be wrong and only one option may be right; a difficulty target is assigned; a reviewer pass checks naturalness, Canadian context and ambiguity. Items enter the bank only when all gates agree. After publication, response statistics keep watching: an item that strong candidates get wrong, or that does not separate strong from weak, is quarantined automatically and replaced.

Strategy lessons are made the same way. An AI session writes each one from the published structure; a separate AI editor session, which never sees the writer's instructions, checks it against our rules: no test-provider material, no wording to reuse, no number that is not in the published format, and nothing about scoring beyond the published dimension names. A lesson that fails gets one rewrite, and it is published only when the editor passes it.

Vocabulary sets go through the same writer and editor. Each entry is a word or a short collocation with its meaning, a usage note where one helps, and a prompt to use it in a sentence of your own — never an example sentence to copy. The editor also checks that every meaning and usage note is accurate, and that no amount, deadline or rule is stated as fact, because those differ between provinces and change.

AI scoring is an estimate

Writing and Speaking are scored on the published performance-standard dimensions — Content/Coherence, Vocabulary, Readability, Task Fulfillment for Writing and Content/Coherence, Vocabulary, Listenability, Task Fulfillment for Speaking. Level-by-level band descriptions are not published, so the level layer is our own inference and is labelled that way on every screen: “Estimated level, based on the published CELPIP Performance Standards dimensions. Our scoring is not yet equated to the score scale the test itself uses, so read it as a range. Not a CELPIP score.”

Each response is scored in several independent passes and the median is reported as a range, for example CLB 8 (7–9). Deterministic checks — word count, time used, format — are done in code, not by a model. A wide spread is escalated to more passes; if it stays wide, the range is shown wide.

What we measure

The trust panel reads these values from the evaluation output. If a value is missing, the panel shows “not yet measured”.

Market review, September 2026

We reviewed 22 active preparation products (web, iOS and Android) and read more than 200 public reviews on app stores and review sites. We recorded prices, stated claims and complaint themes. Findings are published in aggregate only: no product is named in a negative context, and quotes are anonymised. Limits: reviews are self-selected, store ratings can be gamed, and our sample is one point in time.

What we do not claim