Methodology
What we build, how we check it, and what we do not claim.
Content: generated, then gated
Every item is written by an AI model from the published format — the component, part or task type, question counts and time limits — and never from the provider's own materials. No test-provider items, passages, audio or sample answers are copied, paraphrased or used as training data, including free samples.
Before an item can be used it passes independent gates: an independent solver must reproduce the answer key; every wrong option must have a defensible reason to be wrong and only one option may be right; a difficulty target is assigned; a reviewer pass checks naturalness, Canadian context and ambiguity. Items enter the bank only when all gates agree. After publication, response statistics keep watching: an item that strong candidates get wrong, or that does not separate strong from weak, is quarantined automatically and replaced.
Strategy lessons are made the same way. An AI session writes each one from the published structure; a separate AI editor session, which never sees the writer's instructions, checks it against our rules: no test-provider material, no wording to reuse, no number that is not in the published format, and nothing about scoring beyond the published dimension names. A lesson that fails gets one rewrite, and it is published only when the editor passes it.
Vocabulary sets go through the same writer and editor. Each entry is a word or a short collocation with its meaning, a usage note where one helps, and a prompt to use it in a sentence of your own — never an example sentence to copy. The editor also checks that every meaning and usage note is accurate, and that no amount, deadline or rule is stated as fact, because those differ between provinces and change.
AI scoring is an estimate
Writing and Speaking are scored on the published performance-standard dimensions — Content/Coherence, Vocabulary, Readability, Task Fulfillment for Writing and Content/Coherence, Vocabulary, Listenability, Task Fulfillment for Speaking. Level-by-level band descriptions are not published, so the level layer is our own inference and is labelled that way on every screen: “Estimated level, based on the published CELPIP Performance Standards dimensions. Our scoring is not yet equated to the score scale the test itself uses, so read it as a range. Not a CELPIP score.”
Each response is scored in several independent passes and the median is reported as a range, for example CLB 8 (7–9). Deterministic checks — word count, time used, format — are done in code, not by a model. A wide spread is escalated to more passes; if it stays wide, the range is shown wide.
What we measure
- Difficulty deviation: median estimated level on our tests against the published mean per component.
- Scoring consistency: each response in our internal set is scored from scratch several times, each time the way the product scores it (several passes, a second model when they disagree, the median); we report the spread between the highest and lowest estimate — what you would see if you submitted the same response again.
- Raw pass spread: the same responses scored in ten single passes, before anything is combined; spread between the highest and lowest pass. Stricter than what you see, and published next to it so the combining step is not hiding anything.
- Answer-key accuracy: share of items where an independent solver agrees with the key.
- Reported items: user reports, verification outcome, and time to correction.
- Cohort accuracy: beta members share the level they receive after their test; we compute the mean absolute error of our estimate. This cannot be rushed — it waits for real tests.
The trust panel reads these values from the evaluation output. If a value is missing, the panel shows “not yet measured”.
Market review, September 2026
We reviewed 22 active preparation products (web, iOS and Android) and read more than 200 public reviews on app stores and review sites. We recorded prices, stated claims and complaint themes. Findings are published in aggregate only: no product is named in a negative context, and quotes are anonymised. Limits: reviews are self-selected, store ratings can be gamed, and our sample is one point in time.
What we do not claim
- We are not affiliated with the test provider and do not use its materials.
- Our estimate is not a CELPIP® score and cannot predict one for an individual.
- We do not sell memorised answers; the provider warns that non-original responses can void a score.
- Figures we cite from the provider or from IRCC are quoted with their source; where a link is still being verified we say so.