CELPIP · CLB 9 PATHWAYA CELPIP trainer built to show applied AI, data analysis and automation.
For learners: spaced phrase recall, AI writing feedback on the official criteria, and a quiz built from your own mistakes. For reviewers: a working multi-user product on Cloudflare Workers + D1 with Claude at its core.
AI application
Claude as a calibrated rater and coach, wrapped in code that checks its work.
- Essay grading on the four official CELPIP criteria (content, vocabulary, readability, task) with a CLB estimate per criterion and a minimal-edit revision shown as a word diff.
- Instant sentence check for every phrase card: verdict, rule-level corrections, alternatives, and collocations that can be added to the deck in one click.
- Structured output only: every call is forced through a JSON schema (tool use), so the app never parses free text.
- Guardrails after generation: errors must quote the learner’s text verbatim, words the model invented in a rewrite are flagged, unclear-meaning answers are marked low-confidence and kept out of the quiz.
- Prompts personalized from the learner’s settings: explanations in their first language, calibrated to their current and target CLB and exam date.
Data analysis
Every answer becomes a data point; the product turns them into decisions.
- Normalized study data in SQL: each review stores rating, recall latency and the interval before it; essays store per-criterion scores; every AI call stores tokens, latency and outcome.
- Recurring-error analysis: a fixed 21-tag error taxonomy (model tags, code counts) ranks each learner’s patterns across essays and compares this week with last week.
- Scheduling signals: slow-but-correct recalls are downgraded, and a 7-day review-load forecast tells the learner whether to add new cards today.
- Insights dashboard: observed forgetting curve with 95% Wilson intervals, memory half-life per number of prior successes (maximum-likelihood fit), recall-speed distribution, top error tags, essay score trend, AI cost per feature.
- Estimator validated against ground truth: a simulated learner (the app’s real scheduler plus a memory model with known parameters) can be loaded into any empty account, and the dashboard recovers the model’s true recall rates and stability growth.
Automation
The routine work runs itself — for the learner and for the operator.
- Adaptive daily queue: due cards first, a personal new-card budget, look-alike words kept 7 days apart, rest days rolled over automatically.
- Mistakes Quiz generated only from the learner’s own errors; items retire after two correct answers; a differing rewrite is re-graded by Claude automatically.
- Operations: one-command deploy (database provisioning, config generation, secret upload, bundle secret scan, test gate), self-applying schema migrations, nightly cleanup job.
- Batch localization pipeline (Message Batches API): Claude translates 1,500 cards into each supported language, code checks catch structural errors, a second Claude pass grades every translation, and low scores are re-translated with the reviewer’s note.in progress
Stack
Next.js App Router on Cloudflare Workers (vinext) · Cloudflare D1 (SQLite) · Anthropic Claude Messages API · TypeScript end to end · guest-first accounts with optional Google sign-in (OAuth 2.0 + PKCE) · unit-tested scheduling, quota and account-merge logic.