Put the entry point where you can't miss it
That raised yellow "Build a paper" button in the middle of the tab bar is the only element in the whole app that breaks out of the tab row.
A WeChat Mini Program for independent A-Level / IB / AP tutors. It takes the three things that drain a tutor's week —— building papers, marking homework, and answering to parents —— and turns them into one production line that compounds.
Backend is live, six subjects are open, and the whole paper-building path runs end to end for real.

Here is what it actually takes for an independent tutor to build one targeted worksheet for one student: dig out a dozen past-paper PDFs → hunt page by page for questions on that topic → screenshot → paste into Word → go find the matching answers in another file → lay it out → export. All of that, for four questions. And the next student needs a different topic, so the whole thing starts over.
"I want a set of P1 integration questions, not too hard, not too easy, about 30 minutes." —— an actual sentence typed into the product
WMA11/01 · Unit P1 · 9709 Paper 3 · chapter codes · session codes
That coding is the exam board's view —— built for setting and marking papers, filed against the specification. A teacher works in the teaching view: how you explain this kind of question, which method it needs, where students usually get stuck.
Below is a frame-by-frame replay of one real session on 2026-07-30. Left is what the system is doing, right is what the teacher sees, and at the end you can open the PDFs that run actually produced.

Look at difficulty → medium (percentile within unit) in the second group: difficulty is not guessed from question numbers, it's a percentile inside the same unit. Ask twice with the same sentence and you get the same difficulty spread. This one comes back later, because it has an honest tail.
This isn't a toy —— it is levels 1–2 of the same parsing funnel that ships in the product: intent lexicon + catalogue lookup + regex, with no model call at all. Around 70% of real conversation turns end right here, spending zero tokens.
All real device screenshots. The phone changes screen as you read —— and the yellow note under each step is the product judgement behind that screen.
That raised yellow "Build a paper" button in the middle of the tab bar is the only element in the whole app that breaks out of the tab row.
"Natural language (recommended)" sits next to "By chapter". The second one is an honest three-level filter: board → subject → unit.
It doesn't generate straight away. It restates what it understood —— "Edexcel Maths P1 integration practice, moderate difficulty, completable within 30 minutes" —— and tells you 30 questions are available in that range.
Swap, move up, move down, delete. Each one carries its full provenance: original question, January 2023 session · Pure Mathematics P1 (WMA11) · Q3.
A chapter distribution bar: Integration, 4 questions, 27 marks, 100%. Plus question count, total marks, estimated time, chapters covered.
Rendering the PDF takes a few seconds to a dozen; the screen shows a live progress bar throughout.
When it finishes you get two separate entry points: view question paper / view mark scheme. Plus "generate a parallel paper" and "assign to student".
Standard exam layout: TIME / PAPER REFERENCE(S) / QUESTIONS / TOTAL MARKS, source session and original question number above each question, watermark across every page, provenance in the footer.
The list separates "draft" from "generated", and papers can be duplicated or deleted. Duplicate one, swap two questions, and that's the next student's paper.
That unremarkable line in the replay —— "30 questions available in this range" —— is the design rule I've held onto hardest in this product.
Before selecting anything, the system checks stock and tells the teacher the real number. If a topic only has 3 questions, it won't pad the count, won't quietly borrow from an adjacent topic, and definitely won't repeat the same question to fill the gap. It says "there are only 3", and offers alternatives.
This rule makes the product look less capable at certain moments. What it buys is the thing that matters: the teacher can trust that every question on the paper is one they asked for.
Padding to the requested count is easier to build, and it looks better on the dashboard —— 100% generation success.
But a teacher only has to find one irrelevant question slipped into a paper, and they will never use the feature again. That trade-off isn't a technical question. It's a decision only someone who understands the usage context makes.
Not mockups —— what came back after tapping "Generate paper PDF" on a real device. Open them and page through.
I cut them out of past-paper PDFs myself, classified them one by one, ran them through QA, then loaded them in. All six subjects are open, and essentially every question carries its original figure.
Retrieval resolves to the unit, not the subject. The numbers below come straight out of the unit picker inside the product.
Every line below was deliberately cut from the original concept. None of them is "didn't get around to it".
In the body text, not hidden in footer small print.
In the "edit paper" screenshot above, the top-right corner of each question reads "difficulty not set" —— I didn't retouch it. The definition (percentile within unit) is settled, but backfilling runs in batches and this batch of P1 questions hasn't come up yet. Feature design and data coverage are two different things, and shouldn't be reported as one.
Classification quality has a full QA pipeline. But "is this the paper the teacher actually wanted" currently rests on constraint satisfaction and human review —— there is no formal eval set. Building one is the next step, even if it's only 20 papers scored by teachers.
The claim that "official taxonomy doesn't work for teaching" comes from working with front-line teachers and rebuilding the taxonomy to a teacher's standard —— but I never turned that into systematic interviews, so the sample is small.
The backend is live and stable; the Mini Program is still on the trial channel. Also, only the paper-building side is wired to a real conversational AI —— homework and report generation still run the default mock. Growth work: none at all.