Technical Deep Dive

One sentence into one paper,
across four levels of funnel

And past a gate that keeps unqualified questions out of the production bank.
Upstream is turning thousands of past-paper PDFs into a bank that is retrievable, classifiable and QA-able question by question. Downstream is collapsing one natural-language sentence into a set of executable retrieval constraints. This page is how both were built —— and what I fell into along the way.

api.teacheros.cn Express + Prisma + PostgreSQL 16 routes · 38 models · 65 test files ~70% of turns cost zero tokens
Edit-paper screen: each question carries its source and difficulty label
Architecture

A monolith. Enough. Replaceable.

No microservices at MVP stage. The selection criterion was exactly one: can a single person deploy and roll this back reliably?

Runtime path

WeChat Mini Program (native, 44 pages / 5 tabs)
  ↓ HTTPS
Nginx :443 · Let's Encrypt auto-renew
  ↓ 127.0.0.1:3000
PM2 · teacher-os-api
  ↓
Express 4 + TypeScript 5
  ↓ Prisma 6
PostgreSQL 14+ · daily backups

Three provider abstractions

So the MVP could close the loop first and swap in real implementations one at a time:

AImock / openai-compatible
paper building is wired to DeepSeek; homework/report generation is still mock
Storagelocal / cos / wechat-cloud
the last two are placeholders, no real SDK yet
RepositoryPrismaRepository / MemoryRepository
tests don't need a database
0backend route modules
0Prisma models
0database migrations
0Vitest files
0Mini Program pages
Slot Filling

If a rule can settle it, the model doesn't get it

The easy way to build natural-language paper generation is to throw the whole sentence at an LLM and ask for JSON back. I didn't, because that means every single turn costs money, adds latency, and can be unstable. What ships is a four-level funnel: whatever an earlier level can catch never reaches the next.

L1Intent lexicon —— direct hits on board / subject / action intentzero tokens
L2Catalogue lookup —— resolve unit and topic against a subject-scoped catalogue snapshotzero tokens
L3Regex —— strongly formatted slots: duration, count, markszero tokens
L4LLM fallback —— only when the first three still can't fill a critical slotcalls DeepSeek
Result About 70% of turns stop at L1–L3 and spend zero tokens. That's a cost decision and a stability decision at once —— rule-layer behaviour is deterministic, model-layer behaviour isn't. The parser you can type into on the product page is a simplified L1–L3.
Build-by-chapter screen: three-level filter over board, subject and unit, with question counts next to each unit
The "By chapter" route that I kept is essentially L1–L3's three slots turned into a manual UI: board, subject, unit. The natural-language route just compresses those three steps into one sentence —— underneath, both hit the same catalogue snapshot.

Why catalogue lookup has to be subject-scoped

Topic lookup supports bare topics —— a teacher says "integration" without naming a unit, and the system still has to find it. Which means retrieval must be able to resolve globally, with no unit context.

And once global resolution is allowed, the P0 below becomes inevitable.

P0 · five stacked root causes

A physics teacher got an economics question

Symptom

Intermittent, no pattern, and the tests asserted "pass". The front end showed nothing wrong either —— the unresolved-slots list was empty, so the UI looked perfectly healthy.

Root cause (five layers)

  1. Different subjects share the same bare unit codes (U1/U2/U4)
  2. subject was dropped while assembling parameters
  3. Ordering in the retrieval layer made one subject always win (a stable wrong answer, not a random one)
  4. Empty unresolved list → the front end could not notice
  5. Test assertions stopped at "ask for clarification" and never reached actual generation

Fix

  • Unit identity is scoped by subject —— the qualifier travels with the identifier
  • Catalogue snapshot is filtered by subject before retrieval
  • Added a zero-question guard: if nothing can be selected, raise —— never silently return an empty paper
  • Extended the tests all the way to real generation
The rule it left behind Anywhere a code is shared across dimensions, the qualifier must travel with the identifier. Filtering at the retrieval layer alone is not enough —— drop it anywhere upstream and downstream fails, consistently.
Difficulty

Turning "a bit harder" into a reproducible number

Early approach: infer from question number

Assume "higher number means harder". Two problems: the assumption was never validated, and the same request generated twice produced a drifting difficulty spread. When a teacher asked "on what basis is this one hard", there was no answer.

Now: deterministic percentile within the unit

Difficulty is the question's percentile inside its own unit. Same request twice, same difficulty spread; and every band is explainable as "this band means this percentile range within the unit".

Edit-paper screen with 'difficulty not set' in the top-right corner of each question card
I did not retouch this screenshot. Every question card reads "difficulty not set" —— the definition is settled, but backfilling runs in batches and this batch of P1 questions hasn't come up. How complete the feature design is and how complete the data coverage is are two different things; reporting them as one is inflation.

Why percentiles instead of absolute difficulty

Absolute difficulty isn't comparable across A-Level units in the first place —— "hard" in P1 and "hard" in FP3 are not the same thing. When a teacher says "P1, a bit harder", they mean harder within P1, not on some cross-unit absolute scale. A percentile expresses exactly that.

The side effect: newly loaded questions have no difficulty value until that unit's distribution is recomputed. That is the direct cause of the "difficulty not set" above, and the price this design charges.

Data Governance

The harder half is upstream: was each question classified correctly?

Whether paper building is usable at all comes down to one tedious thing —— whether every question was filed under the right topic. Having a teacher review thousands one by one isn't realistic; handing it entirely to a model isn't trustworthy.

rulesOfficial specification as the baseline constraintdraws the legal classification space
↓
humanTeacher-view topic ontology + teacher method tagsthe standard is set by people; models take no part in defining it. The physics ontology is on v3, converged from 44 nodes to 27
↓
modelLLM classifies question by questionthe grunt work that needs scale goes to the model
↓
modelTwo-agent independent review: solver and reviewer kept separateso no single agent grades its own work
↓
stronger modelSampled calibration by a stronger modelspecifically to offset same-source bias
↓
gateA / B / C triageproduction · review_queue · quarantine
↓
humanTeachers settle disagreements · taxonomy rebuilt to the teacher's standardfinal judgement always stays with people
↓
releaseTiered release: high-agreement ships first, disputed goes to internal testingneeds_review never reaches production
humans set the standard and make the call models label and review rules and gates

The A / B / C gate

A

production

High confidence, both agents agree → straight into the production bank, retrievable.

B

review_queue

Disagreement or middling confidence → into the review queue for a teacher to settle. Never ships to production.

C

quarantine

Low confidence or clearly anomalous → isolated; doesn't even enter the review queue until it's rerun or a human steps in.

Why the sampling check has to use a stronger, different model

I started out using the same model to both answer and review. The agreement rate looked great. Then I realised what it was: same-source bias. A model naturally agrees with its own judgement, so a high agreement rate only proves it is self-consistent —— not that it is right.

The fix was to bring in a stronger model from a different generation for sampled calibration, and to archive every disagreement as its own document. Two independent sets of data show why this was necessary.

Errors that come in clusters are invisible question by question

After rebuilding the physics ontology I ran a full review: across 637 accepted questions, the independent review agent agreed with the pipeline 85.9% of the time, and of the 90 disagreements 79 were ruled pipeline misclassifications. The important part is that those 79 were tightly clustered on 4 knowledge boundaries —— standing waves, emf and circuits, scattering and accelerators, vectors and SUVAT. Systematic errors in clusters like that never surface in spot checks; only a full cross-check brings them out.

The measured cost of same-source bias

Later, while raising quality across chemistry / biology / economics / psychology, I made stronger-model calibration a fixed step specifically to offset that bias. It caught 20 over-corrections by the review agent —— psychology's over-correction rate reached 25.8%, and one unit was ruled against 16 out of 16. Without switching models, those 20 would have been written into the bank as fixes. 164 confirmed corrections were written back in the end.

Definition statement · please read this with the numbers The "agreement rates" I quote (e.g. ~98% on FP2 and ~81% on FP3 in one subtopic relabelling round) are internal metrics: the denominator is the batch of questions being labelled, and the comparison is between the labelling model and an independent review agent —— not against human ground truth. It measures how much two independent judgements agree, and cannot be read as accuracy. A percentage without a stated denominator carries no information in an evaluation context —— which is why the definition sits next to the number.
Corpus

What's in the bank today

Six Edexcel IAL subjects, essentially all carrying their original figures. CIE and AQA have a classification pilot running and are labelled "in progress" inside the product.

SubjectQuestionsStatus and notes
Mathematics2,649includes the teacher method-tag dimension for FP1/FP2/FP3 (decoupled from chapter, used only by natural-language building)
Chemistry1,0560.9% residual contamination after the quality pass
Physics556ontology now on v3 (44 → 27 nodes)
Economics353essay questions need their dotted answer lines stripped before cropping
Biology317—
Psychology308—
Total5,239plus a dedicated multiple-choice set and a quick-drill module

Why build a teacher-view ontology at all

Official chapters are the exam board's view, filed against the specification. Teachers organise by "how I teach it, which method it needs". So alongside the chapter dimension I added teacher method tags as a second building dimension —— decoupled from chapter, serving natural-language building only. Because teachers often don't build a paper by chapter; they build it by method ("let's see whether they can do substitution").

An ontology can absolutely be too fine-grained

The main move in physics ontology v3 was convergence: 44 nodes cut to 27. Teachers won't use a taxonomy that's too fine either —— they don't carry that many boxes in their head. After the rerun, only one unit's review rate genuinely dropped and the rest held flat, which says the extra granularity was producing noise, not precision.

Rendering

From question crops to a paper you can print

What the bank stores are per-question image fragments. Turning those into a directly printable paper takes every step below, all of it hand-built.

Per-question cropping

Cut image fragments out of the source PDF question by question; questions spanning pages have to be stitched. Essay subjects like economics need their dotted answer lines stripped first, or the crop comes out as a field of dots.

Renumbering

Original numbers (Q3 / Q1 / Q10) are renumbered 1–4 on the new paper; multiple-choice and structured sections are numbered independently; the original number stays in the source line.

Flattened watermark

The watermark is composited into the page at render time rather than layered on top, so it can't be trivially removed.

Noise stripping

Source footers, rough-work areas and "rough working" bands are filtered out at crop time and never reach the new paper.

Two-file output

Question paper and mark scheme render into two separate PDFs. The mark scheme carries official mark points (B1 / M1 / A1).

Traceable footer

U:e3e17e38 · WS:65efa3b3 · B:5c9d8b42 · UTC —— user, paper ID, build number, generation time.

First page of the question paper PDF

Question Paper

2 pp · 4 questions · 27 marks · WMA11/01, WMA11/01A
New tab
First page of the mark scheme PDF with the marking grid

Mark Scheme

7 pp · official mark points B1 / M1 / A1
New tab
PDF If your browser won't embed it, use "New tab"
Engineering & Security

Evidence that it actually shipped, and isn't a prototype

Auth and permissions

  • WeChat code2Session → JWT
  • Every endpoint touching user data runs a relationship-based ownership check server-side
  • Tiered rate limits: login 5/min · general 100/min · AI 10/hr · files 30/hr
  • Zod input validation · Helmet · CORS · Morgan
  • Admin console gated by openid, with audit log and role impersonation

Delivery discipline

  • Production migrations only via prisma migrate deploy; db push is banned
  • Fixed change path: edit locally → rsync → build on the server → pm2 restart. Never edit code on the server
  • 65 Vitest files cover the P0 happy paths, authorisation and security, rate limiting, concurrency, API contracts, and paper building; a batch of known flaky tests remains, root cause not yet settled (below)
  • Daily database backups
One security pass before launch Parent binding changed from "whoever holds the invite code is bound" to student confirmation; ownership checks on course writes were tightened; booking now requires an existing teacher–student binding. Changes like these make the flow slower, but if the binding is wrong, every report and every learning record built on it is wrong too —— the error compounds all the way down.
War Stories

Symptom → root cause → response → the rule it left

The cross-subject leak is covered above; here are three more. The third one still isn't fixed, and I'm writing it down as it stands.

"412 duplicate questions" —— none of them were duplicates

Symptom: the dedup script reported 412 duplicates.
Root cause: dedup compared full_question_text, which is a page-level field —— several sub-questions on the same page all get the same block of text. The crops were completely correct. I nearly reran the entire pipeline over this.

Rule: a field's granularity is its semantics. Before deduping, confirm whether the field describes a page or a question.

"Request failed" wasn't a network problem

Symptom: "request failed" everywhere on a real device.
Root cause: I was testing through the admin role impersonation feature, and the impersonated identity used an ID with no real entity behind it —— so every server-side ownership check rejected it, exactly as designed.

Rule: device regression testing must use real accounts; impersonation is for UI preview only. The surface error and the real root cause frequently live on different layers.

Tests go red intermittently (unresolved)

Symptom: the same test batch is sometimes all green, sometimes red, and the victim file differs every time.
My original diagnosis: supertest spins up a temporary server per request, and port reuse creates a race. I built the "persist one global server" fix on that theory, and everything went green at the time.
Later overturned: vitest.config.ts says fileParallelism: false in plain sight —— test files don't run in parallel, so the race story doesn't hold. And that fix was never merged into the current branch. Seven independent reruns still reproduced it 3 times. It now looks more like cross-file leakage from a vi.doMock('../config/env.js') somewhere.

Rule: a plausible root cause is not a correct root cause; without falsification you haven't located anything. Also, log intermittent and deterministic failures separately —— with no local Postgres, dev-login fails every single run, and mixing that in with the flaky ones destroys your ability to judge a test run at all.

Limits & Next

What doesn't hold up yet

No golden set for generated papers

Classification quality has gates, review, calibration and internal agreement rates. But "is this the paper the teacher wanted" still rests on constraint satisfaction plus teacher review —— there is no formal eval set. Next step: a 20-paper teacher-scored set, to move "effectiveness evaluation" from intention to practice.

Inter-teacher agreement isn't quantified

Different teachers will genuinely disagree on how to classify the same question. Today that goes through review_queue for a teacher to settle, but I've never computed IAA / Kappa.

Difficulty labelling doesn't cover the bank

The percentile definition is settled, but backfilling runs in batches. Newly loaded questions have no difficulty value until that unit's distribution is recomputed.

Infrastructure is still thin

No CI/CD; single-point database; file storage is still local with COS / WeChat cloud as placeholders; the video path isn't deployed; the Mini Program hasn't been submitted for WeChat review.