Product Portfolio · Teacher OS / 可见课堂

One sentence from a teacher,
one paper they can print

A WeChat Mini Program for independent A-Level / IB / AP tutors. It takes the three things that drain a tutor's week —— building papers, marking homework, and answering to parents —— and turns them into one production line that compounds.
Backend is live, six subjects are open, and the whole paper-building path runs end to end for real.

api.teacheros.cn backend live 6 subjects · 5,239 questions sentence → PDF real output, end to end
Natural-language paper building inside the Mini Program: the teacher describes what they want in one sentence
The Problem

What a teacher asks for is not what a question bank understands

Here is what it actually takes for an independent tutor to build one targeted worksheet for one student: dig out a dozen past-paper PDFs → hunt page by page for questions on that topic → screenshot → paste into Word → go find the matching answers in another file → lay it out → export. All of that, for four questions. And the next student needs a different topic, so the whole thing starts over.

The sentence in the teacher's head
"I want a set of P1 integration questions, not too hard, not too easy, about 30 minutes." —— an actual sentence typed into the product
How the bank is organised

WMA11/01 · Unit P1 · 9709 Paper 3 · chapter codes · session codes

That coding is the exam board's view —— built for setting and marking papers, filed against the specification. A teacher works in the teaching view: how you explain this kind of question, which method it needs, where students usually get stuck.

Conclusion Two languages that don't line up, with one blunt consequence: make a teacher search by specification codes and they simply won't use it. So the entry point to this product isn't a filter. It's a sentence.
Live Replay

From that sentence to that PDF —— what happened in between

Below is a frame-by-frame replay of one real session on 2026-07-30. Left is what the system is doing, right is what the teacher sees, and at the end you can open the PDFs that run actually produced.

Session Replay Real session · not a live call · 2026-07-30 13:39 UTC
The Mini Program screen during the replay

Look at difficulty → medium (percentile within unit) in the second group: difficulty is not guessed from question numbers, it's a percentile inside the same unit. Ask twice with the same sentence and you get the same difficulty spread. This one comes back later, because it has an honest tail.

This replay ran on a student account Every screenshot on this page, and both PDFs, come from a real session under a student identity —— that's why the generated PDF footer says student, and I didn't retouch it out. Paper building is one engine behind two entry points: teachers use it to build a paper and then "assign to student" (that button is right there on the preview screen), students use it to drill by chapter on their own. The product still centres on the teacher; student self-practice is the same capability reused —— one set of retrieval constraints, two use cases. That is exactly why I built paper generation as a standalone capability instead of burying it inside the homework flow.
Try It

Type a sentence yourself

This isn't a toy —— it is levels 1–2 of the same parsing funnel that ships in the product: intent lexicon + catalogue lookup + regex, with no model call at all. Around 70% of real conversation turns end right here, spending zero tokens.

Slot Filling front-end rule demo, not wired to the live bank
The Flow

Nine screens, one complete path

All real device screenshots. The phone changes screen as you read —— and the yellow note under each step is the product judgement behind that screen.

STEP 01

Put the entry point where you can't miss it

That raised yellow "Build a paper" button in the middle of the tab bar is the only element in the whole app that breaks out of the tab row.

Judgement: this is something the user does every day; it cannot live two levels deep. Where you put an entry point is your bet on how often it gets used.
STEP 02

Two routes —— nobody is forced to talk to the app

"Natural language (recommended)" sits next to "By chapter". The second one is an honest three-level filter: board → subject → unit.

Judgement: not everyone is comfortable describing what they want to an app. Keep the old way intact, mark the new way as recommended —— let users migrate themselves instead of forcing them.
STEP 03

One sentence, then the system reads it back

It doesn't generate straight away. It restates what it understood —— "Edexcel Maths P1 integration practice, moderate difficulty, completable within 30 minutes" —— and tells you 30 questions are available in that range.

Judgement: restating hands responsibility back to the human. If the AI misread the request, the teacher stops it here, instead of discovering it after holding a useless paper.
STEP 04

Every question can be overruled

Swap, move up, move down, delete. Each one carries its full provenance: original question, January 2023 session · Pure Mathematics P1 (WMA11) · Q3.

Judgement: an AI-built paper has to be rejectable question by question, or no teacher will dare send it to a student. Overrulable matters more than more accurate.
STEP 05

See what it actually tests, before printing

A chapter distribution bar: Integration, 4 questions, 27 marks, 100%. Plus question count, total marks, estimated time, chapters covered.

Judgement: the teacher has to explain to students and parents what this paper drills. This chart isn't for the system —— it's the thing they'll paraphrase.
STEP 06

Slow operations need visible progress

Rendering the PDF takes a few seconds to a dozen; the screen shows a live progress bar throughout.

Judgement: a wait with no progress bar reads as "frozen", and the user quits. This is not a place to save effort.
STEP 07

Paper and mark scheme must be two files

When it finishes you get two separate entry points: view question paper / view mark scheme. Plus "generate a parallel paper" and "assign to student".

Judgement: put them in one file and the teacher forwards the answers along with the questions. This is the kind of trap only someone who has actually used it walks into.
STEP 08

The artefact itself has to hold up

Standard exam layout: TIME / PAPER REFERENCE(S) / QUESTIONS / TOTAL MARKS, source session and original question number above each question, watermark across every page, provenance in the footer.

Judgement: the teacher prints this and hands it to a student. It cannot look like "something a tool exported". It has to look like a paper.
STEP 09

Papers you've built stay around

The list separates "draft" from "generated", and papers can be duplicated or deleted. Duplicate one, swap two questions, and that's the next student's paper.

Judgement: one-shot tools don't compound. Only when the output is reusable does the second use cost less than the first.
Student home screen, with the paper-builder entry at the centre of the tab bar Choosing how to build: natural language or by chapter The assistant restates the request and reports how many questions are available Editing the paper: swap, delete and reorder each question, every one showing its source Paper preview with chapter distribution and paper theme colour Progress bar while the PDF is being rendered Done: question paper and mark scheme as two separate entry points First page of the generated question paper PDF Paper list separating drafts from generated papers
Design Decision

If the bank holds 3, it must not pretend it can give you 10

That unremarkable line in the replay —— "30 questions available in this range" —— is the design rule I've held onto hardest in this product.

Before selecting anything, the system checks stock and tells the teacher the real number. If a topic only has 3 questions, it won't pad the count, won't quietly borrow from an adjacent topic, and definitely won't repeat the same question to fill the gap. It says "there are only 3", and offers alternatives.

This rule makes the product look less capable at certain moments. What it buys is the thing that matters: the teacher can trust that every question on the paper is one they asked for.

Why this has to be a product call

Padding to the requested count is easier to build, and it looks better on the dashboard —— 100% generation success.

But a teacher only has to find one irrelevant question slipped into a paper, and they will never use the feature again. That trade-off isn't a technical question. It's a decision only someone who understands the usage context makes.

The Artefact

These are the two files that run produced

Not mockups —— what came back after tapping "Generate paper PDF" on a real device. Open them and page through.

First page of the question paper PDF

Question Paper

2 pp · 4 questions · 27 marks · 35 min
New tab
First page of the mark scheme PDF with the marking grid

Mark Scheme

7 pp · official mark points B1 / M1 / A1
New tab
PDF If your browser won't embed it, use "New tab" above
Per-question provenanceAbove each question: Jan 2023 · WMA11/01 · Q3 · 5 marks. The teacher can judge where it came from and how much it's worth, instantly.
Watermark on every pageThe watermark is flattened into the page, not layered on top —— bank content shouldn't leak out without a trace.
Traceable footerThe footer records the user, paper ID, build number and UTC time. If a paper leaks, you can trace where from.
RenumberingOriginal numbers (Q3, Q1, Q10) are renumbered 1–4 on the new paper, while the original number stays in the source line.
Consistent layoutTIME / PAPER REFERENCE(S) / QUESTIONS / TOTAL MARKS —— matching the information structure of a real exam paper.
Answers kept separateQuestion paper and mark scheme are two files, so a teacher forwarding one doesn't send the answers with it.
Scale

These questions weren't bought from an API

I cut them out of past-paper PDFs myself, classified them one by one, ran them through QA, then loaded them in. All six subjects are open, and essentially every question carries its original figure.

0questions live (six subjects)
0subjects open
0carry the original figure %

Six-subject split

Mathematics2,649
Chemistry1,056
Physics556
Economics353
Biology317
Psychology308

Unit-level granularity in maths

Retrieval resolves to the unit, not the subject. The numbers below come straight out of the unit picker inside the product.

P1227
P2208
P3173
P4171
FP1230
FP2152
FP3152
S1225
S2209
S3144
M1–M3Mechanics
D1Decision
Build-by-chapter screen: pick board, subject and unit, with the question count printed next to each unit
Those per-unit counts aren't a separate tally I ran —— they're what the "By chapter" screen shows in the product. When you pick a unit, the number next to it is how many questions it holds right now.
Being straight about this Only Edexcel is fully open today. In the product, CIE and AQA are explicitly labelled "in progress" (visible in the first row of the screenshot). I didn't make them greyed-out but tappable —— letting a user tap in and find nothing is worse than saying "not yet" up front.
Why the count sits on the option itself When a teacher picks a unit, what they most want to know is how many questions they have to choose from. Printing stock directly on the option means they don't have to tap in to find out. This and the "30 questions available in this range" line in the chat are the same principle landing twice: show people the constraint before they commit to a decision.
What I Cut

What I removed says more about my judgement than what I added

Every line below was deliberately cut from the original concept. None of them is "didn't get around to it".

AI writes student feedback → AI reviews it → publish → the teacher writes and publishes it
Education is a trust-sensitive setting. Letting an AI speak directly to students and parents carries far more risk than the time it saves. I pulled AI back from speaking for the teacher to doing work for the teacher —— it stays in paper building and lesson-recording summaries, where it saves time and mistakes have a fallback, and stays out of anything that makes a statement to users on a human's behalf. That's a call about capability boundaries, not about what the technology can do.
Two-way IM + subscription push → one-way system notifications
WeChat Mini Programs have qualification and capability limits around social features, private messaging and subscription messages. Rather than build an IM the platform could cut off at any time —— and that hands users unread-badge anxiety —— just guarantee that what matters gets delivered. Cutting scope under a platform constraint beats fighting it.
Local recording upload + hosted transcription → external links
Hosting large video plus transcription is expensive, and "how many teachers actually use recording summaries" hasn't been validated yet. Close the loop with links first, get usage data, then decide whether to spend that money.
Inferring difficulty from question number → deterministic percentile within the unit
"Higher question number means harder" is an unvalidated assumption, and the same request generated twice would drift. With percentiles, difficulty became reproducible and explainable —— when a teacher asks "on what basis do you call this one hard", I have an answer.
Community / content feed → cut, placeholder page only
The MVP keeps the core loop and nothing else. A community with no content supply ships as a blank page.
Limits

What doesn't hold up yet

In the body text, not hidden in footer small print.

Difficulty labelling isn't finished

In the "edit paper" screenshot above, the top-right corner of each question reads "difficulty not set" —— I didn't retouch it. The definition (percentile within unit) is settled, but backfilling runs in batches and this batch of P1 questions hasn't come up yet. Feature design and data coverage are two different things, and shouldn't be reported as one.

No eval set for paper quality

Classification quality has a full QA pipeline. But "is this the paper the teacher actually wanted" currently rests on constraint satisfaction and human review —— there is no formal eval set. Building one is the next step, even if it's only 20 papers scored by teachers.

Small teacher-interview sample

The claim that "official taxonomy doesn't work for teaching" comes from working with front-line teachers and rebuilding the taxonomy to a teacher's standard —— but I never turned that into systematic interviews, so the sample is small.

Not yet submitted for WeChat review

The backend is live and stable; the Mini Program is still on the trial channel. Also, only the paper-building side is wired to a real conversational AI —— homework and report generation still run the default mock. Growth work: none at all.