A complete mastery-based education system — curriculum graph, generated materials, fluency assessment, and a seven-year-old's real results — built at home in three days by one parent and his AI.
Zack educates Penelope (7) and Cora (5) at home — the kids run a morning school block before the day gets going. On July 13 a podcast on mastery learning crystallized three questions he'd been circling. Within about 36 hours, this system existed and is in daily use — the packet press had already been running a day ahead of it. This report covers why it was built, what it is, how it was built, and what happened when a real kid used it. Everything in here is real, and most of it is downloadable, including the packets themselves.
The trigger was an education podcast on mastery learning. By the end of it, three questions had sharpened into a spec.
How do we get mastery into our system — the packets and the kids' app?
How do I take a more deliberate approach to structuring a progressive curriculum graph, instead of picking topics ad-hoc?
How do I give Penelope — and ultimately Cora — agency in how they traverse that graph?
One clarification shaped everything that followed: mastery here means fluency, not just retention. Zack's reference point is second grade — timed arithmetic tests, where the bar was fast and accurate, not "remembers it eventually." The reasoning is straight out of cognitive load theory and the precision-teaching literature (Lindsley; Haughton's fluency aims; Binder): automaticity frees working memory, and fluent number facts are what make multi-step reasoning cheap later. Most systems stop at "got it right." This one measures correct-per-minute.
From the design contract, near-verbatim: packet topics were chosen ad-hoc — whatever gap surfaced that week. Mastery state was fragmented per-pack, with no global skill ledger. The kid had no structured say in sequencing, which wastes the strongest engagement lever there is: autonomy. And "mastery" meant retention only. Four defects, one root cause — there was no graph underneath any of it.
Cash buys showing up.
Mastery pays differently.
One loop. The parent runs it from a single console; the AI does everything else.
Every item below is real and appears later in this report.
150 nodes. 232 edges — 155 hard prerequisites, 77 soft transfer edges. Four domains, 29 strands, 14 fluency tiers.
The unit of the graph is a skill — not a lesson, not a worksheet. Skills connect two ways: requires edges are hard prerequisites that gate the frontier, and helps edges are soft transfer links that carry no gating but inform sequencing — clock quarters ↔ coin quarters ↔ fraction quarters; count-up change-making ↔ elapsed time. Fourteen nodes — the fundamentals where automaticity compounds — carry explicit fluency tiers with rate aims.
Know from memory all sums of two one-digit numbers (add/sub facts to 20)
Each domain has a different shape and different traversal rules — and this maps directly onto the kids' wall map.
| Domain | Shape | Traversal | v1 density |
|---|---|---|---|
| Math | mountain — steep prerequisite ridgelines | constrained frontier choice | dense — 90 nodes, counting through pre-algebra |
| Reading & Writing | city/library — decoding → fluency → comprehension → composition | semi-ordered | 21 nodes |
| Thinking | gear the agent carries, not territory | leveled through use inside other domains | 12 nodes |
| World Models | plains & wilds — breadth-first, weakly ordered | free-roam, interest-driven | 27 nodes |
Thinking skills are gear, not territory. Critical thinking taught as an abstract subject transfers poorly (Willingham, "Critical Thinking: Why Is It So Hard to Teach?"). So reasoning, metacognition, and inquiry nodes exist in the graph, but their resources are woven into other domains — a fractions packet carries an "explain why" item; readers carry evidence-hunt questions. Never a standalone logic curriculum for a seven-year-old.
World-models free-roam is a feature. Where sequencing matters least, the kid's agency is widest.
Adapted from Rick Gorman's LATTICE — a friend's parallel project, and credit to him for the frame — each node can carry up to four representations of the same skill: Concrete (manipulatives), Graphical (ten-frames, bar models, number lines), Symbolic (the number sentence), and Narrative (the story the skill lives in) — each written as a teaching move, a check question, and an entry prerequisite. It's the Concrete-Pictorial-Abstract spine that Singapore-style curricula run on, made machine-readable. The packet generator weights the rungs by the kid's current level: just entering → concrete-heavy; consolidating → pictorial into symbolic; fluency push → symbolic and timed.
Sources fanned out across the Common Core progressions documents, Singapore/Dimensions scope and sequence, Beast Academy, Khan Academy's prerequisite graph, Core Knowledge (Hirsch) for world models, and the precision-teaching fluency literature; node granularity was calibrated against Math Academy's live module taxonomy — the closest existing artifact, mastery plus spaced repetition over a knowledge graph. Then an adversarial review panel with named lenses walked the whole graph before it shipped: a prerequisite-gap walk, fluency-aim sanity against the literature, a measurability audit (every gate checkable with instruments that actually exist in this house), age-band sanity, and a maintenance-burden check. The panel demoted edges, split nodes, and flagged two fluency aims as ungrounded — those are marked UNGROUNDED in the source rather than silently kept.
The graph, live
This is the actual export from the system — not an illustration. Filter by territory, paint a kid's mastery over it, hover a node to trace its prerequisite chain, click for the full record.
Mastery
Frontier
Edges
Live data, not an illustration — Penelope's actual mastery state as of July 14, painted over the graph she traverses every morning.
Five levels, every gate measurable with instruments that actually exist in this house.
| Level | Name | The gate (measurable) |
|---|---|---|
| L0 | unseen | — |
| L1 | introduced | a teaching resource completed (packet finished, chapter done) |
| L2 | capable | ~90%+ accurate, untimed, on fresh generated items — never the taught examples |
| L3 | competent | retained: passes a spaced-repetition gate (3+ spaced reps, 4+ lifetime correct) holding over expanding intervals |
| L4 | proficient | automatic: rate-based — X correct/minute sustained across 2–3 separate days. Only for designated fluency nodes. |
Two design points do most of the work here. First, prerequisites unlock at L3 — retained — not L4. Fluency is a stretch goal, never a gate that blocks the path. Second, most nodes cap at L3; only 14 fundamentals — number facts, place-value operations, decoding rate, digit writing — carry L4 aims, because automaticity there pays compounding dividends everywhere else. And every ledger entry carries an evidence string — a human-readable justification citing the actual work:
L4 math · add & subtract within 5 — "Jul 14 fluency probe: 33 correct, 0 errors in one minute against a 20/min aim — clears the aim decisively."
Single-run today; the formal rule requires sustaining the rate on 2 of 3 separate days.
A six-probe printed kit. One minute per probe, parent-administered with a phone timer. Penelope ran it on July 14 at the kitchen table. The real results:
| Probe | Result (1 min) | Aim | Verdict |
|---|---|---|---|
| Add & subtract within 5 | 33 correct, 0 errors | 20/min | Fluent — locked clears the aim decisively |
| Add & subtract within 10 | 21 correct, 1 error | 25/min | Close, not yet re-test scheduled Jul 21 |
| Skip counting 2s/5s/10s | to 66 by 2s · 180 by 5s · 350 by 10s | 60/min | Building re-test Jul 21 |
| Writing digits | full sheet, 15 s to spare | — | Fluent hand speed is not the bottleneck |
| Letter sounds | full sheet, 2 s to spare | — | Fluent automatic |
| Trick words (sight words) | all, 24 s to spare, 1 error | — | Fluent the error: read "her" for "here" at speed |
The skip-count aim of 60/min is probably set too high for age 7 and is flagged for recalibration. Several probe sheets weren't dense enough to find her true ceiling — she finished with time to spare — so the next kit revision packs denser sheets. And these are single-run results; the aims formally require sustaining across 2 of 3 separate days. The instrument gets audited as hard as the kid.
That was Zack's explicit product bar — he asked for it "idiot proof — me being the idiot here." Scoring a fluency result auto-computes the next step: decisive pass → locked; borderline → confirm in 2 days; under the aim → re-test in 7 days. The console leads with a "Fluency re-tests" panel showing exactly what's due and when, with a one-click button that opens the kit. Telegram reminders fire on the due date. The parent never has to remember anything.
July 14, Penelope's 15-minute writing block. The prompt: "What will your life be like when you are 25?" Zack photographed her two handwritten pages and dropped them into the system.
"My job will be to be on stage. And I went to callegeg. My faveret book's will be the searees of the name of this book is secret. And I would like Art. I would like P.E and playing with my sister. And I will like burgers. And I like rollcosters. And I will still Be friend with Millie and Olivia and Penny and Gabby."
What the system extracted from one photographed freewrite:
One artifact, three ledger updates, zero extra work for the kid.
The kids' home app — "gumball," self-built, running on the family LAN — delivers a daily plan: a message from dad, academic modules, jokes, missions. Every packet ships with a matching SM-2 review pack; when the kid finishes the physical packet, its review pack activates, and the app reinforces the material on an expanding-interval schedule — dense for days, then weekly, then long-tail refreshers. Penelope's spelling program: 127 words, 52 mastered through the SM-2 gate as of this week. Mastery state flows both ways: her app work feeds the ledger — a July 14 analysis of her review responses moved silent-e/vowel-team spelling to L3 (retained) and identified the irregular/multisyllable tail as her next spelling edge.
Materials are compiled artifacts, generated from the graph on demand — not a content library. When a node reaches the frontier, the system can print for it.
The node's S/G/C/N rungs plus the kid's current level drive the page sequence — entering → concrete-heavy; consolidating → pictorial into symbolic; fluency → timed. A misconception inventory yields one honest "trap" item each, plus a counter-trap so the trap doesn't become the new rule. Production ≥ recognition: she constructs — draws the hands, writes the fraction — not just recognizes. Practice volume is allocated by difficulty, not page symmetry. And the nothing-used-before-taught walk: every technique a page needs must be taught earlier in the packet.
Any figure whose geometry or quantity is the lesson comes from a deterministic generator with verification wired into the build — clock faces with proportional hands, true-to-scale coins, exact fraction partitions, verified map routes. The image model never draws pedagogical content; it only does decorative B&W line art in the house style.
Every exercise is data. The figure, the prompt, and the answer key derive from the same datum — the key literally cannot disagree with the artwork.
Three independent AI reviewers — pedagogy; accuracy, which blind-reads every figure off the rendered page and then cross-checks; and production, covering pencil ergonomics, print survivability, and copy ambiguity. Findings are adjudicated, not auto-applied — reviewers have been measurably wrong, so each finding is classified: data bug vs. rendering vs. reviewer error. Across the first six packets the panel caught real blockers every single time: untaught notation, a count-up skill gap between adjacent pages, self-answering pages, answer boxes too small for a seven-year-old's "100", gray dashes that drop out on a laser printer, a "drawn true to size" claim that wasn't.
Three PDFs: the duplex packet, the answer key (never visible to the kid), and a keepsake certificate that never gets exercises printed on its back.
Each packet ships with its review pack for the kids' app — rung order matches the packet's own teaching sequence, content strictly limited to what the packet taught.
Kids love continuity, so the packets share one: the town of Maplewood; Time Bureau Press "Field Manual No. N" series numbering; a recurring villain, Shortchange Sam, who cheats with coins and fakes order tickets; and every packet ends CASE CLOSED with a named certificate — "Master of the Mint." A chapter-book kid notices when a cover promises a case that never happens, so narrative promised = narrative delivered is a review-panel rule.
These are the real files. Print one — they're built for a duplex laser printer. Cost per packet is on the order of single-digit dollars of AI compute (estimate) and roughly an hour of unattended generation and review; parent time ≈ pressing print.
The graph the parent sees as data, the kid sees as territory.
The Maplewood Master Map — one tabloid poster per kid — renders the graph as terrain: math as mountain ridgelines, reading as the city and its library, world models as plains and wilds, thinking skills as the gear the agent carries. Mastered skills are claimed territory. The frontier shows up as marked quests. Locked regions wait behind prerequisites. Version 1 is deliberately paper on the wall — stickers and stamps for claims — because paper-first is a house directive for the kids; screens have to earn their keep.
Each morning the kid chooses her next quest from the unlocked frontier. In constrained domains she gets 2–4 unlocked choices; in free-roam domains the field is wide open. The choice is real — the frontier guarantees anything she picks is productive — and it's the status currency doing the motivational work, not cash. There's proto-validation here: before any of this existed, Zack printed about five packets and let Penelope pick the order. She engages with the choice. The map just makes the implicit visible.
Five tabs, one loop: choose → deliver → score.
Technical note: the console is a set of ~15 typed capabilities (vita.education.*) — the same surface the UI, the AI agents, and the Sunday review all consume. React/FastAPI, inside vita, Zack's personal AI system.
Honest framing first: the system is three days old, and this is n=2 kids in one family with a motivated parent. Early signal, not a study.
Her own frontier on the same graph — 11 unlocked plus free-roam — with 6 ledger entries (3 retained), her own map poster, and 2 packets built for her so far (compare-numbers and digraphs). Phonics runs through reading.com and read-alouds; her track is screen-capped and read-aloud-forward.
Whether engagement holds at week 6. Whether the map's status economy stays motivating once novelty fades. Whether the fluency aims are calibrated right for this age. Whether the parent loop — scoring plus probe administration — stays under ~15 minutes a day at steady state. The instrumentation to answer each of those is in place.
The field log, dated. All of it is real.
Six packets in one day; the packet pipeline hardened into a skill. Every rule in the pipeline doc is a real finding from an adversarial review.
The podcast → the three questions → a 170-line structure contract pinning the load-bearing decisions: the ladder, the graph shape, the two currencies, the map.
The AI orchestrated the full build: research fan-out across the curriculum literature → domain authors drafting nodes in parallel → an adversarial review panel with named lenses → hardening. 150 nodes with citations. Same day: the capability layer, the teacher console, both kids' ledgers seeded conservatively from demonstrated evidence, and ten more packets.
An autonomous run — "imagine the system in daily use, find every gap, fill the highest-value ones with extreme care" — produced the wall-map posters, console packet generation, the Sunday-review agent rewired to the graph, and a documented (not half-built) integration plan for the kids' app.
Artifact ingestion shipped and the first artifact processed; the fluency kit generated, administered at the kitchen table, and processed; the idiot-proof re-test scheduler.
This is the part that generalizes. One AI orchestrator holds the vision and the judgment. Fleets of subagent "hands" do research, authoring, and implementation in parallel. And — critically — adversarial review is a standard build step, with findings adjudicated rather than auto-applied. The human's role compresses to what only the human can do: set values ("mastery means fluency," "paper first," "never cash on mastery"), supply ground truth about the learner, administer the one-minute probes, and teach. Everything else — research, authoring, figure generation, review, layout, scheduling, bookkeeping — is AI labor that runs while the family sleeps.
Context: all of this runs inside vita, Zack's personal AI operating system — memory, agents, schedulers, dashboards. The education OS is one subsystem of it. This report was itself produced by that system.
Everything referenced in this report, in one place. All files are the real artifacts the system produced and the kids used.
All 17 packets and their answer keys are downloadable from the gallery in §05.