# Curriculum Graph — Structure Contract

*Drafted 2026-07-13 (CLI session with bio-zack, post alpha-school podcast episode).
This doc is the INPUT to a dedicated research/orchestration session that fills the
graph in robustly (deep research, adversarial review). It pins the load-bearing
structural decisions and marks what the research session must settle. Forge spark 064;
sibling spark 063 (review packs — SHIPPED as gumball review-pack machinery 2026-07-13).*

Owner: bio-zack. Operating agent: Learning agent (`data/education/working-memory.md`).
Kids: Penelope (b. Mar 2019, profile 3), Cora (b. ~2021, profile 2).

## Why (the problem)

Packet topics are chosen ad-hoc (whatever gap surfaces). Mastery state is fragmented
per-pack in gumball with no global skill ledger. The kid has no structured say in
sequencing — which wastes the strongest engagement lever (autonomy). And "mastery"
currently means retention only; there is no notion of FLUENCY (fast + accurate), which
is where fundamentals pay compounding dividends.

## Pinned decisions

### 1. Mastery is a ladder, not a bit

Per (skill, kid), levels with measurable gates:

| level | name | gate (measurable) |
|-------|------|-------------------|
| L0 | unseen | — |
| L1 | introduced | a teaching resource completed (packet finished, chapter done) |
| L2 | capable | accurate untimed on FRESH generated items (~90%+, no time pressure) |
| L3 | competent | retained: passes the SM-2 spaced gate (3+ spaced reps, 4+ lifetime correct — the shipped review-pack gate), holds over expanding intervals |
| L4 | proficient | automatic: rate-based — X correct/min sustained across 2–3 separate days. ONLY for fluency-tier nodes. |

- **Fluency tier is a per-node attribute.** Most nodes cap at L3. A deliberately small
  set of fundamentals (number facts, place-value operations, decoding→reading rate,
  number writing/typing) carry L4 targets with explicit rate aims.
- Rationale: automaticity frees working memory — fluent number facts are what make
  multi-step reasoning cheap later (cognitive load theory; precision-teaching
  literature). "Capable → competent → proficient" is bio-zack's framing; keep it.
- Rate aims come from the precision-teaching / fluency literature (Haughton aims,
  Binder), age-adjusted. **Research session pins the numbers per node.**
- Instrumentation: (a) gumball timed-sprint module type (60s, count correct,
  digits/min, personal-best curve — kids love PB curves); (b) printed timed sheets via
  packetlib (the classic 100-problem page), parent enters the score. Both feed the same
  rate history. Response latency capture in gumball responses is a prerequisite — audit
  what `response_payload` already records.

### 2. Graph shape

- **Node = skill** (not a packet, not a module). Fields: id, domain, title, description,
  target tier (L3 or L4), fluency aim (rate spec, nullable), assessment spec (which
  generative item types measure each level), rough age/stage band, source citations.
- **Edges**: `requires` (hard prerequisite), `helps` (soft transfer — e.g. clock
  quarters = coin quarters = fraction quarters; count-up change = elapsed time. The
  ed-packet panels discovered these empirically; they're first-class edges),
  `parallel` (interleavable).
- **Resources attach to nodes**: ed packets, gumball generators/modules, Beast Academy
  chapters, Singapore units, books/read-alouds, and real-world activities ("visit the
  water treatment plant" is a resource). A node can have zero resources (frontier =
  build/buy signal).
- **Ledger per (kid, node)**: current level, evidence trail (responses, packets, dates),
  rate history for fluency nodes. Per-kid frontiers over a shared graph.
- **Frontier query**: unmastered nodes whose `requires` edges are satisfied at the
  required level. This is what drives next-packet selection and the kid's choice set.

### 3. Multi-domain from the start, filled at different densities

Domains are different TERRAIN with different traversal rules (this maps directly onto
the kid-facing map — see §5):

| domain | shape | traversal rule | density at v1 |
|--------|-------|----------------|---------------|
| math | mountain: steep prereq ridgelines | constrained frontier choice (2–4 unlocked) | DENSE (~60–90 nodes, ages ~6–9 band) |
| reading/writing | city/library: decoding→fluency→comprehension→composition | semi-ordered | skeleton (~15–25 nodes) |
| thinking skills (logic, argumentation, question-asking, metacognition) | GEAR the agent carries, not territory | embedded: exercised inside other domains' materials, leveled by use | skeleton (~10–15 nodes) |
| world models (physics intuitions, how-things-work, civics, water cycle, money/economy) | plains/wilds: breadth-first, weakly ordered | free-roam, interest-driven | skeleton (~15–25 nodes) |

- **Thinking-skills caveat (research-grounded, do not violate):** critical thinking
  taught as an abstract subject transfers poorly (Willingham). These nodes exist in the
  graph, but their RESOURCES are woven into other domains — a fractions packet carries
  an "explain why" production item; readers carry inference/evidence-hunt questions
  (they already do). Tag exercises with the thinking skills they exercise; level up
  through cross-domain use. Never a standalone "logic curriculum" for a 7-year-old,
  with the narrow exception of puzzle-type resources she already enjoys.
- World-models domain: weak ordering is a FEATURE — sequencing matters less, so her
  agency can be widest here. Natural resource types: readers, read-alouds, Macaulay-style
  "how things work" books, field trips.

### 4. Storage & system boundaries

- **Graph is data in vita, versioned in git**: `data/education/curriculum/` (yaml or
  json — research session picks conventions and documents them here).
- **Gumball consumes a projection**: a `skills` table; review-pack concepts map onto
  skill nodes (many-to-one allowed); mastery rolls up from the existing
  `ReviewPractice` records. Sync like review packs (admin endpoint / startup).
- Vita side (ed-packet skill, Learning agent) authors against the graph; the Sunday
  weekly review reports the frontier per kid instead of ad-hoc gap-picking.

### 5. Kid-facing layer: the map + two currencies

- Render the graph AS the Maplewood / Time Bureau universe (the shared universe from
  six packets). Territory claimed on mastery; regions locked until prereqs met; she
  picks the next case from the unlocked frontier. Terrain per domain per §3.
- **v1 is a printed wall map** (paper > devices directive; mapgen exists) with
  stickers/stamps for claimed territory. Gumball renders the same data digitally.
  Cora gets the same world, read-aloud + phonics territories, her own frontier.
- **Two currencies, never mixed**: cash pays for INPUTS (the chore list — 15 min of
  math, etc. — unchanged). Mastery pays in AGENCY and universe status (the pick, the
  certificate, the territory, the PB curve). No cash on mastery events —
  overjustification risk lives exactly there.
- Note: proto-version already validated — bio-zack printed ~5 packets and let Penelope
  pick order; she engages with choice. The map makes the implicit visible.

### 6. Gumball recognition layer (near-term, independent of the graph)

Buildable NOW against shipped review-pack machinery; do not block on the graph:
- Packet completion (pack activation) → next day's plan opens with a case-closed
  celebration in-universe; the review slot framed as "agent maintenance training."
- Mastery events (concept mastered, pack fully mastered, fluency PB) → plan-level
  congratulations + map/territory update when the map exists.
- Route: spec → Fable flow, same path as the spelling Word Architect upgrade.

## Deliverables for the research session

1. Schema + authoring conventions, documented in this directory.
2. **Math domain filled**: ~60–90 nodes for the ages ~6–9 band, prereq + transfer
   edges, fluency tiers marked with age-adjusted rate aims, assessment specs per node,
   resources attached (existing 6 packets, gumball generators, Beast L2–L3 chapters,
   Singapore units).
3. Skeletons for the other three domains (coarse, expandable).
4. Per-node source citations (so future sessions extend without re-research).
5. **Seeded ledgers** for Penelope + Cora from known state (six packets, gumball
   response history, Singapore book 2 position, spelling mastery data, Cora CVC stage).
6. A frontier-query script + a map-poster render script (mapgen-style) proving the
   data is usable end-to-end.
7. Adversarial review panel with named lenses: prereq-gap walk (the ed-packet panel
   skill, applied to the graph), fluency-aim sanity vs literature, measurability audit
   (every gate checkable with our actual instruments), age-band sanity,
   maintenance-burden check (will this stay true at our curation capacity?), scope check.

**Non-goals for that session**: no gumball code changes, no packet generation, no
weekly pacing plans (Learning agent owns pacing).

## Sources to consult (starting set — research session expands)

- Math spines: Common Core progressions documents; Singapore Primary Mathematics scope
  & sequence; Beast Academy L1–L5 scope; Khan Academy prerequisite graph; **Math
  Academy's knowledge-graph writeups ("The Math Academy Way")** — the closest existing
  artifact to what we're building (mastery + fluency + spaced repetition over a graph).
- Fluency: precision teaching (Lindsley), Haughton fluency aims, Binder "Behavioral
  Fluency: Evolution of a New Paradigm."
- Retention: Bloom mastery learning / two-sigma; retrieval practice; interleaving
  (Rohrer); SM-2 (shipped).
- Thinking skills: Willingham "Critical Thinking: Why Is It So Hard to Teach?";
  Right Question Institute (question formulation); Philosophy for Children literature.
- World models: **Core Knowledge Sequence (Hirsch)** — free, grade-by-grade "what
  should a kid know" spine; David Macaulay "The Way Things Work."

## Open questions (research session settles)

- ~~Node granularity~~ RESOLVED 2026-07-13 (schema.md "Node granularity"): node ≈ one
  Math Academy module / 2–5 topic cluster / packet-rung-sized; split only on strategy
  change or fluency rung. Grounded in the live MA taxonomy scrape.
- Exact L2/L4 thresholds and per-node rate aims.
- Yaml-vs-json + file layout conventions.
- How thinking-skill "gear levels" are measured (tagged-exercise accumulation? parent
  observation? both?).
- Whether Beast Academy L2 becomes the math daily core (separate live decision, Learning
  agent owns it — graph should not assume it).
