Evidence of iteration
What changed along the way, and why
This proposal did not arrive in one pass. The versions below record what was tried, what was dropped, and what a question from a reader forced me to redo. They are my own working notes on the concept, not a foundation review process.
How the concept changed
An open chat assistant for families
What it was
A general-purpose assistant a caregiver or child could ask anything, with the foundation's programs as background context.
What changed
Dropped entirely. The scope narrowed to one short, fixed activity with a beginning and an end.
Why
An open-ended assistant with children in scope carries the wrong risk profile, and there is no honest way to evaluate whether an unbounded answer was correct.
A scripted, bounded activity with an evidence step
What it was
One reviewed passage, a small set of questions, and a fixed set of answers — no free-form generation in the family path.
What changed
Added the step where the child has to point at the sentence in the passage that supports their answer, rather than simply being told they were right.
Why
A fixed passage makes correctness reviewable, and the evidence step is where the reading skill and the AI judgment skill turn out to be the same skill.
The activity plus a mistake to catch, a lab, and an architecture
What it was
The bounded activity on its own, described in prose.
What changed
Added a deliberate model mistake for the family to catch, an expandable evaluation lab with proposed test cases and an A/B instruction comparison, and a labeled architecture showing channels, consent, routing, and output checks.
Why
Reviewers asked what would be tested and how it would run, and prose alone could not answer either question.
How the Coach wording changed
The demonstration is scripted, so every line is a deliberate choice. These three were rewritten more than once.
The answer prompt
“Great answer! Maya planted a bean seed. What do you think happens next?”
“Before we go on — which sentence in the passage shows that? Read it out loud to your grown-up.”
Why
The first version rewarded any answer and taught nothing. The second makes the child return to the text, which is the behaviour the activity is actually for.
The mistake prompt
“Maya's bean sprouted after two days.”
“I think Maya's bean sprouted after two days. Check the passage — am I right? It is fine to tell me I am wrong.”
Why
Stating a wrong fact flatly risks the child simply absorbing it. Framing it as a claim to check turns the error into the lesson, and gives the child explicit permission to correct the machine.
The unsupported-inference prompt
“Maya was probably feeling excited and proud of her little garden.”
“The passage does not say how Maya felt. We could guess, but let's be clear that it would be a guess — what in the story makes you think so?”
Why
The earlier wording invented a fact and presented it in the same voice as the real ones. Naming the gap plainly is the single most transferable habit in the whole activity.
Questions that changed the design
Each of these was raised while the work was being read and reviewed, and each one produced a specific change you can still see on the site.
Is this claiming to be an approved foundation program?
What it changed
Added the concept-brief label and the prepared-by line at the top of the proposal, and caution notes on every page where something proposed could be mistaken for something built.
What about families without a smartphone or a data plan?
What it changed
Added printed and facilitated alternatives with equal support, seated and low-movement play by default, and later a proposed spoken channel for caregivers who cannot easily read a screen.
What would this cost?
What it changed
Removed every invented figure. The pilot page now offers a blank scenario calculator and named cost categories so a reviewer enters their own numbers and sees their own result.
One long page does not scroll, share, or search well.
What it changed
Split the proposal into separate pages, each with its own title, description, and structured summary, and kept the old anchor links working.
How would anyone know whether it worked?
What it changed
Added proposed targets, a 0–4 rubric agreed before the pilot starts, explicit continue / change / stop decision steps, and an evaluation lab that shows the test cases without claiming any results.
What stayed, and what got cut
The child has to point at the sentence that proves it
It survived every version, and everything else was arranged around it. A fixed passage exists so the proof can be found. The deliberate mistake exists so there is something worth checking. The rubric scores the child's own work, not the Coach's, because the evidence step is the thing being learned.
Anything that made correctness unreviewable
The open assistant, free-form generation in the family path, praise without evidence, and open-web retrieval all went for the same reason: if no one can say afterwards whether an answer was right, there is nothing honest to evaluate and nothing safe to put in front of a child.
These are the author's working notes on how the concept developed. No foundation staff reviewed or approved these iterations, the questions above are not quotations from named people, and no evaluation results are claimed anywhere on this page.