Skip to main content
EAT. LEARN. PLAY.Family AI Coach

Evidence of iteration

What changed along the way, and why

This proposal did not arrive in one pass. The versions below record what was tried, what was dropped, and what a question from a reader forced me to redo. They are my own working notes on the concept, not a foundation review process.

How the concept changed

v1

An open chat assistant for families

What it was

A general-purpose assistant a caregiver or child could ask anything, with the foundation's programs as background context.

What changed

Dropped entirely. The scope narrowed to one short, fixed activity with a beginning and an end.

Why

An open-ended assistant with children in scope carries the wrong risk profile, and there is no honest way to evaluate whether an unbounded answer was correct.

v2

A scripted, bounded activity with an evidence step

What it was

One reviewed passage, a small set of questions, and a fixed set of answers — no free-form generation in the family path.

What changed

Added the step where the child has to point at the sentence in the passage that supports their answer, rather than simply being told they were right.

Why

A fixed passage makes correctness reviewable, and the evidence step is where the reading skill and the AI judgment skill turn out to be the same skill.

v3

The activity plus a mistake to catch, a lab, and an architecture

What it was

The bounded activity on its own, described in prose.

What changed

Added a deliberate model mistake for the family to catch, an expandable evaluation lab with proposed test cases and an A/B instruction comparison, and a labeled architecture showing channels, consent, routing, and output checks.

Why

Reviewers asked what would be tested and how it would run, and prose alone could not answer either question.

How the Coach wording changed

The demonstration is scripted, so every line is a deliberate choice. These three were rewritten more than once.

The answer prompt

Before

Great answer! Maya planted a bean seed. What do you think happens next?

After

Before we go on — which sentence in the passage shows that? Read it out loud to your grown-up.

Why

The first version rewarded any answer and taught nothing. The second makes the child return to the text, which is the behaviour the activity is actually for.

The mistake prompt

Before

Maya's bean sprouted after two days.

After

I think Maya's bean sprouted after two days. Check the passage — am I right? It is fine to tell me I am wrong.

Why

Stating a wrong fact flatly risks the child simply absorbing it. Framing it as a claim to check turns the error into the lesson, and gives the child explicit permission to correct the machine.

The unsupported-inference prompt

Before

Maya was probably feeling excited and proud of her little garden.

After

The passage does not say how Maya felt. We could guess, but let's be clear that it would be a guess — what in the story makes you think so?

Why

The earlier wording invented a fact and presented it in the same voice as the real ones. Naming the gap plainly is the single most transferable habit in the whole activity.

Questions that changed the design

Each of these was raised while the work was being read and reviewed, and each one produced a specific change you can still see on the site.

Question raised

Is this claiming to be an approved foundation program?

What it changed

Added the concept-brief label and the prepared-by line at the top of the proposal, and caution notes on every page where something proposed could be mistaken for something built.

Question raised

What about families without a smartphone or a data plan?

What it changed

Added printed and facilitated alternatives with equal support, seated and low-movement play by default, and later a proposed spoken channel for caregivers who cannot easily read a screen.

Question raised

What would this cost?

What it changed

Removed every invented figure. The pilot page now offers a blank scenario calculator and named cost categories so a reviewer enters their own numbers and sees their own result.

Question raised

One long page does not scroll, share, or search well.

What it changed

Split the proposal into separate pages, each with its own title, description, and structured summary, and kept the old anchor links working.

Question raised

How would anyone know whether it worked?

What it changed

Added proposed targets, a 0–4 rubric agreed before the pilot starts, explicit continue / change / stop decision steps, and an evaluation lab that shows the test cases without claiming any results.

What stayed, and what got cut

Stayed

The child has to point at the sentence that proves it

It survived every version, and everything else was arranged around it. A fixed passage exists so the proof can be found. The deliberate mistake exists so there is something worth checking. The rubric scores the child's own work, not the Coach's, because the evidence step is the thing being learned.

Cut

Anything that made correctness unreviewable

The open assistant, free-form generation in the family path, praise without evidence, and open-web retrieval all went for the same reason: if no one can say afterwards whether an answer was right, there is nothing honest to evaluate and nothing safe to put in front of a child.

These are the author's working notes on how the concept developed. No foundation staff reviewed or approved these iterations, the questions above are not quotations from named people, and no evaluation results are claimed anywhere on this page.