← Blog

AI as scaffolding, never shortcut: the rule that makes STEMBuddy different

Every other AI-for-kids product has to answer this question, and almost none of them answer it clearly. We will: when your child is using STEMBuddy, the AI does NOT write your child's hypothesis. It does not write the conclusion. It does not produce the explanation your child should be able to give. The kid does the reasoning. The AI helps the kid get unstuck. Those are different things, and the difference is the entire product.

We call this the load-bearing rule: "No system where AI does the reasoning and the kid approves it." When we have had to make a hard product decision — what a button does, what a coach can say, what an autocomplete suggests — this rule has been the tiebreaker every time. Below is what the rule actually looks like inside the product, and the five scaffolding patterns we use to keep the kid driving while still making the AI useful.

What we will not ship

It is easier to define this rule by what we refuse to build. We have shipped none of the following, and we won't:

Each of these is a small, plausible, well-meaning feature. Each one, shipped, would do the same thing: the AI does the structural work, the kid signs off, the kid learns nothing. The reason most "AI tutor" products end up that shape is that it's the easiest UX to build. We deliberately did not build it.

The five scaffolding patterns we use instead

1. Mad-lib slot fill

The structure of the sentence is given. The substance is the kid's. For a hypothesis: "If I change [], I think [] will happen because [____]." The AI never fills the brackets. It can offer the kid a list of kinds of things that might go in the third bracket (an explanation? a prediction? a comparison?), but the actual content is something the kid types. This is how a 9-year-old learns the SHAPE of a hypothesis without being told what to think.

2. Sentence starters

A kid staring at a blank journal entry will often quit before they begin. A sentence starter — "Today I noticed that…" or "The thing that surprised me was…" — is a runway, not a destination. The AI never finishes the sentence. The kid takes the runway and writes the actual observation. This pattern is borrowed almost directly from how good elementary-school teachers handle a stuck kid; we just made the runway available at the moment of stuckness, not the next morning.

3. AI co-draft from kid inputs

Sometimes the kid has the ideas but can't arrange them. In those cases, the kid writes their messy thoughts ("water went up, plant got bigger, leaves greener"), and the AI offers ONE pass at organizing them into a paragraph — visibly labeled as a draft, never as a final answer. The kid then revises it, in their own words. The AI is the messy-draft-to-organized-draft helper; the kid is the editor and the final voice. If the kid never wrote the messy thoughts, the AI has nothing to organize.

4. Iterative review

When the kid finishes a step, the AI doesn't say "correct" or "incorrect." It asks one or two structural questions: "what is the variable you changed?" "what would have to be true for your claim to hold?" "could you have gotten this result if your hypothesis were wrong?" These questions teach the kid to do the questioning themselves over time. After several investigations, kids start asking the questions before the AI gets to. That is the whole goal.

5. Progressive escape hatches

If a kid is really stuck — not just sticky but actually stuck — the scaffolding gets more generous in clearly-marked stages. First more sentence starters. Then concrete examples (drawn from public Codex entries, so the kid sees how another kid handled it). Only at the very last step, with a clear "this is an example, not your answer" label, does the system show a worked variant of the same kind of question. The kid still writes their own. The escape hatches exist so a stuck kid does not give up; they do not exist so an unstuck kid can skip the work.

Why the principle name matters

We are careful about something else parents may not notice at first: the principle a kid is practicing is named the same word on every surface. If the kid is doing Fair Test Design, the kid sees "Fair Test Design" in the investigation, the parent sees "Fair Test Design" in the parent hub, the Field Report says "Fair Test Design," and the analytics event the platform records is "Fair Test Design." We refuse to use whimsical marketing names ("Test It Cleanly!"). Real principles have real names, and a family conversation is only possible when everyone is using the same vocabulary.

When a kid earns a principle (we call it "minting"), the trigger is deterministic — they did the observable behavior, the platform recorded it, the principle is theirs. We don't use free-form AI inference for this. If the AI guesses wrong about whether a kid practiced Belief Revision, parents lose trust in every principle the platform ever credits, including the real ones. So we use rules a parent can audit.

The kid's work, in the kid's voice

The artifact a kid finishes with — a Field Report — reads like a 10-year-old wrote it. Because a 10-year-old did. The AI helped them find a structure. It asked clarifying questions. It offered some sentence starters when they got stuck. It pointed at what an example might look like when they were really stuck. It never wrote the hypothesis. It never wrote the conclusion. It never gave an answer the kid then approved.

There are real trade-offs. A kid using STEMBuddy will sometimes spend more time on a single investigation than a kid using a "let the AI write it" tool. That is not a bug. That is the work. The kid who did the work has the thinking; the kid who pressed accept does not. Over a year of investigations, the gap is enormous, and it shows up in the only place it really matters — in what the kid can do when no app is open.