Back to Blog

What I Learned Researching AI in Schools Around the World

By Orly Izhaki2026-10-056 min readEnglish

I've spent a lot of time lately thinking about what a school needs before it adds another AI tool. That question led to a different one: has anyone actually tried the alternative, a school built around AI from the start, and does it work?

So I went looking. Not for marketing pages. For the independent analysts, the randomized trials, the government reports that came out after the pilot programs did.

What I found didn't match either the hype or the panic.

I expected a success story or a cautionary tale. I found a function-by-function breakdown.

The question everyone wants answered is whether AI can replace a school. As of late 2026, nothing shows that it can, at scale, independently verified.

What AI actually does is take the school apart, function by function. Content delivery, practice and feedback, and lesson preparation are automating the fastest. Motivation, supervision, socialization, and credentialing are still squarely in the school's hands.

That reframes the whole debate. The useful question stopped being "will AI replace school" and became "which specific jobs a school does are already shifting, and which ones no one has figured out how to automate yet."

I expected the flagship AI schools to have the strongest evidence. They have the weakest.

Alpha School and its platform TimeBack are the most visible AI-school model in the US. The company claims students grow "2.6x faster than peers" on national MAP tests, with "most" students in the 99th percentile.

Independent analysts pulled the number apart. Kelsey Piper, writing in The Argument in August 2026, showed how the 2.6x figure gets built: Alpha divides each student's MAP gain by the median expected gain for their level. In high school, the expected gain is close to zero, so a reading student in the 90th percentile who's expected to gain about 0.23 points can register as "9x growth" from normal statistical noise. A parent at the school called the number "miscalculated and meaningless." A "gifted" data point was reportedly based on five children with two tests each.

The one public, non-selective test of the model is Unbound Academy, a public charter in Arizona that opened in 2025 using the same platform. Alpha told the state to expect 65% proficiency in English and 60% in math. The actual first-year results were 28% and 10%.

That's not a rounding error. That's a school that, in its own terms, missed its own target by more than half.

I expected the research on AI tutors to be mixed. It turned out to hinge entirely on whether students actually used the thing.

The largest, longest randomized trial of an LLM tutor is Oreopoulos and Low's two-year study of Khanmigo (NBER working paper 35620). 96% of students tried it at least once. The median student sent it a message on only a third of practice days, and used it on just 17% of the sessions where they got something wrong.

The result: gains of 0.06 to 0.08 standard deviations, no better than using Khan Academy without the AI at all.

Access isn't the same as learning. That line, more than any single statistic, is the finding underneath every other finding in this research.

I expected guardrails to be a nice-to-have. They turned out to be the entire story.

Tutors that give hints and scaffolding instead of answers preserve or improve learning. Chatbots that just answer the question raise practice scores while the tool is in front of the student, and can hurt learning once it's taken away. In one study, unrestricted access to GPT-4 dropped test scores by 17% after the access ended.

Where the gains were real, there was always a structure doing the work. A Nigeria program with tight guardrails produced 0.31 standard deviations of improvement. Tutor CoPilot, which coaches the teacher in real time rather than replacing them, raised student mastery by 4 percentage points for about $20 per teacher per year. LearnLM's automated feedback loop improved how teachers responded to student ideas by 13%.

None of these are chatbots for kids. They're structured tools aimed at the adult in the room.

I expected AI to save teachers a lot of time. It saves them some.

A Gallup and Walton survey found teachers who use AI weekly estimate they save about 5.9 hours a week. That's self-reported, and self-reports of time savings run high.

A controlled trial by the UK's Education Endowment Foundation, with 259 teachers across 68 schools, found an actual savings of 25.3 minutes a week on lesson planning, about 31%, with no measurable difference in lesson quality. A separate study of 311 AI-generated lesson plans found 90% of the activities stayed at a basic level of thinking.

The clearest value showed up where AI supported the teaching itself rather than replacing the teacher's prep work, tools like Tutor CoPilot and LearnLM, not generic lesson-plan generators.

I expected countries to be racing toward AI-taught classrooms. Every one I looked at backed away from that, fast.

South Korea spent more than 1.2 trillion won, roughly $850 million, on AI textbooks. After one semester, it reclassified them as "supplementary material." China now bans unsupervised use of open generative AI by elementary students. The UAE, Estonia, India, and Israel are all expanding structured, supervised programs rather than open ones.

The pattern across every national rollout I found was the same: teach AI literacy, restrict unsupervised use for young children, keep the teacher accountable, and roll out in phases. Full mandates for AI-led teaching, like Korea's, failed first.

I expected the money to be chasing the hype. It's actually chasing the safest bets.

EdTech venture funding hit $2.4 billion globally in 2024, 89% below the 2021 peak and the lowest since 2014. Most of what's left is going to teacher assistants, tutors, and language learning. Peer learning, credentialing, generative worlds, and audio-based learning are still underfunded, which probably means they're underexplored, not that they don't work.

What I'd tell someone building for this space

If the anxiety is that AI will replace teachers, the evidence doesn't support it, not yet, and not by accident. The schools and tools that actually show gains all have an adult, or a tight structure, standing between the student and the model.

If the hope is that AI will finally fix education on its own, the evidence doesn't support that either. The gains are real, specific, small, and entirely dependent on design choices nobody automates away: what the tool is allowed to say, who checks its work, and whether a child using it alone looks any different from a child using it with a teacher who still has the final word.

The numbers show the gap between the claim and the evidence.

The guardrails show where the real work still is.


Sources

  • Kelsey Piper, The Argument, August 2026, analysis of Alpha School's growth claims
  • Oreopoulos & Low, NBER Working Paper No. 35620, two-year randomized trial of Khanmigo
  • Gallup & Walton Family Foundation survey, March-April 2025 (2,232 respondents)
  • UK Education Endowment Foundation randomized controlled trial (259 teachers, 68 schools)
  • HolonIQ global EdTech venture funding data, 2024
  • Arizona Unbound Academy first-year proficiency results, reported August 2026