Comprendo / insights Calibrate

Analysis

Can AI Tutors Actually Teach? We Measured It

US schools spend roughly $13 billion a year on education software, and AI tutors are a fast-growing slice. But can they teach and does anyone check? We built a test to find out.

August 10, 2026

Technical line drawing of a vernier caliper measuring a speech bubble containing the equation 3x + 2 = 17, drawn in dark ink on cream drafting paper.

What buyers say, and what they do

Ask district leaders and they’ll tell you evidence matters. 49% say research evidence is their top influence when judging instructional products.1 Watch what they buy and a different picture appears. In a survey of 515 school and district leaders responsible for edtech adoption, 11% said they demand peer-reviewed research; about half conceded research plays little or no role in their decisions.2 The one observational study analyzing 54 districts’ actual purchases found decisions made on informal pilots and peer word-of-mouth, with needs assessments “rarely conducted.”3

49%of district leaders say research evidence is their top influence
11%actually demand peer-reviewed research before buying
3 in 4of the 100 most-used K-12 tools published no ESSA-aligned research (2023)

This isn’t surprising to anyone who has been around K-12 edtech. Quality gets a lot of attention but doesn’t sell. Of the 100 most-used K-12 tools in 2023, roughly three in four published no ESSA-aligned research at all and sold fine.4 As of 2026, only 7% of edtech tools meet either of the two rigorous federal evidence tiers.5

The federal standard doesn’t do much, either. When researchers audited how one large district’s Title I spending held up against ESSA’s evidence requirement, over 95% of funds technically complied while less than 60% of the money went to practices the full research base actually supports.6 So it’s largely an exercise in compliance.

Most edtech tools have no identifiable evidence behind them share of purpose-built tools by highest ESSA evidence tier, 2026
0% 20% 40% 60% 80% 100% Tier I — strong Tier II — moderate Tier III — promising Tier IV — rationale only No identifiable evidence 2% 5% 14% 19% 60%
Source: Instructure 2026 EdTech Evidence Report, based on LearnPlatform usage telemetry across thousands of districts. Tier IV requires a written rationale and a study planned but not results.

AI widens the gap

The cost of producing an edtech product is falling toward zero. In its most basic form, an AI tutor is a frontier model plus a prompt. There will be many more products and they will change constantly. The underlying models, upon which nearly all tutors will be built, change at the same pace as AI, which already feels incredibly fast. Even a careful one-time evaluation will be stale within a few months. There will be more products, changing faster, with no new way to validate them.

A 2026 survey found 55% of districts require vendors to provide product-safety information. Far fewer ask for evidence of effectiveness, and most district leaders were unfamiliar with the field’s own quality framework.7 State AI guidance, now issued by more than half the states, focuses on acceptable use and AI literacy; almost none of it requires vendors to demonstrate their tools have been evaluated for accuracy or instructional quality.8

Student data privacy has real compliance requirements in the form of signed data-privacy agreements, state laws like New York’s Ed Law 2-d, and SOC 2 audits. Federal and state laws put liability on districts, and districts responded by building controls. Quality, on the other hand, shows up mostly in marketing copy. The procurement process doesn’t need to ask the question how well does it work?

So we measured it

Tutoring happens in conversation, and the standard tools don’t work there. A static benchmark can’t test a dialogue, because the student’s next line depends on what the tutor just said. The test has to talk back. The funder and research community are beginning to address this gap by building better benchmarks for multi-turn taxonomies of tutoring ability.9 But it’s very different from a traditional benchmark like the kind you see frontier models reporting on. When the student is in the loop, you have to judge three things at once. Did the tutor diagnose the misconception (if there was one), was the response pedagogically sound, and was it at grade level. Those are difficult and subjective tasks.

A test that talks back

We simulated 8th-grade students with planted misconceptions. Each misconception was documented in the math-education literature and tagged to standards10. The students then talked with the AI tutor for several turns. Because we planted the error, the open-ended conversation has an answer key. Did the tutor notice the mistake? Did it diagnose the misconception? Did it say anything false? Did it hand over the answer?

Our tutors were six frontier LLMs, not edtech products. That’s a deliberate choice, partly for practical reasons since we don’t have access to six edtech tutors. Nevertheless, most AI tutoring products are or will be built on a frontier LLM. Even fine-tuned or custom models need external quality validation. As those continue to evolve and improve, the vendors will need to adapt, measure regressions, change prompts, ensure guardrails, update content, etc. It makes sense to test the foundation models that most tutoring products will leverage. Not to mention that models are increasingly being marketed directly to teachers and students.

We tested each model two ways. We used a naive prompt a student might produce and a standard, but simple, tutoring prompt. Each of these prompts was run across 10 misconceptions and 3 control scenarios resulting in 779 conversations, roughly 3.7 million tokens.

Naively prompted chatbots don’t tutor. Without a tutoring prompt, every model handed the student the answer in 97% of conversations. This was true across all misconceptions, all six models alike. The models provided a correct diagnosis, made no false statements, and held back the answer in 0 out of 298 conversations. That’s not surprising. Models are trained to give answers, not to tutor students.

In 298 conversations with a struggling 8th grader, naively prompted frontier chatbots produced zero clean tutoring wins.

A tutoring prompt helps significantly, and models differ under it. The same prompt on the same scenarios cut answer-leaking significantly, but compliance varied between models (between one and nine leaked answers per 65 conversations) and restraint also has a price. The prompted models correctly named the student’s actual misconception slightly less often, with a wider spread by model (88% to 100%).

Clean tutoring wins, with a tutoring prompt narrow: didn't state the answer strict: didn't do the student's thinking
0% 25% 50% 75% 100% DeepSeek V4 Pro Grok 4.5 Kimi K3 GPT-5.6 Terra Qwen3.8 Max Claude Sonnet 5 Naive prompt 54 52 42 38 30 28 74 76 80 86 86 66 0
Correct diagnosis + nothing false + restraint, misconception scenarios, 65 conversations per model. Strict applies the Learning Commons Withholding Answers rubric, which also fails a tutor that performs the student's reasoning for them without literally stating the answer. Naive-prompt condition: 0% under both definitions. Judge and rubric were both calibrated against an expert labeling pass and err conservative, so read these as lower bounds.

No model did well at both diagnosis and restraint. The two restraint definitions nearly reverse the ranking. The models that name the error most reliably score worst on strict restraint, because they hold their accuracy by doing the student’s thinking for them and leaving only the arithmetic. The models that genuinely hold back pay for it in missed diagnoses. So there is still opportunity for models to have high diagnostic accuracy and high restraint.

The empty corner diagnosis recall vs. strict restraint, prompted condition
85% 90% 95% 100% 30% 40% 50% 60% 70% 80% Diagnosis recall (named the planted misconception) → Strict restraint pass rate → opportunity Grok 4.5 DeepSeek V4 Pro Kimi K3 GPT-5.6 Terra Qwen3.8 Max Claude Sonnet 5
Recall: of conversations where the student had the planted misconception, the share where the tutor named that error. Restraint: share of conversations passing the Withholding Answers rubric. Which trade-off is right is a judgment / pedagogical call for funders and buyers.

Quality has almost nothing to do with price. The cheapest model in the test, which cost about $0.002 of API cost per conversation, had the best strict clean-win rate. The most expensive spent roughly 12× more per conversation, mostly on long chains of reasoning tokens, for an average result. The spread is 16x across all models.

Strict clean wins per dollar strict clean-win rate ÷ API cost per conversation, prompted condition
0 100 200 300 Strict clean wins per $1 of API cost → DeepSeek V4 Pro GPT-5.6 Terra Grok 4.5 Claude Sonnet 5 Qwen3.8 Max Kimi K3 307 186 149 59 40 19
Cost is the measured API price of the tutor's calls in each conversation (OpenRouter, August 2026). Model prices and versions change monthly, suggesting an ongoing need for measurement.

In some cases the models were excellent. Outright false statements were rare. We found 3 in 779 conversations, each human-confirmed. In 779 conversations, no model ever told a correct student they were wrong. Also, the “students-and-professors” reversal, the misconception the literature predicts is hardest to catch,10 turned out to be the hardest for every model. We never tuned for that, so the match with human studies suggests the harness is measuring something real.

How to make quality matter

Buyers won’t force vendors to care, and they likely can’t do it on their own. Without a check, buyers will keep relying on relationships and incumbents for signals of quality, and adoption of high-quality tools will stay slow. Funders, however, have the motivation. They are already investing heavily in AI and AI tutoring to build expert-labeled tutoring datasets, benchmarks, and evaluator tools, committing hundreds of millions of dollars. One recent $8M fund backs an open tutoring model whose stated rationale is that current AI tutors “give answers too quickly, talk too much, miss signs of student motivation.”11 That is part of what our baseline just measured.

These tools are great. We used the Learning Commons Evaluators rubrics in this experiment. However, the risk is that the research and tools don’t get used. A funded benchmark or evaluator only changes the market if someone keeps running it, and right now no vendor needs to and no buyer is asking. But funders affect both supply and demand. They can make independent evaluation a condition of the grant, or point this kind of measurement at their own portfolio and rerun it as models change. There’s precedent with England’s National Tutoring Programme which paid subsidies only to tutoring vendors that passed an independent quality-assurance review.12

No one has an incentive to keep measuring quality as models change, but an ongoing quality signal can change behavior. When EdReports began publishing free, philanthropy-funded reviews of math curricula in 2015, publishers attacked the methodology. Within a few years they were revising their materials to meet it, and today 43 states reference its reviews.13 Nothing technical is blocking this path.

If you fund, build, or deploy AI tutoring and want it measured, we’d like to talk.


Methodology

We used six frontier models: GPT-5.6 Terra, Claude Sonnet 5, Grok 4.5, DeepSeek V4 Pro, Kimi K3, and Qwen3.8 Max. Each was tested with a naive prompt and with a tutoring prompt. We ran five repetitions per scenario, 779 conversations in all. The 13 scenarios used a specific Algebra 1 misconception documented in the math-education research literature10 alongside control students whose work is actually correct. These were included to catch the opposite failure in the case where a tutor tried to “correct” a right answer.

Each transcript is scored against the answer key provided by our misconception. Did the tutor notice the error, diagnose the misconception the student actually has, say anything false, give the answer away? We scored restraint two ways. We defined the narrow check as “did the tutor state the answer”. The strict check is an external rubric based on Learning Commons’ Withholding Answers, from their Productive Coaching work with Quill.org and Leanlab Education14. It scores every tutor turn on three features. Does it point toward evidence without supplying it, does it prompt revision without rewriting the student’s work, and does it leave the core thinking to the student. Each turn passes or fails, and a conversation counts as restrained only if every turn passes. We are trying to detect if the tutor is doing the student’s thinking for them even when it doesn’t directly provide the answer.

Every transcript was scored by an AI judge that we then audited by using human-labeled conversations over the same transcripts. We used samples from exactly the places the judge was most likely to be wrong. The judge’s errors run consistently conservative. Most were false alarms and under-credited diagnoses, almost never missed problems. Thus the rates presented here should be interpreted as floors instead of averages. Caveats we know about: the strict rubric was built for writing feedback and borrowed for math dialogue14 (it errs strict, agreeing with the human expert 77% of the time on the adversarial sample), and our simulated student runs on a model that shares a family with one of the tutors under test.

Sources

  1. EdWeek Market Brief, "In Judging Instructional Products, Educators Have This Message: Show Us Your Evidence," Dec 2018. marketbrief.edweek.org
  2. EdTech Efficacy Research Academic Symposium (UVA / Digital Promise / Jefferson Education), survey of 515 adoption decision-makers, 2017. wcet.wiche.edu
  3. Morrison, Ross & Cheung, "From the market to the classroom: how ed-tech products are procured by school districts," Educational Technology Research & Development, 2019. link.springer.com
  4. LearnPlatform by Instructure, inaugural EdTech Evidence Report, March 2023. instructure.com
  5. Instructure, 2026 EdTech Evidence Report, Jan 2026. instructure.com
  6. Ginsberg, Hollands et al., "Does ESSA Assure the Use of Evidence-Based Educational Practices?" Educational Policy, 2024. journals.sagepub.com
  7. EdWeek Market Brief, "Security, AI Guidelines Are Top School District Ed-Tech Concerns," May 2026; SETDA 2025 EdTech Quality Indicators Guide. marketbrief.edweek.org · setda.org
  8. ExcelinEd, "State K-12 AI Policy in 2026: Milestones," May 2026; CDT, "Advancing Responsible AI Adoption and Use in K-12." excelined.org · cdt.org
  9. Maurya, Srivatsa, Petukhova & Kochmar, "Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors," NAACL 2025. arxiv.org
  10. Clement, Lochhead & Monk, "Translation Difficulties in Learning Mathematics," The American Mathematical Monthly 88(4), 1981 - the original documentation of the students-and-professors reversal, replicated repeatedly since. jstor.org · scholarworks.umass.edu (open access)
  11. K-12 AI Infrastructure Program ($26M donor collaborative; $8M education AI model RFP). k12-ai-infrastructure.org
  12. Education Endowment Foundation, National Tutoring Programme: Tuition Partners quality-assurance framework. educationendowmentfoundation.org.uk
  13. EdReports, ten-year impact summary; Gates Foundation Q&A on EdReports; EdWeek coverage of the 2015 launch and publisher response. edreports.org · gatesfoundation.org · edweek.org
  14. Learning Commons, open tutoring-quality rubrics (Withholding Answers, from the Productive Coaching work with Quill.org and Leanlab Education). learningcommons.org · marketbrief.edweek.org