An AI Tutor Matched Expert Humans at 1/918th of the Cost. It Was the Cheapest Model Tested

Surya Pratap
By Surya Pratap

October 2, 2026

9 min read

AI & Technology
A two-part diagram. On the left, the StudentBench design: 2,383 participants randomly assigned to AI tutoring, expert human tutoring or no tutoring for one hour between two parallel GRE tests, with pooled AI tutoring statistically equivalent to the human tutors and both improving on no tutoring. On the right, cost per percentage point of learning gain: $4.81 for an expert human tutor at $75 an hour against $0.0052 for Gemma 4 31B, the cheapest AI tutor tested at $0.067 a session, which passed the equivalence test; a note shows that the most expensive AI tutor cost $21.24 a session, and that Gemini 3.5 Flash matched GPT-5.5 Pro's gains at a twentieth of the price.Priced per point, not per tokenHover to explore
The cheapest model tested matched expert human tutors. Price per session told you nothing about which tutors would; only measuring the learning did.

Most comparisons of AI models are about what the models can do. This one is about what the people using them learned, measured before and after, against a control group and against expert human teachers. Its most useful result for founders is not about education at all.

Everything here comes from "StudentBench: AI and human tutoring yield equivalent GRE learning gains" by Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner and Jonas Mueller of Handshake AI Research, a preprint first posted on 23 September 2026 and revised on 30 September. The de-identified data is open source. Handshake also runs the StudentBench tutoring platform the study used. The reading from section 3 onward is mine.

1. What was tested

The study recruited 2,383 adults, most of them aged 18 to 24, to prepare for the GRE, a graduate admissions test. Each session had the same shape:

  1. A pre-test of 27 questions.
  2. One hour in a randomly assigned condition: tutoring from one of 13 AI tutor configurations, tutoring from an expert human over live video, or no tutoring.
  3. A post-test of the same length, on a parallel form.

The learning gain was the post-test score minus the pre-test score, in percentage points. The human tutors were former GRE question writers or tutors with at least five years' experience, and they were not allowed to see the post-test questions.

2. What it found

AI tutoring, pooled across tutors, was statistically equivalent to expert human tutoring on GRE learning gains (p = .015, within a margin of a quarter of a standard deviation). The average difference between AI and human was −0.58 percentage points. Both improved on no tutoring: AI tutoring added 6.15 points over the control group.

Six of the AI tutors individually passed the equivalence test against the human tutors. In five of the seven GRE topic areas, the best AI tutor did better than the human tutors on average.

Then the cost.

  • The AI tutors ranged from $0.067 per session for Gemma 4 31B to $21.24 for GPT-5.5 Pro — more than 300 times apart.
  • Gemma 4 31B, the cheapest, passed the equivalence test (p = .044).
  • Per percentage point of learning gain, Gemma cost $0.0052. An expert human tutor at $75 an hour cost $4.81. That is the 918-fold difference in the headline.
  • Gemini 3.5 Flash produced a similar learning gain to GPT-5.5 Pro at about a twentieth of the cost.

The price of a model told you almost nothing about how much a student would learn from it. Only measuring the learning did.

3. Choose models by cost per outcome

This is the lesson that travels beyond tutoring. The authors did not compare models on a benchmark or on price per token. They compared them on cost per unit of the outcome the product exists to produce — here, dollars per percentage point of learning.

Seen that way, the ranking of models changes. A model that costs 300 times more per session is only worth it if it produces 300 times more of what you are selling, and in this study the expensive end did not.

Most AI products are priced and built the other way round. The team picks a capable, expensive model because it feels safe, measures cost per call, and never measures what each call achieves for the user. Without the outcome, you cannot know whether a model twenty times cheaper would do the job just as well.

4. Speed was part of the product

One secondary result is easy to overlook. On the quantitative sessions, faster AI replies were associated with students sending more messages, more messages with more correct practice problems, and more correct practice with larger learning gains (all p < .002).

It is a correlation, not a controlled experiment, and the authors present it that way. But it points at something founders routinely underrate: in an interactive product, latency is not a technical detail. If a slower, smarter model means fewer turns of practice, the user may end up learning — or getting — less, not more.

5. Copy the study design, not just the result

The most valuable thing about StudentBench for a founder is how it measured, and most of the design is cheap to reproduce.

Measure before and after. Define the outcome your product changes and measure it at both ends of a session. Satisfaction scores and time-in-app are not outcomes.

Keep a no-AI control group. Without one, you cannot tell how much of the improvement is your product and how much would have happened anyway. Here, no tutoring still produced some gain.

Compare models on the same users. Randomly assign users to two or three models and compare outcomes, not just costs. A cheaper model that is equivalent is a gross-margin decision waiting to be made.

Divide cost by outcome. Report cost per unit of result — per resolved ticket, per correct answer, per point learned — and make model decisions on that number.

6. What this means for AI training

For teams that buy or run AI training, the result is encouraging in a narrow way. For drill-style practice with clear right answers, an AI tutor can match an expert human in one sitting, at a tiny fraction of the cost. That makes it sensible to use AI tutors for practice between live sessions.

It does not show that AI replaces a human teacher. The study measured one hour and an immediate post-test. Whether the gains last, and whether AI tutoring works as well for open-ended skills, are questions it did not ask.

7. What I would not claim

It is a preprint from the company that runs the platform. Handshake AI Research ran the study on Handshake's own StudentBench service. The data has been released openly, which helps, but it has not been peer-reviewed.

The human arm was small. About 140 human-tutored sessions against more than 2,100 AI-tutored ones. Equivalence was established for AI pooled across tutors; individual comparisons have less behind them.

Verbal did not establish equivalence on its own. The pooled result combines quantitative and verbal sections. On verbal sections alone, equivalence was not shown, and the human tutors had the highest average gains in all three verbal areas.

The gains are immediate. No delayed test was run. Learning that shows up an hour later may not show up a month later, for either group.

The human cost is a reference rate. The $4.81 figure uses a published rate of $75 an hour from a 2018 source. Actual tutoring prices vary widely, which changes the multiple but not the direction.

I have not verified every model's individual result. The paper reports some figures only in charts. I have used only the results it states in text.

The honest summary

StudentBench is one of the more careful comparisons of AI and human teaching so far: randomised, controlled, with expert human tutors and open data. For one hour of GRE practice, AI tutoring was as effective as an expert human, and the cheapest model tested was among the ones that matched them.

For founders, the more durable lesson is in the method. The authors chose models by cost per unit of the outcome they cared about, and measured that outcome against a control. That ranked a $0.067 model alongside expert humans, and showed that the most expensive option was not the one to pick.

Before you choose a model on price or reputation, measure what each one actually produces for your users, and divide the cost by that.

Source: Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner and Jonas Mueller, "StudentBench: AI and human tutoring yield equivalent GRE learning gains", arXiv:2609.28470, v1 23 September 2026, v2 30 September 2026 — the 2,383 participants, the pre-test, one-hour condition and parallel post-test design, the 13 AI tutor configurations per section and expert human tutors, the pooled equivalence result (p = .015), the −0.58-point difference, the 6.15-point gain over control, the six individual equivalence passes, the five-of-seven domains result, the $0.067 and $21.24 session costs, Gemma 4 31B's equivalence (p = .044) and $0.0052 against $4.81 per percentage point at $75 an hour, the Gemini 3.5 Flash and GPT-5.5 Pro comparison, the latency and engagement correlations, the arm sizes, the verbal-only result and the stated limitations are all as reported there. De-identified data and code: StudentBench on GitHub. The reading in sections 3 to 6 is mine. For why smaller models often win narrow tasks, see most of what your agent calls an LLM for is a yes or a no. For the time problem in corporate AI training, see AI training more than doubled; 56% get no time to do it.

IdeaToMVP Academy

Want to build with AI — not just read about it?

4-week live cohort for founders. Learn to ship AI agents, scope MVPs, and automate your business — taught by the same team that writes these guides.

Explore the Academy →
Share this post :