AI Respondents Said 98% Knew the Answer. 52% of Real People Did. Don't Validate Your Startup With Synthetic Users

Surya Pratap
By Surya Pratap

October 1, 2026

9 min read

MVP Strategy
A two-part diagram. On the left, pairs of bars comparing real American survey respondents with AI digital twins built from their own profiles: on a question about the First Amendment, 52 percent of real people answered correctly against 98 percent of AI respondents; 43 percent of Hispanic adults said they were following the World Cup against 97 percent of AI respondents; and 25 percent had heard about data centers in the news against 3 percent of AI respondents. On the right, the overall picture from Pew Research's study of about 300 questions: an average error of 12.4 points, more than 15 points on 28 percent of questions, real people choosing 'not sure' about four times as often, and nearly half of questions with at least one answer no AI respondent chose.Confident, moderate, and wrongHover to explore
AI stand-ins knew more than real people, were surer than real people, and answered more like the stereotype than the person. That is the customer research they will give you too.

A growing number of founders validate ideas without talking to anyone. They describe their target customer to a language model, generate a few hundred "synthetic users", and ask them whether they would buy. The answers come back fast, cheap and confident. Pew Research has now tested the same idea against real people, carefully, and the confidence is the problem.

Figures come from two Pew Research Center reports published on 30 September 2026 by Athena Chapekis, Arnold Lau, Samuel Bestvater, Sono Shah, Andrew Mercer and Aaron Smith: "Can AI stand in for human survey takers? Not really" and "How synthetic polling results change based on the AI model". Pew studied public opinion surveys, not product research. Applying the findings to founders is my reading, from section 3 onward.

1. What Pew tested

Pew runs the American Trends Panel, a large panel of real US adults it surveys repeatedly. For this study it built a "digital twin" of each panel member: an AI model was given that person's demographic details and their answers to an earlier political typology survey, then asked to answer new surveys as that person would.

The twins answered three real surveys from early 2026, about 300 questions in total, using the same wording and instructions as the humans. Pew then compared the twins' answers with what the real people had said. The main model was Anthropic's Claude Opus 4.6. A second report repeated the exercise with OpenAI's GPT-5.1.

This is close to the best case for synthetic respondents. Each twin was built from a real person's own detailed answers, not from a short persona description.

2. How far off the answers were

Across the three surveys, the AI answers missed the real ones by 12.4 percentage points on average. On about 28% of questions the error was larger than 15 points, and differences of 20 to 40 points on individual answers were common.

A few examples show the shape of the misses:

  • Heard about data centers in the news: 25% of real people, 3% of AI twins.
  • A question on the First Amendment: 52% of real people answered correctly, 98% of AI twins.
  • A question on NATO: 56% of real people correct, 99% of AI twins.
  • Hispanic adults following the World Cup: 43% real, 97% AI.

Pew's conclusion: AI models are "not an adequate replacement" for surveying real people on topics of broad importance, and their results "differ in ways that are often unpredictable."

The twins were built from the real people's own answers, and still disagreed with them by twelve points on average.

3. Five failures, and what each one does to customer research

Pew found consistent patterns in how the AI got things wrong. Each one maps onto something founders are tempted to ask synthetic users.

They know too much. AI twins answered knowledge questions correctly far more often than real people — 98% against 52% on one. A synthetic customer will understand your category, your jargon and your value proposition. Your real customers mostly will not, and that gap is where products fail.

They are rarely unsure. Real people chose "not sure" about four times as often as the AI. In customer research, "I don't know, I've never thought about it" is often the most important answer you can get. Synthetic users almost never give it.

They skip the extremes. On nearly half the questions, at least one answer choice was never picked by a single AI respondent. The people who would never buy, and the small group who would buy tomorrow, are exactly the ones a synthetic sample leaves out.

They play the stereotype. AI twins answered as the caricature of their group: 97% of Hispanic twins following the World Cup against 43% in reality, 95% of Republican twins favourable to Israel against 58%. Describe your ideal customer, and the model will answer as the cliché of that customer.

The fifth failure matters most for anything new. The largest misses were on recent events — the data centers question was off by 22 points — even though the model had access to earlier trends. A new product category is, by definition, something real people have not encountered before. That is the situation in which synthetic answers are least reliable.

4. Change the model, change your market

Pew's second report ran the same digital twins on two different models. The two disagreed with each other as much as with reality.

  • Dissatisfied with the state of the country: 69% of real people, 70% of Claude twins, 100% of GPT twins.
  • Trump voters approving "very strongly": 65% real, 46% Claude, 88% GPT.
  • Choosing a middle option on scaled questions: 45% of real people, 56% of Claude twins, 31% of GPT twins.

One model made everyone moderate. The other made everyone extreme. Pew's summary: each synthetic sample "painted a very different picture of the American public," and "neither truly reflected actual public opinion."

For a founder, this means the "market" a synthetic user study describes depends partly on which model you happen to use. Switch models, and your willingness-to-pay curve moves.

5. What synthetic users are still good for

Pew is not saying AI is useless in research. It names two places where it helps: categorising real answers to open-ended questions, and writing the code to analyse survey results. In both, the AI processes what real people said rather than inventing it.

For founders, the same rule gives a useful list:

  1. Drafting your interview guide.

    Ask a model to suggest questions, then cut the leading ones. It is quick, and nothing it produces is treated as evidence.

  2. Pre-testing your wording.

    Run your survey past a model to find confusing or double-barrelled questions before real people see it.

  3. Rehearsing objections.

    Role-play a sceptical buyer to practise your pitch. Treat it as sparring, not as a forecast of what buyers will say.

  4. Coding the transcripts.

    After you have talked to real people, use a model to tag themes across your notes. This is the use Pew endorses.

What it cannot do is tell you whether people want what you are building, how many of them, or what they would pay.

6. The cheaper alternative that works

Five real conversations with people who have the problem will tell you more than five hundred synthetic ones, because they can surprise you. They can say "not sure", "never heard of that", or "we solved that with a spreadsheet years ago" — the answers Pew found the AI almost never gives.

If you have not set those conversations up before, our guides on running a discovery call and on turning community interest into a first paid customer cover the mechanics.

7. What I would not claim

Pew studied public opinion, not product demand. The questions were about politics, current events and knowledge. I am applying the patterns to customer research because the mechanism is the same, but Pew did not test product questions.

Two models were compared. Claude Opus 4.6 was the main model, and the comparison added GPT-5.1. The headline 12.4-point figure is the average across three surveys; the model comparison reports 11.4 points for Claude and 13.3 for GPT on the questions it covered. Other models, prompts or methods could do better or worse.

Better methods may narrow the gap. Pew says so itself: newer models and new ways of building synthetic samples could produce smaller errors. This is a measurement of now, not a permanent limit.

Your synthetic users are probably built on less. Pew's twins were built from each person's own detailed survey answers. Most founder personas are a paragraph of description. That suggests founder-built synthetic users would do worse, but it is my inference, not something Pew tested.

The honest summary

Pew built the most careful version of synthetic respondents most founders could hope for — AI twins of real people, built from their own answers — and they still missed by about 12 points on average. They knew more than real people, were surer than real people, avoided the extremes and answered like the stereotype of their group. They did worst on new topics. And changing the model changed the picture.

Those are not edge cases for a startup. They are exactly the errors that would make a product look more wanted, better understood and more broadly appealing than it really is.

Use AI to prepare for customer conversations and to analyse them. Don't use it to replace them.

Sources: Athena Chapekis, Arnold Lau, Samuel Bestvater, Sono Shah, Andrew Mercer and Aaron Smith, "Can AI stand in for human survey takers? Not really", Pew Research Center, 30 September 2026 — the digital-twin method, the three early-2026 American Trends Panel surveys and about 300 questions, the 12.4-point average error, the share of questions over 15 points, the data centers, First Amendment, NATO and World Cup examples, the four-times "not sure" gap, the answer choices no AI respondent selected, the stereotype examples, and Pew's conclusions on what AI can and cannot do. The same authors, "How synthetic polling results change based on the AI model", Pew Research Center, 30 September 2026 — the Claude Opus 4.6 and GPT-5.1 comparison, the 11.4- and 13.3-point errors, the middle-option rates, the examples of model disagreement and the caveat about future methods. The application to customer research in sections 3 to 6 is mine.

IdeaToMVP Academy

Want to build with AI — not just read about it?

4-week live cohort for founders. Learn to ship AI agents, scope MVPs, and automate your business — taught by the same team that writes these guides.

Explore the Academy →
Share this post :