oosioo
Concept image for AI synthetic consumer panels: a row of person icons in which a few highlighted real people are mixed among outlined synthetic respondents.
SeriesActually One Click · Ep. 8

Running AI Consumer Research Without Hiring Panel Respondents — What AI Synthetic Panels Are Good For, and the Self-Cannibalization the Survey Industry Fears

· SYSOP · 5 views

I ran a consumer survey with a few lines of AI code. "Ask a woman in her thirties who works full-time how she sees this new product concept." A few minutes later, out came a single report complete with respondent profiles, a distribution of purchase intent, price sensitivity, and open-ended responses. Except no actual person answered that question. The respondents weren't people — they were 30,000 fake consumers built in advance.

This installment is a bit different in character from the previous ones. Until now I've covered one-off things that were done in a single pass, but this one is a tool built over several months, and the key point for the user is that it works with one line, one click. In other words, it's more like a solution-type product. And beyond that, I also looked into what the traditional survey industry is going through as it watches cases like this unfold. Isn't that already interesting?

What I Asked

What I actually entered was a sentence like this.

Show a group of women in their late thirties who work full-time in the Seoul metropolitan area a new low-sugar yogurt concept, and ask whether they'd be willing to buy it, at what price they'd start to feel it's too expensive, and why they hesitate.

The question itself is fairly simple, but I freely included various pieces of information such as the product concept, name, and price. Narrowing down the target consumer to some degree can also be a key point.

What Came Out

200 respondents are automatically selected, and a report comes out based on each one reading the concept description and answering. Purchase intent is captured as a distribution on a 5-point scale, "at what price does it start to feel expensive" is shown as a curve by price band, and the reasons for hesitation are grouped from the open-ended responses into a few clusters. Cross-tabs broken down by gender, age, and region are attached as well.

Up to this point, it looks exactly like a real survey report. There's exactly one difference: the respondents aren't people. That difference is the subject of this entire piece.

Who Answers Instead of People — NeoPersona

The ones answering are fake people built in advance. In this service I called these fictional individuals NeoPersona. There are about 30,000 of them, and each one isn't defined by surface attributes like age, gender, region, and occupation alone — each also carries personality (five axes such as extraversion and openness), political leanings, values, media habits, and consumption style. When running a survey, people matching the given conditions are selected from among these 30,000, questions are posed to each, and the answers are collected.

Why aren't surface attributes enough? If you only give the information "a 35-year-old woman in Seoul," the language model invents on the spot the most average possible 35-year-old woman in Seoul. Then even if you ask 200 of them, all 200 give similar answers. The premise itself strongly generates bias, so the problem is that only predictable results come out. But in real research, people under the same conditions generally still diverge — some like a new product and some are suspicious of it. What creates that texture is attributes like personality and values, and the point of this exercise is enabling responses that reflect that.

Where Did the 1 Million Come From — Nemotron

The 30,000 were built one by one using a language model, which costs both money and time. To go broader, we needed base material, and the Nemotron-Personas-Korea dataset that NVIDIA released in April 2026 became that material.

Nemotron is a set of profiles for one million fictional Korean people. The key point is that they weren't simply invented, but built to match the population distribution from Statistics Korea. The way it's built is interesting: first, a statistical device called a probabilistic graphical model draws, in line with actual population statistics, the probability that "a person of this age in this region has this occupation, education level, and marital status," and on top of that skeleton, a language model (from Google's Gemma family) drapes a Korean-language narrative — that is, who this person is as a person. The population distribution was based on Statistics Korea data, and the name distribution on Supreme Court data. The license is also open (CC BY 4.0), so it can be used as long as the source is credited.

What we gained from this comes down to two things. One is scale — from 30,000 up to 1 million. Each individual even comes with up to seven variants differing in flavor — job, family, travel — so the real scale is even larger. The other is regional resolution. Because Nemotron matches its distribution down to the level of individual si/gun/gu administrative districts (about 252 of them), even narrowing to something like "people in their forties in Haeundae-gu, Busan" produces people that match that specific area's population makeup. It narrows down to the neighborhood level, not just a lumped-together metropolitan city.

To sum up, Nemotron provides a statistically accurate population skeleton, and NeoPersona adds the flesh of personality and values on top of it. It's a structure where statistical validity is borrowed from someone else's well-built data, while we add the psychological depth that scatters the answers. For fun, I also generated 40 million of these separately, and I keep them around to ask this and that from time to time.

Synthetic consumer panel structure diagram: a Nemotron population skeleton of 1 million matched to Statistics Korea's population distribution, layered with NeoPersona's psychological flesh of personality and values, to create synthetic respondents.

What's Happening in the Industry Right Now

But we're not the only ones doing this interesting thing. Over the past two years, an entire market has been forming under the name "synthetic respondents." It looks like a jumble, but it actually breaks into three branches.

ApproachWhat it doesExamples
Pure syntheticAI invents the answers with no real people involvedSynthetic Users, Aaru
Statistical augmentationA small amount of real, actually-collected responses is statistically expandedFairgen
AI moderatorAI interviews real people at scale (not synthetic)Outset, Genway

Simply put, it's the pure synthetic side that dominates. That's where the money is flowing. Synthetic Users runs user interviews with AI participants instead of recruiting real ones. However, an independent review (MeasuringU) found that while synthetic responses track the broad direction of real attitudes, they diverge significantly from deeper behavior. Aaru simulates voter and consumer groups using thousands of AI agents, and in a December 2025 Series A round it was valued at a "headline valuation" of $1 billion. That flashy number, though, doesn't guarantee accuracy — the company predicted a Harris win (53–47) in the 2024 US presidential election, and Trump actually won.

Big research companies have jumped in too. Qualtrics launched a synthetic panel product (Edge Audiences) in March 2025, and NIQ (formerly Nielsen) released a synthetic respondent tool (BASES AI Screener) in April 2025 that screens new product ideas using its own panel data. Toluna unveiled more than 1 million synthetic personas built from anonymized data drawn from its 79-million-person proprietary panel. YouGov went as far as acquiring a synthetic research company (Yabble) outright in August 2024.

The second row, statistical augmentation, has a somewhat different character, and there's a reason to note it here. A company like Fairgen does not invent answers that don't exist. When real responses have already been collected but a particular subgroup's sample is too small, it trains only on that real data and expands the sample statistically. That means it can be "verified against real data," and there are five third-party verification results out, including ones from Google and L'Oréal. This distinction between pure synthetic and statistical augmentation becomes important again later.

What the Traditional Survey Industry Fears — Self-Cannibalization

This is where things get interesting. The big research companies welcome this and fear it at the same time.

The reason for the fear is simple: they'd be breaking their own rice bowl. A survey company's core business is asking people questions and getting paid for it. But if synthetic respondents can knock out a large chunk of that work in minutes, at a fraction of the cost, the company ends up cheapening what it sells with its own hands. This is called cannibalization. So the industry's response splits into two layers.

On the surface, almost everyone says it's "a complement, not a replacement." Patrick Comer, CEO of panel company Cint, states flatly, "Our foundation is the people that we all work with every day in panels: that's not going away." Ipsos, in a white paper, openly called unvetted newcomers "snake oil salesmen" and wrote: "This technology is not magic – it is math." It stresses that the quality of synthetic data depends "entirely on real human data." Kantar ran its own experiments early on with GPT-4 and concluded that today's synthetic samples are "just not good enough to use as a supplement for human sample." It found that GPT-4 answered too positively, that repeated prompting produced only formulaic answers, and that variance and nuance were lacking.

But look beneath the surface, and the story is different. In a survey Qualtrics ran of more than 3,000 researchers across 14 countries (the 2025 Market Research Trends Report), seven in ten (71%) said that within three years, synthetic responses would account for more than half of all data collection. That's the size of the gap between the official position of "just a complement" and a forecast in which more than half the work actually goes to synthetic. Industry outlet GreenBook even ran a piece titled "The Death of the Survey?" Though the marketing scholar quoted in that piece, Mark Ritson, is actually on the skeptical side — his point being that any synthetic respondent, in the end, still leans on data made by real people.

The big companies' real motive is closer to "if we can't stop it, let's swallow it." While keeping their distance from fully synthetic approaches as risky, they use the real data they already own as a moat and sell synthetic products built on top of it. That's the direction behind YouGov buying a synthetic company, and Toluna and NIQ building synthetic personas from their own panel data.

Where It Goes Wrong

Under the rules of this series, this section has to be the longest. Synthetic respondents are plausible enough that it's hard to tell when they're wrong. Unlike a photo, you can't judge it at a glance. So knowing where it goes wrong is close to a prerequisite for using this thing. Fortunately, scholars have dug into this quite a bit already.

First, answers cluster toward the center. When political scientists tested this against ANES, a well-known polling dataset (Bisbee et al., 2024), synthetic responses lacked variance. Answers clumped excessively around the mean, collapsing differences between groups — yet the apparent precision actually looked higher than the real thing. That's the trap. Because it looks precise, it leads people to underestimate the sample size actually needed. It manufactures confidence that isn't warranted. Earlier I said that embedding personality scatters the answers, but even so, they don't scatter as much as real people do.

Second, it tilts toward the majority and erases minorities. Language models pull answers toward whatever they saw most often in training. Multiple studies point out that synthetic responses skew male and highly educated, and underrepresent minority opinions (Santurkar et al., 2023, and others). In market research, the most valuable information is often the signal that "a minority strongly dislikes this" — and that's exactly the spot that gets flattened.

Third, it gives different answers to the same question. When the same political scientists re-ran the identical prompt over a three-month span, nearly half (48%) of the regression coefficients diverged significantly from the real data, and of those that diverged, a third (32%) had their sign flipped entirely. What was reported as liking something turned into disliking it. If it doesn't reproduce, it isn't research.

Fourth, it eats away at itself — a deeper self-cannibalization. Where the earlier cannibalization was a business story, this one is a technical story. A 2024 study published in Nature (Shumailov et al.) showed that repeatedly retraining AI on AI-generated data causes model collapse. With each generation, the tails of the distribution — that is, the rare, extreme cases — disappear first. If synthetic responses get fed back into training, and that model then generates more synthetic responses, the loop structurally erases minority opinions over time. In effect, the second problem (erasing minorities) worsens on its own as time passes.

So is there no encouraging research at all? There is. A Stanford research team (Park et al., 2024) interviewed 1,052 real Americans for two hours each and built agents that mimicked each person, and those agents answered quite similarly to the originals (normalized accuracy of about 0.85). That's the figure commonly cited as "85% accurate." Two things need noting, though. One, that 0.85 is measured against a baseline where perfect (1.0) means "the same person answering again two weeks later." Even real people don't answer exactly the same as themselves two weeks on. Two, achieving that accuracy required a real two-hour interview per person. If you use personas built only from surface attributes, accuracy drops sharply. In other words, how much real data you feed in is everything. This is exactly the point behind Ipsos's line that "it's math, not magic."

So Where Does It Get Used

Listing out where it goes wrong doesn't mean it's useless. There are places it belongs. The dividing line comes down to one thing — how hard it would be to undo the decision if it turns out wrong.

Where it's fine to use (low-risk, upstream)Where it shouldn't be used (high-risk, irreversible)
Pre-checking a questionnaire before actually fielding itDecisions with money on the line, like pricing or launch/no-launch
Narrowing down which of twenty concepts to keepRegulated industries such as healthcare and finance
Forming hypotheses, setting directionResearch where the depth of lived experience or emotion is the whole point
Simulating hard-to-reach groups like executives or expertsSomething genuinely new that isn't in the training data

The key is the sequence: narrow with synthetic, confirm with real. You build hypotheses with synthetic respondents, cut the candidates from twenty down to three, and then only ask real people about those three to finalize the decision. This is the approach the industry is settling into.

Hybrid verification funnel that narrows with synthetic and confirms with real: roughly 20 concepts are pre-screened by synthetic respondents down to 3 candidates, and only those three are put to real people to finalize.

And one more thing: the report must always disclose where the synthetic part ends and the real part begins. In fact, in its 2025 revision, the ICC/ESOMAR international code of conduct added a new clause requiring that, when synthetic data is used, clients be informed of the method, the data source, and the limitations. The point is not to blur synthetic data into looking like real people.

What It Can't Do

It can't produce reactions to something genuinely new. A synthetic respondent is a thing that reflects back the world it saw in its training data. How people would react to a concept that never existed in the world is, by definition, not in the training data. All it can do is imitate "how people reacted when something like this came out before."

It cannot substitute for a firsthand account. In the words of someone who has been sick, suffered a major loss, or experienced discrimination, there's a specificity that can't be invented. Synthetic output can only generate an average-sounding sentence in that spot.

You can't tell immediately when it's wrong. In the photo and video work from earlier installments, mistakes were visible. Synthetic responses are the opposite. Because they produce plausible-looking tables and charts, confirming whether it's wrong ultimately requires checking it against real data. So the most honest use of this thing isn't as "a tool that eliminates real research," but as "a tool that aims real research more precisely."

That's also why we built 30,000 personas and layered on a million-person supply of raw material. It's not about replacing people, but about cheaply figuring out first what to even ask people. Getting an answer with a single click, and also knowing where to trust that answer and where not to — that's what this whole undertaking was about.

There's honestly more here I can't disclose in this piece — plenty of deeper, more interesting refinements — and right now we're also applying it to solving fairly substantial real problems. I think that's a genuinely interesting thing.

Sources and Further Reading

All figures and quotations were confirmed against the primary sources below. Accuracy claims that companies make about their own products are mostly self-reported, so this piece includes only what could be verified.

Datasets

Papers

Companies and Organizations

  • Qualtrics, "Synthetic data in market research" — qualtrics.com · 2025 Market Research Trends Report — prnewswire.com
  • Ipsos, Synthetic Data: From Hype to Realityipsos.com (PDF)
  • Kantar, "What is synthetic sample – and is it all it's cracked up to be?" (2023) — kantar.com
  • Cint, "Why everything we thought we knew about sampling might change" — cint.com
  • YouGov, "YouGov acquires Yabble" (2024) — yougov.com
  • NielsenIQ, "BASES AI Screener" (2025) — nielseniq.com
  • Toluna, "One million synthetic personas" (2025) — tolunacorporate.com
  • ICC/ESOMAR International Code — iccwbo.org

Startups and Industry Media

  • Fairgen, Google synthetic data validation study — fairgen.ai
  • Synthetic Users, independent review (MeasuringU) — measuringu.com
  • Aaru, "headline valuation" Series A (TechCrunch, 2025) — techcrunch.com · 2024 election forecast (Semafor, 2024) — semafor.com
  • GreenBook, "The Death of the Survey?" — greenbook.org

Read this series from the start: Actually One Click.