
The People Who Got 5,000 AI Voters to Cast Ballots — Synthetic Voters That Run the Election in Advance
More companies and researchers are asking AI-generated virtual voters how they would vote instead of real people. This piece follows Aaru, which came within 371 votes of the actual result in a New York primary but missed the presidential race; GPT models that skewed toward the Greens when questioned in German; and a Korean firm that tried predicting local elections — then shares impressions from running the experiment myself with NeoPersona, and why none of this should be called an opinion poll.

Election season brings a flood of opinion polls. The numbers showing who leads by how many points change daily, and each polling firm's results differ slightly. To get a single one of those numbers, pollsters place thousands of phone calls. These days, hardly anyone picks up.
But a few years ago, people began trying to gauge voter sentiment in a completely different way. Instead of asking real people, they ask virtual voters created by AI. They generate thousands of fake people whose age, region, occupation, and political leanings are matched to the actual population makeup, show them the candidates and the issues, and ask who they would vote for. No phone calls are made, and results come back in minutes.
In the last installment (#8), I wrote about using this method for consumer research. This time, it's about people who pointed the same tool at elections. If a new yogurt flavor test gets it wrong, you can just make another batch, but an election happens only once, and getting it wrong causes trouble for a lot of people. That makes this both far more interesting and far more delicate.
It Started With a Single Paper
The starting point was a single paper published in 2023 in the American political science journal Political Analysis. A team of researchers at Brigham Young University (Argyle et al.) gave it a provocative title right from the start: "Out of One, Many."
Here's the method. They took the backgrounds — age, gender, race, education, party affiliation, religion, and so on — of thousands of real respondents from ANES, the most authoritative election survey in the United States, wrote them out as sentences, and fed them into GPT-3. Something like: "I am a white man in my fifties living in Georgia, I attend church regularly, and I consider myself conservative." Then they asked who this person would have voted for in the presidential election. The researchers called this group of AI-generated respondents a silicon sample.
The results were striking. The correlation between the AI's voting choices and real people's actual votes reached 0.90 for the 2012 presidential election, 0.92 for 2016, and 0.94 for 2020. Given nothing but a background sketch, the language model guessed the person's vote quite plausibly.
But look at the appendix and a crack appears. When you isolate pure independents — people leaning toward neither party — the correlation drops to 0.31 for 2012 and all the way to 0.02 for 2020. In effect, the model couldn't predict them at all. Guessing that a Republican votes Republican isn't hard. Elections are always decided by the people in the middle, and that's exactly where the AI performed worst.

Companies Selling Votes
The paper was quickly followed by companies. The most famous is a New York startup called Aaru.

In September 2024, the American outlet Semafor reported that Aaru had come within 371 votes of the actual result in that June's New York Democratic House primary. The company said it used about 5,000 AI respondents instead of real people, that each survey took a little over a minute, and that the cost was less than a tenth of a human survey. As it happens, that's 5,000 people too.
Then, two months later, came the presidential election. Of the seven battleground states, Aaru forecast Harris leading in Michigan, Nevada, Pennsylvania, and Wisconsin. Everyone knows how that turned out. Trump took all seven. The day after the election, a co-founder said: "A coin flip is a coin flip. 53–47 isn't statistically much different from 48–52. We were within the margin of error, so we did fine."

Getting it wrong wasn't the end of the story. In a Series A round in December 2025, Aaru was valued at $1 billion for a partial stake (per TechCrunch). Its client list includes names like Accenture and EY alongside political campaigns. In a project with EY, it simulated in a single day a survey of 3,600 investors across some 30 markets that would normally take six months, and reported a median rank correlation of 0.90 with the real responses.
Similar companies keep appearing.
| Company | Where | What it does |
|---|---|---|
| Aaru | United States | Forecasts elections and markets using AI respondents; clients include political campaigns |
| Simile | United States | Individual-level digital twins. Founded by Stanford researchers who interviewed 1,000 Americans for two hours each to build AI that mimics each person; raised $100 million in February 2026 |
| Electric Twin | United Kingdom | A "synthetic audience" blending real surveys with language models; founded by a former data advisor to the UK Prime Minister's Office |
| Expected Parrot | United States | Open-source tool for running surveys and experiments with AI agents; founded by an MIT economist |
Few of these companies focus solely on elections. For most, the core business is selling corporate clients an answer to "how will people see our ad," and elections are simply the most visible showcase. That's because an election is one of the rare test grounds where the correct answer is revealed a few months later. Get it right, and it becomes an advertisement; get it wrong, and you have to explain yourself, like in the interview above.
It's already an open secret that American campaigns run focus groups using synthetic voters. In August of this year, TechPolicy.Press covered this trend — and its risks — in an article titled "Here Come the Synthetic Voters," describing how campaigns test speeches and ad copy on virtual voters before showing them to real people.
When You Ask in Korean
Everything so far has been about the United States. American elections have a clear two-party structure, English-language text floods the internet, and language models were raised on that text. Does this approach work elsewhere?
A team of German researchers (von der Heyde et al.) tested exactly this. They created virtual figures with the same backgrounds as respondents to the 2017 German federal election survey and asked GPT-3.5 for their voting choices in German. The short answer was: it couldn't predict them. In particular, the model skewed toward the Greens and the Left Party. The researchers pointed out that less than 5% of internet text is in German, while more than half is in English. When the same team broadened the scope to the 2024 European Parliament election, redoing the exercise with the backgrounds of 26,000 voters, the conclusion was similar. Predictions mostly failed, and even the accuracy that did exist varied wildly by country and language.
What about Korea? A paper posted in May of this year by a Korea University researcher (Dynamo-K, pre-peer-review) is almost the only one to take on this question directly. It built 5,000 virtual voters from domestic social survey data and had four language models re-run Korean elections from 2017 to 2025 in Korean. Here too, the number is 5,000. It correctly called the winner in all three presidential elections, and even the 2022 race — decided by a margin of just 0.73 percentage points — was tracked with an average error of 2.1 points. The two general elections, however, both got the largest party wrong. But getting even these results required fixing one thing first. Left unadjusted, the model's virtual voters skewed progressive. Before correction, 97% of centrist virtual voters chose the progressive candidate. Once this skew was explicitly corrected, that figure fell to 59%, and the average error dropped to a fifth of what it had been. It's the same ailment as the German model's tilt toward the Greens. One more thing worth remembering: this study, too, retroactively reproduced elections whose results were already known — it did not predict anything in advance.

Some have tried predicting in advance. The Korean AI company Post AI published, after polls closed on this June 3's local elections, the results of predictions it had made using its own synthetic voters for six metropolitan and provincial chief executive races. It explicitly stated this was "not an exit poll or an opinion survey, but an AI-persona-based simulation." It correctly called the winner in five of the six races, missing only Seoul.

Where It Goes Wrong
Laid side by side, these cases fail in almost the same places.
Answers cluster toward the middle. When American political scientists (Bisbee et al., 2024) ran the same experiment on ChatGPT, the average was close to real survey results, but the spread of responses was much narrower. Real people give wildly different answers even under identical conditions, whereas virtual voters answer similarly to one another. This makes the results look more precise than they actually are — a false sense of certainty.
Ask the same question again and you get a different answer. In the same study, running the identical question again in April, June, and July 2023 produced different distributions, because the model gets quietly updated behind the scenes. You can't compare last month's result to this month's and say "public opinion shifted" — you don't know whether it was public opinion that shifted, or the model.
It leans a certain way. A study published in February of this year (Parikh et al.) placed several language models' answers side by side with actual polling from the 2024 U.S. presidential campaign. Every model rated Harris's favorability 10–40% higher than the polls did. Trump was also rated higher, but only by about 5–10%.

It's the same root cause as the tilt toward the Greens in Germany and toward progressives in Korea. The internet text that language models were trained on simply doesn't represent the full electorate evenly.
It's easy to look like it got it right. Reproducing an election that's already over isn't prediction. Once you know the answer, there's room to turn the dials until the model fits it. If an announcement comes out after an election saying "we ran our model and it matched," the first thing to check is whether that result was published before the election.
AI Won't Cast the Vote for You
There's also an interesting wall here. If you keep asking today's major language models "who would this person vote for" about real, living candidates, they quite often dodge the question or refuse to answer.
That's because the companies that built these models deliberately trained them to be cautious about elections. Ahead of the 2024 election, Anthropic barred Claude from being used in political campaigns, and in April of this year it stated as a principle for election-related answers that it should help users reach their own conclusions without being steered toward a particular viewpoint. OpenAI's model spec likewise states that "the model should not push the user toward a particular side in pursuit of its own agenda." Even when speaking through the mouth of a virtual voter, it's ultimately the model choosing that vote, so this runs straight into that principle.
For those trying to run election simulations, this is inconvenient, but it makes sense when you think about it. If an AI used by hundreds of millions of people leans even slightly toward a particular candidate, its influence dwarfs that of any opinion poll. And as we've seen, these models are already tilted a little to one side.
I Tried It Myself
Actually, I tried this myself. Using the NeoPersona introduced in installment #8 — a pool of virtual figures built to match statistical distributions — I drew voters by region and had them cast ballots ahead of this past June's local elections.
The results were fairly interesting. Races that were clearly decided came out clearly decided, close races stayed close, and the overall picture pointed roughly in the same direction as the polling trends of the time. What was most interesting, though, wasn't the numbers but the process. I ran headfirst into the wall described above — the AI's reluctance to pick a vote on behalf of a real candidate — and I could watch, within minutes, how the whole picture shifted each time I changed a single assumption. Being able to throw out a question like "what happens if this issue blows up in this region" without gathering thousands of real people was, in itself, a remarkable experience.
That said, I won't be putting those results down here as numbers. The reason follows below.
Don't Call It a Poll
The American Association for Public Opinion Research (AAPOR) drew a clear line on this issue when it revised its code of ethics in June of this year. AI-generated responses are not research participants. And because the terms "poll" and "survey" refer to data obtained from actual people, that name should not be attached to AI-generated data. A task force report from the same association offered this sentence as an example disclosure to attach to any result using synthetic responses: "Margins of error cannot be calculated for synthetic responses." That's because a margin of error is a value that can only be computed when real people are randomly sampled.
Korea already has a standard for this too. Ahead of the general election, the National Election Commission issued operating guidelines on generative AI in August 2023. Publishing predictions of support ratings or election outcomes generated by AI is permitted as long as it discloses that the result comes from AI analysis and may therefore have limited reliability. But the guidelines specifically warn that presenting such results in a way that could be mistaken for actual polling could run afoul of the Public Official Election Act's ban on false commentary and reporting. That's exactly why Post AI, mentioned earlier, went out of its way to add the line "this is not an opinion poll."

That's why this article doesn't call the results of virtual-voter experiments a "survey," and why I haven't included the numbers from my own experiment. The moment such numbers spread, they're exactly the kind that get read as "an AI poll shows candidate X leading by Y%."
So What Is It Actually Good For
Does that mean it's useless? Not at all. It just has a specific place where it belongs.
Synthetic voters are less a prediction tool than a device for testing assumptions. You feed in a premise — "if this issue grows," "if this candidate runs on this pledge," "if turnout is this low" — and quickly and cheaply scan where that premise leads. Then you take only the handful of things that seem genuinely important and ask real people about those. The sequence described in installment #8 — "narrow down with synthetic data, confirm with the real thing" — applies just the same to elections. It's ultimately why American campaigns run speeches past virtual voters first.
Conversely, the moment you present results under a headline like "the winner, as picked by AI," this tool moves into its most dangerous role. The model's inherent bias gets dressed up as public sentiment, and a number with no margin of error is placed side by side with numbers that have one.
I'll probably run this again come the next election season — not to get it right, but to see what assumptions the race is standing on.
Sources and Further Reading
Papers
- Argyle et al., "Out of One, Many: Using Language Models to Simulate Human Samples," Political Analysis 31(3), 2023 — arxiv.org/abs/2209.06899
- Bisbee et al., "Synthetic Replacements for Human Survey Data? The Perils of Large Language Models," Political Analysis 32(4), 2024 — cambridge.org
- von der Heyde, Haensch, Wenz, "Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion," 2024 — arxiv.org/abs/2407.08563
- von der Heyde et al., "United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections," 2024 — arxiv.org/abs/2409.09045
- Parikh, Cen, Podimata, "Do LLMs Track Public Opinion?" 2026 — arxiv.org/abs/2602.06302
- Kang, "Diagnosing Korean-Language LLM Political Bias via Census-Grounded Agent Simulation," 2026 (pre-peer-review) — arxiv.org/abs/2605.18395
Companies and News Coverage
- Semafor, on Aaru's New York primary forecast (2024-09-20) — semafor.com · on its presidential election forecast (2024-11-04) — semafor.com · on its post-election explanation (2024-11-06) — semafor.com
- TechCrunch, on Aaru's Series A round (2025-12) — techcrunch.com
- EY, on its AI-simulated survey case study (2025-10) — ey.com
- SiliconANGLE, on Simile's $100 million raise (2026-02-12) — siliconangle.com
- UKTN, on Electric Twin's funding round (2026-02-12) — uktech.news
- Expected Parrot — expectedparrot.com
- TechPolicy.Press, "Here Come the Synthetic Voters" (2026-08-16) — techpolicy.press
- Post AI, on its AI predictions for the June 3 local elections — post-ai.com
Regulations and Policy
- AAPOR, Code of Professional Ethics and Practices (revised June 2026) — aapor.org
- AAPOR, "Responsible AI Integration in Survey Research" (2026-05) — aapor.org (PDF)
- National Election Commission, "Stepping Up the Response to AI-Generated Disinformation" (2023-08-31) — nec.go.kr
- Anthropic, "Preparing for global elections in 2024" — anthropic.com · "An update on our election safeguards" (2026-04) — anthropic.com
- OpenAI, Model Spec — model-spec.openai.com
Photos (all from Wikimedia Commons)
- Polling booth: "대한민국 21대 대통령선거 기표소 1" by revi, CC BY-SA 2.0 KR — commons.wikimedia.org (cropped; usable under the same license)
- Ballot box: "Korean ballots box" by revi, CC BY 2.0 KR — commons.wikimedia.org
- Stickers: "Pile of "I Voted" stickers" by Funknendai, CC0 — commons.wikimedia.org
Related reading (installment 8): Consumer Research With AI Synthetic Panels