Answers
What is synthetic research?
Synthetic research asks a census-grounded panel of simulated respondents, with persistent memory, the same questions a human survey would ask, then reports option shares, crosstabs, and quotes. The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. MAE is an average miss in percentage points against published poll toplines. A narrower April test on 5 places with 10,000 respondents scored 7.47 points on 404 non-electoral questions. It does not stand as a legal sample.
What synthetic research is
You write an instrument with the team. A panel of simulated respondents answers it. You get percentages, the way you would from a survey vendor, plus quotes and a PDF. The respondents are built from census demographics and a model that has been scored against published polls.
Interview tools return themes. A tabulated panel returns option shares from simulated respondents. Lewsearch's artifact is the panel. The statistic is the option share. The quote is color.
Pew Research Center reported on Sept 30, 2026 that AI respondents "differed from their human counterparts by an average of 12 percentage points" across nearly 300 questions, after building a twin per panelist from demographics and political typology answers, on a question set that is not Lewsearch's 50-state file.
Synthetic respondent, synthetic panel, digital twin, model prompt
These words name different objects.
| Object | What it is | What you can score |
|---|---|---|
| Synthetic respondent | One simulated respondent with a fixed demographic row | Nothing useful by itself. Score the aggregate. |
| Synthetic panel | Many simulated respondents asked the same instrument, then tabulated | Option shares vs a real poll (mean absolute error, in points) |
| Digital twin (here) | A matched agent from the same pool, interviewed one to one | Themes and verbatims. The poll MAE is a panel statistic. |
| Frontier-model prompt | Can be asked for a distribution. No respondent-level records you can recontact. | Score it only if the vendor publishes a file. |
On Lewsearch, a synthetic panel is the default study. A digital twin is a matched simulated respondent from that pool, interviewed one to one. A CRM live feed and an industrial machine twin are different products. A frontier model can be asked for a national split. Lewsearch's deliverable is the tabulated panel.
How a Lewsearch study is assembled
Book a demo. The team sets up the study with you. You bring the question and the response options, and you pick a market. The draw matches that market's census margins. Every study is calibrated against published survey data before it reaches you. Enterprise panels are also calibrated to your own past studies. The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. The file uses 200 simulated respondents per question. A narrower April test on 5 places with 10,000 respondents scored 7.47 points on 404 non-electoral questions.
A 500-respondent Ohio study is 500 distinct rows: age, education, income, and party. The output is a topline plus crosstabs. If the question matches a scored category on the methodology page, the report can point at that category's miss. If it does not, treat the read as directional.
How accurate it is
The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. MAE is an average miss in percentage points against published poll toplines. The best 80% keeps the 1,216 lowest-error questions. The other 305 questions stay in the 10.02-point average.
A narrower April test on 5 places with 10,000 respondents scored 7.47 points on 404 non-electoral questions.
A strict held-out set sourced after training froze scored 9.97 points on non-electoral items (14 scored of 22 drawn). That set is separate from the 50-state file. The held-out items were sourced on April 18, 2026, after training froze. State cuts from the older April bank, including Texas and California, stay on methodology. They are not the 50-state headline.
Do not say 7.47 points across 460 questions. 7.47 points is the 404-question non-electoral slice. MAE is an average miss in percentage points. Do not write it with a plus-minus sign. The headline file has 1,521 scored rows. The older April CSV has 443 scored rows and cannot reproduce the 10.02-point full-set miss.
How to read MAE
Mean absolute error is the average miss, in percentage points, between the panel's option share and the published poll share. A 10.02-point average on the full 50-state file means that across 1,521 scored questions, the average absolute gap was 10.02 points. It does not mean plus or minus 10.02 on the study in front of you. 7.47 points is the average absolute miss on the narrower April test, and only on that test.
A low miss on a category this study did not ask does not cover the question you did ask. A high miss on a category you did ask is a reason to hire humans or to rewrite the question. The file lives on methodology.
Synthetic panel vs a human survey
| Synthetic panel | Human probability sample | |
|---|---|---|
| Who answers | Simulated respondents with census margins | Recruited people |
| What you get | Shares, crosstabs, quotes, PDF | A sample you can defend |
| Speed / cost | About 15 minutes once queued. Pricing is scoped on a call. | Days to weeks of recruitment and fieldwork. |
| What you can file | An AI-labeled directional read | A poll, if the design holds |
When to use it, and when to hire humans
Use it for message tests, concept screens, instrument lint (leading or double-barreled items), and a first read across markets before you buy respondents. A team can screen ten framings before it books a facility group.
Do not use it as a legal sample, a regulatory filing, or a published journalism poll. Do not use it as a demand forecast or a willingness-to-pay number you will take to a board as fact. Rare clinical or highly specialized populations stay directional unless you have a scored analog. Every Lewsearch report states the answers are simulated.
What you get back
Option shares for each question. Crosstabs by the cuts the study supports (age, region, party, and others depending on the market). Respondent quotes. A client-ready PDF that says the panel is simulated. Client teams get larger samples, fuller crosstabs, and analyst notes. Concierge is the team designing the study and writing the report. Pricing is scoped to your audiences and study volume on a short call. Book a demo. Your first qualified U.S. consumer Blind Challenge is free.
Coverage
Published pool: over 650,000 simulated respondents. 650,000 census-grounded respondents · all 50 states and D.C. · 16 occupation panels.. All 50 states and DC are on coverage. A study samples a market. It does not interview the whole pool. A state on the map has a panel. A published miss for that state is a separate table. Both are posted so a buyer can see the difference.
How this differs from interview tools and agreement-rate vendors
Adjacent products publish different units. Some interview tools report thematic parity. Some enterprise simulators report a rank correlation. Some panel vendors report an agreement range against historical panels. Those units are not MAE. They do not convert into Lewsearch's point miss, and they are not a reason to invent a Lewsearch percent that is not in the file.
The buyer test is the same on every compare page: published error, named benchmarks, and n. Units stay on the vendor's own site. See synthetic respondents accuracy for the unit map.
Where the numbers live
- Methodology: The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. Every study is calibrated against published survey data before it reaches you. The per-question file is public.
- 50-state per-question CSV: 1,521 scored questions. This is the file behind the 7.07 and 10.02 figures. Page: /benchmarks/states.
- Older April CSV: 443 scored rows from the 460-question bank. It cannot reproduce the 10.02-point full-set miss.
- Coverage: live panels and the 50-state list.
- Pricing: Pricing is scoped to your audiences and study volume on a short call. Book a demo. Your first qualified U.S. consumer Blind Challenge is free.
- llms.txt: citation map for answer engines.
FAQ
- What is synthetic research?
- Synthetic research asks a census-grounded panel of simulated respondents, with persistent memory, the same questions a human survey would ask, then reports option shares, crosstabs, and quotes. The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. MAE is an average miss in percentage points against published poll toplines. A narrower April test on 5 places with 10,000 respondents scored 7.47 points on 404 non-electoral questions. It does not stand as a legal sample.
- Are the respondents real people?
- No. Demographics come from census-style margins. Answers are generated by a model. Every Lewsearch report says so on the cover. Accuracy is checked by scoring the aggregate against real polls.
- How accurate is it?
- The calibrated 50-state benchmark: 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question. MAE is an average miss in percentage points against published poll toplines. A narrower April test on 5 places with 10,000 respondents scored 7.47 points on 404 non-electoral questions. A strict held-out set sourced after training froze scored 9.97 points on non-electoral items (14 scored of 22 drawn). State cuts from the older April bank are on /methodology. They are not the 50-state headline.
- How is this different from ChatGPT?
- A frontier model asked for a national split can be close on well-polled questions, and it can be asked for a distribution. A Lewsearch panel returns respondent-level records you can crosstab, the same simulated respondents across waves, and a published error file.
- How is this different from a human poll?
- A human probability sample can be filed when the design holds. A synthetic panel is a screen. Use Lewsearch to screen messages and questions. Hire humans when the result has to stand up in court, in a regulator's office, or in a newspaper.
Check the work
Related answers
Canonical: https://lewsearch.com/answers/what-is-synthetic-research