The proof, with receipts
The product grade is the calibrated 50-state benchmark. 7.07-point average miss on the best 80% of 1,521 scored questions (1,216 retained). 10.02 points across every scored question (200 simulated respondents per question, calibrated panels). The market note below is a separate check.
68% YES. We said 1%. It resolved NO.
Where no human survey exists, the only external check on a synthetic panel is a market that resolves. So we ran the one test we could. On the single topic where real-money markets exist on opinion, Trump approval, we asked whether our panel could read outcomes before contracts resolved. It's real. It's also not alpha.
98
resolved contracts, pre-registered
70.4%
our direction accuracy vs. 72.4% market
63 / 98
cells where we were the closer estimate
0
settlement leak · odds-only, pre-resolution
The cell that started it
“Will Trump's approval hit 40% in 2025?” At decision time, Polymarket priced 68% YES. Our panel said ~1%. It resolved NO. That wasn't a fluke. The same pattern held across overpriced longshots. The market kept overweighting tail outcomes; the panel kept reading the underlying approval level and calling them out.
Where the market overpriced, the panel held the line
Six clean, pre-resolution examples. Market price vs. our forecast at the decision date. All resolved NO.
The scoreboard
We froze every resolved Trump-approval contract we could map (n=98), ran a nationally weighted panel of 2,000 simulated respondents on each, and compared our forecast to the Polymarket mid-price at a pre-registered decision date: odds, not the eventual outcome. Nothing was scored against settlement.
We were closer on 63 of 98 cells, and our aggregate Brier is worse (0.26 vs 0.16). The classic signature of winning many small cells and eating a few confident misses. Not bankable.
Can we forecast before resolution?
Mostly yes. ~70% direction, strong on brackets, and several large, clean longshot reads where the market was badly mispriced.
Do we beat Polymarket overall?
No. They edge us on direction. We win brackets, lose thresholds. Call it a near-tie that leans the market's way.
Could you trade on us?
No evidence of that. When we disagree with the market we're right about half the time. That's a coin flip, not alpha, and we ship no trading feature.
What the test actually proves
The test does not mean “use Lewsearch instead of Polymarket.” It does not mean “fade the market when we disagree.” What it proves is the thing that matters for our customers: the panel engine isn't hallucinating on opinion-shaped questions. It tracks real-world political outcomes before they resolve, roughly as well as a market full of people betting real money.
Your actual questions don't have markets. “Will this ad move suburban women in Ohio?” “How does Message A test against B with our segment?” The market test is the closest public check for questions like those.
Forecasts are AI-generated estimates, not interviews with human respondents and not probability samples. All 98 cells used pre-resolution market prices only; nothing was scored against settlement. Full write-up and underlying run are linked above.