Evidence

Public study records

These rows are compiled from pages and files Lewsearch already publishes. Types are not pooled into one accuracy score. A date missing here is unknown. No completed Blind Challenge records are listed.

JSON · Vendor evidence dataset · Methodology

RecordTypeMetricResult
As-shipped 50-state public-poll MAE

lewsearch-as-shipped-50state

Retrospective benchmarkMAE (percentage points)About 10 pp overall (10.02 exact on methodology). 10.18 pp on 1,169 non-electoral questions.
April 2026 five-place calibrated benchmark

lewsearch-april-calibrated

Retrospective benchmarkMAE (percentage points)7.47 pp on 404 non-electoral questions of a 460-question bank.
Held-out set sourced 2026-04-18

lewsearch-heldout-2026-04-18

Held-out testMAE (percentage points)9.97 pp non-electoral; 10.68 pp overall.
Public per-question CSV (April bank subset)

lewsearch-public-csv-443

Retrospective benchmarkMAE (percentage points)About 7.44 pp on 389 non-electoral rows. Do not replace with the 50-state book.
Issue 01: NRF back-to-school

lewsearch-report-issue1-nrf

Prospective forecastmean K-12 spend miss (percent of the NRF figure, as published on the issue)Mean K-12 spend within 1%.
Issue 04: 225 pre-registered predictions

lewsearch-report-issue4-lock

Prospective forecastmixed (see issue)See the issue. This record does not invent a pooled accuracy score.
Issue 05: September Consumer Sentiment, graded

lewsearch-report-issue5-ics

Baseline-anchored forecastICS miss (index points)Guess #2 nowcast 52.1 vs 47.8 (plus 4.3). Guess #1 was 55.4 (plus 7.6).
Public Pew forecasts

lewsearch-predictions-pew

Prospective forecastvaries by item (see page)See /predictions. Do not fold these into the 10 pp MAE.
Polymarket scoreboard

lewsearch-proof-polymarket

Ungraded editorial pulsemarket comparison (see page)Credibility check only. Historical disclosure, not a product claim.