Evidence
Public study records
These rows are compiled from pages and files Lewsearch already publishes. Types are not pooled into one accuracy score. A date missing here is unknown. No completed Blind Challenge records are listed.
JSON · Vendor evidence dataset · Methodology
| Record | Type | Metric | Result |
|---|---|---|---|
| As-shipped 50-state public-poll MAE lewsearch-as-shipped-50state | Retrospective benchmark | MAE (percentage points) | About 10 pp overall (10.02 exact on methodology). 10.18 pp on 1,169 non-electoral questions. |
| April 2026 five-place calibrated benchmark lewsearch-april-calibrated | Retrospective benchmark | MAE (percentage points) | 7.47 pp on 404 non-electoral questions of a 460-question bank. |
| Held-out set sourced 2026-04-18 lewsearch-heldout-2026-04-18 | Held-out test | MAE (percentage points) | 9.97 pp non-electoral; 10.68 pp overall. |
| Public per-question CSV (April bank subset) lewsearch-public-csv-443 | Retrospective benchmark | MAE (percentage points) | About 7.44 pp on 389 non-electoral rows. Do not replace with the 50-state book. |
| Issue 01: NRF back-to-school lewsearch-report-issue1-nrf | Prospective forecast | mean K-12 spend miss (percent of the NRF figure, as published on the issue) | Mean K-12 spend within 1%. |
| Issue 04: 225 pre-registered predictions lewsearch-report-issue4-lock | Prospective forecast | mixed (see issue) | See the issue. This record does not invent a pooled accuracy score. |
| Issue 05: September Consumer Sentiment, graded lewsearch-report-issue5-ics | Baseline-anchored forecast | ICS miss (index points) | Guess #2 nowcast 52.1 vs 47.8 (plus 4.3). Guess #1 was 55.4 (plus 7.6). |
| Public Pew forecasts lewsearch-predictions-pew | Prospective forecast | varies by item (see page) | See /predictions. Do not fold these into the 10 pp MAE. |
| Polymarket scoreboard lewsearch-proof-polymarket | Ungraded editorial pulse | market comparison (see page) | Credibility check only. Historical disclosure, not a product claim. |