U.S. Midterms · November 3, 2026
Lab notes · Note 01Sat 03 Oct 2026

Six models, four tests, one shared blind spot

Before we forecast a single race of the 2026 midterms, we wanted to know one thing: can a language model stand in for a group of voters and react to the news the way that group really would? So we asked six of them to try, and scored them against what actually happened.

The test bench ran on 27 and 28 September. It cost $1.52 in total. What it taught us shaped every part of the forecast you will see from 12 October.

We asked every model the same four questions, in the same words, with the same voter personas.

Does a football score move a Senate race? It shouldn't. We showed each model 20 stories that shouldn't change anyone's vote, from sports results to a new seasonal croissant, and checked whether it stayed still.

If you swap the parties in a story, does the reaction swap with them? We told 16 stories twice, once with a Democrat at the centre and once with a Republican, and looked for any lean left over.

Does it move the way people actually moved? We used 19 real events since 2012 whose effect on opinion was measured at the time, and asked whether each model saw the right direction, and the right size.

Does it know how groups of Americans voted in 2024? We compared its answers with the real vote, group by group.

In detail

The six models were Jev 1.13, Kev 4B before any training of ours, GLM-5.3 Flash, MiMo-V2.6 Flash, DeepSeek V4.1 Flash and GPT-6 Luna, all reached through OpenRouter. Every question was asked twice, with the answer scale in both orders, for reasons we explain in Lab notes 03. The voter personas came from real respondents to the 2024 Cooperative Election Study, chosen to span party, race and education.

Four questions we asked every model: does a football score move a Senate race; if the parties swap, does the reaction swap; does it move the way people really moved; does it know how groups voted in 2024.

No model was good at everything. Each was good at one thing, or at nothing.

Jev, a small decision model, was the best at ignoring news that doesn't matter: it said "no change" 99% of the time. GLM-5.3 Flash was the best at reading which way a group moves: it got the direction right on 13 of 13 real events that shifted opinion.

None of them got the size of a reaction right. GLM's sizes barely tracked reality, with a correlation of 0.26 between what it predicted and what was measured.

And on the most basic question, how groups voted in 2024, plain statistics beat every model we tried.

In detail

On that last test we measured the gap between each model's answers and the real vote across demographic groups. A simple statistical model built on the same survey answers was about twice as close as the best language model. An untuned Kev, the model we later trained ourselves, reacted to about 1 irrelevant story in 4, which is why it got no job in this round.

Some got the direction right, none got the size: Jev ignores irrelevant news 99% of the time; GLM gets the direction right on 13 of 13 events; no model gets the size right; plain statistics know how groups voted better than every model.

The most useful finding was a shared one. The models didn't just make mistakes; they made the same mistakes.

All of them were too calm about real shocks: the early weeks of COVID, January 6, the fall of Kabul. All of them were too excited about media spectacles, like the Access Hollywood tape and the presidential debates. A story that dominated television moved them more than a crisis that changed how people lived.

Because the errors line up, averaging several models doesn't cancel them out. So in our forecast, the size of every reaction comes from shifts that were actually measured in past events, never from a model's guess.

In detail

Across the 19 events, the models' errors correlate between +0.88 and +0.97 with each other. Averaging any two of them never beat GLM alone. We use the models for what they do well, the direction of a reaction, and real data for everything else.

They all got the same things wrong: too calm about real shocks like COVID and January 6, too excited about the debates. So the size of every reaction comes from how people really moved.

So this is how the forecast is put together.

Statistics set each race's starting line: past results, the economy and the poll average. How each group of voters usually splits between the parties comes from Kev, a model we trained on real survey answers, which came out slightly closer to the real 2024 vote than the plain statistics on three states it had never seen.

Then, every day, the voter personas react to each day's news, and the model's job is narrow: which way does this story push each group? How far it pushes comes from measured history. Every week, the latest poll average decides how much of that shift to keep.

Tomorrow in Lab notes: what happened when we told the same story with a different newspaper's name on it.

Real history sets the size of every reaction. Statistics set where each race starts, the voter personas react to each day's news, and how far each group moves comes from past events.