October 2026 · Research notes

The weather market learned to read the forecast

Kalshi lets you trade tomorrow's high temperature in seven US cities. The National Weather Service publishes its forecast for free, hours before most of the trading happens. So does the market actually use it? We spent a few weeks finding out, and the answer changed while we were measuring it.

Every day, for every city, Kalshi lists a ladder of 2°F buckets and exactly one of them settles YES. That makes each ladder a small classification problem with a price on every class, which is about the cleanest setup there is for asking whether a model's probabilities beat the crowd's. We pulled every settled ladder back to 2021, about 8,200 city-days and 23,000 ladder snapshots, every trade since mid 2023, 17 million prints, and twelve years of archived NWS forecasts. All of it from free public APIs.

The model is in the spirit of Jev, TypeSafe's new decision model: you give it a state and a typed question, and it gives back calibrated probabilities instead of text. Ours is called isotherm. It puts one probability distribution over the outcome and reads every answer off that, so "which bucket wins" and "will it be above 85" can never contradict each other. Under the hood it is small. A learned log pool of the market and the forecasts, plus an MLP correction, an ensemble of five, trained on log loss with recent days weighted more heavily. It starts out identical to the market and only moves away from the crowd where the outcomes pay for it.

The evaluation is where most of the work went. Everything is walk-forward. Every forecast is the version that was actually public at the moment of the read, so there is no look-ahead. Every model is refit each quarter on the past only, and the last three months were sealed away until the very end. We score with proper scoring rules (log loss, RPS, Brier), always against the market's own implied probabilities at the same timestamp, never against zero. And there is a control: the same network trained on labels sampled from the market itself. It can only beat the market if the pipeline leaks the answer. It scored within 0.002 nats of the market.

fig. 1: log loss gain over the market at the 8am read, by half year. above zero means better than the crowd
halfsimple poolisotherm
2023 H2+0.182+0.227
2024 H1+0.113+0.179
2024 H2+0.089+0.098
2025 H1-0.012+0.022
2025 H2-0.004+0.016
2026 H1-0.021+0.006

First result: there used to be a lot of edge here. In late 2023, blending the market with the free forecast beat the market by 0.18 nats per ladder, which is a huge number for a liquid market. By 2025 the same blend was at zero, then slightly negative. Somebody started reading the forecast. isotherm holds up better and stays positive in every half year, but it decays too, down to about 0.01.

Second result, and our favourite. The NWS puts out two main public forecast products. NBM is the blend behind the forecast on weather.gov, the one everybody looks at. GFS MOS is older and plainer and mostly lives in text bulletins. Over the last twelve months, adding NBM to the market price makes it worse. Adding GFS MOS still makes it better, at every read time, with confidence intervals clear of zero. The crowd priced in the forecast it could see, and not the one it couldn't.

fig. 2: last twelve months, log loss gain from pooling the market with each forecast. whiskers are 95% intervals, resampled by date
readNBMGFS MOS
4pm day before-0.004+0.020
8am-0.021+0.012
noon-0.000+0.009

Better probabilities are not the same thing as money. Every trade crosses a spread and pays Kalshi's fee, and that is where most of the edge went. With the simple blends as a taker, the strategy made about one cent per contract, and one extra cent of slippage turned it negative. As a maker the spread came back to us, but our resting orders mostly got filled right when the price was moving against us. Textbook adverse selection. The model's information cut that loss by two thirds compared to quoting blind. It did not cut it to zero.

The version that worked trades at 4pm the day before, only when the GFS-adjusted price disagrees with the market by at least four cents, sized with fractional Kelly over the whole ladder at once. Out of sample, with each quarter's settings chosen only on earlier quarters, it made $22k on a $10k paper bankroll over three years.

fig. 3: cumulative paper PnL, nested walk-forward, then the sealed test from july 2026. weekly, $10k bankroll
Sharpe, annualised (95% bootstrap CI)2.05 [0.81, 3.27]
Newey-West t on daily PnL3.17
max drawdown$5,195
hit rate71%
deflated Sharpe ratio (28 configs tried)0.85
probability of backtest overfitting (CSCV)0.26
its main config, with +2¢ slippage+$4,516

It is positive in six of seven cities, and the configuration it settles on most often survives two cents of slippage and ten times the size. Most of the profit came in the last year, so this isn't the 2023 edge that already decayed. Where the money comes from is worth a look, because it is the clearest picture of the whole problem: the model's picks beat the mid by a lot, and costs eat about half of it.

fig. 4: where the money goes, day-before strategy, out of sample
componentdollars
alpha vs mid+43,561
spread paid-12,434
fees-9,005
net+22,122

We tried 28 strategy configurations to get there, though, so the raw Sharpe is not the number to trust. The deflated Sharpe ratio, which corrects for how many things you tried and for fat tails, came out at 0.85. The probability of backtest overfitting came out at 0.26. Our bars were 0.95 and 0.20. Close, but not over.

So we ran the one test we could only run once. We wrote down the frozen strategy and the pass rule, pushed them to the repo, and only then scored July to October 2026, which nothing had touched. It passed: +$1,851 over 95 days, 80% of trades winning, a t of 1.77 against a bar of 1.645, and still positive with an extra cent of slippage. Trading noise with the same turnover lost $2,703 over the same days. Then we looked at it by month.

fig. 5: sealed test, paper PnL by month. october is three days
monthdollars
July+1,029
August+517
September+246
October (3 days)+60

That is the honest headline. The edge is real, it survived a test it could not have been fit to, and it is shrinking about as fast as the 2023 edge did. Three months can't tell us whether that's renewed decay, the August switch in Kalshi's settlement source, or noise.

Next we gave the model more to look at on the day itself: five-minute station observations, NBM's forecast for the rest of the afternoon, and a new 2pm read. It helped exactly where you'd expect. At 2pm the model beats the market by 0.017 nats over the last year, its biggest recent gain at any hour of the day. It did not help where it counts. By mid afternoon the order books are thin, and there were only 322 trades worth taking in three years. Resting orders still got picked off too. Pulling our quotes every time a new GFS run lands sounded clever, and the backtest almost never chose it.

Then came the best looking result of the whole project. Kalshi started listing daily low temperatures in December 2025. Same pipeline, same family of strategies, and the backtest cleared every bar we had set: a deflated Sharpe of 0.98, a probability of backtest overfitting of 0.06, and $13k in seven months. We sealed off September and October before looking, wrote down the rule, and scored it once. It made $36.

fig. 6: daily lows, cumulative paper PnL. walk-forward backtest, then the sealed weeks from september 2026

The model's probabilities were still better than the market's on those weeks, by 0.019 nats with an interval clear of zero. The money was gone. The tell was the placebo. In the backtest even random trades made a little money, because a brand new market was badly calibrated. In the sealed weeks random trades lost. The young market had grown up, and our edge had mostly been its growing pains.

We also ran the identical pipeline on MLB game winners, the other end of the liquidity spectrum: 4,484 games, Elo plus starting pitchers, one cent spreads and around a million contracts a game. The best blend landed within 0.001 nats of the market. Nothing there. Efficient markets exist, and that is one of them.

We also tried the last piece of the Jev recipe, pretraining on synthetic data. We built 92,000 fake ladders from 55 weather stations Kalshi doesn't list, with a simulated market tuned until its spreads, overround and accuracy looked like the real one. Pretraining doubled the model's edge when we only gave it a tenth of the real history, and made it slightly worse when we gave it all of it. Handy for a market that just launched. Not for this one.

What we'd tell anyone trying this. Score against the market, not against zero. Use proper scoring rules, and train on them too. Keep a control that can only win if your pipeline leaks. Count your trials and deflate your Sharpe. Be suspicious of young markets, because their early mistakes look a lot like skill. And lock some data away before you start, because you only get to be surprised once.

The strategy is now running live in shadow mode. Every day at 4pm local, a scheduled job scores the next day's ladders and commits its paper trades before the outcome exists, so the timestamps are the proof. It carries a decay alarm: a rolling 30-day score against the market that flips to STOP if two windows in a row sit clearly below zero. The model also sits behind a small API with a model card, so you can ask it for the chance Chicago tops 85 tomorrow and get a calibrated number back. What's left is mostly waiting. If the alarm stays quiet for a month, the next step is real execution. If it fires, we'll write that up too.

Explore every result interactively on the isotherm dashboard. Every number here is reproducible from physicalreasoning/isotherm, with the code, raw results and the full findings record, including the bugs we caught and the gate we had to amend. The live paper ledger is on the shadow-ledger branch.

References. G. W. Brier, Verification of forecasts expressed in terms of probability, Monthly Weather Review, 1950. T. Gneiting, A. E. Raftery, A. H. Westveld, T. Goldman, Calibrated probabilistic forecasting using ensemble model output statistics, Monthly Weather Review, 2005. T. Gneiting, A. E. Raftery, Strictly proper scoring rules, prediction, and estimation, JASA, 2007. H. R. Glahn, D. A. Lowry, The use of model output statistics in objective weather forecasting, Journal of Applied Meteorology, 1972. J. P. Craven, D. E. Rudack, P. E. Shafer, National Blend of Models, Journal of Operational Meteorology, 2020. J. L. Kelly, A new interpretation of information rate, Bell System Technical Journal, 1956. L. R. Glosten, P. R. Milgrom, Bid, ask and transaction prices in a specialist market with heterogeneously informed traders, Journal of Financial Economics, 1985. D. H. Bailey, M. López de Prado, The deflated Sharpe ratio, Journal of Portfolio Management, 2014. D. H. Bailey, J. M. Borwein, M. López de Prado, Q. J. Zhu, The probability of backtest overfitting, Journal of Computational Finance, 2017. The typed question design follows TypeSafe AI's Jev; isotherm is an independent open implementation, not affiliated with TypeSafe. Full list in the repository.