← All posts
·9 min read

We Tested Kalshi's Markets for Mispricing. Mostly, There Isn't Any.

Every service selling prediction-market picks claims to find edges everywhere. That should be suspicious. Markets with real money and real volume are usually priced about right, and the honest version of this business is admitting where you can't beat them.

So here are our results across four Kalshi markets — including the ones that didn't go our way, and the model we built, measured, and then refused to launch.

Bitcoin and Ethereum: no edge

Kalshi's crypto ladders let you bet on where BTC or ETH closes at a given hour. They carry enormous open interest — over 400,000 contracts on a single Bitcoin event — and that popularity is exactly the problem.

We priced every rung of the live ladders against our own model. Across 18 actively quoted contracts, our fair value landed a median of 1.1 probability points from the market's. The single best opportunity we could find was worth about 0.7% expected value — against a 4% threshold we require before publishing anything.

That near-agreement is good news about our model and bad news about the opportunity. These are the most heavily arbitraged contracts on the exchange. Anyone telling you they consistently beat Kalshi's Bitcoin book is describing something we could not reproduce.

Fed decisions and inflation: nothing free

Some mispricings need no forecast at all. A Fed meeting has exactly one outcome, so the prices of all possible outcomes must add to $1. An inflation ladder is nested — "CPI above 3.0%" must always be at least as likely as "CPI above 3.5%". When a market violates its own arithmetic, the profit is locked the moment you enter.

We checked 1,002 contracts across 15 ladders, at executable prices and after fees. We found zero violations. Fed outcome sets summed to between $1.01 and $1.03 — that's the bid-ask spread, not free money.

We also stopped publishing inflation forecasts entirely. We had been extrapolating the last CPI print forward, which sounds reasonable and is worthless: it contains no information the market doesn't already have. A guess dressed as a signal is worse than nothing.

Baseball: we built the model, then didn't launch it

Kalshi's MLB game markets are liquid and tight — 13.6 million contracts of open interest at roughly a 1.9¢ spread. We have a genuine baseball model built on pitcher fundamentals, so this looked like the obvious place to point it.

We backtested it on 881 settled Kalshi games over 38 days. Two details mattered for making the answer trustworthy: every feature was pulled as of the day before each game, so the model never saw its own future; and the train/test split was chronological, never random. Scores are Brier — lower is better.

Brierlog-loss
Kalshi market0.24380.6807
Our model (refit)0.24480.6827
Our model (original)0.24610.6853
Home-field baseline0.25200.6972

Our model beats a naive home-field baseline comfortably, so it knows something real. It also loses to Kalshi's closing price. It captures roughly 88% of the market's skill and never exceeds it.

We then simulated actual bets on the held-out games, paying the real spread and Kalshi's fees. At a 2% edge threshold the picks returned +5.3%. At 4%, −1.2%. At 6%, +11.5%. At 8%, −2.9%.

That pattern is the whole lesson. A real edge gets better as you demand more from it. Ours flips sign at random, on samples as small as thirty bets. If we had shown you only the 6% row, we could have launched a product on it — and it would have been noise.

Weather: the exception, and why

The one place we do find repeatable mispricing is daily temperature. The reason isn't that our math is better. It's that pricing these markets correctly requires unglamorous work most traders skip.

Each contract settles on the official reading from one specific weather station — Chicago's market settles at Midway, not O'Hare; Houston's at Hobby, not Bush. Get the station wrong and you have a systematic error that feels like an edge. The contracts settle on whole degrees, so the real boundary of "above 96°" sits at 96.5. Kalshi's fee is quadratic and peaks in the middle of the price range, which is precisely where most apparent edges live.

None of that is clever. It's just work someone has to do, and it doesn't exist in a Bitcoin order book, where everyone already has the same price feed you do.

We're still building the public record on weather, and we'll publish the calibration curve — how often outcomes land at the probabilities we quoted — when there's enough settled history to make it meaningful. Until then we'd rather show you the tests that failed than make a claim we can't back yet.

Then we tested our one good model, and it failed too

Kalshi isn't the only venue running daily temperature markets. Polymarket lists them as well, in the same bucket format, across roughly 143 open events. If our weather model is genuinely good, a second exchange with a mostly-politics audience seemed like the obvious place to find it paying off.

First we checked which weather station each market actually settles on, because that's the detail that quietly ruins everything. Three of ten differed from Kalshi's: Polymarket settles Chicago at O'Hare where Kalshi uses Midway, Denver at Buckley Space Force Base rather than Denver International, and Dallas at Love Field rather than DFW. Two markets with identical-sounding titles can resolve to different numbers on the same day.

On the seven cities where the stations do match, we compared our probability to theirs on every quoted contract:

contractsmedian gap90th pct
Polymarket530.0730.176
Kalshi610.0940.202

We agree more with Polymarket than with Kalshi. There was no edge to collect — the opposite of what we went looking for.

The more useful result was the one we weren't looking for. On the single most likely outcome in each city, our probability was lower than the market's in six cases out of six — averaging 0.27 where both exchanges averaged 0.40. Two independent markets, with different traders and different settlement sources, both saying our next-day forecasts are roughly 1.5× too uncertain.

When two independent markets disagree with you in the same direction, the parsimonious explanation is that you are wrong. An over-wide forecast doesn't just lose precision — it manufactures apparent edges, because every outcome the market considers unlikely looks underpriced to you. It is entirely possible that some share of the next-day edges we've been publishing are artifacts of that, and we're treating them as unproven until the settled record says otherwise.

We could have closed the gap in an afternoon by tuning our uncertainty down until the numbers matched. We didn't, and won't. A model adjusted to agree with the market is just a slower way of reading the market. The fix has to come from measuring our own forecast errors against what actually happened, which takes weeks of recorded history — so that's what we're doing.

What to take from this

If you trade prediction markets, the useful conclusion isn't "Betstinct is honest." It's that efficiency tracks attention. The markets everyone watches — crypto, rates, major sports — are priced by people with the same data you have and better execution. The edges live where pricing requires information somebody has to go out and collect.

Ask anyone selling you picks which markets they tested and failed to beat. If the answer is none, they either haven't checked or aren't telling you.

See the markets that did pass

We email the Kalshi contracts our model thinks are mispriced — free right now, and every pick is graded publicly, win or lose.

No spam. Unsubscribe anytime. 21+. Please gamble responsibly.

See today's weather edges → · Our full track record →