The question
Prediction markets are new, volatile, and growing fast. The question I wanted to answer was simple: could a solo trader make money on price mismatches between the two biggest prediction-market exchanges?
Fees, latency, liquidity, resolution conditions, and the risk of only one leg filling all get in the way, so I weighed the question against every one of them. The result was no: none of the pairs I tested stayed profitable once real fees and available order-book depth were accounted for — even with the low-latency setup I tested, which ran capture from servers a few miles from the exchanges’ own to trim round-trip time off every quote.
What I built
A trading system that became a research platform
It started as a real trading bot: it could authenticate against Kalshi, price a trade, and place live orders. Once the results came back negative, I retired the order path and kept everything underneath it as a research platform that reads markets and never trades.
I split the system into three independent services connected through a small message bus. They talk through versioned messages instead of reaching into each other’s state, so one service can restart or fall behind without corrupting the others, and a malformed message is quarantined instead of replayed. That boundary is what kept a one-person project predictable under real market timing rather than turning into one fragile script.
Architecture: outbox → router → inbox.
- Concurrent Kalshi and Polymarket order-book capture over REST and WebSocket feeds on asyncio, sustaining roughly 2.9M rows per day at 52 ms median cross-venue leg skew.
- RSA-PSS signed Kalshi authentication and a kill-switch-gated order path, since retired to read-only, paper-only operation.
- Contract-equivalence review, quote normalization, freshness and skew gates, and depth-aware price calculations, with pgvector-backed similarity supporting the equivalence screen.
- Fee, VWAP, slippage, collateral-carry, and failed-leg modeling, so a displayed spread is never mistaken for money I could actually take.
- State and the pair registry persisted in PostgreSQL through SQLAlchemy 2.0, with append-only capture and a Parquet and DuckDB layer for analytical queries over order-book history.
- Testing and deployment: mypy --strict across 26K lines, 879 tests (532 unit, 347 PostgreSQL integration) run as lint, then types, then tests, and a multi-stage non-root Docker image with healthchecks.
The speed limit
Every price in this project had to travel. Polymarket’s order books are served from London, Kalshi’s from the Chicago area, and the system comparing them ran in New York. Nothing travels faster than light, and light in optical fibre is about a third slower than in a vacuum: it covers roughly 127 miles every millisecond. That puts a hard floor under every quote, however fast the code is.
The speed limitLight in fibre ≈ 127 mi per ms
Route by route, the floor looks like this:
- London → New York: 3,460 miles, at least 27 ms one way (18.6 ms even in a vacuum). Acting on a price there means a round trip, at least 55 ms.
- Chicago area → New York: 710 miles, at least 6 ms one way and 11 ms round trip.
- London ↔ Chicago: 3,950 miles between the exchanges themselves, at least 31 ms.
For an arbitrage, that floor does two kinds of damage. The two prices never arrive together: even at the limit, a Polymarket quote reaches New York more than 20 ms after a Kalshi quote from the same instant. And acting on a gap takes a round trip, so by the time an order reaches London, the price it was based on is at least 55 ms old. Anyone with a server next to the exchange has had that whole window to take the gap first.
Moving the computer doesn’t remove the limit; it only moves it. You can put a server beside each exchange, as the low-latency setup I tested did, but the decision still needs both prices in one place, and the exchanges are 3,950 miles apart. Even a computer floating in the middle of the Atlantic would see one of the two prices at least 15 ms late. Real networks are slower still: cables don’t follow the great circle, and every hop on the way adds time.
That’s why the platform never treats two quotes as simultaneous. It measures the gap between the two legs it compares, a median 52.3 ms in the live soaks, and its freshness and skew gates reject any pair whose prices are too far apart in time to trust.
Multiple 24-hour-plus live soaks averaged 2.9 million order-book rows per day across 97 approved cross-market pairs, with both legs of every trade staying within 52.3 ms of each other at the median. Even at that level of precision, no approved pair cleared a one-cent profit once real fees and order-book depth were factored in.
The long story
This started as a do-first, think-later project. The first step was scraping every public resource I could find of people attempting something similar and pulling out anything useful, which turned out to be not much.
From there, the work turned into research: learning exactly how Kalshi and Polymarket operate, down to their fee structures, their APIs, rate limits, and settlement mechanics.
Only after that did the real problem show up: finding contracts on the two venues that were actually equivalent. Plenty of pairs looked nearly identical, but small differences in their resolution rules could have left me holding two losing positions instead of a hedged trade. Getting this right mattered more than anything else in the project, so pair selection ran through both an AI-assisted screen and a manual review of the actual contract language. Out of hundreds of candidates, 97 pairs survived that review and were cleared for testing.
verified_equivalent), strategy eligibility (excluded), and scan status (approved_for_paper). A pair can be genuinely equivalent and still be excluded from every reported metric — that separation is what keeps a similar-looking contract from quietly entering the results.Testing almost lied to me. An early run looked genuinely profitable, until I traced it back to a bid/ask normalization defect that was quietly inflating the spread. I fixed the live data path, built an idempotent backfill to correct the rows it had already touched, and hardened regression coverage to 879 tests (532 unit and 347 PostgreSQL integration) so the same bug couldn’t come back unnoticed. Once the fix was in, the attractive result disappeared, which was the correct outcome. The project was never supposed to find a number I liked; it was supposed to find the true one.
That bug is a fair summary of the whole project. The exciting result was wrong, and the boring, correct one was no. I’d rather build something that tells me the truth than something that tells me what I want to hear. The bot didn’t survive contact with real fees, latency, and fine print, but the platform underneath it did, and that’s the part worth carrying forward.
What’s public, and what isn’t
Three things on this page point at the same project, and they are not the same thing:
- The original research system. Private. The full 97 reviewed pairs and the live-order path that could place real trades. Every number on this page comes from here.
- The open-source release. Public, MIT, paper-only. A reduced build that ships 23 reviewed pairs, all of them excluded from strategy metrics, so a scan records real order-book data and deliberately produces no strategy result. Every screenshot on this page comes from here.
- The interactive diagram. A read-only map of the open-source code, pinned to a commit, where each box links to the file it describes.
Where I’d take it next
The arbitrage result was a no-go, but most of the infrastructure is reusable: the capture pipeline, the fee and depth models, and the contract-equivalence review don’t care which strategy sits on top of them. I could reuse the same market-data pipeline for ideas that don’t need two simultaneous fills — mean reversion on a single contract, correlated markets that tend to move together, or tracking where attention and volume shift inside one market. Those all carry real directional risk instead of a locked-in spread, the opposite trade-off from this project, which is exactly what makes them interesting to test next.
Look at how it works
You don’t have to take my word for any of this. The research platform is open source, and the diagram behind the button below maps it as it runs: every box links to the code it describes, pinned to a specific commit, so each citation opens the exact file and line on GitHub.
It follows one path end to end: two public order books arrive, get paired with an explicit capture-skew measurement, get priced through the real per-venue fee engine, and become a depth-adjusted edge row. The parts that make the negative result trustworthy are on it too — the pair registry whose equivalence screen can only flag and never clear, and the mode gates and kill switch that keep the whole thing read-only.
See the repo as an interactive diagram →Diagram generated with archify by tt-a1i (MIT).