Research project · Open source

I Built a System to Test Prediction-Market Arbitrage

I built a cross-venue trading system to find out whether price gaps between Kalshi and Polymarket were actually exploitable. After fees, liquidity, latency, and contract differences, they weren’t. The project ended up being less about finding arbitrage and more about building something I could trust to tell me when there wasn’t any.

The research platform is now open source. github.com/frahank/prediction-market-arbitrage

See the repo as an interactive diagram → Read the source on GitHub →

Diagram generated with archify by tt-a1i (MIT).

97 Reviewed pairs
2.9M Order-book rows / day
52.3 ms Median leg skew
879 Tests
3 Independent services
No-go No viable arbitrage found

The question

Prediction markets are new, volatile, and growing fast. The question I wanted to answer was simple: could a solo trader make money on price mismatches between the two biggest prediction-market exchanges?

Fees, latency, liquidity, resolution conditions, and the risk of only one leg filling all get in the way, so I weighed the question against every one of them. The result was no: none of the pairs I tested stayed profitable once real fees and available order-book depth were accounted for — even with the low-latency setup I tested, which ran capture from servers a few miles from the exchanges’ own to trim round-trip time off every quote.

What I built

A trading system that became a research platform

It started as a real trading bot: it could authenticate against Kalshi, price a trade, and place live orders. Once the results came back negative, I retired the order path and kept everything underneath it as a research platform that reads markets and never trades.

I split the system into three independent services connected through a small message bus. They talk through versioned messages instead of reaching into each other’s state, so one service can restart or fall behind without corrupting the others, and a malformed message is quarantined instead of replayed. That boundary is what kept a one-person project predictable under real market timing rather than turning into one fragile script.

Diagram of the platform's data path: Kalshi and Polymarket order books are captured, paired behind a skew gate, and priced through fees and depth. None of the 97 reviewed pairs cleared one cent.
The data path, simplified. Quotes that look like an edge (the moving blue dot) go into the fee and depth model and come out without one; none of the 97 reviewed pairs cleared one cent.

Architecture: outbox → router → inbox.

  • Concurrent Kalshi and Polymarket order-book capture over REST and WebSocket feeds on asyncio, sustaining roughly 2.9M rows per day at 52 ms median cross-venue leg skew.
  • RSA-PSS signed Kalshi authentication and a kill-switch-gated order path, since retired to read-only, paper-only operation.
  • Contract-equivalence review, quote normalization, freshness and skew gates, and depth-aware price calculations, with pgvector-backed similarity supporting the equivalence screen.
  • Fee, VWAP, slippage, collateral-carry, and failed-leg modeling, so a displayed spread is never mistaken for money I could actually take.
  • State and the pair registry persisted in PostgreSQL through SQLAlchemy 2.0, with append-only capture and a Parquet and DuckDB layer for analytical queries over order-book history.
  • Testing and deployment: mypy --strict across 26K lines, 879 tests (532 unit, 347 PostgreSQL integration) run as lint, then types, then tests, and a multi-stage non-root Docker image with healthchecks.
The access and safety tab, showing paper runtime mode, real orders disabled, a clear kill switch, and storage-only credential fields.
The safety controls: runtime mode, all three real-order gates, the global kill switch, and a credential store that is storage-only and wired to no trading path. Placing a real order requires three independent gates, and is denied by default whenever any one of them is missing.

The speed limit

Every price in this project had to travel. Polymarket’s order books are served from London, Kalshi’s from the Chicago area, and the system comparing them ran in New York. Nothing travels faster than light, and light in optical fibre is about a third slower than in a vacuum: it covers roughly 127 miles every millisecond. That puts a hard floor under every quote, however fast the code is.

The speed limitLight in fibre ≈ 127 mi per ms

Where each price starts and how fast it can possibly reach me. Distances are great-circle; the times are the floor for light in fibre, which real cables only add to. Hover or tap a stop to see its route.

Route by route, the floor looks like this:

  • London → New York: 3,460 miles, at least 27 ms one way (18.6 ms even in a vacuum). Acting on a price there means a round trip, at least 55 ms.
  • Chicago area → New York: 710 miles, at least 6 ms one way and 11 ms round trip.
  • London ↔ Chicago: 3,950 miles between the exchanges themselves, at least 31 ms.

For an arbitrage, that floor does two kinds of damage. The two prices never arrive together: even at the limit, a Polymarket quote reaches New York more than 20 ms after a Kalshi quote from the same instant. And acting on a gap takes a round trip, so by the time an order reaches London, the price it was based on is at least 55 ms old. Anyone with a server next to the exchange has had that whole window to take the gap first.

Moving the computer doesn’t remove the limit; it only moves it. You can put a server beside each exchange, as the low-latency setup I tested did, but the decision still needs both prices in one place, and the exchanges are 3,950 miles apart. Even a computer floating in the middle of the Atlantic would see one of the two prices at least 15 ms late. Real networks are slower still: cables don’t follow the great circle, and every hop on the way adds time.

That’s why the platform never treats two quotes as simultaneous. It measures the gap between the two legs it compares, a median 52.3 ms in the live soaks, and its freshness and skew gates reject any pair whose prices are too far apart in time to trust.

The result was a no-go.

Multiple 24-hour-plus live soaks averaged 2.9 million order-book rows per day across 97 approved cross-market pairs, with both legs of every trade staying within 52.3 ms of each other at the median. Even at that level of precision, no approved pair cleared a one-cent profit once real fees and order-book depth were factored in.

The long story

This started as a do-first, think-later project. The first step was scraping every public resource I could find of people attempting something similar and pulling out anything useful, which turned out to be not much.

From there, the work turned into research: learning exactly how Kalshi and Polymarket operate, down to their fee structures, their APIs, rate limits, and settlement mechanics.

Only after that did the real problem show up: finding contracts on the two venues that were actually equivalent. Plenty of pairs looked nearly identical, but small differences in their resolution rules could have left me holding two losing positions instead of a hedged trade. Getting this right mattered more than anything else in the project, so pair selection ran through both an AI-assisted screen and a manual review of the actual contract language. Out of hundreds of candidates, 97 pairs survived that review and were cleared for testing.

The pairs registry tab, listing candidate market pairs with badges for settlement equivalence, strategy eligibility, and scan status.
Pair review. Each pair carries three independent states: settlement equivalence (verified_equivalent), strategy eligibility (excluded), and scan status (approved_for_paper). A pair can be genuinely equivalent and still be excluded from every reported metric — that separation is what keeps a similar-looking contract from quietly entering the results.

Testing almost lied to me. An early run looked genuinely profitable, until I traced it back to a bid/ask normalization defect that was quietly inflating the spread. I fixed the live data path, built an idempotent backfill to correct the rows it had already touched, and hardened regression coverage to 879 tests (532 unit and 347 PostgreSQL integration) so the same bug couldn’t come back unnoticed. Once the fix was in, the attractive result disappeared, which was the correct outcome. The project was never supposed to find a number I liked; it was supposed to find the true one.

That bug is a fair summary of the whole project. The exciting result was wrong, and the boring, correct one was no. I’d rather build something that tells me the truth than something that tells me what I want to hear. The bot didn’t survive contact with real fees, latency, and fine print, but the platform underneath it did, and that’s the part worth carrying forward.

What’s public, and what isn’t

Three things on this page point at the same project, and they are not the same thing:

  • The original research system. Private. The full 97 reviewed pairs and the live-order path that could place real trades. Every number on this page comes from here.
  • The open-source release. Public, MIT, paper-only. A reduced build that ships 23 reviewed pairs, all of them excluded from strategy metrics, so a scan records real order-book data and deliberately produces no strategy result. Every screenshot on this page comes from here.
  • The interactive diagram. A read-only map of the open-source code, pinned to a commit, where each box links to the file it describes.
The paper-trading tab of the research cockpit, showing scanner controls, cadence, and a counter reading 23 scannable with 0 counting toward strategy metrics.
The paper-trading tab in the open-source release. The eligibility counter reads 0 count toward strategy metrics: every bundled pair was rejected during review, so a scan captures real market data and reports no edge by design.

Where I’d take it next

The arbitrage result was a no-go, but most of the infrastructure is reusable: the capture pipeline, the fee and depth models, and the contract-equivalence review don’t care which strategy sits on top of them. I could reuse the same market-data pipeline for ideas that don’t need two simultaneous fills — mean reversion on a single contract, correlated markets that tend to move together, or tracking where attention and volume shift inside one market. Those all carry real directional risk instead of a locked-in spread, the opposite trade-off from this project, which is exactly what makes them interesting to test next.

Look at how it works

You don’t have to take my word for any of this. The research platform is open source, and the diagram behind the button below maps it as it runs: every box links to the code it describes, pinned to a specific commit, so each citation opens the exact file and line on GitHub.

It follows one path end to end: two public order books arrive, get paired with an explicit capture-skew measurement, get priced through the real per-venue fee engine, and become a depth-adjusted edge row. The parts that make the negative result trustworthy are on it too — the pair registry whose equivalence screen can only flag and never clear, and the mode gates and kill switch that keep the whole thing read-only.

See the repo as an interactive diagram →

Diagram generated with archify by tt-a1i (MIT).