BlockRun

PLAYBOOK · UPDATED 6 OCT 2026

The prediction-market agent playbook (2026), with Claude Code and BlockRun

Thirteen chapters for anyone building an agent that researches or trades prediction markets — Polymarket, Kalshi, Limitless, Opinion, Predict.Fun. Each starts from a mistake that costs real money, then gives the play, a prompt to run in your agent, and the rule to keep, and for every data step, what happened when we ran it.

The data steps were run on one market: Polymarket's "Will no Fed rate cuts happen in 2026?", on 6 Oct 2026. One research dossier, 13 paid calls, $0.100 in total — including the two calls we wasted, which are in here too.

This is method, not advice. Nothing here is a recommendation to trade, and complying with a venue's terms in your jurisdiction is yours to check.

SET UP ONCE, IN CLAUDE CODE, CODEX OR CURSOR
set up BlockRun - https://blockrun.ai/llms.txt

For agents: this page as markdown is at /prediction-market-agent.md.

Where to start

  • Your agent found an "arbitrage" that wasn't one — chapters 2 and 3
  • Fills come in worse than the price the agent saw — chapter 4
  • The smart-money signal says something different every time you ask — chapter 5
  • The wallet you copied stopped winning the week you copied it — chapters 6 and 7
  • The agent explains every move with a story you can't check — chapter 9
  • The model's verdict changes when you rerun it — chapter 10
  • Nobody can say what the agent is allowed to spend — chapters 11 and 13
  • You don't know what one decision costs — chapter 12

CHAPTER 1

What does a prediction-market agent need in 2026?

An agent (Claude Code, Codex, Cursor, Hermes, your own loop). A market-data layer that covers every venue: listings, prices, order books, trades, candles, wallets, smart-money positioning and the oracle state that decides resolution. A research layer for the "why": web and news search with dates. A model to summarise evidence. An execution path with hard caps. And a ledger of what every step cost.

BlockRun is one base URL for all of it: cross-venue market data at $0.0085 a call, web search from $0.011, 113 models, and a Polymarket execution tool with a per-order cap — paid per call from the agent's own wallet, no accounts or API keys to collect.

PROMPT

Set up BlockRun from https://blockrun.ai/llms.txt. Then, without placing any order, list the open markets across every venue for "<your topic>" and tell me the five with the most volume.

Rule: One data layer for every venue, or your agent will compare numbers that were never measured the same way.

CHAPTER 2

How do you find the same market on Polymarket, Kalshi and the rest?

Search every venue at once, then read the questions. Two markets with the same words in the title are often different bets, and an agent that matches on title will find arbitrage that isn't there.

PROMPT

Search all prediction-market venues for "<topic>", status open. Group the results by the exact event each one resolves on, not by title. For each group, say whether the questions are equivalent, and why.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, one call to markets/search for "Fed rate" ($0.0085): 5 venues searched, 16 open markets back.

Three different questions under one phrase. Polymarket asks how many cuts happen in 2026; Kalshi asks how many hikes; Limitless and Predict ask "another hike in 2026?". None of them is the other's hedge.

Lifetime volume ranged from $88.56 (Kalshi, "at least 4 hikes") to $8.6M (Polymarket, "no cuts"). Polymarket rows came back without a price; Kalshi and Predict rows carried one.

Rule: A market matches across venues only when the event, the resolution source and the deadline all match. Otherwise it is related, not the same.

CHAPTER 3

Read the resolution rules before you read the price

The price is the market's opinion of the rules as written, not of the headline. Fetch the full description and the oracle state, and make the agent quote the clause it is relying on.

PROMPT

For condition <id>, fetch the market and its UMA oracle record. Quote the resolution clause, the deadline, and anything that counts that a reader might not expect. Report the oracle state, any disputes, and the bond.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, two calls ($0.017). The rules: emergency cuts outside scheduled meetings count; a 50 bp cut counts as two; a cut of 1–24 bp counts as one; the market stays open until 31 Dec 2026, 11:59 PM ET.

Oracle: state pending, not flagged, not paused, 0 resets, no events, proposal bond $500, reward $5.

Rule: The agent quotes the resolution clause it is betting on. If it cannot quote one, it has not read the market.

CHAPTER 4

Why does my bot fill worse than the price it saw?

The listed price is the last trade, or the midpoint. What you pay is the ask at your size. Preview the order at the size you would actually trade — on BlockRun that preview is free and signs nothing.

PROMPT

Preview, without placing, a market buy of $<size> on the <outcome> side of <condition>. Report the best quote, the worst fill price, and how far that is from the listed price.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026. Listed: YES 0.9581, NO 0.0419. Previews of a $10 and a $25 buy of NO both filled at 0.043, worst fill 0.043, minimum size 5 shares — the book was deep enough ($1.49M liquidity) that our size did not move it. Free.

The other end: Kalshi's "exactly 2 Fed hikes in 2026" listed at 0.67 on $165.78 of lifetime volume. A price on that little volume is a quote, not a consensus.

A history call to the order-book endpoint over a 10-minute window came back with zero snapshots and still cost $0.0085. For "what would I pay now", use the free preview; keep the history endpoint for backtests.

Rule: Preview the fill at your real size before deciding. Skip any market where your size moves the price by more than your edge.

CHAPTER 5

Does smart-money tracking work on Polymarket? It depends on the filter you set

"Smart money" is not a fact the data contains. It is a filter you choose, and the answer changes with it. Write the criteria as numbers in the prompt and report the wallet count next to the percentage.

PROMPT

For condition <id>, get smart-money positioning twice: once with min_trades 100, once with min_trades 100, min_realized_pnl 50000 and min_win_rate 0.55. Report wallet count and net-buyer share for each, and what changed.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, two calls ($0.017). Loose filter (100+ trades): 10,121 wallets, 75.8% net buyers. Strict filter (also $50K+ realized PnL and 55%+ win rate): 56 wallets, 58.9% net buyers.

Same market, same hour: the loose signal reads as a crowd leaning hard one way; the strict one is close to a coin flip. The strict cohort's average win rate in this market came back as 48%, under the 55% we filtered on — the filter and the average are not measured over the same trades, so read the field definitions before you trust either.

Rule: Every smart-money claim carries its criteria and its wallet count. "75% of smart money" without both is not a signal.

CHAPTER 6

Polymarket leaderboards rank paper gains — rank by realized PnL instead

A leaderboard sorted by total PnL on an open market is mostly unrealized marks. Positions that have not closed have not been proven right.

PROMPT

Get the top 5 wallets for condition <id> by total PnL. For each, show realized PnL, total PnL and positions closed. Re-rank by realized PnL among wallets with at least 3 closed positions.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, one call ($0.0085). Of the top 5 by total PnL, 4 had closed zero positions in this market. Rank 1 showed $146,342 total PnL and −$382 realized.

Only one of the five — rank 4 — had closed positions (3 of 3 won, $47,064 realized).

Rule: Rank wallets by realized PnL with a minimum number of closed positions. Unrealized PnL is the market's opinion, not the wallet's record.

CHAPTER 7

Check a wallet across time before your agent copies it

One good market does not make a forecaster. Pull the wallet's whole record and read the windows side by side: all-time, 30 days, 7 days. ROI says more than PnL; a high-volume trader can show millions in profit on a thin edge.

PROMPT

Profile wallet <address>: realized PnL, ROI, win rate and positions closed for all-time, 30 days and 7 days, plus trade count and wallet age. Say whether the recent windows agree with the all-time record.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, one call ($0.0085), on that rank-4 wallet. All-time: $2.51M realized on $150M volume — 2.1% ROI over 606,193 trades and 72,380 closed positions, 56% win rate, wallet age 1,464 days.

Last 30 days: −$109,549 realized, profit factor 0.58. Last 7 days: +$5,746. A high-volume trader with a thin edge that is currently negative — not someone whose next trade tells you much.

Rule: Copy a wallet only when its 30-day and all-time records agree, and judge it on ROI and closed positions, never total PnL.

CHAPTER 8

Wallet clusters are transfer graphs, not identities

Cluster data links wallets that moved money to each other. That is a lead about behaviour, not a statement about who owns what — and it can be old.

PROMPT

Get the transfer cluster for wallet <address>. Report how many linked wallets came back, what kind of link each is, and the date the cluster was computed.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026, one call ($0.0085). 50 linked wallets, every one at confidence 100, most linked by a single inbound transfer. The cluster was computed on 6 Jul 2026 — three months before the run.

Rule: Check computed_at before using a cluster, and treat a transfer link as context, never as the same owner.

CHAPTER 9

No dated source, no "why"

Agents are fluent at explaining price moves after the fact. Make every explanation carry a source and a date, and check that the date lines up with the move in the candles.

PROMPT

Pull daily candles for <condition> for the last 30 days and find the largest one-day move. Search the news for that date. Report the move, the dated source that explains it, or "no source found".

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026. One candles call ($0.0085, interval 1440 minutes): YES closed 0.936 on 15 Sep, 0.948 on 16 Sep, 0.954 on 17 Sep, 0.959 on 6 Oct.

One neural news search ($0.011) returned the Federal Reserve's own statement of 16 Sep 2026, the Chair's press-conference transcript, and Reuters and CNBC reports of the same day: the FOMC raised the target range to 3.75%–4.00%. The move and the source share a date.

Rule: No source link, no explanation. No date that matches the move, no causation.

CHAPTER 10

Let the model summarise the evidence, not pick the trade

A cheap model is good at turning a dossier into structured evidence for and against. It is not a reason to trade. Tell it so in the system prompt, ask for JSON, and leave enough room for the answer to finish.

PROMPT

System: you never recommend a trade; state what the evidence supports and what would change it. Max 3 items per list, 20 words per item. JSON: market_implied, evidence_for, evidence_against, what_would_move_it, confidence_in_data. User: <the dossier from chapters 2–9>.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026 on gpt-5.6-luna, $0.002 a call. The first attempt, capped at 600 tokens with no length limit in the prompt, stopped mid-list (finish_reason: length) and was billed anyway. The second, with the length rule, finished.

The answer: market-implied 0.958; for — the September hike and a hawkish statement, the price trend; against — two scheduled meetings left and emergency cuts count, only 59% of the strict cohort net-buying, other venues' contracts not equivalent; confidence in the data: medium.

Rule: The model writes evidence, in a schema, with a length budget. A person — or a rule you wrote down — decides.

CHAPTER 11

Put the money rules where the agent can't argue with them

Three ceilings, enforced by the tools rather than by the prompt: a budget per research run, a cap per order, and a refusal before payment when a request is malformed. Prompts get ignored under pressure; a refused call does not.

PROMPT

Delegate a $5 budget to agent id "pm-research" and keep the per-order cap at $25. Then run the dossier for <condition> under that agent id and report spend at the end.

WHAT HAPPENED WHEN WE RAN IT

Recorded 6 Oct 2026. A $500 order preview was refused by the $25 per-order cap before anything was signed. The research ran under a $5 delegated budget and used $0.100 of it.

Three malformed data requests — a candle interval written as "1d" instead of minutes, a smart-money call with no criterion, an order-book call with no time window — were refused before payment, each with the fix in the error and "No payment was made".

Rule: Every agent runs under a budget, every order under a cap, and nothing goes autonomous until both have been hit at least once in testing.

CHAPTER 12

What does one prediction-market research dossier cost?

Count it per decision, not per month. Here is the whole bill for the dossier in this playbook — one market, search to summary.

WHAT HAPPENED WHEN WE RAN IT

10 market-data calls × $0.0085 = $0.085 (search, market, order-book history, smart money ×2, leaderboard, wallet, cluster, oracle, candles).

1 news search = $0.011. 2 model calls × $0.002 = $0.004. Previews and the refused requests: free.

Total: $0.100 for one market, of which $0.0105 was waste we could have avoided (the empty order-book window and the truncated model reply). A hundred markets a day is about $10.

Rule: Know the cost of one decision before you scale to a thousand, and log the waste separately from the work.

CHAPTER 13

Roll it out like software: shadow, small size, human gate, then autonomy

Run the agent in shadow first: it writes the dossier and the decision it would have made, with the price, and you compare against resolution. Then small size with a person approving every order. Autonomy last, and only inside the caps from chapter 11.

Log per decision: the market, the clause quoted, the fill preview, the smart-money criteria and count, the sources and dates, the model's JSON, and the cost. That log is what tells you which step earned its money.

Rule: Nothing trades on its own until it has a shadow record you can score against resolution.

Every recorded run behind this playbook

  • 6 Oct 2026 — one dossier on "Will no Fed rate cuts happen in 2026?": 13 paid calls, $0.100, plus 3 free order previews and 4 requests refused before payment.

Prediction-market agent questions: data, bots, smart money, cost

Can an AI agent trade on Polymarket?
Yes. BlockRun's MCP gives an agent Polymarket market data, free order previews and an execution tool that signs from the agent's own wallet, with a per-order cap enforced by the tool rather than the prompt. Whether it trades unattended is a separate decision; the playbook argues for shadow mode first.
What data does a prediction-market bot need?
Listings across every venue, prices and order books, trades and candles, wallet records and smart-money positioning, and the oracle state that decides resolution, plus dated news for the why. One data layer for all of it keeps the numbers comparable.
Is Polymarket smart-money data reliable?
It is as reliable as the filter you set. On the market in the playbook, a loose filter showed thousands of wallets leaning hard one way, and a strict one showed dozens close to a coin flip, in the same hour. Report the criteria and the wallet count with every percentage.
Can my agent find arbitrage between Polymarket and Kalshi?
Only between markets that resolve on the same event, source and deadline. In our run, three venues phrased Fed questions alike and asked three different things: cuts, hikes, and another hike. Matching on the title finds arbitrage that is not there.
Why does a prediction-market leaderboard mislead an agent?
Sorted by total PnL on an open market, it mostly ranks unrealized marks. Most of the top wallets in our run had closed nothing in that market. Rank by realized PnL with a minimum of closed positions, and read the recent window next to the all-time record.
How much does it cost to research one prediction market with an agent?
The full dossier in the playbook, from search to a model summary, cost about ten cents in USDC, priced per call with no subscription. The bill is broken down call by call, including the two calls that were wasted.
Do I need a Polymarket or Kalshi API key for my agent?
Not for the data: the agent pays per call from its own wallet, with no accounts or keys. Trading on Polymarket through the BlockRun tool uses a deposit wallet the tool sets up from the same key.
Should I use a prediction-market data API instead of scraping each venue?
If your agent compares venues, yes. Each venue shapes prices, volumes and wallets differently, and an agent that stitches scrapers together compares numbers that were never measured the same way. One cross-venue layer, paid per call, is cheaper than maintaining the scrapers.