Favorite-Longshot Bias in Prediction Markets: How to Test Your Own Tail Purchases
Test favorite-longshot bias in your resolved prediction-market purchases: compute returns by price tail, compare three aggregation methods, and see what the evidence can and cannot show.
The favorite-longshot bias is a return pattern: purchases at low prices, which pay a lot when they win, earn lower average returns than purchases at high prices. To look for it in your own resolved prediction-market purchases, record the price you paid per token, the number of tokens, and whether the token you bought paid $1 or $0. Then group purchases by price tail, such as under 10 cents and 90 cents or above, and compute the return per dollar spent in each tail. Repeat the calculation under more than one averaging rule, because the answer depends on whether you weight dollars, separate markets, or parent events equally. In a large 2026 Polymarket preprint, the same resolved purchases showed a longshot loss under two of those rules and a longshot gain under the third, and the largest loss was statistically imprecise. Your record can describe your own purchases; it cannot show why they behaved as they did.
This article is about that price-tail test. It assumes you already know how to compare stated forecasts with realized frequencies, which prediction market probability calibration covers, and it does not repeat that method. It also does not decide whether a given contract can be entered or exited at a workable price, whether a position is sized sensibly, or how a contract’s resolution wording could change the outcome.
How do you test for a favorite-longshot bias in your own contract record?
- Log each purchase’s price per token, number of tokens, side (YES or NO), category, and parent event when you enter it.
- After resolution, record the payoff of the token you bought: 1 if that token paid $1, 0 otherwise. A NO purchase scores 1 when NO wins, not when the underlying question resolves YES.
- Fix your price tails, your resolution cutoff, and your averaging rule before looking at results.
- For each tail, compute the pooled return per dollar (total profit divided by total cost, before fees) and the probability error (share of tokens that paid $1 minus the average price paid).
- Repeat the return under per-market and per-parent-event averaging, and report purchase, token, market, and event counts beside each figure.
- Treat the observed return in a thin tail as a valid descriptive result, but treat it as inconclusive evidence about a persistent underlying pricing bias, and keep any conclusion about your own record separate from the market-wide evidence.
What is the favorite-longshot bias, and what does the evidence show?
The bias is a long-standing regularity from horse-race betting, where long shots were overbet and favorites underbet relative to how often each actually won.1 The Polymarket study used in this article defines it in return terms: low-priced purchases earn lower average returns than high-priced ones.2 Three related statements are easy to conflate:
- Return pattern. Under a stated averaging rule, low-priced purchases return less per dollar than high-priced purchases.
- Probability error. The share of tokens that ended up paying $1, minus the average price paid, in percentage points.
- Two-sided pattern. Longshots return below zero and favorites return above zero.
Return and probability error are linked but not interchangeable. For one pooled set of purchases they share a sign, because return is probability error divided by average price (the identity below). Averaging across markets breaks that link: percentage returns weight low-priced markets heavily, and percentage-point errors do not. With each child market weighted equally, the Polymarket longshot return was −6.30% while the longshot probability error was +0.24 percentage points.2 A negative average return therefore does not by itself mean tokens paid off less often than their prices implied.
Published explanations are commonly grouped into misperceived probabilities, a preference for risky payoffs, differences in beliefs or information among bettors, and features of how the market is built;2 Ottaviani and Sørensen’s chapter gives an overview of the main ones.3 Snowberg and Wolfers tested two of these against horse-race pools that included compound bets and found evidence favoring misperception of probabilities, as prospect theory suggests, over risk-love.1 The prospect-theory reasoning is that people weight small probabilities more heavily than their size warrants and are risk-seeking toward low-probability gains, so a long shot looks more attractive than its odds justify.4 That is a proposed mechanism from betting data, not a measured fact about any individual prediction-market trader.
Prediction markets have since produced their own evidence. An analysis of Kalshi data documented a favorite-longshot bias and modeled how it differs between traders who post offers and traders who accept them.5 A September 2026 Polymarket preprint, not yet peer reviewed, used a dataset of 588,287,492 observed transactions by about 2.48 million wallets between November 11, 2022 and March 29, 2026. A wallet is not necessarily one person. After excluding invalid records, markets with more than two outcomes, and purchases by buyers with highly concentrated counterparties (a wash-trading screen), 586,123,163 purchases remained. Only the 560,909,624 purchases in markets resolved by March 29, 2026 enter the return calculations, covering 591,187 child markets in 250,307 parent events.2 In that resolved sample, pooled purchases below 10 cents lost 19.35 cents per dollar, with a 95% confidence interval from −46.99% to +8.30%, while purchases at or above 90 cents earned 0.83%. These are what each purchase would have earned if held to resolution, before fees, not any wallet’s realized profit.2
Two boundaries on the evidence matter here. First, these studies describe aggregate purchase data from a platform, not what your own account will show. Second, a price pattern alone cannot tell you why it exists; the Polymarket authors say their results narrow the possible explanations without identifying a single one.2
How is a tail-price test different from a calibration review?
A calibration review groups your forecasts by the probability you stated and asks whether that number matched how often those forecasts came true. It tests your estimate. The tail-price test groups purchases by the price you paid and asks whether tokens bought at that price paid off often enough to cover it. It tests the payoff of a price region, whatever your estimate was.
The two can disagree. A trader could be well calibrated across stated probabilities and still lose money on contracts priced under 10 cents, for example on purchases where the stated probability sat below the price, because the price paid, not the stated probability, sets the break-even success rate. The reverse also happens: no visible tail loss alongside miscalibration elsewhere in the range. Neither result substitutes for the other, which is why a record that keeps only stated probabilities cannot run this test. It needs the price paid and the token quantity for each purchase.
What does a tail purchase have to win to break even?
For each purchase i, define the price paid per token p_i, the number of tokens bought q_i, and the payoff of the token you bought y_i, which is 1 if that token paid $1 and 0 otherwise. For a YES purchase, y_i is 1 when YES wins. For a NO purchase, y_i is 1 when NO wins. Then:
Cost_i = p_i × q_i
Payout_i = y_i × q_i
Profit_i = q_i × (y_i − p_i)
r_i = (y_i − p_i) / p_i (return per dollar; needs p_i > 0)
For a set of purchases B, such as one price tail:
PooledReturn_B = Σ q_i(y_i − p_i) / Σ p_i q_i
= (total payout − total cost) / total cost
AveragePrice_B = Σ p_i q_i / Σ q_i
SuccessRate_B = Σ y_i q_i / Σ q_i
ProbabilityError_B = SuccessRate_B − AveragePrice_B
Because both the numerator and denominator can be divided by Σ q_i, the pooled return equals the probability error divided by the average price:
PooledReturn_B = ProbabilityError_B / AveragePrice_B
This holds when both statistics use the same pooled purchases and these weights. It does not carry over to an average of separate market-level returns and an average of separate market-level errors, which is exactly where the two can disagree in sign.
Keep the units apart: a return is a percentage of money spent, a probability error is in percentage points of tokens, and profit is in dollars. In the paper’s pooled longshot row, a −19.35% return and a −0.23 percentage-point probability error describe the same purchases.2
The break-even success rate before fees is the average price. A tail purchase near 8 cents needs its tokens to pay $1 about 8% of the time. Because the price is small, a modest shortfall produces a large percentage loss: if the success rate is 5%, the return is 0.05 ÷ 0.08 − 1, or −37.5%. A gap of the same three points on the favorite side is far smaller in percentage terms: at an average price of 92 cents, an 89% success rate returns about −3.3%, and 95% returns about +3.3%. That asymmetry is why the tails get their own test. The figures above are arithmetic on assumed success rates, not observed data, and they exclude fees and the difficulty of getting the quoted price, which trading slippage and execution costs covers separately.
The formulas apply to ordinary binary tokens with a known $0 or $1 payoff. Purchases in markets that have not resolved are neither wins nor losses; leave them out and use one resolution cutoff for the whole review. Contracts that were canceled, voided, refunded, or settled under nonstandard terms do not fit the $0/$1 formula, so report them separately instead of forcing them in. How a venue defines resolution is outside this article.
Why does the result change with the unit you average over?
The Polymarket study’s most useful contribution for a reviewer is a demonstration that one set of resolved purchases can give different answers under different weightings.2 A child market is one separately traded question, and a parent event is the group of related questions listed together, such as every candidate market under one election.
| Averaging unit | Longshot (under 10¢) | Favorite (90¢ and above) |
|---|---|---|
| Each child market weighted equally | −6.30% [−8.38%, −4.22%] | +0.277% [+0.242%, +0.312%] |
| All purchases pooled, weighted by dollars paid | −19.35% [−46.99%, +8.30%] | +0.83% [+0.60%, +1.07%] |
| Each parent event weighted equally | +4.09% [+1.29%, +6.90%] | +0.392% [+0.345%, +0.440%] |
Returns are per dollar paid, with 95% confidence intervals in brackets. The authors cluster the intervals for the first two rows by parent event.2 Each rule builds its figure differently:
- Child market. Pool the eligible purchases inside each child market, calculate that market’s dollar-weighted return, then give every represented child market equal weight. This is not a simple average of individual trades or tokens.
- Pooled. Pool every eligible purchase in the tail and divide total profit by total cost.
- Parent event. Pool the eligible purchases inside each parent event, calculate that event’s dollar-weighted return, then give every represented event equal weight.
A child market can appear in both tails because purchases are classified by price, not markets: a NO token bought at 8 cents and a YES token bought at 92 cents in the same market land in different tails and are kept separate.
Read the intervals before the point estimates. The pooled longshot interval includes zero, so that −19.35% is large but is not statistically established at the conventional two-sided 5% level; the authors’ own bootstrap p-value for it is 0.184.2 The other longshot intervals exclude zero, but they point in opposite directions: a loss under child-market weighting and a gain under parent-event weighting. All three favorite intervals exclude zero and are positive. The three calculations do not all establish longshot losses. On longshot probability errors the three rules gave +0.24, −0.23, and +0.20 percentage points, so the sign of the return and the sign of the error can differ.2
The authors trace the reversal to how spending and market counts are distributed. Parent events with many child markets tended to have lower longshot returns and pull down the equal-child-market figure, while events attracting the most longshot spending tended to have lower returns and pull down the pooled figure.2 Category also matters: Crypto and Politics showed longshot losses and favorite gains under both primary calculations, with all four intervals excluding zero, while Sports was the clearest exception, with positive longshot point estimates whose intervals include zero.2
The lesson for your own review is not that one weighting is correct. Each answers a different question. The pooled figure describes hypothetical held-to-resolution performance under the stated purchase set, not necessarily how your account fared if you exited positions early; the per-market average describes a typical market, and the per-event average describes a typical event. If you pick after seeing which one looks best, you have turned a diagnostic into a rationalization. Decide which question you are asking, write it down, and report the others alongside it.
How do you test your own tail purchases?
Start with the record, since the test cannot be run on data you did not capture. At purchase, log the purchase timestamp, contract identifier, parent-event identifier, side, actual execution price, token quantity, category, and the scheduled resolution date if one is known. After resolution, add the actual resolution date, the final payoff of the token you bought, and the resolution status. The actual resolution date is not known at entry, and a scheduled date can change. The parent-event identifier is easy to omit, and the aggregation result above shows why it matters.
Then set the review rules in advance:
- Bands. The research used under 10 cents and 90 cents or above as the tails. You can use coarser or finer bands, but choose them first, and prefer wider bands when a narrow one would leave a handful of purchases.
- Resolution cutoff. The paper’s return sample includes only markets with known payoffs by March 29, 2026, and recent purchases had less time to resolve, so later periods describe markets that resolved quickly.2 Pick one cutoff date for your review, and never score an unresolved purchase as a loss or a win.
- Held-to-resolution return. Compute each tail’s return as if every purchase had been held to settlement, and note separately any positions you actually exited early, since the two numbers answer different questions.
- Three averages. Report the pooled return, the per-market average, and the per-event average side by side. If they disagree in sign, the disagreement is itself the finding.
- Longshots against favorites. A complete personal review compares longshot returns with favorite returns, using the same averaging rule and the same observation period for both. The two-sided pattern asks whether longshots return below zero and favorites above zero, which is a different question from whether longshots return less than favorites. A negative longshot return alone does not establish the relative pattern. A formal claim about the longshot-favorite difference needs uncertainty estimated for the difference itself; do not infer it from two separate confidence intervals.
- Counts. Report the number of purchases, total tokens, distinct child markets, and distinct parent events for each tail. For scale, the paper’s 560.9 million purchases sit in 250,307 parent events, and its purchase counts measure trading activity, not positions.2
- What would change your mind. Decide in advance what evidence would change your behavior, and apply the discipline in trading data statistical reliability so that you do not try many bands until one looks meaningful.
Grouping by parent event helps you notice dependence, but it does not settle it. The paper notes that some parent events group alternative answers to one question, while others group related questions that can both be true, so you cannot assume every market in an event is mutually exclusive, and you cannot assume different events are independent of each other. Purchase counts and event counts are useful descriptive measures, but neither is an effective sample size, and a count of 19 parent events is not 19 independent observations. Treating correlated purchases as independent can misestimate uncertainty, and the direction and magnitude depend on the covariance structure: positive covariance can increase the variance of an average, while negative covariance can reduce it.
The study clusters its primary intervals by parent event.2 Event clustering suits your own data only when its assumptions and grouping structure are defensible. With very few events, conventional large-sample cluster-robust inference can be unreliable, and dependence across events can invalidate simple event clustering. Where a formal interval is justified, use an appropriate small-cluster method and disclose the assumptions. Where it is not, report descriptive estimates and counts without an interval.
A personal record can show whether your own tail purchases lost money. It cannot show why. A tail loss may reflect the market-wide pattern, your own selection of which longshots to buy, or plain variance. The Polymarket authors found that the top decile of wallets ranked by past longshot buying earned about the same negative return on later purchases as other wallets did (−21.96% versus −20.33%), which weighs against the idea that the losses belong to a distinctive group of habitual longshot buyers.2 Your record is evidence about your process, not a test of the market, and a market-wide pattern is not evidence that you have a profitable edge.
A worked example: a longshot loss in a small sample
Suppose a trader reviews five months of resolved purchases. Thirty-four purchases, YES or NO, were made below 10 cents, averaging 7 cents, across 19 parent events. The token bought paid $1 in two of them. For simplicity, assume every purchase was for the same number of tokens, so the simple average price equals the quantity-weighted average price and the success rate is the quantity-weighted one.
- Success rate: 2 ÷ 34 ≈ 5.88%.
- Total cost, buying one token per purchase: 34 × $0.07 = $2.38. Total payout: $2.
- Pooled return: (2 − 2.38) ÷ 2.38 ≈ −15.97%.
- Probability error: 5.88% − 7.00% ≈ −1.12 percentage points, and −1.12 ÷ 7.00 ≈ −16%, matching the identity above.
Now weigh how much that shows. If purchases were independent with a true 7% success rate, the standard error of a 34-purchase success rate would be √(0.07 × 0.93 ÷ 34), about 4.4 percentage points. The observed shortfall is about 1.1 points, roughly a quarter of that, and two wins against 2.38 expected is a difference of less than one purchase. This is an illustrative standard-error calculation, not a formal hypothesis test: with only 2.38 expected successes, a normal approximation should not be treated as a reliable standalone significance test. Independence is a strong assumption here, because purchases in the same event can be correlated. Positive covariance would widen the real uncertainty and negative covariance could narrow it. The 34 purchases and 19 events are descriptive counts, not an effective sample size.
The same formula shows why small tail samples are noisy: at a 7% rate and independent purchases, about 651 observations would bring the standard error to one percentage point. That is a precision calculation under an independence assumption, not a required sample size. It does not establish 80% power, does not guarantee that a one-point difference would be detected, and does not hold without qualification when purchases within an event are dependent.
The observed hypothetical held-to-resolution return is approximately −16%. This descriptive result does not establish a persistent negative expected return. Without a favorite comparison sample from the same period under the same averaging rule, it cannot establish a favorite-longshot return differential, as described above. The useful next step is to keep logging under the same rules and re-run the review after more events have resolved.
Common ways the test goes wrong
Scoring NO purchases by the underlying YES outcome. The payoff variable belongs to the token you bought. Using the question’s YES result for a NO purchase reverses its win and loss.
Treating one longshot hit as proof of skill. A single win at 6 cents pays a large multiple and feels like validation. A single large payoff does not establish a persistent pricing advantage or a reliably profitable purchase process.
Treating correlated purchases as independent. Contracts under one parent event may be linked: some events group mutually exclusive answers, and others group related questions whose outcomes can both be true. Treating correlated purchases as independent can misestimate uncertainty; the direction and magnitude depend on the covariance structure. Purchase and event counts describe your record, but neither equals an effective sample size.
Choosing the weighting after seeing the results. With three defensible averages that can differ in sign, picking the flattering one is easy and hard to notice.
Forcing unresolved or canceled contracts into the formula. Unresolved purchases are not losses, and voided or refunded contracts do not follow the $0/$1 payoff.
Treating a market-wide bias as your own edge. The bias existing on average does not mean a rule built around it is profitable after fees, spreads, and the price you can actually get. Minimum edge after trading costs covers how much margin a strategy needs before costs consume it.
Reading held-to-resolution returns as realized profit. If you sold early, or fees applied, your account’s outcome differs from the calculation.
Inferring the cause from the price pattern. A tail loss does not tell you whether misperception, risk preference, or market structure produced it, and aggregate purchase records cannot diagnose an individual trader’s psychology.
Where Costante fits
Costante’s relevance here is limited. Its core workflow is built for discretionary intraday traders: planning rules before a session, making guardrails visible, logging decisions, and reviewing behavioral drift afterward. A prediction-market trader could apply the review discipline in this article, including a rule for tail purchases and a scheduled review of resolved contracts, but would need to keep the price, token quantity, side, payoff, parent event, and resolution record separately, whether in a spreadsheet or in a journal app whose fields they can define (the trading journal app guide covers what to look for). Costante does not compute favorite-longshot bias or the other statistics in this article, provide a dedicated prediction-market record, track prediction-market contracts, import resolution outcomes, connect to prediction-market exchanges, execute trades, or recommend contracts. The tails, the averaging rule, and every decision about whether to buy a longshot remain the trader’s own.
Frequently asked questions
What is the favorite-longshot bias in prediction markets?
It is a return pattern: under a stated averaging rule, low-priced purchases earn lower average returns than high-priced purchases. The Polymarket study also reports a separate probability-error measure, and the two can differ in sign once results are averaged across markets. Studies of Kalshi and Polymarket data have reported the bias, but its size depends on the market, the category, and how results are averaged.
Does the favorite-longshot bias mean longshots always lose money?
No. In the Polymarket study, longshot purchases lost money when each child market was weighted equally (−6.30%, interval excluding zero) and when dollars were pooled (−19.35%, interval including zero, so not statistically established), but gained when each parent event was weighted equally (+4.09%). Sports showed positive longshot point estimates under both primary calculations. Results also exclude fees and describe held-to-resolution returns, not any buyer’s realized profit.
How many resolved contracts do I need to test for it?
There is no fixed number. Under an independence assumption at a 7% success rate, about 651 purchases would bring the standard error to one percentage point, but that is a precision calculation, not a required sample size or a power guarantee. Whether a personal record is large enough depends on its actual sample size, how dependent the purchases are within and across events, and the size of the effect being tested; a small record or a narrow price band may leave the result inconclusive.
Is this the same as checking whether my probabilities are calibrated?
No. Calibration groups forecasts by the probability you stated. This test groups purchases by the price you paid and checks whether that price was covered by how often the purchased tokens paid $1. A trader can pass one and fail the other.
Sources
Costante provides educational workflow tools, not financial advice. Trading involves risk.
Footnotes
-
Snowberg, E., & Wolfers, J. (2010). Explaining the Favorite–Long Shot Bias: Is it Risk-Love or Misperceptions? Journal of Political Economy, 118(4), 723–746. ↩ ↩2
-
Cardozo, M., & Rivero-Wildemauwe, J. I. (2026). The Favorite–Longshot Bias in Prediction Markets: Evidence from Polymarket. arXiv:2609.12878v1, submitted September 11, 2026 — preprint, not peer reviewed. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17
-
Ottaviani, M., & Sørensen, P. N. (2008). The Favorite-Longshot Bias: An Overview of the Main Explanations. In D. B. Hausch & W. T. Ziemba (Eds.), Handbook of Sports and Lottery Markets, Elsevier, 83–101. ↩
-
Tversky, A., & Kahneman, D. (1992). Advances in Prospect Theory: Cumulative Representation of Uncertainty. Journal of Risk and Uncertainty, 5(4), 297–323. ↩
-
Bürgi, C., Deng, W., & Whelan, K. (2025). Makers and Takers: The Economics of the Kalshi Prediction Market. CEPR Discussion Paper DP20631, first published September 8, 2025, revised January 31, 2026. ↩