Probability Revision Quality: Did Your Update Follow the Evidence?
Review whether a probability revision matched the evidence you had at the time, judged by direction, size, timing, and rationale rather than by the outcome.
A probability revision was well supported if its direction, size, timing, and rationale fit the evidence available when you made it. That is different from process compliance (did you follow your predefined procedure?) and from predictive performance (did the method forecast well across many observations?). Following a poorly designed rule proves consistency, not normative correctness; departing from a rule is a review flag, not automatic proof of irrationality. Evidence can remain uncertain even when documentation is complete.
The result does not settle the quality of one revision. A sound revision can precede a NO, and a sloppy one can precede a YES. Outcomes still matter for ex-post validation of an updating method across many resolved forecasts, while this article focuses on ex-ante process review: what was knowable and defensible at the time.
The reader who needs this is anyone who states a probability, sees new information, and changes the number: a trader working a prediction-market contract, or a discretionary trader who has written “I think this level holds about 65% of the time” and later moved it to 80% after a headline. The stated probability is the object under review. A revision is a change in that probability; a revision decision is the recorded response to evidence, including updating, retaining the prior, or deferring judgment.
How should you review a probability revision?
- Record the forecast or contract and its prior probability before new evidence is considered; never overwrite it.
- Log each potentially relevant evidence event, including events that produce no update.
- Record the decision—update, no update, or deferred judgment—and the contemporaneous rationale.
- At a fixed review boundary, assess direction, magnitude, timing, and rationale before looking at the outcome.
- Look for patterns across many revisions and independent resolved forecasts, not a verdict on one.
What is a probability revision, and what is revision quality?
A prior is your probability before the new information. Evidence is the specific new information. A revised forecast or updated probability is the number you hold afterward; calling it an exact Bayesian posterior is unwarranted unless the likelihood assumptions and inputs were actually estimated. Revision quality is how well the step from prior to updated probability fits the evidence: whether it moved in the direction the evidence points, by an amount the evidence can defensibly support, at a reviewable time, for a reason you can state.
The tidy version of this is Bayes’ rule in odds form: posterior odds equal prior odds multiplied by a likelihood ratio, where the ratio compares how likely the evidence is if the event will happen against how likely it is if it will not. Starting from a 40% prior (odds of about 0.67 to 1), the table shows what evidence of different diagnostic strength would justify.
| Evidence is this many times more likely if YES than if NO | Revised probability from a 40% prior |
|---|---|
| 1.5 (weak) | 50% |
| 3 (substantial) | 67% |
| 10 (very strong) | 87% |
These figures are illustrations, not a recommended scale. A discretionary trader usually does not know a precise likelihood ratio, and false precision is its own hazard. For multiple evidence items, multiplying individually estimated ratios is valid only when the items are conditionally independent under both competing hypotheses. Otherwise, later ratios must account for evidence already observed. An official announcement, a news report repeating it, and the market reaction to it may all reflect the same underlying information; three headlines or sources do not automatically make three independent signals. The same LR of 3 moves a 10% prior to 25%, a 40% prior to about 67%, and a 70% prior to 87.5%, showing why fixed percentage-point expectations are misleading.
How is revision quality different from calibration?
The two reviews use different evidence and answer different questions, so a trader can pass one and fail the other.
| Calibration review | Revision-quality review | |
|---|---|---|
| Question | Do my stated probabilities match realized frequency? | Did each change follow the evidence? |
| Unit of analysis | A bucket of many resolved forecasts | One revision, then a pattern across revisions |
| Needs resolutions? | Yes | No |
| Typical failure found | Overconfidence in a probability range | Overreacting to weak evidence, reacting late, or following the price without a documented or justified rationale |
Prediction market probability calibration covers how to choose an observation convention for a probability that changes before resolution and how to read calibration buckets. This article starts one step earlier, at the revision itself, and assumes you already keep a revision history.
What are the four dimensions of a revision?
Direction. Did the full documented information set point the way you moved? A revision toward YES after considering evidence that, in aggregate, is more likely under NO may be a direction error; one isolated item is not enough if other relevant evidence was incorporated simultaneously. A price move may itself be the identified signal, so distinguish independent information, an explicitly used market signal, mechanical price copying, and an unexplained change with no documented trigger.
Magnitude. Was the change proportionate? Two opposite errors are documented. In experiments on detecting regime shifts, Massey and Wu found that underreaction was most common in unstable environments with precise signals, and overreaction was most common in stable environments with noisy signals. Their explanation, the system-neglect hypothesis, is that people react mainly to the signals they observe and only secondarily to the system that produced them.1 For a trader, the review question is therefore not “did I move too much?” in general, but “did the size of my move take into account how noisy or reliable this kind of evidence usually is?”
Timing. Did the revision follow a defensible information and decision timeline? Record public release time, the time you actually accessed the information, and the time you revised. The publication-to-revision interval is not a direct measure of your delay unless access time is known. Market reaction time is not automatically your access time; verification can justify a delay, and an immediate update is not automatically superior. A late update should be judged against an explicit review standard, not merely against a price move.
Rationale. Is there a contemporaneous sentence naming the evidence and why it matters? A rationale reconstructed afterward is unreliable. In a classic study that asked people to state probabilities before a set of events and recall them later, Fischhoff and Beyth found that remembered probabilities generally drifted toward what participants believed had happened.2 The study tested recalled probabilities, not recalled rationales, so this is an inference rather than a tested result: if remembered probabilities drift toward the outcome, a rationale written at the review is exposed to the same drift, which is why the prior and the rationale are recorded at the time of the revision.
How do you grade the evidence that arrived?
Grade evidence before deciding how far to move, and write the grading rule in advance. Four questions do most of the work:
- Is it new to this forecast? Previously incorporated evidence should not be counted twice. Information reflected in the market price may still be absent from your own forecast and may justify an update.
- Is it independent? A news item and the price move it caused overlap heavily, so do not count them as two independent pieces of evidence.
- Is it credible and diagnostic? How reliable is the source, and would this evidence look very different if the event were going to happen than if it were not? A dramatic headline that would appear under either outcome is weak evidence however striking it is.
- How much is behind it? One anecdote and a large data release are different weights.
The credibility and amount-of-evidence questions echo a finding by Griffin and Tversky: people tend to focus on the strength or extremeness of evidence, such as how warm a letter is, with insufficient regard for its weight, such as the credibility of the writer or the size of the sample. Their result was overconfidence when strength is high and weight is low.3 Applied to revisions, that plausibly means a dramatic but thin piece of evidence is a case where a large revision looks justified and may not be. That is an application of their confidence finding, not something they tested for revision size.
A practical convention is to grade the trader’s documented information set weak, moderate, or strong under a written definition and use illustrative maximum review thresholds: flag a change exceeding 10 percentage points for weak evidence, 20 points for moderate evidence, or 35 points for strong evidence. No minimum change is required by any category. These are user-defined audit triggers—not validated optimal sizes, Bayesian likelihood-ratio equivalents, or universal prescriptions. A justified revision can exceed a threshold, and a revision inside one can still be poorly supported. When several evidence events contribute, assess their combined diagnostic effect and dependence rather than applying one isolated item’s threshold mechanically.
There is also research on update size itself. Atanasov and colleagues studied forecasters in a four-year geopolitical forecasting tournament and distinguished the frequency of updates, their magnitude, and a tendency to confirm the initial judgment. They reported that the most accurate forecasters made frequent, small updates, while low-skill forecasters were prone to confirm initial judgments or make infrequent, large revisions, and that update magnitude mediated the causal effect of training on accuracy.4 The evidence concerns geopolitical questions rather than trading, and it does not say small steps are always right. Massey and Wu’s result is the counterweight: underreaction, meaning steps that are too small, was what they found most common in unstable environments with precise signals. Update size should follow evidence weight, not a fixed habit.
Was the market price the evidence?
In a prediction market, the price can be information, but it is not automatically an independent proof or an exact probability. Spreads, fees, market structure, and other frictions matter. The review question is what role the price played in your forecast. Distinguish three cases:
- Independent evidence: “Revised from 40% to 50% after the committee minutes; weak-to-moderate evidence.”
- Market as an explicit signal: “Used the market move from 41¢ to 61¢ as one input; revised from 40% to 50% after checking why it moved.”
- Mechanical copying: “Revised from 40% to 61% because the price rose from 41¢ to 61¢; no forecasting rationale.”
Cases two and three should not receive the same judgment. Deliberately using a market-implied probability as a baseline can be legitimate if documented; mechanically copying a price without an articulated rationale is a different process finding. Price agreement is not independent confirmation that the forecast is correct.
How do you preserve the prior and the revision history?
Use two linked, append-only records. An evidence event is information documented independently of whether a probability decision follows. A revision decision records how the trader responded to one or more observed evidence events.
Evidence event record:
Evidence ID: unique evidence identifier
Forecast/contract: identifier and precisely defined event
Description: evidence description and source
Public time: publication time, when known
Observed time: when the trader actually accessed it, if known
Optional evidence fields include dependence notes, an initial evidence assessment, a market-price snapshot, and verification status. The event can exist even when no decision follows.
Revision decision record:
Decision ID: unique decision identifier
Forecast/contract: identifier
Evidence ID(s): one or more referenced events, or "none identified"
Prior: probability and prior forecast reference
Decision: update / no update / deferred
Decision time: timestamp
Updated probability: when applicable
Rationale: contemporaneous explanation
Optional decision fields include evidence-strength category, trigger type, review-threshold flag, and a follow-up deadline for a deferred decision. An update may reference several evidence events, and one evidence event does not require exactly one revision. If revisions are written back over the prior, the review can no longer tell whether you moved late, moved far, or moved at all. These prediction-market-specific records are user-maintained unless a product explicitly provides them.
Use precise findings: documented no update means the evidence was reviewed and the prior was retained; missing required reassessment record means the trader observed relevant information, a predefined logging rule required a decision by a deadline, that deadline passed, and no decision was recorded; confirmed missed reassessment requires additional evidence that the trader did not actually reassess; unknown actual reassessment means the record cannot establish whether private reassessment occurred; unsupported or delayed revision means a numerical change lacks sufficient justification or occurred later than the applicable timing standard. An evidence event without a decision is a candidate for investigation, not automatically a process failure. Unknown exposure and a valid deferral whose deadline has not passed must not be classified as missing.
Writing update conditions before the trade is a related but separate step. Trading decision-making covers naming, in advance, what information would justify reconsidering an action. This article covers the review afterward: whether an update actually followed the conditions you set.
How do you review revisions at a fixed boundary?
- Fix the boundary in advance. A number of resolved contracts, a calendar interval, or the end of a session. Do not choose it after seeing results.
- Review with the outcome hidden if you can. Cover the resolution column or review revisions before looking up how the contract settled. This protects the review from hindsight drift; it is a suggested practice, not an established protocol.
- Separate the review outputs. Record process status as followed documented procedure, documented deviation, or cannot determine. Record evidence-support status as supported by the available record, questionable or insufficiently documented, contradicted by the available evidence, or cannot determine. A procedural deviation can coexist with a defensible update, and a compliant update can still be poorly supported. Missing documentation is not proof of irrationality.
- Grade the four dimensions. Assess direction, magnitude, timing, and rationale using only what was knowable at the time. An evidence event without an associated decision is a candidate for investigation. A missing required decision record requires observed information, an applicable rule, a passed deadline, and no corresponding decision; unknown exposure or an unexpired deferral remains uncertain. A missing evidence entry can be identified only by comparing the log with an independent evidence timeline or other reliable source; if no such source exists, report evidence completeness as unverified. An externally available event is not proof that the trader observed it, so unknown exposure is not a behavioral failure.
- Only then look at outcomes. Predictive validation requires multiple resolved forecasts, a consistent observation convention, an appropriate probabilistic scoring method, and a defined baseline when data exist. Revisions on one contract are not independent outcomes: 30 revisions across two contracts are not 30 independent resolved questions. Keep this brief; calibration owns the outcome-based review.
Sample size matters here exactly as it does for calibration. A handful of revisions can show a habit worth watching but cannot establish that the habit is stable. Treat an early pattern as a reason to keep logging.
A worked example
This example is hypothetical, and it reads contract prices as probabilities while ignoring spread and fees. A trader forecasts whether a central bank will cut its policy rate at the next meeting. Their written convention flags changes exceeding 10 points for weak evidence, 20 for moderate evidence, and 35 for strong evidence; no category requires a minimum change. A single committee member’s speech counts as weak evidence. For review, when a market reaction appears to reflect that same speech and no independent information has been identified, the trader provisionally treats the speech and price reaction as one dependent evidence bundle. That is a documentation convention, not proof that the price contained no additional information. The trader also writes an illustrative personal review rule: “After observing a material inflation release, reassess the forecast and record an update, no-update decision, or deferral by the end of that day.” This is an illustrative, trader-defined documentation deadline, not a universal optimal trading rule, and it does not require the numerical probability to change.
- Day 0. Prior is 35%. Rationale: recent inflation prints were above target and committee language was cautious.
- Day 2. One committee member gives a speech leaning toward a cut. The contract price rises from 33¢ to 48¢. The trader revises from 35% to 60% within the hour. The rationale field reads: “price jumped.”
- Day 6. An inflation release comes in clearly below expectations, evidence the trader would grade as strong. The evidence-event log records its public release time, that the trader accessed it that day, and that it had not already been incorporated. By the end of Day 6, the trader’s rule required a reassessment record, but the deadline passed without any corresponding update, no-update decision, or deferral being recorded. The price rises from 52¢ to 71¢.
- Day 8. The price is 78¢. The trader revises from 60% to 78%. The rationale field reads: “matches market,” without saying whether the trader used the market as an explicit signal or copied it mechanically.
- Resolution. The bank cuts, and the contract resolves YES.
Read by outcome alone, this looks like a success: the final 78% was on the right side. The revision review reads differently, using only what was known at each step.
The Day 2 revision moved 25 points on the provisionally bundled weak evidence, exceeding the trader’s 10-point review threshold. That is a review flag requiring closer inspection—not a binding maximum and not, by itself, a procedural violation. The total incorporated information set is insufficiently documented because the rationale says only “price jumped.” Working the odds backward, (0.60 / 0.40) / (0.35 / 0.65) is approximately 2.786, or roughly 2.8. This describes the probability change’s implied evidential strength under the hypothetical single-bundle interpretation; it is not the measured likelihood ratio of the speech or the market move. The process status and evidence-support status therefore remain review questions, not automatic findings that 60% was wrong.
The Day 6 event is visible because it is in the evidence log, and the example establishes that the trader accessed it, had not already incorporated it, and was required to record a reassessment record. The demonstrable finding is therefore a missing required reassessment record, not proof that the trader failed to reassess privately and not a missed numerical update. A rational no-update decision remained possible under the rule. On Day 8, the trader revises from 60% to 78%, two days after the release, but “matches market” does not establish that the inflation release caused the change. Its relationship to Day 6 remains unclear; the review must distinguish independent re-estimation, an explicitly documented market signal, and mechanical copying.
What the review cannot say is what the mathematically correct number should have been on any day. It can say whether a documented process deviation or missing required decision record occurred, and whether the available information supports the stated rationale. If the contract had resolved NO, none of those process findings would change. Resolution YES does not validate the process; resolution NO would not invalidate it. One sequence on one contract shows a pattern worth watching, not a stable behavioral pattern.
What are common failures in revision review?
Judging the revision by how the contract resolved. Letting the result stand in for the quality of the decision is outcome bias, and it applies to a probability update as much as to a trade. The outcome is not evidence about whether the update followed the information available at the time.
Treating compliance as correctness. A trader can follow a poorly designed rule perfectly. Compliance is evidence about process consistency, not proof that the rule or the revision was normatively correct.
Treating small steps as a universal rule. The tournament finding that frequent, small updaters were most accurate does not override the finding that underreaction is common when signals are precise and the environment is unstable.
How does this differ from confirmation bias and anchoring?
Those are named biases; revision quality is the measurable step that a bias can distort. Confirmation bias concerns how a trader searches for and interprets evidence in ways that favor an existing belief, so it can corrupt which evidence you grade and how you grade it. Anchoring concerns a salient number, such as your first estimate or an entry price, exerting more influence than it deserves. Anchoring may contribute to insufficient adjustment away from an initial probability, but a small revision alone cannot establish anchoring; weak evidence, dependent evidence, appropriate uncertainty, or other updating mechanisms can produce the same pattern. A revision review can identify patterns consistent with insufficient adjustment or problematic evidence selection, but it cannot establish anchoring, confirmation bias, or systematic misweighting from revision magnitude or an incomplete evidence record alone. Both articles cover their own diagnostics.
Where Costante fits
Costante’s relevance here is adjacent and concerns the discipline around a decision, not the probability arithmetic. Costante is built for discretionary traders who want a process observable before, during, and after a decision. A trader can use its session planning to state conditions before a session, its low-friction logging to capture the reason for a decision, and its structured review afterward. The trader still owns the probability, the grading rule for evidence, and every revision.
Costante does not compute Bayesian updates or likelihood ratios, grade evidence, estimate event probabilities, detect overreaction or underreaction automatically, or tell a trader what a revised probability should be. It does not provide a dedicated prediction-market probability record, connect to prediction-market exchanges or brokers, or execute trades. The evidence log, the weight bands, and the review judgments are the trader’s own.
A compact review checklist
- Was the forecast and event definition fixed?
- What new evidence was actually observed?
- Was it already incorporated or dependent on other evidence?
- Was the update, no-update, or deferred decision documented?
- Did direction and magnitude have a defensible rationale?
- Was the decision timely relative to information access?
- Does the record support a process finding without relying on the outcome?
Frequently asked questions
How much should a probability change after new evidence?
There is no universal amount. The size of a justified change depends on prior probability and diagnostic strength. A written grading rule with maximum review thresholds can flag large changes, but it does not determine the correct posterior: weak changes over 10 points, moderate over 20, and strong over 35 are heuristic flags with no required minimum. A justified change can exceed a threshold, and an in-threshold change can still be unsupported.
Is updating toward the market price a bad revision?
Not automatically. The information may be in the market but not in your own forecast, and you may deliberately use the market as a documented signal. Distinguish that from mechanically copying the price without a forecasting rationale. Record the trigger type and whether the information was already incorporated. A price is not necessarily an exact probability because spreads, fees, and market structure matter.
Can a probability revision be good if the contract resolved against it?
Yes. A revision is judged by whether it fit the evidence available at the time, using direction, size, timing, and rationale. A well-founded 70% that fails to happen is consistent with sound updating, just as a poorly founded 70% that succeeds is not evidence of skill.
How is a revision review different from a calibration review?
Calibration compares stated probabilities with realized frequencies across many resolved forecasts. A revision review examines whether each change in a probability followed the evidence, and it does not need any contract to have resolved. The two can disagree: a trader can be well calibrated on final probabilities and still update late or in step with the price.
What if new evidence arrives but my probability should not change?
Record the evidence event and a no-update decision with the reason. The information may be weak, redundant, dependent on evidence already incorporated, or not strong enough to justify a numerical change. A no-update decision is distinct from a deferred decision, which should include a follow-up condition or deadline. An event may be publicly available yet absent from the trader’s log; what may remain unknown is whether the trader observed it, recorded it, incorporated it, reassessed it, or numerically revised. Do not infer irrationality from unknown exposure.
Sources
Costante provides educational workflow tools, not financial advice. Trading involves risk.
Footnotes
-
Massey, C., & Wu, G. (2005). Detecting Regime Shifts: The Causes of Under- and Overreaction. Management Science, 51(6), 932–947. ↩
-
Fischhoff, B., & Beyth, R. (1975). “I knew it would happen”: Remembered probabilities of once-future things. Organizational Behavior and Human Performance, 13(1), 1–16. ↩
-
Griffin, D., & Tversky, A. (1992). The Weighing of Evidence and the Determinants of Confidence. Cognitive Psychology, 24(3), 411–435. ↩
-
Atanasov, P., Witkowski, J., Ungar, L., Mellers, B., & Tetlock, P. (2020). Small Steps to Accuracy: Incremental Belief Updaters Are Better Forecasters. Organizational Behavior and Human Decision Processes, 160, 19–35. ↩