Decision Load and Late-Session Trading Execution: How to Test the Interaction
Test whether accumulated decision load and elapsed session time interact to degrade trading execution, how to screen for the pattern, and what the evidence can and cannot establish.
A count × time interaction is interaction-consistent when the relationship between cumulative decision load and execution quality is materially different later in the session than earlier in the session, after obvious trigger and context differences are accounted for. That is different from either variable acting alone: a high decision count reached early in a session, a long session with few decisions, and a high count reached late in a session are three different states, and only comparing across all of them — not looking at the high-count/late-time state in isolation — can tell them apart. The trading-specific evidence cited here does not establish that this interaction exists; the sources reviewed consist of adjacent theory, vigilance research, mixed decision-fatigue evidence, and trader-specific evidence for alternative mechanisms such as post-loss risk shifts — not direct evidence for the proposed interaction — plus whatever pattern a trader can find in their own session history. This article treats the interaction as a hypothesis to screen for, distinguishes a coarse practical heuristic from a stronger decision-level analysis, and is explicit about which is which.
What is the decision-load × elapsed-time interaction hypothesis?
The broader decision-fatigue literature uses several proxies for accumulated load — decision order, number of prior decisions, time into a work period, time of day — and the empirical evidence behind them is mixed. In this article, accumulated decision count is used as one operationalization of decision load for the trading-specific diagnostic framework, checked against the same decision-making checkpoints that framework reviews. Late-session timing is a separate variable: how much session time has elapsed, independent of how many decisions produced it. Each can move without the other — a trader can rack up a dozen decisions in the first twenty minutes of a volatile open, or sit through four quiet hours with two setups.
This is one of several distinct mechanisms that can produce late-session performance decline; the question this article covers is narrower: does the relationship between accumulated decision count and execution quality change as elapsed session time increases — equivalently, does the relationship between elapsed time and execution quality differ at different levels of accumulated decision count? An interaction is not the same claim as “decision count is high and it’s late” — that combined state can show up for reasons that have nothing to do with count and time compounding each other, including several confounders described below. Interaction-consistent evidence requires comparing count and time across each other, not describing one session that happens to sit in the high-count/late-time corner.
Redefining the interaction correctly
Put in plain language: an interaction is present when accumulated decision load matters more (or less) for execution quality at some elapsed-time levels than at others. If decision count is associated with the same degree of execution decline whether it’s reached in the first hour or the fourth, there’s no interaction — the pattern is more consistent with a decision-load-dominant association. If elapsed time is associated with the same degree of decline regardless of how many decisions were made along the way, the pattern is more consistent with a duration-dominant association. An interaction is the specific claim that these two do not act independently — and, like the associations that make it up, it is itself an observational pattern, not a demonstration of which variable causally produced the decline.
A simple way to represent this is a four-condition comparison, not a single high-count/late-time observation:
| Earlier / shorter elapsed time | Later / longer elapsed time | |
|---|---|---|
| Lower cumulative decision load | baseline condition | duration-heavy condition |
| Higher cumulative decision load | decision-load-heavy condition | combined high-load/high-duration condition |
The bottom-right cell by itself is not sufficient evidence of an interaction — a session can land there simply because a busy, long session is more common than a busy, short one. The relevant question is whether the deterioration associated with higher decision count is materially different at later elapsed times than at earlier elapsed times, after accounting for plausible confounders. That requires seeing all four cells, or a comparable stand-in for them, not just the combined one.
Why an interaction is theoretically plausible
Two lines of research make a shared-capacity interaction theoretically plausible without demonstrating that one exists in trading.
Hockey’s compensatory-control framework proposes that performance can initially be protected under high demand through additional effort, while costs emerge elsewhere — in strain, fatigue, subsidiary activity, less efficient strategies, or delayed effects — rather than showing up immediately in the visible output.1 Applied cautiously to a trading session, this is a theoretical reason not to assume that process deterioration under combined load and duration must appear immediately or follow a smooth, linear path. It is not a demonstrated finding that a trading session’s decline actually follows this shape, and it does not by itself imply that decision count and elapsed time must interact rather than simply add.
Separately, a review of vigilance research found that sustained-attention tasks impose a substantial subjective and cognitive workload on their own, and that workload/performance consequences are sensitive to processing demand — event rate, discrimination difficulty, and task duration.2 That is vigilance-task evidence, not trading-specific evidence, and it does not establish the shape of an interaction in a trading session. It does make a shared-capacity interaction theoretically plausible — if sustained attention already draws heavily on a limited resource, a high rate of decisions on top of it could plausibly draw on that same strained capacity — but it does not establish that such an interaction exists in trading, or that it is super-additive.
Decision-fatigue evidence: what current research supports
The broader decision-fatigue literature this article draws on is less settled than a single mechanism story suggests. A 2025 systematic review examined decision fatigue in healthcare professionals’ medical decision-making and found that operationalizations of decision load varied widely across studies — some used decision order, some used caseload, some used time-on-shift — and that support for a decision-fatigue effect across those studies was mixed rather than consistent.3 That review is healthcare evidence, not trading evidence, and its operationalizations do not map cleanly onto a trading session; it is cited here to show that even within a single, well-studied professional domain, the underlying construct is not settled science, not as evidence that traders experience the same pattern.
A separate large-scale, preregistered field study using real-world healthcare data found no evidence of decision fatigue in that setting.4 It reinforces the same point from a different angle: current evidence for decision fatigue as a general construct is genuinely mixed, not merely under-studied in trading specifically. Neither source should be read as transferable evidence about trading execution — they establish that the broader construct this article borrows a proxy from is itself contested, which is a reason to hold the trading-specific interaction hypothesis loosely rather than treat it as inherited support from adjacent fields.
Distinguishing the interaction from other patterns
The table below separates patterns that can look similar from the outside but require different evidence and different responses.
| Pattern | Decision count | Elapsed session time | What discriminates it |
|---|---|---|---|
| Decision-load-dominant pattern | High | Not necessarily late | Decline tracks decision count even in a compressed, short session, and does not grow disproportionately at later elapsed times |
| Duration-dominant pattern | Not necessarily high | Late | Decline (if present) appears with few decisions made, from elapsed time alone, and does not grow disproportionately at higher decision counts |
| Interaction-consistent pattern (this article) | High | Late | The count–execution relationship differs materially across elapsed-time conditions, or the time–execution relationship differs materially across decision-count conditions — not simply that both happen to be high in one observation |
| Loss-triggered shift (tilt) | Irrelevant | Irrelevant | Change follows a specific loss or triggering event, not accumulated count or duration |
| Documentation-only deterioration | Irrelevant | Often late | Record completeness declines while the underlying trade decisions show no corresponding change in rule adherence — a logging-friction pattern, not an execution-quality pattern |
The fourth row matters because a real, trading-specific finding can look like this pattern from the outside without being it. A study of Chicago Board of Trade proprietary futures traders found that traders who had lost money in the morning session took on above-average risk in the afternoon, consistent with loss-averse “recovery” behavior rather than gradual accumulated-load decline.5 That is genuine evidence of a within-session behavioral shift in real traders — but it is scoped to a discrete morning loss driving afternoon risk-taking, not to decision count and elapsed time compounding in the absence of a triggering loss. The study does not establish a fixed duration for that association, so the post-trigger window used to handle this confound should be defined by the analyst and reviewed separately; see the loss-trigger handling below.
The fifth row matters because the most commonly available session metric — completed-record rate — is a process-adherence proxy, not a direct measure of trading execution. The next section covers why that distinction matters and what to track instead where possible.
The outcome-measurement problem: documentation is not execution
A decline in the share of decisions with a fully completed record is easy to measure and is used throughout the decision-fatigue framework this article builds on. But record completion is a process-adherence or documentation proxy, not automatically a measure of trading execution quality. A trader can log a thinner record late in a session for reasons that have nothing to do with worse trading: documentation fatigue, lower willingness to journal in the moment, time pressure to close out the session, or general logging friction. None of those prove the trade itself was executed worse.
Where possible, distinguish two kinds of evidence:
- Documentation-adherence evidence: whether the record was completed — the proxy this article and its sibling framework rely on most often, useful as a screening signal but not proof of execution decline on its own.
- Execution-quality evidence: whether the trade itself followed the trader’s predefined rules — entry-condition adherence, position-size adherence, stop or exit-rule adherence, re-entry-rule adherence, session-cutoff adherence, or adherence to a pre-defined execution checklist, and more broadly planned-versus-actual execution and rule adherence generally.
Costante supports structured trade logging and review of recorded trading behavior. Where a trader has captured rule-adherence or execution-related evidence in that workflow, those records can provide inputs for comparing patterns across sessions rather than reconstructing them entirely from memory afterward.
How to screen for the interaction in your own history
A coarse comparison can point toward the pattern, but it should not be mistaken for identifying it. Dividing sessions into high-count and low-count groups and comparing first-third-versus-last-third completion-rate decline is a practical screening heuristic, not a causal or statistical interaction estimate. It does not control for the fact that total session decision count is often correlated with volatility, market regime, setup frequency, opportunity frequency, decision density, session length, time of day, prior P&L, loss events, strategy type, or unusual news and event sessions. Any of those can produce a pattern that looks like a count × time interaction without one being present.
What to record
A stronger screen starts at the decision level, not the session level. For every decision or eligible execution event, record:
- cumulative decision number reached at that point in the session;
- elapsed minutes from session start;
- decision timestamp;
- whether a known triggering loss or event occurred beforehand;
- the relevant execution-quality outcome for that decision (not only whether the record was completed);
- session identifier.
Level 1 — practical screening
With decision-level records, compare materially similar observations across the four count × time states from the table above, rather than only comparing session-level group averages. Prefer comparisons such as:
- a similar cumulative decision count reached at different elapsed times across different sessions;
- a similar elapsed time reached with different cumulative decision counts across different sessions.
The pattern should repeat across multiple comparable sessions before it counts as screening evidence, and any decision inside a flagged post-trigger window (see below) should be set aside or reviewed separately rather than folded into either comparison.
Level 2 — stronger interaction analysis
A more formal decision-level model would estimate something conceptually like execution quality as a function of cumulative decision count, elapsed time, their interaction, and session or context controls. This does not automatically establish causality — it is still an observational estimate, sensitive to which controls are included and how the outcome is defined.
Repeated decisions within the same session are not independent observations, so a serious version of this analysis should account for session-level clustering, fixed effects, random effects, or another repeated-measures structure rather than treating every decision as if it came from a separate, unrelated session. If the outcome is binary — adhered to the rule or did not — an interaction term in a logistic model does not translate directly into a simple additive probability effect; that detail matters for anyone actually running the analysis, but the article stays at the conceptual level here rather than working through the statistics.
Most traders will only ever run Level 1. Level 2 is described so the practical heuristic is not mistaken for something more rigorous than it is.
Fixing the loss-trigger confound
Coval and Shumway’s finding — Chicago Board of Trade proprietary traders taking above-average afternoon risk after a morning loss — is evidence that a prior trading outcome can be associated with behavior later in the same trading day.5 The study does not demonstrate a fixed post-loss duration, persistence for a specific number of decisions, decision fatigue, or a count × time interaction. Treating a broader post-trigger window as worth flagging, rather than excluding only the single next decision, is an analyst-defined choice for handling this confound, not a duration estimated by Coval and Shumway.
Instead:
- flag the post-trigger period following a known loss or event, not only the single next decision;
- exclude that flagged window from the count × time comparison, or analyze it separately rather than folding it into either group;
- control or stratify by prior same-session P&L where the record allows it;
- test whether the apparent count × time pattern still holds in sessions, or session segments, that contain no identifiable trigger at all.
A pattern that only appears in sessions with a triggering loss is more consistent with the loss-triggered-shift row of the table above than with the count × time interaction this article is testing for.
Worked example (hypothetical)
The following uses illustrative, hypothetical numbers to demonstrate the logic of an interaction — it is not a report of observed trading data, and none of the figures are thresholds. Rule-adherence rate is used as the outcome, compared across four conditions built from many comparable decision segments rather than two individual sessions:
| Condition | Illustrative rule-adherence rate |
|---|---|
| Low count, early elapsed time | 92% |
| High count, early elapsed time | 90% |
| Low count, late elapsed time | 86% |
| High count, late elapsed time | 68% |
This comparison is shown on an illustrative additive percentage-point scale, not a fitted statistical model. Read across the rows: at early elapsed time, a higher decision count is associated with almost no change in the adherence rate (92% → 90%, a −2 percentage-point difference). Read down the low-count column: elapsed time alone is associated with a modest decline (92% → 86%, a −6 percentage-point difference). Under a simple additive percentage-point comparison, combining those two differences gives a no-interaction reference of roughly 84% (92% − 2pp − 6pp) for the high-count, late-elapsed-time cell. The illustrative value shown for that cell, 68%, is substantially below that reference. This does not imply that an additive probability scale is the uniquely correct scale for modeling interaction — an actual analysis run as a logistic model on binary rule-adherence outcomes would need its interaction term interpreted on that model’s own scale, not read off as a simple percentage-point subtraction. The shape is interaction-consistent, not proof of a causal interaction — it would need to hold across many comparable sessions, with post-trigger windows excluded and plausible confounders checked, before it supports more than a working hypothesis for that trader.
A single session with stable performance throughout — or a single session showing decline — does not establish or rule out a count × time interaction on its own; the comparison only means something once it is repeated across multiple comparable conditions.
Decision response: if the interaction is present
A session shutdown boundary based only on decision count or elapsed time does not explicitly represent a count × time interaction. If repeated, comparable observations consistently show the combined state is associated with worse execution, a joint boundary may be a more targeted risk-control rule than two independent limits — a count ceiling that tightens once elapsed session time passes a point the trader’s own history has shown to matter, rather than two limits that each have to be hit on their own. For illustration, not as a universal threshold: a trader whose own review repeatedly shows the pattern starting around decision 10 in sessions running past three hours might test a joint rule that conditions the allowable count on elapsed time, rather than assuming either “stop after 15 decisions” or “stop after four hours” alone represents the interaction.
This is a boundary-setting input, not a proven treatment. An interaction-consistent observational pattern can justify testing a joint boundary going forward — it does not prove the boundary will improve trading performance, only that it targets the specific combination the trader’s own history flagged. The trader still decides, in advance, what the joint threshold is and what happens when it is reached. The full decision rule for converting confirmed fatigue evidence into a shutdown trigger — including when the evidence is not yet sufficient to change anything — is a separate question from establishing the interaction itself. If the observed deterioration is broader than this count-by-time pattern, trading performance under pressure provides the wider planned-versus-actual execution review. And if the decline does not reset with a normal night’s rest — persisting across many sessions rather than within one — that is a different, cross-session pattern; see trading burnout for how to recognize it and set a recovery boundary.
Review and prevention
Feed cumulative decision count and elapsed session time into review as a pair, tagged with rule-adherence outcomes and any post-loss flag, not as disconnected session-level totals. A review that logs “12 decisions” and “3.5 hours” without linking them to timestamped, decision-level outcomes cannot later show whether a bad late decision belongs to the interaction, to decision count alone, to duration alone, to a loss trigger, or to documentation friction rather than an execution problem at all. Over several sessions, this record is also what lets a trader revisit and tighten the boundary in the previous section as more evidence accumulates, instead of setting it once from a guess and leaving it fixed.
Common failure modes
Treating the high-count/late-time cell alone as evidence. A single session with a high decision count late in the day is one point in one corner of the four-condition comparison. Without the other three corners — or a comparable stand-in for them — it cannot distinguish an interaction from decision-load-dominant or duration-dominant decline.
Reading completion-rate decline as execution decline. A thinner record late in a session can reflect documentation fatigue rather than worse trading. Check rule-adherence or execution-outcome evidence where it exists before attributing a completion-rate drop to the interaction.
Excluding only the single decision after a loss. Coval and Shumway do not identify a fixed post-loss duration. Define a post-trigger window as an analyst-defined choice, then exclude or separately analyze it.
Setting only one type of boundary after seeing an interaction-consistent pattern. A single-variable boundary can still reduce exposure to the combined state, but it does not explicitly condition the allowable level of one variable on the other. If repeated review shows an interaction-consistent pattern, a joint rule may target that specific combination more directly.
Treating a screening comparison as a statistical finding. The grouped, first-third-versus-last-third comparison is a starting heuristic. It does not control for volatility, regime, setup frequency, or the other confounders listed above; a repeated, decision-level comparison across matched conditions is the stronger version.
Frequently asked questions
Is this the same as trading decision fatigue?
Not exactly. This article operationalizes decision load partly through cumulative decision count, then asks whether its association with execution changes as elapsed session time increases. Decision-fatigue research uses broader and inconsistent operational definitions, so the count × time pattern here should be treated as a specific diagnostic hypothesis rather than the definition of decision fatigue itself.
How is this different from a loss-triggered risk shift?
A loss-triggered shift, like the afternoon risk-taking documented in Chicago Board of Trade proprietary traders after a morning loss, follows a specific triggering event and can appear at any point in a session.5 The interaction covered here is defined by the absence of a specific trigger — it is the hypothesized joint pattern of accumulated decision count and elapsed time, tested only after a discrete trigger has been ruled out or its window excluded.
Does this mean every long, decision-heavy session degrades?
No. The interaction is a pattern to screen for, not a guaranteed outcome. Some sessions with high count and long duration will not show a materially different count-execution relationship than shorter or lighter sessions; the screening process in this article exists specifically to tell those sessions apart from the ones that do.
Can a single time limit or a single decision-count limit prevent this?
Either limit can reduce exposure to the combined state, but neither explicitly conditions one boundary on the level of the other. A joint rule is more targeted when repeated review suggests that the relevant association depends on the combination — set from that trader’s own screening history, not assumed in advance.
Where Costante fits
Costante supports structured trade logging and review of recorded trading behavior. Where a trader has captured rule-adherence or execution-related evidence in that workflow, those records can provide inputs for comparing patterns across sessions rather than reconstructing them entirely from memory afterward.
Costante does not record a formal decision-by-decision timestamp log, calculate or flag a count × time interaction automatically, diagnose its cause, or enforce a joint session boundary. It does not generate a strategy or guarantee that adjusting a session boundary will improve results. Identifying the pattern in review, and deciding what boundary to set from it, remains the trader’s responsibility.
Sources
Costante provides educational workflow tools, not financial advice. Trading involves risk.
For the broader process framework connecting late-session decisions to rule adherence, see trading discipline.
Footnotes
-
Hockey, G. R. J. (1997). Compensatory control in the regulation of human performance under stress and high workload: A cognitive-energetical framework. Biological Psychology, 45(1–3), 73–93. ↩
-
Warm, J. S., Parasuraman, R., & Matthews, G. (2008). Vigilance requires hard mental work and is stressful. Human Factors, 50(3), 433–441. ↩
-
Maier, M., Powell, D., Murchie, P., & Allan, J. L. (2025). Systematic review of the effects of decision fatigue in healthcare professionals on medical decision-making. Health Psychology Review, 19(4), 717–762. ↩
-
Andersson, D., Lindberg, M., Tinghög, G., & Persson, E. (2025). No evidence for decision fatigue using large-scale field data from healthcare. Communications Psychology, 3, Article 33. ↩
-
Coval, J. D., & Shumway, T. (2005). Do behavioral biases affect prices? The Journal of Finance, 60(1), 1–34. ↩ ↩2 ↩3