Published September 13, 2026

Trading Setup Performance: Measure One Setup Without Hiding the Conditions

Measure one trading setup accurately by separating setup identity, sample eligibility, session and regime conditions, and execution quality before judging its results.


A single trading setup’s results are interpretable only when four things stay separate: what the setup actually is, whether the sample of trades judging it is large and eligible enough to mean anything, what session or market conditions those trades came from, and whether they were executed as planned. Collapse any of these into the others and the same setup can look strong or broken for reasons that have nothing to do with the setup itself — a thin sample read as a trend, a good regime read as a good setup, or a badly executed version of the setup read as the setup failing. This sits one layer below the broader trading performance diagnosis triage: once that framework points at a specific setup rather than discipline, risk, or market context generally, this is the checklist for evaluating that one setup on its own terms.

Define the setup before measuring it

A setup is not “a good-looking trade.” It needs a written entry trigger, the market condition it’s meant to occur in, and the invalidation that ends it — recorded before the trade, not reconstructed afterward from trades that happened to work. Without that definition, “how did this setup perform?” has no fixed target: every review implicitly redefines the setup around whichever trades are being looked at, which makes any expectancy or win-rate figure describe the sample instead of the setup.

Two setups that share an indicator or chart pattern are not the same setup if their entry trigger, context, or invalidation differ. Treat them as separate setups with separate samples until a specific comparison justifies merging them.

The metrics, and what each one actually tells you

MetricWhat it measuresWhat it doesn’t tell you
Win rateShare of trades that closed positiveNothing about size of wins vs. losses — a 70%-win-rate setup can still lose money
Average win / average lossMean size of winning trades and mean size of losing tradesNothing about frequency — pair with win rate, not instead of it
ExpectancyP(win) × average win − P(loss) × |average loss|Whether the sample is large enough to trust that number, or whether execution matched the plan
Profit factorGross winning P&L ÷ absolute value of gross losing P&LSame sample-size and execution caveats as expectancy

Win rate and expectancy answer different questions and neither substitutes for the other. A setup can win often and still have negative expectancy if losses run larger than wins; a setup can win rarely and still be profitable if wins are large relative to losses. Report both, not one framed as if it were the whole answer.

Breakeven trades — outcomes that closed at exactly zero — are neither wins nor losses and contribute nothing to either term above, so P(win) and P(loss) don’t need to sum to 1. Read simply, expectancy is the average outcome per trade in the sample, expressed in a consistent unit such as dollars, points, or R.

Expectancy calculated here describes the setup’s realized results in the sample, not a validated edge. Mistake-adjusted expectancy covers the separate question of what changes when execution deviations are removed from that same figure — useful once execution quality itself is in question, not a replacement for defining the setup and sample correctly first.

Is the sample big enough to say anything?

A setup that has fired only a handful of times has not yet produced a number worth trusting, even if every one of those trades is an unambiguous win. Ordinary variance in a small sample can look identical to a genuine edge or a genuine problem, and adding more observations generally narrows that uncertainty. There’s no universal trade count that makes a sample “enough” — how much evidence is required depends on how variable the setup’s outcomes and payoffs are, how much precision the decision actually needs, whether every trade in the sample was genuinely eligible under the setup’s own definition, and whether the observations reasonably represent the same setup and conditions rather than a mix of them. A thin sample should be reported as provisional, not as a verdict on the setup, and the honest response to “did it work” is often “not enough evidence yet” rather than a premature yes or no.

Watch for a related distortion: repeatedly narrowing the definition after the fact (“this setup, but only on Tuesdays, but only above this size”) until the remaining trades look better. Each narrowing shrinks the sample and increases the chance the resulting figure reflects noise in that subset rather than a real pattern. Decide the eligibility criteria before looking at outcomes, not after.

Segment by session and market condition before judging the setup itself

A setup’s expectancy is not one fixed number — it can shift by session and by market regime, and averaging across very different conditions can hide both a working sub-case and a broken one. This kind of segmentation is most trustworthy when the session or regime being checked was defined and recorded independently of the outcome — decided as part of the setup’s own conditions, not chosen after seeing which slice looks best. Holding time fails that test more than most splits, because a trade’s duration is largely set by which exit it hit; trade duration analysis covers how to compare results by holding time without mistaking stop-outs for a duration effect.

  • By session: the same setup can behave differently in different trading sessions if liquidity, volatility, or typical range differ between them. Split results by session before concluding a setup “doesn’t work,” since it may only be failing in one session’s conditions.
  • By market regime: a trending-market setup tested only through a range-bound stretch (or the reverse) will show results that describe the regime as much as the setup. Where regime is known and recorded, segment by it; where it isn’t reliably classifiable, report that the check couldn’t be run rather than guessing at a regime label after the fact.

Watch the difference between that kind of segmentation and slicing the data until something looks good: if a session or regime split only surfaces after repeatedly narrowing by session, regime, or anything else until one subset looks favorable, treat the result as exploratory rather than confirmed. A positive segment found that way is a hypothesis worth watching, not evidence strong enough by itself to redefine the setup — where practical, check whether it holds on later or otherwise unseen eligible trades before treating it as stable.

This is descriptive segmentation, not statistical proof of a causal effect — checking whether a known condition lines up with a meaningful outcome difference, not isolating why. If a predefined, reliably recorded session or regime segment differs materially from the aggregate and there’s enough evidence to estimate it separately, report that segment on its own rather than letting the aggregate hide the conditional pattern. If the segment was instead discovered by inspecting outcomes first, treat it as exploratory until later or independent observations support it.

What the numbers do and don’t justify

FindingJustifiesDoes not justify
Small or imprecise sampleTreating any expectancy figure as provisional; collecting additional eligible tradesConcluding the setup works or doesn’t
Negative aggregate expectancy, but a predefined, reliably recorded session/regime segment differs materially with enough evidence to estimate it separatelyReporting and evaluating that segment separately from the aggregateClaiming a durable conditional edge — especially if the segment was found by inspecting outcomes first, in which case treat it as exploratory until later or independent observations support it
Positive expectancy across the conditions the setup was defined to trade, with reasonably stable evidenceEvidence is consistent with the current setup definitionAssuming the result is permanent, or that a conditional setup must also work outside the conditions it was defined for
Deviated trades concentrated in this setup (see the aligned-vs-classifiable gap in mistake-adjusted expectancy)Reviewing execution on this setup specifically before judging the setup’s own edgeConcluding the setup itself is the problem before execution is checked

Before trusting any of the rows above, confirm the trades in the sample were actually executed as planned. A setup that looks weak on paper can simply be a setup that was rarely traded as designed — that’s an execution-adherence question, not a setup-performance one, and the execution-quality scorecard is the right tool for classifying which trades were aligned, deviated, or unclassified before this checklist’s numbers are trusted.

A setup that fails this checklist hasn’t been proven to fail — it means the current evidence doesn’t yet support the claim being made about it. Sometimes that’s “needs more trades,” sometimes it’s “underperforms outside the conditions it was defined for,” and sometimes the evidence genuinely doesn’t support the setup as currently defined. Which of those applies depends on which check actually failed; what to do with that finding is a separate decision for the trader to make.

Where Costante fits

Costante’s trade logging records the setup label, session, and outcome for each logged trade, which is the raw material this kind of setup-level review needs. Costante does not calculate expectancy or win rate for you, does not grade setup quality (there is no built-in “A+ setup” score), does not decide whether a setup’s sample is large enough, and does not classify market regime automatically. The trader defines the setup, runs the segmentation, and interprets the result; Costante keeps the underlying records consistent enough that the comparison is possible.

Frequently asked questions

Does a higher win rate mean a better setup?

Not by itself. Win rate and expectancy answer different questions, and a high-win-rate setup can still have negative expectancy if losses run larger than wins. Always report expectancy (or profit factor) alongside win rate, not as a substitute for it.

Can setup quality be graded automatically, like an A+ setup score?

A single automated grade can obscure sample size, session, regime, and execution unless the scoring model explicitly represents each of them — collapsing those distinctions is what makes an automated score misleading, not automation itself. Defining the setup, judging trade eligibility, classifying context, and interpreting the result still require trader judgment. As a product fact: Costante does not generate an automated setup grade or “A+ setup” score.

How many trades does a setup need before I trust its expectancy?

There’s no universal trade-count threshold. How much evidence is enough depends on how variable the setup’s outcomes and payoffs are, how much precision the decision requires, and whether the trades in the sample genuinely represent the same setup and conditions. Treat an estimate from a very small sample as provisional, and keep collecting eligible trades under the same definition rather than redefining the setup to make an early result look more decisive. Trading data statistical reliability covers the underlying sample-size and multiple-comparisons reasoning in more depth.

My setup’s overall numbers are bad, but I think it works in one specific condition. How do I check that?

Segment the existing sample by a meaningful, reliably recorded condition such as session or market regime, and compare that segment’s result with the aggregate. Then check whether you defined that condition before looking at outcomes or found it by slicing the data until a subset looked good: a predefined segment with enough evidence to estimate separately is worth reporting on its own, while a segment discovered after inspecting outcomes is exploratory until it holds up on later or otherwise unseen eligible trades. If the segment is still too small or uncertain to judge either way, the honest answer is that the question isn’t settled yet.

Costante provides educational workflow tools, not financial advice. Trading involves risk.