Every rule was significant. Almost none of them mattered.
We tested five confirmation rules on 1.13 million bars. All five passed at p below 0.0001. All five were worth about half a percentage point.
In short
- At 1,130,850 observations every rule we tested cleared p < 0.0001, and every one was worth between a third and half a percentage point. Significance is cheap at this sample size; effect size is not.
- Trade-with-the-trend and stay-on-the-right-side-of-VWAP improved a momentum read and a mean-reversion read alike, which means they are not confirming the call. They are quietly replacing it.
- Our first pass overstated those two by roughly double, because the matching strata did not include the direction of the call and the market's upward drift leaked in through the mix.
- The one effect large enough that a person could plausibly notice it is not on anyone's list of rules, and it is a warning rather than a setup.
Our last measurement ended on a loose thread. Adding the textbook's context requirement to the engulfing candle — the rule that is supposed to separate a real reversal from a coincidence — made the pattern worse, not better.
That was one filter on one pattern. This is the general version: take a plain directional read, bolt on each of the confirmation rules that get repeated everywhere, and see whether any of them earns its place.
Five rules, 1,130,850 decision points (regular-session bars from 11:10 ET onward; why the opening period is excluded is discussed below, and it is a real limitation). Every one of them cleared p < 0.0001. That turned out to be the least interesting thing about the result.
What we tested#
Two base reads, because a rule that only helps one of them is a fact about that read rather than about the rule:
- follow — call the same direction as the decision bar's body
- fade — call the opposite
Neither is a good read. Over this sample follow hits 50.16% and fade 49.84%, and both are behind
simply calling "up" every time, which lands at 51.76% because price drifted up. They are here as
something to attach a rule to, not as candidates.
Then five rules:
| Rule | As it is usually stated |
|---|---|
volume |
The bar trades at least 1.5× its own 20-bar average — "the move needs volume behind it" |
trend |
The call agrees with where price has been over the last ten bars — "trade with the trend" |
vwap |
Up calls only above the session VWAP, down calls only below — the day-trading rule |
break |
The bar closed beyond the previous bar's extreme in the direction you are calling — "wait for confirmation" |
extended |
The bar closed beyond the previous bar's extreme in its own direction, whichever way you are calling — "it has gone too far" |
The last two sound like the same rule and are not, which matters later. break is defined relative
to your call; extended is defined relative to the bar. When you are following the bar they
select exactly the same set. When you are fading it, they come apart completely.
First, the correction#
Each rule is scored as a treatment: occurrences where it fired, against non-firing bars in the same cell, with expected counts summed across strata. The cells are symbol, year, hour of session, volatility tercile — and the direction of the call.
That last one was not in the first run, and leaving it out overstated two of the rules by roughly double. Price drifts up over this sample, so an up call hits more often than a down call for reasons no rule deserves credit for. Three of these five are direction-aware — "trade with the trend", "long above VWAP" — so switching them on changes the up/down mix of the treated set, and the drift walks straight into the result.
Adding call direction to the strata cut the apparent trend effect from +0.71 to +0.34, and vwap
from +0.84 to +0.42. Half of what we first measured was the market drifting, wearing a rule's
name.
The results#
| Rule | Read | Coverage | Observed | Matched expectation | Difference |
|---|---|---|---|---|---|
volume |
follow | 16.8% | 50.56% | 50.07% | +0.48 |
volume |
fade | 16.8% | 49.44% | 49.93% | −0.48 |
trend |
follow | 59.0% | 50.43% | 50.09% | +0.34 |
trend |
fade | 41.0% | 50.23% | 49.90% | +0.33 |
vwap |
follow | 56.1% | 50.50% | 50.08% | +0.42 |
vwap |
fade | 43.9% | 50.27% | 49.87% | +0.40 |
extended |
follow | 52.6% | 50.47% | 50.04% | +0.43 |
extended |
fade | 52.6% | 49.53% | 49.96% | −0.43 |
break |
follow | 52.6% | 50.47% | 50.04% | +0.43 |
break |
fade | 1.7% | 39.46% | 50.23% | −10.78 |
Every row is significant at p < 0.0001. At 1.13 million observations, almost anything is. That is what a large sample buys you: the ability to detect effects, including effects far too small to be worth detecting.
So the useful question is not whether a rule is real. It is how many observations it would take before a person could notice it.
Between 85,000 and 180,000. Very few individual traders will ever accumulate that many comparable observations. This is the same wall we hit measuring the engulfing candle, and the same wall our own statistics panel runs into from the other side — it stays locked until 85 answers because that is what a large effect costs to detect, and these are not large effects.
The rules are not confirming your call. They are replacing it.#
Look again at trend and vwap. Both improved the momentum read and the mean-reversion read.
That is not what we would expect from a filter that merely strengthens the existing call. The two
treated sets are complements — when trend fires for follow it cannot fire for fade — so a rule
that simply made your existing read more reliable should help one and hurt the other, the way
volume and extended do, in exact mirror image.
Helping both means something else is happening. In each case the call that won was the one pointing the way the trend, or the VWAP side, already pointed. The rule is not adding confidence to your read of the bar. It is substituting its own opinion for it, and the arithmetic credits the improvement to the read.
That is worth knowing independently of the size of the effect, because it changes what the advice is. "Trade with the trend" is not a confirmation step you add to a setup. It is a different setup wearing a confirmation's clothes.
Volume is the one rule here whose effect behaves most like the confirmation story people tell about it. It is direction-agnostic, and it mirrors exactly: it makes following the bar better by 0.48 points and fading it worse by 0.48. "The move needs volume behind it" describes what happened in this sample, on this test. It is also worth half a percentage point.
The one effect big enough to see#
There is a single row in that table that is not like the others, and it is not one of the rules anyone teaches.
When you apply "wait for confirmation" to a fade, you select a strange 1.7% of bars: ones that closed beyond the previous bar's extreme in the direction you want to call, while the body went the other way. A bar that pokes above the previous high and closes red, and you call up anyway.
Calling up there hits 39.46%, against a matched expectation of 50.23%. A 10.78 point deficit. You would need about 167 comparable observations to notice that — not 85,000.
We checked it every way it could be an artefact, because a large effect on 1.7% of the data is exactly where concentration hides:
- All 22 symbols show the effect in the same direction, from −2.4 points to −30.3.
- All five years, from −8.7 to −16.0.
- Every hour of the session in scope, from −9.5 to −15.3.
- A day-clustered bootstrap over 1,082 trading days puts the 95% interval at −11.5 to −10.0. Clustered by day rather than by row because 22 correlated names move together, and a row-level interval would be far too narrow.
So it holds up. And we are still not going to dress it up as a setup. It is a negative result — a description of a call that goes badly — measured on direction alone with no stop, no target and no costs, on a definition we wrote ourselves. The honest summary is that the largest thing in this study is a warning about a trade, not an instruction to take one, and it happens to be the one thing here that a person could actually check against their own records.
What this study cannot see#
One limitation is big enough that it belongs in the body rather than a footnote.
The first hundred minutes of every session are missing. The volume rule needs 20 bars of history and the ATR needs 14, so nothing before 11:10 ET is in scope. The opening period is where a large share of day trading happens, and none of these results speak to it. We measured the open separately in the opening-range piece, and it does behave differently.
Beyond that: direction only, at one fixed horizon, close to close. No stop, no target, no position sizing, no costs. A rule that does nothing for direction can still change the distribution of outcomes in a way that matters to a strategy, and that is a different study from this one. Twenty-two US names on 5-minute bars from 2022 to 2026 is not every market, and each rule's parameters — 1.5× volume, ten bars of trend — are choices we made, not laws.
What we take from it#
Three things, in order of how much they changed our own thinking.
Significance is cheap and effect size is not. A p-value answers "is this real?", which stops being the interesting question the moment your sample is large. Every rule here is statistically detectable. Four of them are tiny on this directional test.
Check whether a filter is a filter. If a rule improves opposite reads, it is not refining your judgement — it is overriding it. That test costs nothing and we had never thought to run it.
The baseline has to be matched, and it has to include the thing you are choosing. Half the effect we first measured was drift arriving through the composition of the treated set.
This is also why the question bank has no rule tags in it. We do not label a question "trend continuation" or "volume breakout", because a label implies the category carries information, and where we have been able to measure that, it carries about half a percentage point of it.
SwipeTA is a training game and a simulation: no real money, no broker, and it does not provide investment advice. The parameters behind our measurements are on the methodology page.
Sources#
- SwipeTA confirmation-rule study, measured 2026-08-16: 1,130,850 decision points across 22 US equities and ETFs, 5-minute buckets from cached 1-minute bars, 11:10-16:00 ET, 2022-03-07 to 2026-06-30. Every rule's coverage, observed rate, matched expectation, difference and p-value, plus the day-clustered bootstrap and the per-symbol, per-year and per-hour cuts, come from this run. Script and output: research/confirmation_rules.py and results/confirmation_rules.json in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
- SwipeTA sample-size calculation, run 2026-08-16: the observations needed to distinguish each measured effect from its matched expectation at 80% power and 5% two-sided. Script: research/coin_vs_skill.py in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
- https://www.swipeta.net/methodology
- https://www.swipeta.net/learn/an-ordinary-bar-goes-up-51-percent-of-the-time