How many calls before your record means anything

If you are right 60% of the time, it takes about 194 calls to show that is not luck. At 55% it takes 783. Here is the arithmetic, and why your own record is harder.

In short

  • Against a 50/50 reference, at 80% power and 5% two-sided, distinguishing a genuine 60% caller from chance takes about 194 calls. A 65% caller takes 85. A 55% caller takes 783.
  • The cost rises with the square of how small the edge is. The cases people actually argue about sit nearer 55-60% than 70%, and those smaller deviations are the expensive ones to establish.
  • Comparing two situations is a different and dearer test - telling a ten-point gap apart takes roughly 392 calls on each side, not 392 in total.
  • A personal trading record is harder than any of this. You chose which setups entered it, the market offers no 50/50 reference, and no number of trades repairs either problem.

If you are right 60% of the time, it takes roughly 194 calls before that can be told apart from luck. At 55% it takes about 783. At 65% it takes 85. Those figures are for a one-sample proportion test against a 50/50 reference, at 80% power and a 5% two-sided significance level.

Everything below is why they are shaped like that, what "194" does and does not mean, and why a personal trading record is a harder case than any of them.

Quick answers#

Each line is complete on its own, because these get read one at a time.

How many calls does it take?#

Against a reference that is genuinely 50/50 — a balanced question bank, a coin, anything with no built-in lean — this is what it costs to establish that you are better than chance, at 80% power and a 5% two-sided significance level throughout:

If you are truly… Calls needed
70% 47
65% 85
60% 194
55% 783
52% 4,904
Answers needed before a directional lean is real
A 70/30 lean 47
A 65/35 lean 85
A 60/40 lean 194
A 55/45 lean 783
How many answers does it take to detect a directional bias? At 80% statistical power and a 5% two-sided significance level, approximately 85 answers are needed to detect a 65/35 lean, 194 for 60/40, and 783 for 55/45 — measured against the question bank's deliberate 50/50 split. The highlighted row is the threshold SwipeTA uses before it will show a direction bias at all; below it the panel stays locked. Calculated for this piece on 2026-08-10, reproducing the figures in the comment at the top of apps/mobile/src/services/statsRepo.ts.

What does "194 calls" actually mean?#

It does not mean that 194 calls prove you are a 60% caller. It means that if your true rate were 60%, a one-sample test against a 50/50 reference at 80% power and 5% two-sided would have about a four-in-five chance of distinguishing it from chance. The number describes the sensitivity of the measurement, not a verdict about the person being measured.

Two consequences follow, and both get lost when the figure travels alone.

If you reach 194 calls and the test comes back inconclusive, that is not proof you have no edge — an 80%-powered test misses a real effect one time in five. And if you reach 194 calls at 61%, that is not confirmation you are a 60% caller; it is consistent with a range of true rates, of which 60% is one.

Why is 55% so much more expensive than 65%?#

Because the cost scales with the square of how small the edge is. Halving the edge quadruples the sample. Going from a fifteen-point edge to a five-point edge is a factor of nine.

That is the uncomfortable part, because the sizes people actually argue about are the expensive ones. Claims around 70% are rare; questions about whether someone is at 55% are common, and 55% against a 50/50 reference is a 783-call question.

What if I want to compare two situations?#

Different test, and dearer. "Am I better on longs than shorts?" is not one proportion against a fixed number — it is two estimates, each carrying its own error.

Telling a genuine ten-point gap apart takes roughly 392 calls on each side, not 392 in total, on a two-sample proportion test at 80% power and 5% two-sided. On the same terms a twenty-point gap takes 97 a side, and a five-point gap takes 1,569 a side.

⚠️ These are easy to confuse with the numbers above and they are not interchangeable. The one-sample figure of 783 and the per-side figure of 392 happen to land near each other for these inputs. That is a coincidence of the arithmetic, not a relationship.

Can twenty answers tell me anything?#

Not about skill, in most cases.

A caller with no skill at all, answering a balanced 50/50 set at random, scores 60% or better in one twenty-answer session out of four. Over a hundred such answers they will hit a five-in-a-row streak 81% of the time, and a seven-in-a-row streak 32% of the time. Those are exact binomial and longest-run probabilities, not estimates.

How often a player with no skill scores 60% or better
In a 20-answer session 25.17%
Over 50 answers 10.13%
Over 100 answers 2.84%
Exact binomial probability that a player answering at random scores 60% or better, by how many answers they have given. A single good session is not evidence of skill: one session in four reaches 60% by chance alone. Volume is the only thing that separates the two. Calculated for this piece on 2026-08-11; script and output at research/coin_vs_skill.py.

So a good week is usually not enough evidence to separate skill from luck at these sample sizes. It is close to what a coin does when a coin is watched for a week, and the only thing that separates the two is volume — which is a design problem before it is a statistical one.

Does this apply to my real trading record?#

Yes, and it is harder there. Three reasons, in ascending order of how much they cost you.

There is no 50/50 reference in a market. Every number above is measured against a balanced baseline. Real prices drift: across 22 US equities and ETFs on 5-minute bars from 2022 to 2026, an ordinary bar is followed by a higher close fifteen minutes later 51.73% of the time. Being right 52% of the time in that universe is therefore not two points of skill, it is approximately none — and distinguishing a genuine one-point deviation from a 51.73% baseline, at 80% power and 5% two-sided, takes about 19,600 observations.

You chose the trades in it. A journal contains the setups you had the confidence to enter, so every comparison inside it is between two groups you selected. That is a different problem from sample size, and no number of trades fixes it — a thousand self-selected trades are still self-selected. Sample size and sample selection are separate failures, and collecting more data only addresses the first.

Your trades are not independent of one another. These formulas assume every call is a fresh coin. Positions taken the same morning, on correlated names, in one market regime are not. Where calls are correlated the effective sample is smaller than the count, so every figure on this page is the optimistic end of a range.

What we do with these numbers#

They are not decoration. They are wired into the app, and they are the reason it says less than it could.

The statistics panel will not report a directional lean until 85 decided answers, because 85 is where a 65/35 lean becomes detectable on the test described above, and anything earlier would be reporting noise as insight. A situation slice is not drawn at all below 12 answers — below that the interval spans most of the axis and says nothing. The confidence label steps at 85 and 400. Until those thresholds the panel shows a countdown rather than a figure it cannot stand behind.

A number at answer 20 would be more engaging. On the arithmetic above, it would also be a coin.

What this does not say#

Not that reaching any of these thresholds transfers to profitability: being right more often than not says nothing about position sizing, exits or costs, and those decide outcomes. Not that a short record is worthless — only that it is not yet evidence about skill. And not that these figures hold outside their assumptions; every one of them is a one- or two-sample proportion test at 80% power and 5% two-sided, and each assumes calls are independent, which real trading records are not.

The single practical consequence is this: before asking whether a record is good, ask whether it is long enough to be readable at all. Many are not, and knowing which case you are in is the cheapest thing on this page.

SwipeTA is a training game and a simulation: no real money, no broker, and it does not provide investment advice. The parameters behind our measurements are on the methodology page.

Sources#

  • SwipeTA coin-versus-skill calculation: one-sample proportion test against p0 = 0.5 at 80% power and 5% two-sided (47 / 85 / 194 / 783 / 4,904 answers for a 70 / 65 / 60 / 55 / 52 percent caller), the two-sample equivalent for separating two situations (97 / 392 / 1,569 per side for a 20 / 10 / 5 point gap), and exact binomial and longest-run probabilities for a player answering at random. Script and output: research/coin_vs_skill.py and results/coin_vs_skill.json in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
  • SwipeTA measurement of the drifting baseline: an ordinary 5-minute bar across 22 US equities and ETFs is followed by a higher close 15 minutes later 51.73% of the time, and detecting a one-point deviation from that baseline needs about 19,600 observations. Scripts: research/engulfing_definitions.py and research/coin_vs_skill.py in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
  • SwipeTA app source: the statistics panel stays locked until 85 decided answers (BIAS_UNLOCK), a situation slice is not drawn below 12 answers (SLICE_MIN), and the confidence label steps at 85 and 400. apps/mobile/src/services/statsRepo.ts, read 2026-08-17.
  • https://www.swipeta.net/methodology