The study that would settle whether this works, and why we cannot run it
We have written eight times that the transfer question is open, and never once what would close it. Here is the protocol, the arithmetic, and why we cannot run it.
In short
- The testable claim is narrow: does practising on a dealt question bank improve directional reading of charts the reader has never seen, against a control group, on a test set not selected on its own outcome.
- The arithmetic is unforgiving. Separating a 5-point improvement between two groups needs about 1,569 observations per side, a 3-point one about 4,360, before correcting for answers clustering within a person.
- We cannot answer it from our existing users or question bank. Self-selected installs cannot supply a valid control group, and a bank selected on outcomes cannot serve as its own test set.
- Until someone runs it, the honest status is unknown rather than promising. An unrun study is not evidence, and neither is the fact that we designed one.
Across eight of the articles on this site there is some version of the same sentence: the transfer question is open, and no study we know of settles it.
That is honest as far as it goes. It is also cheap, because we have never once written down what would settle it. An open question with no stated test is not a caveat — it is a permanent excuse.
So here is the protocol, the arithmetic it would need, and the reasons it cannot be answered from the data this product already produces.
What claim would actually be tested?#
The claim is not a financial one. Nothing here measures money, and a study of that would be a different discipline with different obligations.
The narrow, testable claim is this:
Does practice on a dealt question bank improve a person's directional reading of charts they have never seen — measured against a control group, on a test set that was not selected on its own outcome?
Every clause in that sentence is doing work. Dealt rather than chosen. Charts they have never seen. Against a control group. And the last clause is the one that matters most, for a reason that is inconvenient for us.
The protocol#
Pre-registered. The hypothesis, the sample size, the outcome measure and the analysis written down and time-stamped before anyone collects anything. Otherwise the study joins the literature described in the reading list, where in-sample results were real and did not persist.
Randomised, with an active control. Not practice-versus-nothing, which measures enthusiasm. Practice against a group spending equal time on a plausible alternative — reading about chart patterns, say, or the same charts without the immediate answer. The comparison people care about is against the next-best use of the same hour.
A held-out test set built differently from the training set. ⚠️ This is where our own architecture disqualifies us. Our question bank keeps a decision point only when the next fifteen minutes moved between one and three times ATR, and it skips forward after each accepted question — two selection steps applied with knowledge of the outcome. A test set built the same way would measure whether people got better at our selection, not at reading charts. The test set would have to be sampled without reference to what happened next, which means most of it would be undramatic, and scores would be lower and noisier for everyone.
Blind marking, fixed horizon, no discretion. The outcome is what price did. That part is easy.
A pre-committed publication rule. The result gets published whichever way it lands, and the protocol says so before the data exists.
How big would it have to be?#
This is where the design stops being an aspiration.
Comparing two groups is a two-sample problem, and the sample sizes are brutal. At 80% power and 5% two-sided:
| Improvement to detect | Observations needed per group |
|---|---|
| 10 percentage points | 392 |
| 5 points | 1,569 |
| 3 points | 4,360 |
| 2 points | 9,810 |
⚠️ This is not a participant count, and the difference is not a detail. Those are observations, and the calculation assumes each one is independent. In a between-group study answers cluster within a person: one participant's hundred answers carry far less information than a hundred participants' one answer each. Turning 1,569 observations into a number of people requires the within-person correlation and the number of test charts each person answers, and we have estimated neither, so we cannot honestly convert the table into "X people per group". Read it as a floor on information, not as a recruitment target.
And the effect size worth planning for is small. Everything we have measured at this horizon lands in the low single digits: five confirmation rules each worth about half a percentage point, a market baseline that drifts by 1.73 points before anyone reads anything. A training effect of ten points would be extraordinary. A study powered only for ten points would therefore be a study designed to find nothing.
Why we cannot run it#
⚠️ First, what "cannot" means here, because the heading overstates it. The study is possible. It is not possible as a by-product of running the app: it would have to be built as a study, with recruited participants, random assignment and a purpose-built test set. What follows is why ordinary SwipeTA usage data cannot answer the question, not a claim that the question is unanswerable.
Three reasons, and none of them is budget.
Our existing users cannot supply the control group. Everyone who installs SwipeTA chose to install SwipeTA, and everyone who stays chose to stay. Comparing engaged players with lapsed ones measures who enjoys the product. A real control means randomly assigning recruited participants to something other than the thing being tested — which is a study, not an analytics query.
Our test set is contaminated by construction. Covered above: the bank is selected on outcomes. Building a clean held-out set means building a second question bank on different principles, using it for nothing else, and never letting it leak into the product. That is possible. It is not something we can do incidentally.
We should not be the only party anyone has to trust. We could run the first trial. What we should not do is be the sole source of the only result — a party with a commercial interest running the single trial is the arrangement the data-snooping literature exists to be suspicious of. That is fixable rather than fatal: pre-registration, a raw data release, and ideally an independent replication.
There is a fourth, smaller reason: we do not use player data for questions like this. The statistics the app shows a player are for that player.
What we can run, and what it would not show#
We are not helpless, but the honest labels matter.
We can measure whether a player's hit rate rises with volume inside our own bank. That is a within-subject trend on a selected sample, and it is confounded three ways: practice effects on the format rather than the skill, survivorship — the players still answering at 500 questions are not a random subset of those who answered ten — and regression to the mean.
We can report the distribution of accuracy across players, and we do. It says nothing about transfer.
We can keep the thresholds honest, which is why the app reports no directional read below 85 decided answers and draws no situation slice below 12. That is a statement about our own measurement, not about the world.
None of that is the study. Calling any of it evidence of transfer would be the same move as reporting an in-sample result and stopping.
What a negative result would mean#
Less than people expect, and we would rather say so now than after the fact.
A well-powered null would mean that at this horizon, with this format, over that training period, transfer to an unselected test set was not detectable at that sample size. It would not mean the practice was worthless — an 80%-powered test misses a real effect one time in five, and a training effect could be real, small, and slower than the study. What it would rule out is the large, quick effect that anyone selling practice implicitly promises.
We would publish it. That commitment is easy to make in advance and worth exactly as much as the protocol it is attached to, which is why it is written here rather than in a marketing page.
What this article does not do#
It does not report a result. There is no study, no data, and no pilot. Designing a study is not evidence, and we are not going to let this page be cited as though it were — the honest status of the transfer question is unknown, not promising.
It is not a complete protocol either. A real pre-registration needs the exact test instrument, the training dose, the analysis plan and the exclusion rules, and those are decisions rather than paragraphs.
What it does is close a gap in our own writing. We had been repeating that the question is open while never specifying what closing it would require — which made the sentence sound rigorous while costing us nothing. This is the price list. We are applying to ourselves the same four-part standard we used on thirty years of technical trading research: in-sample pattern, statistical significance, economic usefulness, and out-of-sample validity are different things, and we have not established the fourth one about our own product.
SwipeTA is a training game and a simulation: no real money, no broker, and it does not provide investment advice. The parameters behind our measurements are on the methodology page.
Sources#
- SwipeTA sample-size calculation, run 2026-08-20: two-sample proportion tests at 80% power and 5% two-sided give 392, 1,569, 4,360 and 9,810 observations per side for a 10, 5, 3 and 2 percentage point difference between groups. Script and output: research/coin_vs_skill.py and results/coin_vs_skill.json in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
- SwipeTA question bank construction: candidates are kept only when the 15-minute move falls in a 1-3x ATR band, and the generator advances past a horizon after each accepted decision point. Both are selection steps applied with knowledge of the outcome, which is why the bank cannot serve as its own held-out test set. Measured in research/fade_rule_decomposition.py in the site repository.
- SwipeTA app source: the statistics panel reports no directional lean below 85 decided answers and draws no situation slice below 12, apps/mobile/src/services/statsRepo.ts, read 2026-08-20.
- https://www.swipeta.net/learn/the-best-reading-on-technical-analysis-is-the-record-of-it-being-tested
- https://www.swipeta.net/learn/what-research-says-about-trader-intuition
- https://www.swipeta.net/methodology