You chose every trade in your journal
A trade journal records the setups you were confident enough to enter. That makes it a selected sample, and a selected sample cannot grade the reads you never acted on.
In short
- A journal is a record of the trades you entered, so every comparison inside it is between two groups of charts you picked. That is a selected sample, and it cannot separate your reading from your choosing.
- A dealt question bank supplies the other half - the setups you would have passed on - which is what makes a contrast like above-VWAP against below-VWAP a comparison of one reader rather than of two populations.
- Our situation panel is built so it can return nothing: thin slices are hidden, the directional read stays locked until 85 decided answers, and one verdict branch exists only to report that more time made no difference.
- This is a claim about what each tool can measure, not about outcomes. A journal sees your real executions and money and nerves; we see none of those, and never will.
Our home page has a row of four cards on it: practise, analyse, execute, review. We are in the first slot. A charting platform is in the second, a broker or a paper-trading account in the third, and a trade journal in the fourth.
That row is a claim about scope, and it is worth defending properly rather than leaving as a diagram. The interesting part is not that four tools exist. It is that they differ in who decides what you look at.
Execution and review are built out of decisions you actually acted on: one records them, the other replays them. Analysis ranges far wider — nobody has to trade a chart to study it — but what lands on your screen is still something you went and opened. Practice is the only slot where the situations can be dealt to you.
That difference is not a feature comparison. It changes what each tool is able to measure about you, and the tool most people expect to learn from is the one where it bites hardest.
A journal is a sample you selected#
A trade journal is very good at what it does. It holds your entries, your exits, your sizing, your notes on what you were thinking, and it will happily slice all of that by setup, by time of day, by instrument.
But every row in it got there the same way: you looked at a chart and were confident enough to act. That is not a criticism of journalling software. It is the definition of a journal.
The consequence shows up the moment you try to compare two situations inside it. Suppose you want to know whether you read continuation better than reversal. Your journal has forty continuation trades and six reversal ones — and those six are not a random six. They are the reversals that looked clear enough, on the day, for you to put money on. The ones that were genuinely ambiguous, the ones you stared at and passed, the ones you misread and therefore never entered: none of them left a record.
So the comparison is not "you on continuation versus you on reversal". It is "you on the continuations you liked versus you on the reversals you liked". Two different populations of chart, selected by the same judgement you are trying to grade.
This is an old problem with a plain name — selection — and it is not fixable by writing better notes. The missing rows are missing because of a decision you made, and that decision is the thing under examination.
What the other half looks like#
The practice slot exists to supply the rows a journal structurally cannot have: the charts you would not have chosen.
In SwipeTA the questions are dealt, not browsed. The bank is built as a deliberate 50/50 split between up and down outcomes, you get one question per symbol per session, and you answer what arrives. You can briefly defer a small number of them, but they return before the round is over. There is no permanent "skip this setup" path, because that is exactly the button that would reintroduce the problem.
What that buys is a set of contrasts where both sides were dealt to you in comparable numbers:
- long calls against short calls
- fast answers against slow ones
- traps — where the obvious read fails — against readable setups
- price above VWAP against price below it
- choppy tape against calm tape
- price near the day's high against price near its low
- the answer after a wrong one against the answer after a right one
That last pair is the one we wrote a whole article about, and it is a good illustration of the point. A journal has a hard time holding it too, because after a losing trade a disciplined trader may well choose not to trade again straight away — which leaves the tilt question with no next decision to look at, on precisely the occasions it matters.
Each row is drawn with a 95% interval around it and a line at the halfway mark, so that a gap you cannot yet distinguish from chance looks like what it is. We use a Wilson interval rather than the textbook normal one, because the textbook one misbehaves at small samples and near the extremes — which is where a feature like this lives for the first few weeks of a player's life.
Built so it can return nothing#
A measurement feature has to be allowed to return "nothing yet". Otherwise every thin sample eventually gets turned into an insight, because a panel that reads you is under permanent pressure to say something. Here is what that principle cost us.
A slice under twelve answers is not drawn at all. Below that, the interval would span most of the axis and the row would be decoration.
The directional read stays locked until 85 decided answers. That number is not a product instinct; against a bank that is a deliberate 50/50, detecting a 65/35 lean with reasonable power takes about that many. A 60/40 lean takes about 194. Until then the card shows a countdown rather than a figure it cannot stand behind.
Separating two slices honestly is much more expensive than it looks. Comparing one situation against another is a two-sample problem: both sides are estimates, so both carry error. On a two-proportion test at 80% power and 5% two-sided, a genuine ten-point gap — say 55% against 45% — needs about 392 answers per side before it can be told apart from noise. A twenty-point gap needs 97 a side; a five-point gap needs 1,569.
That is a different test from the one behind the unlock thresholds above, which asks whether a single number sits above the bank's 50%. The two are worth keeping apart deliberately, because their answers land close enough together to be mistaken for one another. Both are in the same committed script.
It is also why the intervals are drawn rather than hidden: they say "not yet" without a lecture, and they narrow on their own.
One branch of the verdict engine exists purely to report a null result. If you have enough fast and slow answers and the two are indistinguishable, the app tells you that taking more time made no difference — which is a real finding, and a less flattering sentence than most products are willing to print.
The daily journal never shows a single day's hit rate. The server returns it. We deliberately do not map it into the client, because a twenty-answer day swings by around sixteen points on luck alone. The weekly line only claims improvement after passing a significance test, which means most days it stays quiet and falls back to reporting volume. That silence is the intended behaviour: a claim of progress that fires every day is not a claim, it is decoration.
The metric we deleted from our own home screen#
The clearest thing we can say about this design is what we removed.
Until 2026-08-03 the journal's headline number was decision time. It was, statistically, the best-behaved figure we had — stable, fast to converge, hard to fool. And it was the wrong one to put on the home screen, because any number a product presents as progress becomes a target, and a tool that rewards a shrinking decision time is teaching a trader to answer faster. Quick is not right.
So the metric that could not be gamed was quietly teaching a bad habit, and we replaced it with hit rate — a worse-behaved number that at least points at the task we are actually asking the player to perform, and that we then had to wrap in a significance test to stop it lying on quiet weeks. Note what it still is not: a single correct call is not evidence you read the chart well, which is the argument we made at length elsewhere.
We mention this because it is the sort of decision the four-card row is actually about. The practice slot is not "the same measurement as a journal, but earlier". It is a different instrument with different failure modes, and most of the work is in refusing to overstate it.
Where the journal wins, and we do not#
This would be a dishonest article if it stopped there, so:
A journal has your real executions. Slippage, partial fills, the exit you moved, the position you sized up because you felt good that morning. We have none of that and never will.
A journal has money in it. The single largest difference between a decision that costs nothing and one that costs something is not information, it is nerve, and nerve does not show up in a question bank. Anyone telling you otherwise is selling something.
A journal has your instruments, your session, your strategy. Ours are 5-minute US intraday bars on a fixed universe, settled at a fixed horizon. If you trade futures at the open on a one-minute chart, our questions are adjacent to your job rather than a model of it.
A journal covers the parts of trading that are not reading. Sizing, holding, exiting, sitting out. We have written before that a replay is a better chart while a question bank is a better test, and the same asymmetry applies here: the narrower the thing you measure, the more you can say about it, and the less of the job it covers.
What this does not claim#
Not that a controlled sample of your reads makes you better at trading. Not that any of our numbers transfer to a live market — that question is open, and we said so in what the research says about trader intuition. Not that you should stop journalling; the four cards on our home page are four slots, not a ranking.
SwipeTA is a training game and a simulation: no real money, no broker, and it does not provide investment advice. Hit rate is performance on our questions.
The claim is narrower. A journal can tell you how your trades went. It has a harder time telling you how you read, because you picked the charts it contains. Filling that gap needs someone else to do the picking, which is the entire reason the first card exists. The parameters behind our questions are on the methodology page.
Sources#
- SwipeTA app source, read 2026-08-14: the situation slices, their per-slice minimum (SLICE_MIN = 12), the 95% Wilson interval drawn on each row, the directional unlock (BIAS_UNLOCK = 85), the per-side split minimum (SPLIT_UNLOCK = 20) and the confidence tiers (85 / 400 decided answers) are all in apps/mobile/src/services/statsRepo.ts.
- SwipeTA app source, read 2026-08-14: the journal's headline metric changed from decision time to hit rate on 2026-08-03, and the significance test that gates any claim of improvement, are in apps/mobile/src/game/journalInsight.ts. The excluded fields and the reason they stay unmapped are documented in apps/mobile/src/services/journalRepo.ts.
- SwipeTA sample-size calculation, re-run 2026-08-14: the per-side figures for separating two situation slices (97 / 392 / 1,569 answers for a 20, 10 and 5 point gap) come from a two-sample proportion test at 80% power and 5% two-sided; the unlock thresholds (85 / 194 / 783) come from a one-sample proportion test against p0 = 0.5. Both are in research/coin_vs_skill.py with output results/coin_vs_skill.json in the public research repository https://github.com/BOHARRY/swipeta-research (MIT / CC BY 4.0).
- SwipeTA question bank construction: 5-minute US intraday bars, a deliberate 50/50 direction split, one question per symbol per session, and a 15-minute settlement horizon. Parameters at https://www.swipeta.net/methodology
- https://www.swipeta.net/methodology