58,000 Simulated Trades, Every Level Family — and Nothing Beats Random Levels
After two studies found weak, in-sample edges around our levels, we did what a sceptical reader would do: threw every level family, every interaction and every conditioning variable at the data, and then asked the only question that matters — does the best thing we found beat the best thing a grid of random levels finds?
The short version
No. Across 460 unconditioned cells and 3,355 conditioned ones, the best real cell had a t-statistic of 2.68. The median best cell of a random-level grid had 2.80, and beat the real one in 14 draws out of 20. With conditioning, 3.79 against 3.81. On 84 sessions, the grid contains nothing distinguishable from chance — including the cells that look spectacular in isolation.
What "everything" meant
- Level families: Call Wall, Put Wall, Gamma Flip, Vol Trigger, both Battle Zone edges, upside and downside clusters, Focus 1 to 3, the levels frozen at the open, the hybrid QQQ+SPY walls, the three Value Zone bands on each side, and the overnight anchor.
- Interactions: fade at contact; acceptance beyond the level; acceptance then a retest of the level that holds; sweep and reclaim (a wick through the level, then a close back — the classic failed breakout).
- Exits: stop 20 / target 40, 40 / 60, 40 / 80, 60 / 120. Entry at the next bar's open, stop on the wick, forced exit 15:55 ET, 1 point of cost.
- Conditioning, after the fact: hour of day, gamma regime, side of the flip, sign of the net-premium slope, sign of the 30-minute GEX change, sign of the hedging pressure, compression flag.
58,112 simulated trades on 78 RTH sessions of NQ 1-minute bars, with the levels the Indicator actually published each minute.
The cells that "stood out"
| Cell | Trades | Win rate | Pts / trade | t-stat |
|---|---|---|---|---|
| Fade the upper VZ2 band approached from below, stop 40 / target 80 | 64 | 59% | +21.4 | 2.6 |
| Same, stop 60 / target 120 | 50 | 60% | +34.3 | 2.7 |
| Same, transitional regime, stop 40 / target 60 | 45 | 76% | +29.0 | 3.8 |
| Sweep and reclaim of the Battle Zone low, stop 20 / target 40 | 61 | 57% | +9.6 | 2.2 |
Any of these would make a convincing landing page. The first one even survives its own checks: positive in June, July and August, robust to the touch distance and to the stop size, confidence interval above zero by day. And it is exactly what the best cell of a random grid looks like.
The arbiter
Judging a cell by its own p-value is how backtests lie: search three thousand cells and a few will clear any threshold. The honest control is to rerun the entire grid on levels that cannot contain information — the levels of another day, shifted onto today's price — and record the best t-statistic each random grid reaches. Twenty grids, the same cells, the same conditioning. That distribution is the bar. A real cell only means something if it clears what the best random cell reaches.
| Real grid | Random grids (median) | Random grids that beat the real one | |
|---|---|---|---|
| Best t, unconditioned cells | 2.68 | 2.80 | 14 / 20 |
| Best t, with conditioning | 3.79 | 3.81 | 10 / 20 |
| Cells with t ≥ 2.5 | 14 | 9 | |
| Cells with t ≥ 3 | 5 | 2.5 |
One practical trap for anyone repeating this: permute the levels without re-centring them on the day's price and the random grids barely trade — prices differ by hundreds of points from one day to the next — which makes the null look weak and the real cells look strong. We caught that on the first run and fixed it before drawing any conclusion.
What survives, and what it means for the levels
Descriptively, the same asymmetries kept appearing across all three studies: continuation after acceptance beats fading at contact; rejections worked at the top of the day's range and not at the bottom; buying the Put Wall at contact lost 17 points per trade; buying an upside break of Focus 1 lost 14. Those are consistent, and they are what the Academy has always taught — read the regime, respect acceptance, never fade against the drift. What none of it is, on this sample, is a mechanical edge.
So the two rules we froze earlier are not "the edge". Their in-sample t-statistics sit inside what the best random cell reaches. They are hypotheses with every parameter written down, measured forward on the level statistics page, judged at 200 trades each. That is the only kind of evidence we will show, and the reason our levels are sold as context.
The levels, with the numbers next to them
The Indicator draws the map; the statistics page shows how it behaved; the Academy teaches the reading. Context, never signals.
Start the 7-day free trial — $9.99/moMethod notes, for the people who will check
- The independent unit is the session, not the trade: bootstrap confidence intervals are drawn by day.
- Colliding cells were de-duplicated before pooling — the same false breakout touches several levels at once, which is how a family can look 71% positive cell by cell and make 2 points pooled.
- Vanna and charm levels, max pain, the hedging band and the vanna wall could not be tested: the historical logs never carried them. They are being sampled now.
- Scripts, data and the full report are kept in our research repository; the numbers on this page were not selected after the fact.