Cascade Indicators · validation study, waves 1–2 · ES, ZB & NQ futures · Nov 2024 – May 2025

The detectors work, the predictions don't, and the events mark calm rather than action.

Two indicators (EQH/EQL and ORB + CVD) were rebuilt in Python and tested on three futures markets. Wave 1 tested direction ("which way?"). Wave 2 tested salience ("does anything happen at all?") and ran the ORB logic on true order-flow data to see whether better data would rescue it. Every claim had a pass/fail line written down before results were seen.

✗ CVD = real order flow, FAIL ✗ Confirmations select better trades, FAIL (2 markets × 2 data engines) ✗ Regime dial predicts, FAIL ◑ ZB fade, real but sub-tick ↓ "Pay attention", INVERTED: events mark quiet ✓ Experiment resolved: it's the concept, not the data
Summary. Two indicator concepts were rebuilt outside TradingView and tested on 5.7 million one-second bars of ES, NQ, and ZB futures (Nov 2024 to May 2025), with pass and fail criteria written before results were viewed. The detection engineering held up. No directional edge was found at 5 to 60 minute horizons. CVD confirmation did not improve breakout selection with estimated or true order flow. The one positive result: on 1-minute charts, resolved equal-high/low levels marked below-normal subsequent movement, an effect that fades by 15-minute charts.
Contents. Attention test · Data-quality experiment · CVD accuracy · Confirmation checks · Regime dial · ZB fade · Limitations and protocol
How to read this page. Returns are in bp, basis points; 1 bp = 0.01% ($10 on a $100k position). Whisker lines show the range of plausible truth after accounting for luck: if a whisker touches the reference line, we cannot claim an effect. Dashed vertical lines are the pre-specified pass gates.

Wave 2 headline, "Does this event mean the market is about to move?"

Movement in the next 30 minutes after an event, compared to ordinary bars at the same time of day. 1.0 = normal.
INVERTED, EVENTS MARK QUIET

The directional tests failed, so the fallback claim was salience: even if the signals don't say which way, maybe they say "pay attention now." Pass gate: movement ≥ 1.25× normal. The answer came back below 1× almost everywhere, several with the whole whisker below normal:

What the inversion means We asked if these moments were louder than ordinary ones. They're quieter, and the effect is strongest exactly where the product implied the opposite: after a ×3+ multi-tested pool resolves, the market moves about 40% less than normal. In hindsight it's obvious: equal highs and lows only form when price keeps returning to the same spot, that is a trading range, and quiet markets tend to stay quiet. The indicator isn't a fire alarm. It's a map of the consolidation you're standing in, and wave 2b showed that map is a 1-minute phenomenon: the calm signature decays smoothly as the timeframe grows (ES all-events: 0.79x @ 1m → 0.91x @ 5m → 1.01x @ 15m) and is gone by 15m. Liquidity-pool behavior is not scale-invariant. At 5m, only sweeps on the index futures stay reliably quiet (ES 0.75x, NQ 0.79x, both whiskers below 1.0).

Method notes: the baseline-matching exclusion knob biases toward flattering the events, so the 1m inversion is, if anything, stronger than shown (narrow-exclusion re-run still queued). ZB at 5m leaned the other way, movement after breaks was mildly elevated (1.28x at 15 min), but that cell was not pre-specified and failed the 30m gate, so like the +14bp before it, it goes to the hypothesis list, not the product. NQ ORB (0.98x) stays the neutral case: breakouts there are simply ordinary.

Wave 2 experiment, "Would perfect data fix the confirmations?"

Same 41 sessions of NQ, same breakouts, run twice: once on the indicator's guessed order flow, once on the real thing.
RESOLVED: IT'S THE CONCEPT, NOT THE DATA

Wave 1 showed the indicator's volume-delta guess only ~70% matches true order flow, leaving an open question: would the confirmations work with real data? The NQ trades file records the true aggressor on every trade, so we ran the identical ORB logic on both:

Guessed flow (production proxy)
−1.4 bp
confirmed minus unconfirmed breakouts, 30 min
True aggressor flow (perfect data)
−1.9 bp
same breakouts, same logic, real order flow
Why data quality is not the bottleneck Giving the confirmation logic perfect information made it no better, marginally worse, in fact. The idea of confirming a breakout with same-moment delta direction is what's empty, not the measurement of delta. This settles whether better data would change the conclusion, and it means the wave-1 data-quality failure barely matters for this product.

Caveat: the NQ trades file covers 41 sessions (a ~2-month subset), one market regime. But the proxy-vs-true comparison is internally controlled, identical breakout set, so the exoneration conclusion doesn't depend on the window.

Claim 1, CVD measures real buying & selling

"When the indicator says buyers are in control, is that actually true?"
FAIL

The indicator guesses trade direction from each second's price tick. Compared against NQ data that records the real aggressor (57,006 minutes):

Correlation with true order flowgate: 0.95
Direction agreement (per minute)gate: 90%
How the estimate should be described Right about 3 minutes out of 4, a decent estimate of the tape. Product pages can say "volume-delta estimate"; they cannot say "institutional order flow." Wave 2's experiment above showed this gap doesn't drive the confirmation failure, but it still governs what the CVD display can claim to be.

Measurement caveat: this benchmark predates the front-month filter fix, so contract mixing biased it downward; the corrected number is queued and can only be higher.

Claim 2, The two confirmation checks earn their place

"Do breakouts that pass the CVD + composite checks make more money than ones that don't?"
FAIL, NOW ON TWO MARKETS

Part A, are they different checks? PASS

The two checks genuinely measure different things, agreement correlation ≈ 0.46–0.50 on ES and NQ (redundancy gate was 0.80).

Part B, does passing them pay? FAIL

ES, 30 minutes after breakout, by which checks fired. Every whisker touches zero; "neither" did best. NQ replicated the null on both data engines.

Why the confirmation filters are optional Two independent gauges that select nothing, on two markets, with both guessed and perfect data. Average ES breakout: +1.1 bp, under the cost of trading it. v2.4 direction is data-dictated: confirmations become optional-off; CVD repositions as telemetry, not gatekeeper.
REFUTED IN WAVE 2 The tempting number, resolved: wave 1's "+14 bp CVD-only" subgroup, flagged then as one of ten scratched lottery tickets, came back −22.5 bp (opposite sign) on the NQ re-test. The post-hoc flag held: the result was unstable and never entered product copy.

Claim 3, The regime dial predicts what works next

"When the dial says 'Reversion', do sweep-fades actually pay better than when it says 'Trend'?"
FAIL

Pass condition: fading sweeps in Reversion beats Trend with separated whiskers on 1m and 5m. Measured on ES:

ES 1m, sweep-fade return, next 30 min

ES 5m, sweep-fade return, next 30 min

What the dial can and cannot say Everything hugs zero; on 1m the ordering is even backwards. The dial truthfully describes the recent past and predicts nothing. In the product it should be labelled a description ("recent resolution mix"), never a forecast. Wave 2 reframes what it's for: combined with the consolidation finding, pool activity is a live chop gauge, a natural input to the CCRD product line.

Claim 4, The one "significant" wave-1 result

"ZB breaks fade with p = 0.000, isn't that the edge?"
REAL, BUT SMALLER THAN ONE TICK

On 30-year Treasury futures, break-continuations lost with near-certain statistics (27% winners at 5 min). But the whole effect fits inside the instrument's minimum price step:

Why this is market plumbing The effect (−1.3 to −1.5 bp) lives inside half of one tick (±2.6 bp), a bar closing upward through a level closes near the ask, and prices settle back half a step. Market plumbing, not a signal: like measuring a doorway with a ruler marked only in feet.

Limitations, implications, and what comes next

The validated product claim that emerged: these events map consolidation.

Future holdout protocol

The existing tests were internally pre-specified but not publicly timestamped. Future studies will follow this process so replications can be called what they are:

  1. Freeze the indicator version and the claims under test.
  2. Specify instruments, dates, timeframes, exclusions, metrics, and pass gates.
  3. Publish or immutably timestamp the protocol before retrieving the holdout data.
  4. Run the holdout once.
  5. Publish all specified results regardless of outcome.
  6. Treat any subsequent tuning as development, requiring a fresh holdout.

An unviewed historical period selected mechanically, with the protocol timestamped before retrieval, qualifies as an out-of-sample holdout. The term preregistered replication is reserved for a genuinely public, timestamped protocol.

Implication for product copy: the tools are sold as structure detection, consolidation mapping, alerting, and labeled telemetry. No edge or urgency claims are made, and the recorded negative results define the limits of what the copy can say.

This study is published in full, including every failed claim, as the evidence base for Cascade Indicators products. Methodology questions: contact@cascadeindicators.com. Trading involves substantial risk of loss; see the risk disclosure. © 2026 Cascade Analytics LLC.