Two indicators (EQH/EQL and ORB + CVD) were rebuilt in Python and tested on three futures markets. Wave 1 tested direction ("which way?"). Wave 2 tested salience ("does anything happen at all?") and ran the ORB logic on true order-flow data to see whether better data would rescue it. Every claim had a pass/fail line written down before results were seen.
The directional tests failed, so the fallback claim was salience: even if the signals don't say which way, maybe they say "pay attention now." Pass gate: movement ≥ 1.25× normal. The answer came back below 1× almost everywhere, several with the whole whisker below normal:
Method notes: the baseline-matching exclusion knob biases toward flattering the events, so the 1m inversion is, if anything, stronger than shown (narrow-exclusion re-run still queued). ZB at 5m leaned the other way, movement after breaks was mildly elevated (1.28x at 15 min), but that cell was not pre-specified and failed the 30m gate, so like the +14bp before it, it goes to the hypothesis list, not the product. NQ ORB (0.98x) stays the neutral case: breakouts there are simply ordinary.
Wave 1 showed the indicator's volume-delta guess only ~70% matches true order flow, leaving an open question: would the confirmations work with real data? The NQ trades file records the true aggressor on every trade, so we ran the identical ORB logic on both:
Caveat: the NQ trades file covers 41 sessions (a ~2-month subset), one market regime. But the proxy-vs-true comparison is internally controlled, identical breakout set, so the exoneration conclusion doesn't depend on the window.
The indicator guesses trade direction from each second's price tick. Compared against NQ data that records the real aggressor (57,006 minutes):
Measurement caveat: this benchmark predates the front-month filter fix, so contract mixing biased it downward; the corrected number is queued and can only be higher.
The two checks genuinely measure different things, agreement correlation ≈ 0.46–0.50 on ES and NQ (redundancy gate was 0.80).
ES, 30 minutes after breakout, by which checks fired. Every whisker touches zero; "neither" did best. NQ replicated the null on both data engines.
Pass condition: fading sweeps in Reversion beats Trend with separated whiskers on 1m and 5m. Measured on ES:
On 30-year Treasury futures, break-continuations lost with near-certain statistics (27% winners at 5 min). But the whole effect fits inside the instrument's minimum price step:
The validated product claim that emerged: these events map consolidation.
The existing tests were internally pre-specified but not publicly timestamped. Future studies will follow this process so replications can be called what they are:
An unviewed historical period selected mechanically, with the protocol timestamped before retrieval, qualifies as an out-of-sample holdout. The term preregistered replication is reserved for a genuinely public, timestamped protocol.
Implication for product copy: the tools are sold as structure detection, consolidation mapping, alerting, and labeled telemetry. No edge or urgency claims are made, and the recorded negative results define the limits of what the copy can say.