Start with an identifiable dataset
The July baseline uses archived Binance spot candles and Alpaca stock data. Every selected audit records its date range, usable candle count, data-quality checks, venue calendar, and fixture identity.
Checks cover duplicate timestamps, ordering, incomplete bars, positive prices, OHLC consistency, and unexplained gaps. A content hash identifies the archived input. This prevents a later dataset revision from quietly changing the meaning of an old result.
Walk forward without resetting reality
The baseline uses six-month training/context windows and one-month test windows. Parameters are fixed; this is not a claim that the strategy is refitted or optimized every month.
The corrected continuous-state engine preserves indicators, open positions, capital, and Turtle skip state across calendar boundaries. Trades belong to the out-of-sample window in which their entry decision occurred. A trade is not counted again merely because it crosses a month boundary.
Full-period diagnostic trade counts are not interchangeable with exact out-of-sample counts. Returns are evaluated on daily equity observations using the appropriate venue calendar.
The formulas in the code
These definitions describe our implementation. Units and assumptions matter as much as the equation.
Daily return
E is the last equity observation of the UTC day for crypto, or an observed stock session. Missing weekends are not inserted as zero stock returns.
Annualized Sharpe
s is sample standard deviation (n − 1). A is 365 for crypto and 252 for stocks. This implementation uses raw returns without a risk-free-rate adjustment. Fewer than two returns or nonpositive dispersion produce 0, which is not evidence of an edge.
Maximum drawdown magnitude
The largest decline from a prior equity peak. The public table displays a positive loss magnitude; some engine functions retain the signed negative drawdown internally.
Profitable out-of-sample windows
A window is evaluated using disjoint entry-attributed trade P&L. This is not the percentage of winning trades.
Deflated Sharpe probability
SR* estimates the best Sharpe expected from the recorded trials under the no-skill benchmark. γ₃ is skewness and γ₄ is kurtosis. Inputs use per-observation Sharpe units. Insufficient observations or an invalid variance estimate return no estimate. The recorded gate is at least 0.95.
A positive result can still fail
The V6-04 multi-horizon trend study evaluated 12 selected ETFs across 102 month-end decisions. It records 2,120 session returns from 2,121 equity valuations, ending 10 July 2026.
The registered constrained portfolio passed 11 of 12 checks but did not beat the static risk-balanced benchmark’s Sharpe. The verdict was baseline-only: useful risk control, without demonstrated benchmark edge.
| Portfolio | Total return | Sharpe | Max DD |
|---|---|---|---|
| Constrained trend | 15.19% | 0.999 | 3.34% |
| Equal weight | 64.63% | 0.869 | 14.35% |
| Static risk balanced | 20.24% | 1.225 | 4.45% |
| Faber monthly | 65.08% | 0.898 | 11.47% |
11/12 checks passed. Benchmark edge failed. Retained as a baseline, not promoted.
Selected ETF slate: BIL, EEM, EFA, GLD, IEF, IYR, PDBC, SHY, SPY, TIP, TLT, UUP. This universe is not point-in-time or survivorship-free.
What this evidence cannot establish
The ETF slate is selected today rather than a point-in-time, survivorship-free universe. Adjusted provider histories can be revised in later snapshots. These limitations remain attached to the result.
The July baseline artifacts predate the later per-trial out-of-sample return-series field required for PBO evaluation. Those old files cannot be treated as passing a modern overfitting gate without new evidence.
No historical simulation proves future returns. We do not combine the paper book with live performance, hide failed trials, or replace missing evidence with zero.