What a backtest can't tell you
A backtest replays a strategy over past prices with fixed settings and reports what would have happened. It is the best evidence available before risking money, and it is easy to read too much into. The problems are well known and mostly come from one source: the past is a single path, and a strategy can be fitted to that path rather than to anything that will repeat.
Overfitting
Give a strategy enough settings and enough tries, and some combination will have done brilliantly on any stretch of history, by chance. The more settings searched and the shorter the history, the more likely the best result is luck. An overfitted setting looks like skill on the data it was chosen on and falls apart on the next stretch.
Signs to look for:
- Oddly exact settings, such as a 37-bar average with a 2.13 band, where round numbers nearby do much worse.
- A result carried by one or two trades.
- A return far above anything else on the same symbol.
In-sample and out-of-sample
The standard defence is to split the history. The settings are chosen on the first part, the in-sample window, and judged once on the later part they never saw, the out-of-sample window. Only the second number says anything about the future.
On daily bars, A9's search does exactly that: it tunes on the first 70% of the history and scores the result once on the last 30%. The nightly shortlist does the same and also shows each calendar year separately, so one lucky year cannot pass for a track record.
Hourly and finer bars on tokenized stocks are a harder case, because most of those tokens have only a few months of history. A 30% slice of a few months is too thin to judge on its own, so these rankings are scored on the whole recent window, with a check that the held-out end of it does not contradict the rest, and the gates below carry most of the weight.
The gates Auto-tune and the shortlist apply
On hourly and finer bars, Auto-tune runs the same recipe as the nightly shortlist, on your symbol and window; on daily bars it runs the 70/30 search above. On the finer bars a setting is offered only if it clears all of these:
- At least three completed round trips. One open position marked to market is an anecdote, not a record.
- A positive return that beats simply holding over the same window.
- A return at least as large as its maximum drawdown (half of it, for a grid, which holds inventory through the dips it buys).
- Real timing. The strategy's in-and-out pattern is shifted through time 149 ways; the real alignment with price has to beat most of its own shifted copies. A strategy that only looked good by being out of the market most of the time fails this.
- Stable neighbours. Each setting is nudged by 20% up and down, one at a time, and the typical result of those neighbours has to clear the same bars. A peak one setting wide collapses here.
- A result that is extraordinary is re-run with the window ending one to seven days earlier. If the return lived in its last few days, it is dropped.
If nothing clears, Auto-tune says so rather than returning the least bad setting. That outcome is information too.
Where the day is cut
A daily bar has to end somewhere. Binance ends it at midnight UTC; OKX at 16:00 UTC, which is midnight in Hong Kong. Neither is more correct, and the choice moves results. Measured on the same hourly prices for XRP in 2026, a 10/30 EMA cross returned -13.6% with the day cut at midnight and +18.5% with it cut at 16:00.
So a daily-bar candidate on the shortlist is re-run with the day cut at six hours, every fourth hour from midnight, and has to make money at the typical cut and at most of them. A result that lives at one cut out of six is a boundary, not an edge.
What the simulation leaves out
- Fees are charged at the exchange's standard tier: on OKX spot 0.08% maker and 0.10% taker, on perpetuals 0.02% and 0.05%. Your account may pay less; it never pays nothing.
- Fills happen at the prices in the bars. The backtest does not model an order moving the market, so the shortlist reports a capacity beside each result: 10% of the typical bar's traded value, the size beyond which an order starts to push the price.
- Funding payments on perpetuals are not charged, and leverage is modelled at 1 times.
- Bar resolution. A grid's resting orders fill whenever the price touches a line, but a backtest only sees each bar's open, high, low and close. The same grid on XSOXL over the same three months read +7.8% on 15-minute bars, +11.7% on 5-minute, +11.4% on hourly and +8.1% on 1-minute bars. A9 therefore backtests grids on 1-minute bars, the bar live grids run on.
A short checklist
- Did the result come from Auto-tune, and what did its neighbours earn?
- How many round trips does it rest on?
- Does it beat buy and hold over the same window?
- Would you sit through its maximum drawdown with real money?
- Is the history long enough to include a market unlike the current one?
Then start small, or on paper, and compare the live record with the backtest after a few weeks.