Markets & On-ChainGuide

How to Backtest a Crypto Strategy: Data, Pitfalls & Overfitting

This article demonstrates each backtesting pitfall on real data rather than describing it. Across a 100-combination grid search, only two configurations beat the market's own drift in the period they were tuned on.

How to Backtest a Crypto Strategy: Data, Pitfalls & Overfitting

Introduction

A grid search over 100 RSI configurations, run on 23,610 daily bars from eight major crypto assets, produced 78 profitable results. Measured against the base rate rather than against zero, only two of them beat simply holding the same assets over the same period. The other 86 looked like strategies and were the market. That gap is what this article is about. Every backtesting pitfall below arrives as a measured number from a real run rather than as a warning. What the base rate does to results, what a parameter search actually finds, how much execution assumptions move the answer, what survivorship removes, and how fast ordinary costs erase the rest. The tests also audit this series' own earlier work where it is vulnerable.

Key Takeaways

  • Holding these assets for 30 days returned 17.92% on average in 2017 to 2021 and 2.40% in 2022 to 2026, so the era dominates any strategy result.
  • Of 88 tested configurations, 78 were profitable in raw return but only 2 beat the base rate; the mean edge was -4.44 percentage points.
  • The best in-sample configuration rested on 79 signals and improved out of sample on 43, which is too few to decide anything.
  • Shifting execution from the signal close to the next close cost 0.295 points, roughly a third of the measured edge.
  • CoinPaprika lists 61,463 coins of which 78.6% are inactive, and all eight assets tested here still trade.

What Does a Backtest Actually Measure?

A backtest measures how a fixed rule would have behaved on a fixed slice of history, applied to a fixed set of assets. That is a much narrower claim than most published results imply. Almost every backtesting failure traces back to forgetting how narrow it actually is.

What a backtest measures

The output is a description of one rule applied to one period on one set of assets. Change any of the three and the number changes, often by more than the supposed edge. A backtest cannot tell you the rule will work again, because it contains no information about periods it did not cover. That limitation is structural rather than a matter of technique. What it can do is rule things out. A rule that failed on the data it was designed for is unlikely to succeed elsewhere. Eliminating candidates is a legitimate and underrated use of the technique.

What it cannot measure

It cannot measure whether the result came from the rule or from the market. That distinction is the subject of the next section and the single largest source of misleading crypto backtests. It also cannot measure anything the price series does not contain. Liquidity at the moment of execution, the emotional cost of a drawdown, or whether the exchange was even reachable during the move. Every experiment below runs on the same 23,610 daily bars from eight major USDT pairs. Each pitfall therefore arrives as a measured number rather than as a warning, including the ones that make this series' own earlier work look weaker.

The first pitfall dwarfs all the others.

Why Does the Base Rate Change Everything?

Holding these eight assets for 30 days returned 17.92% on average between 2017 and 2021, and 2.40% between 2022 and 2026. That single difference is larger than any edge measured anywhere in this series, and it belongs to the market rather than to any strategy.

Seventeen percent versus two

Splitting the sample at the end of 2021 and computing the unconditional forward return gives a stark contrast (Binance Spot API, 2026-09-11) . Over three days the mean was 1.35% in the first window and 0.19% in the second. Over seven days it was 3.32% against 0.47%. Over fourteen days it was 7.22% against 0.98%, and over thirty days 17.92% against 2.40%. Every horizon tells the same story, and the gap widens as the holding period lengthens. The market's own drift was roughly 7.5 times stronger in the earlier period at the monthly horizon. Nothing about a strategy needs to change for its reported returns to collapse by that factor.

Holding Period2017 to 2021 Mean2022 to 2026 Mean
3 days+1.35%+0.19%
7 days+3.32%+0.47%
14 days+7.22%+0.98%
30 days+17.92%+2.40%

Data current as of September 2026.

The era is the result

Any strategy backtested only on 2017 to 2021 inherits that drift. A rule that buys frequently and holds for a month shows a large average return in that window whether or not it contributes anything. Almost any entry did. This is why a published backtest without its base rate is unreadable rather than merely incomplete. There is no way to tell from the headline figure whether a rule contributed anything at all. A companion article on market cycles measures how different those regimes actually were ↗. The practical consequence is direct. Report the unconditional return for the same period and the same assets, or the result cannot be interpreted at all.

Bar chart of unconditional mean returns by holding period in two eras, showing 2017-2021 far above 2022-2026 at every horizon

Searching harder does not solve this.

What Happens When You Search 100 Parameter Combinations?

A grid search over 100 RSI configurations produced 78 profitable results in the period it was tuned on. Measured against the base rate instead of against zero, only two of those eighty-eight beat simply holding the same assets over the same window.

One hundred combinations

The grid varied the RSI lookback across 7, 10, 14, 21 and 28 periods and the oversold threshold across 20, 25, 30, 35 and 40. Holding periods were 3, 7, 14 and 30 days. That is 100 combinations, of which 88 produced at least 30 in-sample signals (Binance Spot API, 2026-09-11) . Judged on raw return, 78 of the 88 made money in-sample and all 88 made money out-of-sample. Those numbers look like a discovery and are nothing of the sort. They are a description of what the assets did, dressed as a description of what the rule did.

Two of eighty-eight

Subtracting the base rate for the matching holding period changes the picture completely. Only 2 of the 88 configurations produced a positive edge in-sample, and the mean edge across all 88 was -4.44 percentage points. In other words, 86 of 88 tested rules underperformed simply holding the same assets over the same window, while appearing profitable in every case. The search did not find an edge and then lose it out of sample, which is the textbook story. The search never found an edge at all. The raw returns that suggested otherwise belonged entirely to the market, and no amount of further searching would have changed that.

ConfigurationIn-Sample EdgeOut-of-Sample Edge
RSI(28) below 30, 7-day hold+0.23 pp+9.05 pp
RSI(10) below 30, 3-day hold+0.07 pp+0.84 pp
RSI(7) below 20, 3-day hold-0.00 pp+1.19 pp
RSI(14) below 35, 3-day hold-0.06 pp+0.66 pp
Mean of all 88 configurations-4.44 pp+1.82 pp

Data current as of September 2026.

Stat cards: 100 configurations tested, 78 of 88 profitable in-sample, but only 2 of 88 beat the base rate

The winner is worth examining closely.

Does the Best In-Sample Setting Survive Out of Sample?

The top configuration improved out of sample, from a +0.23 point edge to +9.05. That sounds like validation, and it is not. The out-of-sample result rests on 43 signals, which makes this the clearest illustration in the article of why sample size decides almost everything.

The winner got better

RSI over 28 periods, below 30, held seven days, produced the best in-sample edge of the 88 tested. In-sample it fired 79 times for an edge of +0.23 percentage points, which is barely distinguishable from nothing. Out of sample it fired 43 times for an edge of +9.05 points (Binance Spot API, 2026-09-11) . A textbook overfitting example has the winner collapse out of sample. This one did the opposite. That is a more useful lesson precisely because it is unexpected, and because it cannot be explained by the usual story.

Forty-three signals is noise

A 43-signal sample of crypto returns can swing enormously in either direction, and a long RSI lookback with a strict threshold fires rarely by construction. The configuration that wins a grid search on such thin samples is largely the one that got lucky. Luck does not know which direction to run in the next period. Nothing here says the rule works. It says the measurement is too imprecise to decide, which is a different and more honest conclusion. Any grid search should report the signal count beside every result. Anything resting on fewer than a few hundred independent events belongs in the undecided column rather than the discovered one.

Execution assumptions cost less than that, but not by much.

How Much Does Look-Ahead Bias Change the Result?

Running the identical rule under three execution assumptions produced mean returns of 2.978%, 2.464% and 2.168%. The spread is 0.81 percentage points across exactly the same 889 signals, which is comparable to the entire edge being measured in the first place.

Three execution assumptions

RSI below 30 with a seven-day hold fired 889 times. Entering at the signal-day close returned 2.464% on average with a 59.73% hit rate. That is the standard assumption and a slightly impossible one, since the close is only known once the day is over. Entering at the next bar's open returned 2.978% with a 60.85% hit rate. Entering at the next bar's close returned 2.168% with a 56.58% hit rate (Binance Spot API, 2026-09-11) .

Execution AssumptionMean ReturnHit Rate
Next bar's open+2.978%60.85%
Signal-day close+2.464%59.73%
Next bar's close+2.168%56.58%

Data current as of September 2026.

Bias does not always flatter

The interesting result is that the realistic assumption was the best one. Entering at the next open beat the impossible same-close entry, because oversold signals often continue falling overnight and hand the buyer a better price. Look-ahead bias is usually described as inflating results, and here it deflated them. What matters is the magnitude rather than the direction. Deferring execution by one bar to the next close cost 0.295 points, roughly a third of the measured advantage over the base rate. An assumption that moves the answer by a third of the edge deserves stating explicitly in every published result. Most do not state it at all, which leaves the reader unable to tell which version they are looking at.

The universe of assets carries a larger distortion still.

What Does Survivorship Bias Cost You?

CoinPaprika lists 61,463 coins and marks only 13,154 of them active. That means 78.6% of every coin it has ever tracked is now inactive. No backtest built from today's asset list contains a single one of them, which quietly removes most of the ways a crypto position can fail.

Four in five coins are dead

The inactive count is 48,309 (CoinPaprika API, 2026-09-11) . Exchange data tells a milder version of the same story, because exchanges list fewer assets and delist them more slowly than the long tail dies. Binance's exchangeInfo returned 742 USDT spot pairs, of which 490 had status TRADING and 252 did not. Fully delisted pairs are dropped from that response entirely rather than marked. The list of assets available to test today is therefore a filtered sample of winners. The filter was applied by the market itself, after the fact, using exactly the information a backtest is supposed to exclude.

This cluster has the same flaw

Every experiment in this article, and every result in the companion pieces on indicators ↗ and on-chain metrics ↗, uses eight assets that all still trade. That is survivorship bias by construction, and it inflates every number in the same direction. Bitcoin, ether, solana, BNB, XRP, cardano, dogecoin and chainlink were chosen because they have long histories. That is precisely the selection criterion guaranteed to exclude the failures, and it was applied knowingly. Naming that limitation is more useful than pretending a corrected sample was available, because for most retail-accessible data it simply is not. Readers can then discount the results accordingly rather than trusting them whole.

Even the surviving data is not perfectly clean.

How Clean Is Exchange OHLCV Data Really?

Across 23,610 bars from one major exchange there were no missing days and no zero-volume bars, which is cleaner than most people expect. There were also 97 bars where the high exceeded the low by more than half. Those are the ones that matter.

Gaps, wicks and zero volume

Auditing the sample directly gives a short list of defects (Binance Spot API, 2026-09-11) . Missing-day gaps: zero. Zero-volume bars: zero. Bars where the high was more than 1.5 times the low: 97, or 0.411% of the sample. The gap between one bar's close and the next bar's open averaged 0.0231%, with a 99th percentile of 0.307% and a maximum of 6.41%. Continuous markets do not gap the way equities do. That removes one class of problem and quietly introduces another. A rule can assume a fill at any price with no visible gap to flag the assumption.

Clean enough is not clean

A 0.411% rate of extreme-range bars sounds negligible until a strategy trades those bars specifically. Stop-loss rules, breakout entries and anything keyed to the high or low concentrate on exactly those bars. They are also the bars most likely to contain a flash move no real size could have transacted at. A companion article covers how those records are built and why a wick can represent a single trade ↗. The practical defence is simple. Check whether a strategy's results depend on a small number of extreme bars, then re-run it with those bars neutralised and see what survives.

Costs then remove what remains.

When Do Trading Costs Eat the Whole Edge?

The best configuration in the whole grid search had an in-sample edge of 0.23 percentage points. A 0.40% round-trip cost turns that into -0.17, and a 1.00% round trip into -0.77. The edge is smaller than the cost of collecting it.

One percent kills it

Applying costs to the winning configuration produces a short and decisive table of outcomes (Binance Spot API, 2026-09-11) . At zero cost the in-sample edge is +0.23 points. At 0.20% round trip, which is roughly two taker fees, it falls to +0.03. At 0.40%, adding modest slippage, it is -0.17. At 1.00%, which is realistic for a mid-cap asset in size, it is -0.77. The entire measured advantage is smaller than the cost of acting on it. That was the best of 88 configurations, not a middling one.

Costs scale with turnover

The damage is proportional to how often a rule trades, which makes short holding periods far more expensive than the raw returns suggest. A three-day rule pays the round trip ten times as often as a thirty-day rule for the same exposure. A companion article measures what execution actually costs on a live order book ↗, including how quickly slippage grows with order size. Any backtest that omits costs is reporting a ceiling rather than an estimate. For high-turnover rules that ceiling sits far above anything achievable in practice.

Bar chart of the best configuration's in-sample edge after round-trip costs, falling from +0.23 pp to -0.77 pp

Splitting the data properly helps, within limits.

What Is Walk-Forward Testing and Does It Help?

Walk-forward testing fits parameters on one window, tests them on the next, then rolls both windows forward and repeats. It is the strongest routine available to a retail tester working alone. It also fixes considerably less than its reputation suggests.

Walk-forward in practice

The method is straightforward. Choose an in-sample window and select parameters on it. Apply those parameters unchanged to the following out-of-sample window, record the result, then advance both windows and repeat. The output is a sequence of genuinely out-of-sample results rather than one, which reduces the chance that a single lucky period carries the conclusion. It also produces something a single split cannot. If the best parameters jump around between windows, that instability is itself evidence, and usually evidence against the rule.

What it does not fix

It does not fix survivorship, since every window uses the same surviving assets. It does not fix regime dependence, because a strategy tuned across 2017 to 2021 windows is still being validated on windows from the same era. It does not fix sample size, which was the binding constraint on the winner above. And it does not stop a tester running the whole procedure repeatedly with different rules until one passes. That reintroduces the search problem one level up. Walk-forward is a real improvement over a single split. It is not a substitute for deciding what you are testing before you start, and for testing few things rather than many.

Which reduces to a short reporting standard.

How Do You Run a Backtest That Tells You Something?

Seven items make a backtest interpretable to somebody who did not run it. None of them requires special software or extra data. A result missing any one of them cannot be evaluated by a reader at all, however impressive the headline number looks.

Seven things to report

State the base rate for the same assets and period, because the comparison is the result. Report the number of independent signals rather than signal days, since threshold rules generate hundreds of overlapping observations from a handful of events. State the execution assumption explicitly. Name the full asset universe including anything that has since been delisted, or admit the survivorship limitation. Include costs at a realistic level. Report the test period, since 2017 to 2021 and 2022 to 2026 are effectively different markets. Then verify on data that was never looked at during development, which is the only step that cannot be faked after the fact.

ItemWhy It Matters
Base rate for the same periodA 55% hit rate against a 50% base is a five-point edge, not a 55% one
Independent signal countThreshold rules inflate hundreds of days from a dozen real events
Execution assumptionShifting entry by one bar moved the result by a third of the edge
Full asset universe78.6% of coins ever tracked are inactive and appear in no test
Realistic costsA 1% round trip turned the best edge from +0.23 pp to -0.77 pp
Test period statedThe 30-day base rate was 7.5 times higher in 2017-2021 than after
Untouched verification dataOtherwise the search has already used every observation

Data current as of September 2026.

Reading someone else's backtest

Applied to published results, that list is brutal. A strategy claiming a large annual return on 2017 to 2021 crypto data, with no base rate, no signal count and no costs, is reporting the era rather than the rule. The honest summary from every experiment here is that the ordinary case is not a discovered edge that decays. It is an edge that was never there, obscured by a rising market. Testing carefully mostly produces negative results, and that is the correct outcome. The value of a backtest lies in what it eliminates, not in what it promises.

Five-step diagram: state the base rate, count independent signals, defer execution, include the full universe and costs, verify on untouched data

Summary

A backtest describes one rule on one period and one asset set, and the period usually matters more than the rule. Split at the end of 2021, the unconditional 30-day return for these eight assets was 17.92% in the first window and 2.40% in the second. That is a factor of 7.5. Any result quoted without its base rate is therefore uninterpretable rather than merely incomplete. A grid search over 100 RSI configurations demonstrated the point directly. Some 78 of 88 were profitable, and exactly 2 produced a positive edge once the base rate was subtracted.

The remaining pitfalls are smaller but comparable to the edges being chased. Moving execution one bar later cost 0.295 percentage points. The realistic assumption happened to beat the impossible one, so the bias does not always flatter. Survivorship is the largest uncorrectable distortion, with 78.6% of tracked coins inactive and every asset in these tests still trading. Data quality was better than expected, with no gaps and no zero-volume bars. Still, 0.411% of bars had a high more than 1.5 times the low. Costs finish the job: a 1.00% round trip turned the best edge of +0.23 points into -0.77.

Conclusion

The useful conclusion from every experiment here is negative, and that is the point. Testing carefully mostly produces results saying a rule does not work. A method that eliminates candidates cheaply beats one that manufactures confidence. The common failure mode is not an edge that decays out of sample. It is an edge that was never present, hidden by a market that rose regardless of what anyone did. Report the base rate, the independent signal count and the execution assumption. Report the asset universe, the costs and the test period, then verify once on data nobody looked at. A backtest that survives that is rare and informative. A backtest that skips those steps is a description of the past dressed as a plan.

Why You Might Be Interested?

If a strategy was backtested on 2017 to 2021 crypto data, its returns include a market that paid 17.92% a month to do nothing. If a result rests on a few dozen signals, it is undecided rather than proven. If costs were omitted, a 1% round trip is enough to flip the best result measured here from positive to negative.

Of 88 tested configurations, 78 were profitable and only 2 beat the base rate.

Quick Stats

  • 17.92% vs 2.40% — unconditional 30-day return, 2017-2021 against 2022-2026
  • 2 of 88 — configurations beating the base rate in the period they were tuned on
  • -4.44 pp — mean edge over the base rate across all 88 configurations in-sample
  • 0.295 pp — cost of deferring execution by one bar, about a third of the edge
  • 78.6% — share of the 61,463 coins CoinPaprika tracks that are inactive
  • +0.23 to -0.77 pp — best configuration's edge before and after a 1% round-trip cost

Data current as of September 2026.

FAQ

?What is a crypto backtest and what can it prove?

It is a simulation of a fixed rule applied to historical price data for a fixed set of assets. It can show that a rule failed, which is genuine information and the technique's most reliable use. It cannot show that a rule will work again, because it contains nothing about periods it did not cover. Nor can it separate the rule's contribution from the market's drift unless the base rate is reported alongside.

?Why does the test period matter so much?

Because the market's unconditional return changes enormously between eras. For these eight assets the mean 30-day return was 17.92% from August 2017 to the end of 2021 and 2.40% from 2022 to September 2026. A strategy tested only on the earlier window inherits a drift 7.5 times stronger. It will report large returns whether or not the rule adds anything.

?What is overfitting in a trading strategy?

It is selecting parameters that fit the noise in one sample rather than any durable behaviour. The usual demonstration has a winner collapse out of sample. In this test something more instructive happened. Across 88 configurations only 2 ever beat the base rate in-sample, so the search never located an edge to overfit. The apparent profitability of the other 86 came from holding the assets.

?How many signals does a backtest need?

More than most report, and the count that matters is independent events rather than signal days. The best configuration here rested on 79 in-sample signals and 43 out-of-sample ones. That is thin enough for a handful of outcomes to move the result substantially. Anything below a few hundred independent events should be treated as undecided rather than confirmed.

?What is look-ahead bias and how large is it?

It means using information the strategy would not have had when it acted, such as entering at a close that had not yet printed. Measured across 889 signals, entry at the signal close returned 2.464%, at the next open 2.978% and at the next close 2.168%. Deferring by one bar cost 0.295 points, roughly a third of the measured edge, so the assumption must be stated.

?How does survivorship bias affect crypto backtests?

Badly, because the failure rate is extreme. CoinPaprika lists 61,463 coins and marks only 13,154 active, so 78.6% are inactive and appear in no test built from today's listings. Every asset used in this series' experiments still trades, which inflates all of those results in the same direction. Most retail-accessible data cannot correct for this, so the limitation should be stated rather than ignored.

?Is exchange OHLCV data clean enough to backtest on?

Mostly. The 23,610-bar sample had no missing days and no zero-volume bars. It did contain 97 bars, or 0.411%, where the high exceeded the low by more than 1.5 times. Those matter disproportionately for stop-loss rules and breakout entries, which key off exactly those extremes. Check whether results depend on a few unusual bars and re-run without them.

?Does walk-forward testing fix overfitting?

It helps and does not solve it. Fitting on one window and testing on the next produces a sequence of out-of-sample results instead of one. Parameter instability between windows is itself useful evidence. It does not fix survivorship, regime dependence or small samples, and it does not prevent a tester from repeating the whole procedure until something passes.

References / Sources

Sources
  • All experiments run first-hand on one Binance daily OHLCV pull covering eight USDT pairs, 23,610 completed bars, earliest 17 August 2017. Survivorship figures taken from CoinPaprika's public coin list on the same day.
  • - Binance: Spot API Kline/Candlestick Endpoint (binance.com, Sep 2026)
  • - Binance: Spot API Exchange Information Endpoint (binance.com, Sep 2026)
  • - CoinPaprika: Coins Endpoint (coinpaprika.com, Sep 2026)
Share this guide
Written by Bartek Hagan
Report a correction
All guides

Cryptocurrencies are highly volatile and involve significant risk. You may lose part or all of your investment.

All information on Coinpaprika is provided for informational purposes only and does not constitute financial or investment advice. Always conduct your own research (DYOR) and consult a qualified financial advisor before making investment decisions.

Coinpaprika is not liable for any losses resulting from the use of this information.