# Your Backtest Isn't an Edge — Until You Correct for How Many Times You Tried

*Published 2026-07-04 · Apex (smartapex.uk) · 5 min read*

## TL;DR

- Most edges that look good in a backtest are **multiple-testing overfitting**. Correct for "I tried many variants and kept the best one" and they collapse.
- I watched this happen **live** for 5 weeks. A strategy that backtested at +0.33% avg PnL died in live tracking at **t = −2.19**.
- You can diagnose your own strategy in 30 seconds. No install, no signup — just your returns array.

## The moment 91.7% becomes 38.9%

Say you have a returns series, a decent Sharpe, and it looks "almost statistically significant" — for example a PSR (probability the true Sharpe > 0) of **91.7%**.

But what if that strategy was the best of **12 variants** you tried? The moment you correct for that fact (the Deflated Sharpe Ratio), 91.7% collapses to **38.9%**. If you roll the dice 12 times, you get a Sharpe that good by pure luck even when the true edge is zero.

This is the most common trap in retail algo trading: **sweep parameters, symbols, and windows; keep the prettiest curve; believe it's a discovery.** Academia has addressed this head-on — Harvey & Liu (2015) argue the effective significance bar under multiple testing is not t≈2 but **t≈3**; Bailey & López de Prado (2014) proposed the Deflated Sharpe Ratio, which discounts Sharpe by the number of trials.

## Live evidence: how a +0.33% backtest died

Theory is theory. I went through this **in real time**.

`stillme` is a DATA ONLY research system that collects 750+ Binance USDT-M futures symbols around the clock (no orders, no live trading — pure validation). A momentum strategy (OI-flow v2) looked promising in backfill backtests: **+0.33% average PnL**.

So I put it behind a live gate and watched for 5 weeks:

| Metric | Backtest (in-sample) | Live |
|---|---|---|
| Avg net PnL | +0.33% | **−0.19% to −0.51%** |
| t-statistic | (looked fine) | **−2.19** (significantly negative) |
| Gate | — | **failed 4/4** |

The backfill's +0.33% was a **pure in-sample illusion** — the SHORT-only rule itself had been chosen *after looking at* the backtest data. Live told the truth, and the truth was t = −2.19.

This is still public, on a dashboard that does not hide its failures:

**→ Live dashboard: https://stillme.net** (shows the 4/4 gate failure as-is)

## Diagnose your strategy in 30 seconds

You can run the same diagnostic on your own returns. No install; a returns array is enough.

```bash
curl -sS -X POST 'https://smartapex.uk/api/tools/deflated-sharpe/run' \
  -H 'Content-Type: application/json' \
  -d '{"returns":[0.01,-0.002,0.004,0.006,-0.001],"n_trials":12}'
```

Put your per-trade returns in `returns` and the number of variants you actually tried in `n_trials`. You get back:

- **Sharpe / t-stat** — surface performance
- **PSR** — naive significance
- **DSR** — probability of a real edge *after multiple-testing correction* ← this is the one that matters
- **MinTRL** — minimum track record length needed for confidence

If DSR is low (say, under 0.5), your "edge" is likely a six rolled on the twelfth die. Knowing that before you fund the strategy is much cheaper than after.

## Why trust this tool — Apex

This diagnostic is one of the verified component cards on **Apex** (https://smartapex.uk) — a trust layer where AI agents discover and run production-validated code capabilities instead of scraping random GitHub repos. The deflated-sharpe tool never exposes raw source; it runs behind a bounded read-only wrapper, with public verification reports and usage reviews. And the stillme dashboard above is this card's **origin production system** — provenance you can check live, not marketing.

If you use an AI agent, it can call Apex directly over MCP:

```json
{ "mcpServers": { "apex": { "url": "https://smartapex.uk/api/mcp" } } }
```

Every anonymous tool run returns a verification receipt (hashes of input/output, checks performed, boundaries). Runs signed with a GitHub-backed agent passport also accumulate reputation — the receipt itself tells your agent how.

## One line

**A pretty backtest can be trial count, not skill. Before real money, ask in numbers whether it's edge or luck.** The curl above is all it takes.

---

### References

- Bailey & López de Prado (2014), "The Deflated Sharpe Ratio"
- Harvey & Liu (2015), "Backtesting" — the t≈3 bar under multiple testing
- Live evidence: https://stillme.net · Tool: https://smartapex.uk

*Disclaimer: educational statistics tooling. Not investment advice; no strategy's future returns are guaranteed.*
