A verdict is only worth as much as the process behind it. Here is ours, including the parts that limit what we can honestly claim.
We test the strategy the way its advocates describe it — their entry signal, their exit rules, their holding period, and if it is an options strategy their deltas and tenors too. Testing a straw-man version and declaring it broken proves nothing. If the claim is vague, we test the most common interpretation and say so.
Entry, exit, position size, and what happens in the awkward cases — a gap straight through the stop, a trading halt, an assignment — are written down first. This is the step that stops a test drifting into an exercise in finding the settings that look best.
Not a favourable window. The full period, including the regimes where the strategy was never going to work. A strategy that only survives when you choose the years hasn't survived.
Average profit per trade after losses are counted at full size. A strategy can win nine times in ten and still lose money — that is the single most common way a bad strategy passes for a good one.
A result measured on the same history the rules were chosen from proves only that the rules fit that history. So part of the data is held back, and the verdict rests on how the strategy did there. Where there are enough trades, we walk it forward across the whole period instead of trusting one split.
Most strategies aren't good or bad, they're good inside a boundary. We look for the holding period, the entry threshold, or the market regime where the answer flips, and name it. "Conditional" is a real verdict, not a hedge.
Every test says what it can't tell you — before anyone asks. If we wouldn't be comfortable with a sceptic reading the caveats first, the test isn't ready.
"Sell puts on names you'd happily own. You either keep the premium or you buy a stock you wanted anyway. It can't really go wrong."
Every strategy arrives already believed. It comes with a story that makes sense, a name people recognise, and someone confident explaining it. None of that is evidence, and none of it is a reason to dismiss it either.
"It wins nine times in ten and the tenth takes back everything. You're picking up pennies."
The objection usually sounds just as reasonable. Both sides argue from stories and from the handful of trades they remember. Neither has counted.
"Neither of you has run it. I have."
That is the entire job of this site. Not to have an opinion about a strategy, but to run it across everything we have and publish what came out — including when the answer is boring, when it is inconvenient, and when it contradicts something we said about the same strategy last year.
See what's been testedPeople argue about "buying the dip", "the golden cross" or "the wheel" as though the name were the thing. It isn't. Two people who both say they buy breakouts can be running strategies with almost nothing in common, and both will be certain the other one is doing it wrong. Before anything can be tested, the claim has to be pinned to numbers.
What that takes depends on what is being traded.
Options need all of the above and more, because the contract itself has to be specified before it can be priced at all. For a credit spread that is four numbers, and all four have to be stated:
Width and credit together are the risk. Neither one alone tells you anything. Sell a spread $5 wide and collect $50, and you are risking $450 to make $50 — a shade over nine to one. Collect that same $50 on a spread $80 wide and you are risking $7,950 for it.
Same delta. Same tenor. Same credit. One of them is a trade nobody would knowingly place. A test that fixes only the credit quietly averages the two together and reports the blend as though it were a strategy.
Arithmetic, not a result. No figure on this page came out of a run — those live on the test pages, once there are any.
When a claim arrives without all of it — and most of them do — we test the most common interpretation, and the parameters table on that test says exactly what we chose. You are entitled to disagree with the choice. You should not have to guess at it.
The expensive mistake in testing is almost never a wrong formula. It is a parameter that quietly fell back to a default while the run carried on looking entirely normal. Ask for wings at 11% of the share price, get a fixed $5 spread on every underlying from $40 to $440, and the machine has answered a question nobody asked — accurately, and at length. The share-trading version is a position-size rule that silently reverts to a fixed number of shares, so the strategy you tested was never the strategy you wrote down.
Nothing about the output looks broken. The trades are real trades, the maths is right, the chart is convincing. It is simply not the strategy that was requested.
Before a verdict is published, the trades that came out are checked against the strategy that went in — one parameter at a time. A run that cannot prove it traded what it was asked to trade does not get a verdict, it gets run again. We would rather publish late than publish a confident answer to the wrong question.
A strategy that wins nine times out of ten sounds unarguable. Here is one that wins nine times out of ten and makes you nothing at all.
Ten trades. Nine of them hit the target and make $50. The tenth goes against you and runs all the way to the stop.
| What happens | Each | Total |
|---|---|---|
| 9 × hits the target | +$50 | +$450 |
| 1 × runs to the stop | −$450 | −$450 |
| Ten trades · 90% winners | $0 |
Nothing there is unusual. It is the shape of any strategy with a tight target and a distant stop — dip-buying on shares, mean-reversion, and very nearly every premium-selling strategy in options. Win small, win often, hand it back in one go. Move that loss to $500 and the same 90% win rate is a losing strategy. Move it to $400 and it is a good one. The win rate did not move at all.
Expectancy is the average profit per trade once losses are counted at full size — the only number that answers the question people actually care about, which is whether the thing makes money. Every verdict on this site leads with it. Win rate is reported, because people want to see it, and it is never what decides the verdict.
This is also the honest reason the site exists. A high win rate is exactly what makes a strategy popular — it feels like winning, every week, for months. Popularity is what puts a strategy on the list to be tested. It is not evidence of anything.
A number can be right and still mean nothing, if it rests on too few trades. Every verdict states how many trades it was measured over, and a thin sample gets said out loud rather than being averaged into confidence. Where the sample is too thin to call, the verdict stays Pending — which is why Pending is a real verdict here and not an empty slot.
Here is the uncomfortable fact about backtesting: if you try enough variations of a strategy, one of them will look excellent by pure chance. Nothing about that result is fraudulent. The trades are real, the arithmetic is right, and the equity curve goes up and to the right. It simply will not do it again.
A backtest that has been adjusted until it passes is not evidence. It is a record of adjusting. So an expectancy number does not become a verdict until it has survived four checks.
The rules are fixed on one part of the history, then run on a part the strategy has never seen. We split roughly seventy-thirty, and we split on trades rather than on dates, so both halves carry a comparable amount of trading rather than one half happening to contain a quiet year.
The out-of-sample half is the one that counts. A strategy that looks strong on the first part and comes apart on the second was fitted to the past, not found in it.
The distance between the in-sample result and the out-of-sample result is itself a number worth seeing, so we report it rather than only reporting the half that flatters. A small gap means the strategy travelled. A wide one means the first number described a particular stretch of history and not the strategy — and a strategy can be profitable in both halves and still fail here, if the drop between them is steep enough.
One split is still one draw, and a single lucky cut of the data can flatter a strategy just as easily as a single unlucky one can bury it. Walking forward repeats the exercise across the whole history: fix the rules on a window, test on the window immediately after it, roll both forward, and do it again to the end.
What comes out is a chain of out-of-sample results from many different market conditions instead of one verdict resting on an arbitrary boundary. It is the closest a backtest gets to having actually traded the thing. It also needs a great many trades to mean anything, so it runs automatically when a test has enough — and when a test does not have enough, the page says walk-forward was not available rather than quietly leaving it out.
If we test forty variations of a strategy and publish the best one, "the best of forty" is a different claim from "this works", and reporting the first as though it were the second is how a testing site turns into a machine for manufacturing impressive flukes. The more variations a test explores, the higher the bar its winner has to clear before it earns a verdict.
Alongside this we shuffle the order of the trades many times over and look at the spread of outcomes. It answers something the headline figure hides — how much of the result depended on the wins and losses happening to arrive in a favourable order, and how bad the worst stretch could reasonably have been rather than how bad it happened to be.
A strategy can be clearly profitable across the full history and still not earn Holds — because it did not survive on the data it had never seen. That is not a contradiction, it is the check working. When that happens the test says so plainly and shows both numbers.
And where there are too few trades to run the checks properly, the answer is not a cautious pass. The verdict stays Pending until there is enough to test, however good the raw figure looks.
Positive expectancy on data the strategy was never fitted to, surviving reasonable changes to the assumptions. Not a recommendation — a statement that the claim stood up when it was checked properly.
Works inside a stated boundary and fails outside it. The boundary is always named. Most honest strategies land here.
Negative expectancy, or positive only under assumptions no one could actually trade. Often paired with a high win rate — that's usually the point.
Everything up to here applies to any strategy on any instrument. This part does not. An option has a defined maximum loss, an expiry and a premium, and each of those introduces a way of counting risk that is wrong often enough to be worth naming. Both of these turn up in published strategy advice, in trading software, and in our own work.
On a debit trade the two are the same thing: the most you can lose is what you paid. On a credit trade they are nowhere near each other. Collect $0.55 on a spread $5 wide and you have staked $55 while risking $445. A stop set at "100%" sounds like it protects the whole position, and it fires at about an eighth of the real risk.
Neither setting is wrong. Believing you have the second one when you have the first is. Every exit rule we test states which of the two it is measured against.
The price can only finish on one side. Both wings cannot be breached by the same expiry, so the risk is one side's width less everything you collected. Wings $5 wide with $100 taken in is $400 at risk, not $900. It is one of the most commonly traded structures there is, and doubling its apparent risk changes every ratio anyone computes from it.
Neither trap has an equivalent in a share strategy, where the most you can lose is what the position is worth and there is no expiry deciding it for you. That is the reason the two are separated on this page rather than blended into one list of rules: a method that pretends every instrument behaves the same way is already wrong before it runs.
| Price data source | — |
|---|---|
| Options data source | — |
| Period covered | — |
| Underlyings | — |
| Pricing basis | — |
| Commissions modelled | — |
A backtest is a record of what a set of rules would have produced against historical prices. It is evidence, not proof. It tells you whether a claim has ever been true, which is a much lower bar than whether it will be true — and a much higher bar than most strategy advice clears.
Historical price series are usually adjusted twice over — once for splits, once for dividends. We adjust for splits, because otherwise a four-for-one split reads as a 75% crash and every condition built on it fires nonsense. We do not adjust for dividends.
The reason is that you would have traded at the number that was actually on the screen that morning, not one rewritten afterwards to account for payouts nobody had received yet. A test that picks its trades off a dividend-adjusted chart and then fills them at real prices is quietly running two different histories against each other. It is a small discrepancy that gets everywhere and is far easier to avoid than to find later. It bites hardest on options, which are priced off the underlying's real level, but it distorts a share strategy's entry signals just as readily.
We will not publish a figure that didn't come out of a run. We will not quietly change parameters until a strategy passes. We will not sell a strategy we graded, and we take no payment from anyone whose strategy we test — not to run one, not to delay one, and not to leave one unpublished. A result that is inconvenient for us goes up exactly like any other.
Every test publishes its parameters. Anyone with the same data and the same rules should reach the same verdict — and if you run it and don't, we want to hear about it.