Research
How to validate a trading edge
A trading edge is validated when its results survive tests designed to break it. Expectancy tells you whether the average trade is positive. Sample size tells you whether that average means anything. Out-of-sample data tells you whether the pattern was discovered or invented. A permutation test tells you whether the same result could have arisen from noise.
A record that has not been through those four tests describes the past and settles nothing about the future.
Expectancy is the place to start
An equity curve rising from left to right is the least informative summary of a strategy. It compresses every decision into one line and hides the distribution underneath. Expectancy is the better starting point, because it states what the average trade produces.
Expectancy is the win rate multiplied by the average win, minus the loss rate multiplied by the average loss. Expressed in units of risk, where 1R is the amount risked per trade, a strategy that wins 45 percent of the time at 2R per win and loses 1R per loss has an expectancy of:
(0.45 x 2) - (0.55 x 1) = 0.90 - 0.55 = 0.35R per trade
That figure is what the strategy is worth per decision before costs. It also shows why win rate alone is meaningless. The same 45 percent win rate at a 1:1 reward-to-risk ratio produces an expectancy of minus 0.10R, and the strategy loses money while winning almost half its trades.
Costs come out of that number directly. If the average round trip gives up a quarter of an R to spread, commission and slippage, the 0.35R expectancy becomes 0.10R. A strategy can be genuinely correct about market direction and still be unprofitable once execution is priced in.
Sample size decides whether expectancy means anything
An expectancy calculated from thirty trades and an expectancy calculated from three thousand trades are different kinds of statement. The arithmetic is identical. The confidence is not.
The reason is that the uncertainty in an average shrinks in proportion to the square root of the number of observations. Quadrupling the number of trades halves the uncertainty in the estimated expectancy. Getting the uncertainty down by a factor of ten requires a hundred times the trades.
This has a practical consequence that is easy to state and uncomfortable to accept. For a strategy whose individual trade outcomes vary widely, a few dozen trades cannot distinguish an expectancy of 0.35R from an expectancy of zero. The record is compatible with both. Traders often present such a record as proof, in good faith. The record simply does not contain enough information to settle the question.
Two habits help here. The first is to report the number of trades alongside every performance figure, so the reader can judge the weight of the claim. The second is to treat any strategy with few trades as a hypothesis awaiting data, whatever its return.
Separate discovery data from confirmation data
Most strategies are found by looking. A researcher tries parameters, filters and conditions until something works. The process is legitimate and it is how most ideas begin. It also guarantees that the resulting performance on the data used for searching is optimistic, because the search selected for whatever happened to work on exactly that data.
The correction is to reserve data that took no part in the search. A common arrangement splits history into a development period and a held-back period, and the held-back period is examined once, after the strategy is finished. Examining it repeatedly turns it into development data, and the protection is gone.
Two failure patterns show up when the held-back period is finally opened:
- The strategy performs far worse. The pattern was a property of the development window rather than of the market.
- The strategy performs similarly but with a different shape, for instance the same return arriving through fewer, larger trades. This is worth understanding before deploying, because the risk profile has changed even though the summary figure has not.
A strategy tested only on the data that produced it has been fitted to that data. Validation requires a period the strategy never saw.
Ask whether noise could have produced the same result
Out-of-sample testing answers whether a pattern generalises across time. It does not answer a different question: whether the specific market structure the strategy claims to exploit is doing the work.
A permutation test addresses that. The method is to break the structure being claimed, while preserving the broad statistical character of the data, then rerun the unchanged strategy many times on the altered series. Each run produces a result. Together they describe what the strategy would earn from data where the claimed effect has been removed.
The real result is then compared against that distribution. If it sits comfortably inside the range, the strategy is consistent with having exploited nothing in particular. If it sits far outside, the claimed structure is doing measurable work.
What to permute depends on the claim. A strategy that trades a relationship between two instruments can have that relationship destroyed by shuffling the instruments independently within each day, which leaves each instrument's own behaviour intact while removing any link between them. A strategy that claims to predict direction can have direction shuffled at a fixed frequency. The test is only as sharp as the thing it destroys, so the permutation has to attack the actual mechanism rather than something adjacent to it.
Check that the exit assumptions are honest
Validation usually concentrates on entries, and exits are where most inflated results are manufactured. The mechanism is subtle enough to survive several rounds of review.
Consider a bar in which price touches both the stop level and the target level. The bar records a high and a low. It does not record which came first. A simulator has to choose, and the choice is worth a large fraction of the reported return in strategies with tight stops. Resolving the bar in the strategy's favour produces results no broker would have filled.
The same issue appears when a position is opened and closed inside a single bar, when an order is assumed to fill at a price that was only touched once, and when a limit order is assumed to fill because price reached its level without trading through it. Each assumption is individually small and defensible. Together they can account for the entire edge.
The direct fix is to simulate at a finer resolution than the decision timeframe, so the sequence inside the bar is observed instead of assumed. Where finer data is unavailable, the alternative is to resolve every ambiguous bar against the strategy and accept the pessimistic number as the honest one.
The order matters
These tests are cheap in a specific sequence and expensive in any other.
- Compute expectancy after realistic costs. If it is not positive, nothing further is needed.
- Count the trades and state the uncertainty. If the sample cannot distinguish the expectancy from zero, the answer is more data rather than more analysis.
- Fix the execution assumptions before running anything else, because every subsequent test inherits them. A permutation test on an optimistic simulator produces a confident answer about a fiction.
- Open the held-back period once.
- Permute the claimed mechanism and compare.
A strategy that survives all five has failed to be broken by the tests that break most strategies. That is the strongest available statement about something that has to work in the future.
Common questions
undefined
undefined
undefined
undefined
undefined
undefined
undefined
undefined
Leplace Capital
Leplace Capital validates trading edge, converts strategies into algorithms and allocates capital progressively. The five stages are set out on the process page.
Read the process