Read it. Rebuild it. Break it.
The method is deliberately boring and it does not change between studies. A method that changes to suit the result is not a method, it is a search for the answer you wanted.
Every study goes through the same door.
- ReadTake the paper at its word and write down exactly what it claims. Which instruments, which period, which rule, which holding time, measured how. If the claim cannot be written down precisely enough to fail, there is nothing to test and we stop here.
- ReproduceRebuild the result from the described data and rules, with no shortcuts and no help from hindsight. If we cannot get near the published number on the published sample, that is the finding, and it happens more often than the literature would suggest.
- Test out of sampleRun it on data the authors never saw, over periods they never tested, with costs and slippage that a real account would pay. The sample they fitted on is the one place the result is guaranteed to look good.
- DiscardThrow away what does not survive, and record that it did not. A negative result is cheap to produce and expensive to rediscover, which makes it worth writing down properly.
What that phrase means here.
It is used loosely almost everywhere, so it is worth being exact. For us it means the rule was fixed before the data was touched. Not tuned on the training period and then shown on the test period. Not re-run with a different parameter because the first pass looked poor. Fixed, then run, then reported, including when the answer is that it does not work.
The honest version of this is uncomfortable. Every time you go back and adjust something after seeing the test set, the test set becomes training data and the number you report stops meaning what you think it means.
The four we check for first.
None of these are exotic. They are just easy to do by accident, and each one produces a beautiful backtest.
- LookaheadUsing a number at a time it was not yet known. Revised data, close prices in a signal computed at the open, an index constituent list as it is now rather than as it was.
- SurvivorshipTesting on the names that are still listed. The ones that went to zero are the whole point and they are missing from most convenient datasets.
- Multiple testingTry two hundred variations and the best one will look excellent for reasons that have nothing to do with markets.
- CostsSpread, fees, slippage and the fact that the size you assumed was available often was not. Many published effects are real and still smaller than the cost of capturing them.
The method is why the products look the way they do.
Everything in QFI Terminal and DeepAlpha that could be mistaken for a prediction is labelled for what it is. A detector that marks a level says it marks a level. A trained model that scores a setup reports a score, not an instruction. Where we know the limits of something, the limits are written next to it inside the product, not buried in a page like this one.