How we judge a strategy
This is the method. There is no secret indicator: the edge, if it exists, comes from the process. Everything below is applied to every hypothesis before anyone looks at a result.
1. Pre-registration: the rule is written before looking
A hypothesis only enters testing if it is falsifiable: entry, exit, invalidation and sizing written without ambiguity, with the fewest free parameters possible (ideally zero). The pass and fail criteria are fixed and saved with a date before anything runs. If a result does not meet what was written, the ground is marked barren. Nothing is renegotiated afterwards.
Each hypothesis also states its economic rationale: why the edge should exist and who is on the other side losing it. If there is no convincing answer, it is born weak and that is recorded.
2. Real costs, always
No result is reported without commissions, slippage and, where it applies, funding or borrow costs. It is also measured with costs doubled: an edge that does not survive twice the friction is not an edge.
3. Out of sample and across time
The fitting period is separated from the validation period, and the effect must show up in the stretch the original author could not see. Anti-decay asks about the agent: once an anomaly is published, someone arbitrages it, and the normal outcome is that it arrives exhausted. An effect that only exists before its publication is not tradable.
4. Against chance: nulls that do not give away the drift
Buying US stocks between 2009 and 2026 validates almost any rule. So each result is compared against shuffled versions of the same strategy that keep the market drift and the calendar composition and destroy only the selection. If the rule cannot be told apart from picking at random within its own universe, the ground is barren. Information is not the same as a trade: a variable predicting something does not mean the prediction can be collected.
5. Multiplicity: every attempt is paid for
Testing many variants and keeping the best one manufactures results. Every hypothesis declares how many tests it ran and which mechanism family it belongs to, and the statistical threshold gets stricter with the accumulated number of attempts in that family, including those from other sessions and other days. A family with several barren members demands more from the next one. Closed families are not reopened without genuinely new data.
6. Survivorship bias: an explicit discount
We measured how much a backtest is inflated by using only the companies that are still alive: in our data, more than half of the apparent Sharpe ratio. That is why universes are rebuilt date by date, including the companies that were delisted, and where that is not possible a 0.6 discount is applied to the Sharpe of the survivors. A result without that discount is a promise, not a measurement.
7. Colliders and filters that are consequences
A filter that improves the fit a lot or flips signs is suspect: it is usually a consequence of the return, not a cause. Every condition in a rule must be measurable before entry and explain itself. It is declared in the pre-registration and reviewed in the field report.
8. One look, one bullet
When a test consumes something non-renewable, such as the first look at virgin data, it runs exactly once, with a date and with guards that prevent spending it by accident. Looking twice at the same virgin data is mining, not science.
9. Nothing is deleted
Discarded hypotheses keep their full record: the numbers, the criterion that they failed and what would bring us back. The record of barren ground is the real control against multiple testing, and it is what this site publishes.