Every line on this page is published finance research. Across forty-six years the literature describes three failures — and they lock together into a deadlock that no amount of disclosure can open. That deadlock is the reason MIZAN exists, and the reason nobody closed it sooner.
Every correction for overfitting needs the number of strategies tried. It is self-reported.
Returns arrive from the person being judged, and the literature shows they bend. Inspection is required.
An edge that is disclosed stops being an edge. Inspection is forbidden. This is a result in economics, not a preference.
Every serious correction for backtest overfitting takes one input: how many strategies were tried before the one being shown. Here is the field arriving at that conclusion across twenty-four years — and stopping there.
The origin of the field's correction machinery: a formal test for whether the best rule found in a large search is better than chance, once the full extent of the search is accounted for.
The Reality Check — and Hansen's sharper Superior Predictive Ability test built on it (2005) — is the statistical answer to data snooping. Every one of these tests takes the size of the search as an input. The machinery has existed for a quarter century; the input has never been checkable.
Characterisation of the works, not a verbatim quotation.
The building blocks of the Sharpe ratio — expected returns and volatilities — are unknown quantities that must be estimated statistically and are, therefore, subject to estimation error.
Lo showed a fund's annual Sharpe ratio can be overstated by as much as 65% through serial correlation alone, and that a monthly Sharpe cannot simply be annualised by √12. The headline number was unreliable before anyone tried to game it.
Because most financial analysts and academics rarely report the number of configurations tried for a given backtest, investors cannot evaluate the degree of overfitting in most investment proposals.
This is the sentence MIZAN exists to answer. Four mathematicians, in the journal of the American Mathematical Society, stating that the one input needed to judge a track record is the one input nobody supplies. They also showed that high simulated performance is achievable after trying only a small number of configurations — and that the more you try, the more certain the overfit.
The same authors then built the correction itself: the Deflated Sharpe Ratio — the probability that a record's Sharpe is genuine once the number of trials, the sample length, and the non-normality of returns are all accounted for.
The DSR deflates a reported Sharpe against the expected maximum of N random trials. It is the correction this entire page converges on — and its decisive input is N, the trial count, which the 2014 AMS paper had just shown nobody reports. A complete cure, missing its one ingredient, for twelve years.
Characterisation of the work, not a verbatim quotation.
"Highly significant" backtested performance is easy to generate by selecting stocks on the basis of combinations of randomly generated signals, which by construction have no true power.
Elsewhere Novy-Marx made the same point with signals drawn from the position of Mars and Saturn, and from sunspots. Statistical significance, alone, carries no information about whether an edge is real.
Most claimed research findings in financial economics are likely false.
Their remedy was to raise the bar: a newly discovered factor should clear a t-ratio above 3.0 rather than the conventional 2.0 — precisely because so many factors were tried before the one published. The correction depends on knowing how many. Nobody knows how many.
The more backtesting a quant has done for a strategy, the larger the discrepancy between backtest and out-of-sample performance.
Not theory — 888 real algorithms, each with at least six months of genuine out-of-sample life. Their other finding is the one that should trouble every allocator: the backtest Sharpe ratio predicted out-of-sample performance with R² below 0.025. The number the industry screens on carries almost no information.
The strongest case for the defence: a Bayesian re-examination arguing that most published equity factors do replicate, and that the field's alarm is overstated.
Included against interest, because this page would be dishonest without the rebuttal. Note what even the defence must assume: a model of how many factors were tried. Both sides of the replication debate deflate against a trial count — they differ on its value, because it is unobservable. The debate itself is evidence for the missing mechanism: with N committed rather than assumed, this argument becomes arithmetic.
Characterisation of the work, not a verbatim quotation.
Previously reported LLM advantages deteriorate significantly under broader cross-section and over a longer-term evaluation.
The same reviews report models reaching 100% in-sample accuracy and roughly 50% out-of-sample — a coin flip, dressed as a result. What took a quant months in 2014 now takes an afternoon. The number of configurations tried has exploded. The mechanism for reporting it has not changed at all.
Marked as characterisation of a body of work, not a single verbatim quotation — unlike every entry above.
Act I assumes the reported returns are themselves honest. The literature does not.
Reported hedge-fund returns are systematically smoother than the economic reality beneath them — and the smoothing inflates Sharpe ratios.
Serial correlation in reported returns traces to illiquidity and discretionary smoothing of marks — losses spread quietly across months. The volatility an allocator sees is lower than the volatility that exists, so the Sharpe arrives pre-inflated before any strategy question is even asked.
Characterisation of the work, not a verbatim quotation.
The number of small gains far exceeds the number of small losses… The discontinuity is absent in the 3 months culminating in an audit.
The most important finding on this page. Reported returns show a kink exactly at zero: small losses quietly become small gains. The kink vanishes in the months before an audit, and the performance reverses afterwards. Present in live funds and dead ones, funds of every age — so it is not a database artifact.
Verification changes behaviour. That was demonstrated on real data in 2009 — and the only thing being verified was the accounting, never the research.
Hedge funds report significantly higher returns in December than in other months — and the December spike concentrates exactly where incentives to manage year-end numbers are strongest.
Independent confirmation of Bollen–Pool by a different mechanism: returns bend around the dates that matter — year-end marks, incentive-fee crystallisation. The pattern strengthens with discretion over valuation. Same conclusion from a second team: what the manager reports responds to what the manager needs.
Characterisation of the work, not a verbatim quotation.
Operational risk increases the likelihood of subsequent poor performance and fund disappearance, but does not influence investors' return-chasing behavior.
Built from actual due-diligence reports. Many funds carry operational problems including limited disclosure of legal and regulatory matters — and the signal predicts failure while investors chase returns anyway. Allocators are not careless. They are working from what the manager hands them.
All commercial hedge fund databases are based on hedge funds' voluntary reporting.
Managers choose whether to report, and typically begin after a strong run — at which point the entire prior history is backfilled in. Funds that die stop reporting. The industry's benchmark data is assembled by the people being benchmarked.
So the allocator needs to see inside. Here is why the quant can never allow it.
There is an equilibrium degree of disequilibrium: prices reflect the information of informed individuals but only partially, so that those who expend resources to obtain information do receive compensation.
If prices fully reflected available information, nobody would be paid for gathering it, and the gathering would stop. An edge has value only while it is not fully revealed — which makes the quant's refusal to disclose a structural feature of functioning markets, not stubbornness.
Stiglitz received the Nobel in 2001. "Prove your edge by showing me your edge" asks a manager to destroy the asset in order to sell it.
Its modern institutional form: every road to being believed demands the strategy — the allocator's diligence team, the verification firm, the platform holding code in custody. The NDA offered in exchange cannot protect a number whose value dies on contact: damages are unprovable, knowledge is irreversible, and no contract un-teaches the reader. So the quants with real edges rationally refuse — and disclosure-based verification adversely selects for the strategies not worth protecting. The examiner sees everything except what matters.
Every mechanism the industry built picks a side. Audits verify the accounting and never touch the research. Managed accounts show what a manager does after you allocate, never how the record that won the mandate was produced. Platforms that see the strategy own it, or compete with it. Disclosure regimes ask politely and receive whatever the manager chooses to write.
The two requirements are jointly unsatisfiable by any disclosure-based mechanism. Not for want of will — the field has been unanimous since 2014 — but because satisfying both at once means proving a statement about data without exhibiting the data.
That is a cryptographic capability, and it did not exist in usable form until recently. The problem survived four decades because the tool to close it had not been built.
A manager commits the entire candidate set before evaluation runs. The trial count becomes the leaf count of that commitment — a property of a Merkle tree, not a figure anyone supplies. The proof forces the presented winner to be the maximum of the committed set, and recomputes the deflation on price data the verifier re-derives independently. Verification takes milliseconds, on the verifier's own machine, trusting no issuer.
Both requirements are met at once. The allocator inspects a proof instead of trusting a claim, so Act II is answered. The strategy never leaves the quant's machine, so Act III is preserved. Neither side concedes anything, because the proof carries the truth and leaves the secret behind.
The literature specified the correction. It could not specify a way to trust the input, or a way to look without taking. VTR-1 (published on SSRN) makes the input checkable without making the strategy visible.
Committing a search proves how many strategies were tried and that the winner was genuinely the best of them. It does not prove the searcher was honest about the future. Only a public time-anchor placed before the data existed can do that — which is why credentials are timestamped, not merely signed.
Nor can it make a bad strategy good. The first verdict this engine ever published was a refusal of its author's own flagship: a Deflated Sharpe Ratio of 0.68 against a 0.95 bar. That result is public, with its proof, and it is the point.