The problem · documented by the field itself

Not our claim.
Their findings.

The argued version — the problem document, with every source openable →

Every line on this page is published finance research. Across forty-six years the literature describes three failures — and they lock together into a deadlock that no amount of disclosure can open. That deadlock is the reason MIZAN exists, and the reason nobody closed it sooner.

I · The count

Every correction for overfitting needs the number of strategies tried. It is self-reported.

II · The allocator

Returns arrive from the person being judged, and the literature shows they bend. Inspection is required.

III · The quant

An edge that is disclosed stops being an edge. Inspection is forbidden. This is a result in economics, not a preference.

§ 01
Act I

The number nobody can check

Every serious correction for backtest overfitting takes one input: how many strategies were tried before the one being shown. Here is the field arriving at that conclusion across twenty-four years — and stopping there.

2000

The origin of the field's correction machinery: a formal test for whether the best rule found in a large search is better than chance, once the full extent of the search is accounted for.

Halbert White · "A Reality Check for Data Snooping"
Econometrica 68(5), 2000 · applied to 7,846 trading rules in Sullivan, Timmermann & White (1999)

The Reality Check — and Hansen's sharper Superior Predictive Ability test built on it (2005) — is the statistical answer to data snooping. Every one of these tests takes the size of the search as an input. The machinery has existed for a quarter century; the input has never been checkable.

Characterisation of the works, not a verbatim quotation.

2002

The building blocks of the Sharpe ratio — expected returns and volatilities — are unknown quantities that must be estimated statistically and are, therefore, subject to estimation error.

Andrew W. Lo · "The Statistics of Sharpe Ratios"
Financial Analysts Journal 58(4), 2002

Lo showed a fund's annual Sharpe ratio can be overstated by as much as 65% through serial correlation alone, and that a monthly Sharpe cannot simply be annualised by √12. The headline number was unreliable before anyone tried to game it.

2014

Because most financial analysts and academics rarely report the number of configurations tried for a given backtest, investors cannot evaluate the degree of overfitting in most investment proposals.

David H. Bailey · Jonathan M. Borwein · Marcos López de Prado · Qiji Jim Zhu
"Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance"
Notices of the American Mathematical Society, May 2014

This is the sentence MIZAN exists to answer. Four mathematicians, in the journal of the American Mathematical Society, stating that the one input needed to judge a track record is the one input nobody supplies. They also showed that high simulated performance is achievable after trying only a small number of configurations — and that the more you try, the more certain the overfit.

2014

The same authors then built the correction itself: the Deflated Sharpe Ratio — the probability that a record's Sharpe is genuine once the number of trials, the sample length, and the non-normality of returns are all accounted for.

David H. Bailey · Marcos López de Prado · "The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality"
The Journal of Portfolio Management 40(5), 2014

The DSR deflates a reported Sharpe against the expected maximum of N random trials. It is the correction this entire page converges on — and its decisive input is N, the trial count, which the 2014 AMS paper had just shown nobody reports. A complete cure, missing its one ingredient, for twelve years.

Characterisation of the work, not a verbatim quotation.

2015

"Highly significant" backtested performance is easy to generate by selecting stocks on the basis of combinations of randomly generated signals, which by construction have no true power.

Robert Novy-Marx · "Backtesting Strategies Based on Multiple Signals"
NBER Working Paper 21329, July 2015

Elsewhere Novy-Marx made the same point with signals drawn from the position of Mars and Saturn, and from sunspots. Statistical significance, alone, carries no information about whether an edge is real.

2016

Most claimed research findings in financial economics are likely false.

Campbell R. Harvey · Yan Liu · Heqing Zhu · "…and the Cross-Section of Expected Returns"
The Review of Financial Studies 29(1), 5–68

Their remedy was to raise the bar: a newly discovered factor should clear a t-ratio above 3.0 rather than the conventional 2.0 — precisely because so many factors were tried before the one published. The correction depends on knowing how many. Nobody knows how many.

2016

The more backtesting a quant has done for a strategy, the larger the discrepancy between backtest and out-of-sample performance.

Thomas Wiecki · Andrew Campbell · Justin Lent · Jessica Stauth (Quantopian)
"All That Glitters Is Not Gold: Comparing Backtest and Out-of-Sample Performance on a Large Cohort of Trading Algorithms"
The Journal of Investing 25(3), 69

Not theory — 888 real algorithms, each with at least six months of genuine out-of-sample life. Their other finding is the one that should trouble every allocator: the backtest Sharpe ratio predicted out-of-sample performance with R² below 0.025. The number the industry screens on carries almost no information.

2023

The strongest case for the defence: a Bayesian re-examination arguing that most published equity factors do replicate, and that the field's alarm is overstated.

Theis Ingerslev Jensen · Bryan Kelly · Lasse Heje Pedersen · "Is There a Replication Crisis in Finance?"
The Journal of Finance 78(5), 2023

Included against interest, because this page would be dishonest without the rebuttal. Note what even the defence must assume: a model of how many factors were tried. Both sides of the replication debate deflate against a trial count — they differ on its value, because it is unobservable. The debate itself is evidence for the missing mechanism: with N committed rather than assumed, this argument becomes arithmetic.

Characterisation of the work, not a verbatim quotation.

2025 – 2026

Previously reported LLM advantages deteriorate significantly under broader cross-section and over a longer-term evaluation.

Characterising the recent literature, incl. "Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?" · arXiv:2505.07078, 2025

The same reviews report models reaching 100% in-sample accuracy and roughly 50% out-of-sample — a coin flip, dressed as a result. What took a quant months in 2014 now takes an afternoon. The number of configurations tried has exploded. The mechanism for reporting it has not changed at all.

Marked as characterisation of a body of work, not a single verbatim quotation — unlike every entry above.

§ 02
Act II

The allocator cannot verify

Act I assumes the reported returns are themselves honest. The literature does not.

2004

Reported hedge-fund returns are systematically smoother than the economic reality beneath them — and the smoothing inflates Sharpe ratios.

Mila Getmansky · Andrew W. Lo · Igor Makarov · "An Econometric Model of Serial Correlation and Illiquidity in Hedge Fund Returns"
Journal of Financial Economics 74(3), 2004

Serial correlation in reported returns traces to illiquidity and discretionary smoothing of marks — losses spread quietly across months. The volatility an allocator sees is lower than the volatility that exists, so the Sharpe arrives pre-inflated before any strategy question is even asked.

Characterisation of the work, not a verbatim quotation.

2009

The number of small gains far exceeds the number of small losses… The discontinuity is absent in the 3 months culminating in an audit.

Nicolas P. B. Bollen · Veronika K. Pool · "Do Hedge Fund Managers Misreport Returns? Evidence from the Pooled Distribution"
The Journal of Finance 64(5), 2257–2288, October 2009

The most important finding on this page. Reported returns show a kink exactly at zero: small losses quietly become small gains. The kink vanishes in the months before an audit, and the performance reverses afterwards. Present in live funds and dead ones, funds of every age — so it is not a database artifact.

Verification changes behaviour. That was demonstrated on real data in 2009 — and the only thing being verified was the accounting, never the research.

2011

Hedge funds report significantly higher returns in December than in other months — and the December spike concentrates exactly where incentives to manage year-end numbers are strongest.

Vikas Agarwal · Naveen D. Daniel · Narayan Y. Naik · "Do Hedge Funds Manage Their Reported Returns?"
The Review of Financial Studies 24(10), 2011

Independent confirmation of Bollen–Pool by a different mechanism: returns bend around the dates that matter — year-end marks, incentive-fee crystallisation. The pattern strengthens with discretion over valuation. Same conclusion from a second team: what the manager reports responds to what the manager needs.

Characterisation of the work, not a verbatim quotation.

2012

Operational risk increases the likelihood of subsequent poor performance and fund disappearance, but does not influence investors' return-chasing behavior.

Stephen J. Brown · William N. Goetzmann · Bing Liang · Christopher Schwarz · "Trust and Delegation"
Journal of Financial Economics 103(2), 221–234, 2012

Built from actual due-diligence reports. Many funds carry operational problems including limited disclosure of legal and regulatory matters — and the signal predicts failure while investors chase returns anyway. Allocators are not careless. They are working from what the manager hands them.

2000 – 2003

All commercial hedge fund databases are based on hedge funds' voluntary reporting.

Established across the database-bias literature · backfill effect estimated at 1.4% (Fung & Hsieh, 2000) and 0.4% (Barry, 2003) · survivorship bias measured above 2% per year

Managers choose whether to report, and typically begin after a strong run — at which point the entire prior history is backfilled in. Funds that die stop reporting. The industry's benchmark data is assembled by the people being benchmarked.

§ 03
Act III

The quant cannot reveal

So the allocator needs to see inside. Here is why the quant can never allow it.

1980

There is an equilibrium degree of disequilibrium: prices reflect the information of informed individuals but only partially, so that those who expend resources to obtain information do receive compensation.

Sanford J. Grossman · Joseph E. Stiglitz · "On the Impossibility of Informationally Efficient Markets"
The American Economic Review 70(3), 393–408, 1980

If prices fully reflected available information, nobody would be paid for gathering it, and the gathering would stop. An edge has value only while it is not fully revealed — which makes the quant's refusal to disclose a structural feature of functioning markets, not stubbornness.

Stiglitz received the Nobel in 2001. "Prove your edge by showing me your edge" asks a manager to destroy the asset in order to sell it.

Its modern institutional form: every road to being believed demands the strategy — the allocator's diligence team, the verification firm, the platform holding code in custody. The NDA offered in exchange cannot protect a number whose value dies on contact: damages are unprovable, knowledge is irreversible, and no contract un-teaches the reader. So the quants with real edges rationally refuse — and disclosure-based verification adversely selects for the strategies not worth protecting. The examiner sees everything except what matters.

The deadlock

Act II demands inspection.
Act III forbids it.

Every mechanism the industry built picks a side. Audits verify the accounting and never touch the research. Managed accounts show what a manager does after you allocate, never how the record that won the mandate was produced. Platforms that see the strategy own it, or compete with it. Disclosure regimes ask politely and receive whatever the manager chooses to write.

The two requirements are jointly unsatisfiable by any disclosure-based mechanism. Not for want of will — the field has been unanimous since 2014 — but because satisfying both at once means proving a statement about data without exhibiting the data.

That is a cryptographic capability, and it did not exist in usable form until recently. The problem survived four decades because the tool to close it had not been built.

46
years of published warnings, from Grossman–Stiglitz to today
888
real algorithms measured — backtest Sharpe explained under 2.5% of out-of-sample variance
0
kink in reported returns during the three months before an audit. It returns afterwards
1
number every correction depends on — typed in by the person being judged
§ 04
What changes

The count stops being a claim

A manager commits the entire candidate set before evaluation runs. The trial count becomes the leaf count of that commitment — a property of a Merkle tree, not a figure anyone supplies. The proof forces the presented winner to be the maximum of the committed set, and recomputes the deflation on price data the verifier re-derives independently. Verification takes milliseconds, on the verifier's own machine, trusting no issuer.

Both requirements are met at once. The allocator inspects a proof instead of trusting a claim, so Act II is answered. The strategy never leaves the quant's machine, so Act III is preserved. Neither side concedes anything, because the proof carries the truth and leaves the secret behind.

The literature specified the correction. It could not specify a way to trust the input, or a way to look without taking. VTR-1 (published on SSRN) makes the input checkable without making the strategy visible.

§ 05
What this does not fix

Committing a search proves how many strategies were tried and that the winner was genuinely the best of them. It does not prove the searcher was honest about the future. Only a public time-anchor placed before the data existed can do that — which is why credentials are timestamped, not merely signed.

Nor can it make a bad strategy good. The first verdict this engine ever published was a refusal of its author's own flagship: a Deflated Sharpe Ratio of 0.68 against a 0.95 bar. That result is public, with its proof, and it is the point.

Sources — White, Econometrica 68(5), 2000 · Sullivan, Timmermann & White, Journal of Finance 54(5), 1999 · Hansen, JBES 23(4), 2005 · Getmansky, Lo & Makarov, JFE 74(3), 2004 · Agarwal, Daniel & Naik, RFS 24(10), 2011 · Bailey & López de Prado, JPM 40(5), 2014 · Jensen, Kelly & Pedersen, Journal of Finance 78(5), 2023 · Grossman & Stiglitz, American Economic Review 70(3), 393–408, 1980 · Lo, Financial Analysts Journal 58(4), 2002 · Bollen & Pool, The Journal of Finance 64(5), 2257–2288, 2009 · Brown, Goetzmann, Liang & Schwarz, Journal of Financial Economics 103(2), 221–234, 2012 · Bailey, Borwein, López de Prado & Zhu, Notices of the AMS, May 2014 (SSRN 2308659) · Novy-Marx, NBER w21329, 2015 · Harvey, Liu & Zhu, The Review of Financial Studies 29(1), 5–68 · Wiecki, Campbell, Lent & Stauth, The Journal of Investing 25(3), 69, 2016 · database-bias estimates from Fung & Hsieh (2000) and Barry (2003) · arXiv:2505.07078, 2025.

Quotations in the 1980–2016 entries are verbatim from the cited works. The 2025–2026 entry is marked as a characterisation of a body of recent literature rather than a single quotation. Nothing on this page is a MIZAN finding.
Read the standard these findings led to →
ENGINE v11 · 3ac3b10b… ● LIVE REGISTRY 77 CREDENTIALS · APPEND-ONLY VERIFY ~81 MS · OFFLINE · TRUSTING NO ONE ANCHOR BITCOIN #962,013 SPEC VTR-1 · FROZEN · CC BY ROOT 88298a2e…c6825a · MERKLE-COMMITTED