The problem · documented by the field itself

Not our claim.
Their findings.

Every line on this page is published finance research. Across forty-six years the literature describes three failures — and they lock together into a deadlock that no amount of disclosure can open. That deadlock is the reason MIZAN exists, and the reason nobody closed it sooner.

I · The count

Every correction for overfitting needs the number of strategies tried. It is self-reported.

II · The allocator

Returns arrive from the person being judged, and the literature shows they bend. Inspection is required.

III · The quant

An edge that is disclosed stops being an edge. Inspection is forbidden. This is a result in economics, not a preference.

Act I

The number nobody can check

Every serious correction for backtest overfitting takes one input: how many strategies were tried before the one being shown. Here is the field arriving at that conclusion across twenty-four years — and stopping there.

2002

The building blocks of the Sharpe ratio — expected returns and volatilities — are unknown quantities that must be estimated statistically and are, therefore, subject to estimation error.

Andrew W. Lo · "The Statistics of Sharpe Ratios"
Financial Analysts Journal 58(4), 2002

Lo showed a fund's annual Sharpe ratio can be overstated by as much as 65% through serial correlation alone, and that a monthly Sharpe cannot simply be annualised by √12. The headline number was unreliable before anyone tried to game it.

2014

Because most financial analysts and academics rarely report the number of configurations tried for a given backtest, investors cannot evaluate the degree of overfitting in most investment proposals.

David H. Bailey · Jonathan M. Borwein · Marcos López de Prado · Qiji Jim Zhu
"Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance"
Notices of the American Mathematical Society, May 2014

This is the sentence MIZAN exists to answer. Four mathematicians, in the journal of the American Mathematical Society, stating that the one input needed to judge a track record is the one input nobody supplies. They also showed that high simulated performance is achievable after trying only a small number of configurations — and that the more you try, the more certain the overfit.

2015

"Highly significant" backtested performance is easy to generate by selecting stocks on the basis of combinations of randomly generated signals, which by construction have no true power.

Robert Novy-Marx · "Backtesting Strategies Based on Multiple Signals"
NBER Working Paper 21329, July 2015

Elsewhere Novy-Marx made the same point with signals drawn from the position of Mars and Saturn, and from sunspots. Statistical significance, alone, carries no information about whether an edge is real.

2016

Most claimed research findings in financial economics are likely false.

Campbell R. Harvey · Yan Liu · Heqing Zhu · "…and the Cross-Section of Expected Returns"
The Review of Financial Studies 29(1), 5–68

Their remedy was to raise the bar: a newly discovered factor should clear a t-ratio above 3.0 rather than the conventional 2.0 — precisely because so many factors were tried before the one published. The correction depends on knowing how many. Nobody knows how many.

2016

The more backtesting a quant has done for a strategy, the larger the discrepancy between backtest and out-of-sample performance.

Thomas Wiecki · Andrew Campbell · Justin Lent · Jessica Stauth (Quantopian)
"All That Glitters Is Not Gold: Comparing Backtest and Out-of-Sample Performance on a Large Cohort of Trading Algorithms"
The Journal of Investing 25(3), 69

Not theory — 888 real algorithms, each with at least six months of genuine out-of-sample life. Their other finding is the one that should trouble every allocator: the backtest Sharpe ratio predicted out-of-sample performance with R² below 0.025. The number the industry screens on carries almost no information.

2025 – 2026

Previously reported LLM advantages deteriorate significantly under broader cross-section and over a longer-term evaluation.

Characterising the recent literature, incl. "Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?" · arXiv:2505.07078, 2025

The same reviews report models reaching 100% in-sample accuracy and roughly 50% out-of-sample — a coin flip, dressed as a result. What took a quant months in 2014 now takes an afternoon. The number of configurations tried has exploded. The mechanism for reporting it has not changed at all.

Marked as characterisation of a body of work, not a single verbatim quotation — unlike every entry above.

Act II

The allocator cannot verify

Act I assumes the reported returns are themselves honest. The literature does not.

2009

The number of small gains far exceeds the number of small losses… The discontinuity is absent in the 3 months culminating in an audit.

Nicolas P. B. Bollen · Veronika K. Pool · "Do Hedge Fund Managers Misreport Returns? Evidence from the Pooled Distribution"
The Journal of Finance 64(5), 2257–2288, October 2009

The most important finding on this page. Reported returns show a kink exactly at zero: small losses quietly become small gains. The kink vanishes in the months before an audit, and the performance reverses afterwards. Present in live funds and dead ones, funds of every age — so it is not a database artifact.

Verification changes behaviour. That was demonstrated on real data in 2009 — and the only thing being verified was the accounting, never the research.

2012

Operational risk increases the likelihood of subsequent poor performance and fund disappearance, but does not influence investors' return-chasing behavior.

Stephen J. Brown · William N. Goetzmann · Bing Liang · Christopher Schwarz · "Trust and Delegation"
Journal of Financial Economics 103(2), 221–234, 2012

Built from actual due-diligence reports. Many funds carry operational problems including limited disclosure of legal and regulatory matters — and the signal predicts failure while investors chase returns anyway. Allocators are not careless. They are working from what the manager hands them.

2000 – 2003

All commercial hedge fund databases are based on hedge funds' voluntary reporting.

Established across the database-bias literature · backfill effect estimated at 1.4% (Fung & Hsieh, 2000) and 0.4% (Barry, 2003) · survivorship bias measured above 2% per year

Managers choose whether to report, and typically begin after a strong run — at which point the entire prior history is backfilled in. Funds that die stop reporting. The industry's benchmark data is assembled by the people being benchmarked.

Act III

The quant cannot reveal

So the allocator needs to see inside. Here is why the quant can never allow it.

1980

There is an equilibrium degree of disequilibrium: prices reflect the information of informed individuals but only partially, so that those who expend resources to obtain information do receive compensation.

Sanford J. Grossman · Joseph E. Stiglitz · "On the Impossibility of Informationally Efficient Markets"
The American Economic Review 70(3), 393–408, 1980

If prices fully reflected available information, nobody would be paid for gathering it, and the gathering would stop. An edge has value only while it is not fully revealed — which makes the quant's refusal to disclose a structural feature of functioning markets, not stubbornness.

Stiglitz received the Nobel in 2001. "Prove your edge by showing me your edge" asks a manager to destroy the asset in order to sell it.

The deadlock

Act II demands inspection.
Act III forbids it.

Every mechanism the industry built picks a side. Audits verify the accounting and never touch the research. Managed accounts show what a manager does after you allocate, never how the record that won the mandate was produced. Platforms that see the strategy own it, or compete with it. Disclosure regimes ask politely and receive whatever the manager chooses to write.

The two requirements are jointly unsatisfiable by any disclosure-based mechanism. Not for want of will — the field has been unanimous since 2014 — but because satisfying both at once means proving a statement about data without exhibiting the data.

That is a cryptographic capability, and it did not exist in usable form until recently. The problem survived four decades because the tool to close it had not been built.

46
years of published warnings, from Grossman–Stiglitz to today
888
real algorithms measured — backtest Sharpe explained under 2.5% of out-of-sample variance
0
kink in reported returns during the three months before an audit. It returns afterwards
1
number every correction depends on — typed in by the person being judged
What changes

The count stops being a claim

A manager commits the entire candidate set before evaluation runs. The trial count becomes the leaf count of that commitment — a property of a Merkle tree, not a figure anyone supplies. The proof forces the presented winner to be the maximum of the committed set, and recomputes the deflation on price data the verifier re-derives independently. Verification takes milliseconds, on the verifier's own machine, trusting no issuer.

Both requirements are met at once. The allocator inspects a proof instead of trusting a claim, so Act II is answered. The strategy never leaves the quant's machine, so Act III is preserved. Neither side concedes anything, because the proof carries the truth and leaves the secret behind.

The literature specified the correction. It could not specify a way to trust the input, or a way to look without taking. VTR-1 makes the input checkable without making the strategy visible.

What this does not fix

Committing a search proves how many strategies were tried and that the winner was genuinely the best of them. It does not prove the searcher was honest about the future. Only a public time-anchor placed before the data existed can do that — which is why credentials are timestamped, not merely signed.

Nor can it make a bad strategy good. The first verdict this engine ever published was a refusal of its author's own flagship: a Deflated Sharpe Ratio of 0.68 against a 0.95 bar. That result is public, with its proof, and it is the point.

Sources — Grossman & Stiglitz, American Economic Review 70(3), 393–408, 1980 · Lo, Financial Analysts Journal 58(4), 2002 · Bollen & Pool, The Journal of Finance 64(5), 2257–2288, 2009 · Brown, Goetzmann, Liang & Schwarz, Journal of Financial Economics 103(2), 221–234, 2012 · Bailey, Borwein, López de Prado & Zhu, Notices of the AMS, May 2014 (SSRN 2308659) · Novy-Marx, NBER w21329, 2015 · Harvey, Liu & Zhu, The Review of Financial Studies 29(1), 5–68 · Wiecki, Campbell, Lent & Stauth, The Journal of Investing 25(3), 69, 2016 · database-bias estimates from Fung & Hsieh (2000) and Barry (2003) · arXiv:2505.07078, 2025.

Quotations in the 1980–2016 entries are verbatim from the cited works. The 2025–2026 entry is marked as a characterisation of a body of recent literature rather than a single quotation. Nothing on this page is a MIZAN finding.
Read the standard these findings led to →