This document states the problem MIZAN exists to solve — precisely, with its literature, its price tag, and the reason every previous attempt failed. Nothing here is our claim alone: every element below was documented by the field itself, across four decades, and never once closed. Our claim, argued in full below: this was the biggest solvable problem in quantitative finance — the laws of markets (non-stationarity, capacity, crowding) cannot be repealed; this deadlock could be, and overfitting, dishonest data and pedigree-based allocation were all downstream of it.
The numbers below this section are real and large. But nobody bleeds $80 billion. Three people bleed — the same three, every day the hole stays open — and the cruelest part is which one the world was built to reward.
Five years. A real edge, a Sharpe they actually earned. They walk into the allocator meeting and get the one question they cannot answer: prove it. Proving it means handing it over — and a disclosed edge stops being an edge. So they say “trust me,” the allocator can't, and the money goes to a warmer introduction with a worse strategy.
They wired real money against a backtest that looked clean — because there was no instrument on earth that could show them the nine hundred strategies tried behind the one they were shown. It reverts to noise. Then the redemption letter, the LP call, the number that was never real. They didn't get unlucky. They got something unprovable, and no way to catch it.
Generated 898 strategies for $40, kept the best-looking one, wrote “Sharpe 2.4, since 2021,” and raised. It worked — because the honest quant and the liar hand over identical-looking PDFs, and nothing could tell them apart. The liar wasn't punished. The liar was rewarded, and the honest one subsidized the search.
The cruelty was never that fraud exists. It was that honesty and fraud became indistinguishable — so the market stopped paying for honesty at all.
Everything downstream — the $80B a year, the diligence tax, the pedigree game, the warm-intro economy — is a whole industry improvising around one question nobody could answer. Here is that question, stated exactly, and the forty-six years it stayed open.
To verify a trading record, the examiner must inspect the research behind it. To survive, the quant can never reveal that research — a disclosed edge stops being an edge. These two requirements are jointly unsatisfiable by any disclosure-based mechanism. That is why the problem stayed open for forty-six years: every prior verifier had to pick a side. Auditors demanded disclosure and got refused by anyone with real IP. Platforms demanded custody of the strategy and became gatekeepers nobody portable trusted. Allocators, unable to check, substituted proxies — pedigree, references, the warm introduction — none of which measure the only thing that matters: whether the record is real.
And the deadlock is not abstract — it is a toll booth every quant stands at today. To be believed, the strategy must be shown to someone. The allocator's diligence team asks for the methodology and the factor exposures. The verification firm asks for the research file. The platform asks for the code itself, held in custody on its servers. Each road ends in the same place: the edge, in someone else's hands — hands that belong, as often as not, to the party best equipped to redeploy it.
The standard defence — the NDA — fails exactly where it matters most. An edge is destroyed by being known, not by being published: damages are unprovable, knowledge is irreversible, and no contract un-teaches a reader what your research taught them. Every quant with something real knows this, which is why they refuse.
— Bailey & López de Prado, 2014; Harvey & Liu, 2015
Every statistical correction for luck — the Deflated Sharpe Ratio, White's Reality Check, Hansen's SPA, Harvey–Liu's multiple-testing haircuts — depends on a single input: N, the number of strategies tried before the one being shown. Understate N and every correction silently evaporates. And N has always been typed in by the person being judged. The Notices of the American Mathematical Society (2014) called failure to report the trial count a form of scientific fraud. It remained unreported — because no mechanism existed to make it reportable.
Because most financial analysts and academics rarely report the number of configurations tried for a given backtest, investors cannot evaluate the degree of overfitting in most investment proposals.
Most claimed research findings in financial economics are likely false.
Generating a compelling backtest once took months of skilled work; the cost was a natural rate-limiter on fabrication. Large language models reduced it to minutes and dollars. One 2026 study generated 898 AI strategies for $40 — zero survived an honest backtest. Published literature now clocks models at 100% accuracy in-sample and a coin flip out of sample. Manual diligence — reading, references, judgment — cannot filter output produced faster than it can be read. The trust problem became an arms race, and trust lost.
| Cost | Derivation |
|---|---|
| ~$80B per year, recurring | If systematic capital ($8T) underdelivers its backtested expectation by just 1% — conservative, given R² < 0.025 — the annual shortfall is $80B, paid by allocators who had no way to check. The assumption is stated so a sceptic can move it; halving it still leaves $40B. |
| The diligence tax | $50,000 manual audits per manager that structurally cannot check the trial count at all — the one number the corrections need. |
| The substitute bill | ODD reviews at $25–75K per fund · GIPS verification at $10–50K a year · performance examinations at $50K · diligence-questionnaire platforms at up to $100K a year — billions a year already paid for verification that stops exactly where the number begins. The pain is not hypothetical; it is a standing budget line, spent today on strictly weaker instruments. |
| The invisible loss | Real edge that goes unfunded because its owner cannot prove it without surrendering it: talent priced at zero for lack of a verification mechanism. Unmeasurable, and every allocator knows it exists. The stake for the excluded is not abstract: a single $100M allocation carries ~$2M a year in fees — careers are priced against a wall that measures pedigree, not truth. |
| The trust ceiling | The industry's answer — "invest in people you know" — caps the market at the size of everyone's rolodex. $23T allocated on handshake-era mechanisms. |
| Attempt | What it verifies | What it structurally cannot do |
|---|---|---|
| Manual audit ($50K) | That reported returns match statements | Cannot see the trial count, the search, or hindsight; annual, not continuous |
| GIPS (1999) | That the reporting process follows standards | Verifies process, not math; silent on overfitting, trial counts and backtests entirely |
| Platform track records | Live performance inside one venue's walls | Records die at the platform boundary; the platform is a player, not a referee; backtests unverified. The custody experiment ran at scale — hundreds of thousands of quants uploaded strategies to platform servers; when the largest venue closed in 2020, every record died with it |
| Disclosure under NDA | The strategy itself, shown to the examiner under legal promise | Cannot un-reveal; damages unprovable; the examiner keeps the knowledge. Real edges rationally refuse — so the process adversely selects for strategies not worth protecting |
| Track-record databases | Self-reported numbers, collected | Garbage in, garbage archived — the input is the thing that needed verifying |
| Trust + pedigree | Where you worked, who vouches | Measures social position, not statistical truth; excludes everyone outside the rolodex |
The number of small gains far exceeds the number of small losses… The discontinuity is absent in the 3 months culminating in an audit.
— forty-six years, fifteen published works, every one openable
The testimony above is not selective. Here is the full record — a half-century of the field documenting the same open problem, in chronological order. Not one of these is a MIZAN claim. Each work is quoted and annotated in full on the evidence page.
| Year | The work | What it established |
|---|---|---|
| 1980 | Grossman & Stiglitz — American Economic Review | Disclosure destroys the edge: markets cannot pay for information once it is given away. The deadlock's first half. |
| 2000 | White — Econometrica | The correction machinery is born — and takes the size of the search as its required input. |
| 2000–03 | Fung & Hsieh · Barry — database-bias literature | The databases themselves overstate: backfill and survivorship bias exceed 2% per year. |
| 2002 | Lo — Financial Analysts Journal | The Sharpe ratio is an estimate — gameable, and overstated by up to 65% through serial correlation alone. |
| 2004 | Getmansky, Lo & Makarov — Journal of Financial Economics | Reported returns are systematically smoother than economic reality — smoothing hides the risk. |
| 2009 | Bollen & Pool — The Journal of Finance | The kink at zero: small losses reported as small gains — and the kink vanishes when an audit approaches. |
| 2011 | Agarwal, Daniel & Naik — Review of Financial Studies | Returns spike in December — exactly where the incentive to manage year-end numbers is strongest. |
| 2012 | Brown, Goetzmann, Liang & Schwarz — Journal of Financial Economics | Operational red flags predict fund failure — and investors, working only from what managers hand them, chase returns anyway. |
| 2014 | Bailey, Borwein, López de Prado & Zhu — Notices of the AMS | Unreported trial counts named a form of scientific fraud; investors cannot evaluate overfitting without N. |
| 2014 | Bailey & López de Prado — Journal of Portfolio Management | The Deflated Sharpe Ratio: the correction for luck exists — but its decisive input, N, stayed self-reported. |
| 2015 | Novy-Marx — NBER | “Highly significant” backtests are easy to manufacture from combinations of random signals. |
| 2016 | Harvey, Liu & Zhu — Review of Financial Studies | “Most claimed research findings in financial economics are likely false.” Verbatim. |
| 2016 | Wiecki, Campbell, Lent & Stauth — Journal of Investing | 888 real algorithms: backtest Sharpe predicts live performance with R² under 0.025. |
| 2023 | Jensen, Kelly & Pedersen — The Journal of Finance | Included against interest — the strongest defence of the literature; both sides still deflate against a trial count nobody can check. |
| 2025–26 | The LLM-strategy literature — arXiv | Models clock 100% in-sample and a coin flip out of sample; fabricating a compelling record now costs ~$40. |
Zero-knowledge proofs became practical — at this cost and speed, only in the last two years. A computation can now be performed inside a cryptographic circuit that proves the result correct without revealing the inputs. Applied to the deadlock: the record can be judged while the strategy stays sealed. The two jointly-unsatisfiable requirements stopped being joint.
MIZAN's construction closes each element of the problem in turn:
| The problem element | The mechanism that closes it |
|---|---|
| The deadlock (§ 01) | The full evaluation — returns, costs, out-of-sample behaviour, the gate verdict — is computed inside a zero-knowledge proof on committed data. Anyone verifies the result offline in ~81 ms; nobody sees the strategy. Not the allocator. Not MIZAN. |
| The self-reported N (§ 02) | The committed trial ledger: the candidate set is Merkle-committed before evaluation; N becomes the leaf count; the winner is proven maximal in-circuit; the Deflated Sharpe is recomputed against that N. The anti-luck correction can no longer be understated by the party it constrains. To our knowledge, the first mechanism of its kind (SSRN 7187038). |
| The AI flood (§ 03) | A locked gate, identical for everyone, enforced in-circuit — fabricated backtests don't argue with diligence, they fail cryptography. And for AI strategies themselves: the sealed path proves a model's record with weights never revealed. |
| Hindsight (the last refuge) | The live chain: strategy hashed and anchored (Bitcoin + RFC-3161) before the data it will be judged on exists, then re-proven on the growing window, every link chained to its parent. “Was this built in hindsight?” becomes a millisecond check. First live chain public — anchored in Bitcoin block 962,013. |
The gate refused Apple buy-and-hold at +272,548% (Sharpe 0.89 vs the 1.20 bar) — REFUSED. It refused the S&P 500 itself, proving in-circuit that the benchmark is 99% market beta — REFUSED. It refused its maker's own +895% AI model, and shipped its maker's own flagship as NOT SIGNIFICANT at the 95% bar. 78 credentials on an append-only public wall, 54 of them refusals — because a gate that would pass everything certifies nothing. The standard behind it (VTR-1: 41 normative requirements, conformance vectors, public governance and errata) is published on SSRN under CC BY 4.0 — free for anyone to implement, including competitors. That is what makes it a standard rather than a product.
For forty-six years: “trust me — I can't show you.”
Now: “check me — I don't have to show you.”
Every era of finance installed a verification standard once the market outgrew handshakes — ratings in 1909, audit in 1923, GIPS in 1999. Systematic finance, the last $23T holdout, gets its standard now. The document half is published. What remains is the requirement — and the wedge producing it is already running.