MIZAN · The trust layer for machine finance · The problem we solved

The most important number in finance was self-reported for forty-six years.

Download as PDF ↓

This document states the problem MIZAN exists to solve — precisely, with its literature, its price tag, and the reason every previous attempt failed. Nothing here is our claim alone: every element below was documented by the field itself, across four decades, and never once closed. Our claim, argued in full below: this was the biggest solvable problem in quantitative finance — the laws of markets (non-stationarity, capacity, crowding) cannot be repealed; this deadlock could be, and overfitting, dishonest data and pedigree-based allocation were all downstream of it.

46years of published
warnings, unclosed
$80Bper year, derived —
assumption movable
<0.025R² of backtest Sharpe vs
reality · 888 real algorithms
898AI strategies for $40 —
zero survived honesty
0mechanisms that could
check the trial count. until now
Before the theory — the cost, in the first person

Every day this stayed open, the honest one paid.

The numbers below this section are real and large. But nobody bleeds $80 billion. Three people bleed — the same three, every day the hole stays open — and the cruelest part is which one the world was built to reward.

The maker

The honest quant, made invisible.

Five years. A real edge, a Sharpe they actually earned. They walk into the allocator meeting and get the one question they cannot answer: prove it. Proving it means handing it over — and a disclosed edge stops being an edge. So they say “trust me,” the allocator can't, and the money goes to a warmer introduction with a worse strategy.

The edge was real. It died unfunded.
The capital

The allocator, paying for a coin flip.

They wired real money against a backtest that looked clean — because there was no instrument on earth that could show them the nine hundred strategies tried behind the one they were shown. It reverts to noise. Then the redemption letter, the LP call, the number that was never real. They didn't get unlucky. They got something unprovable, and no way to catch it.

Not a bad bet. A blind one — by force.
The fraud

The liar, funded first.

Generated 898 strategies for $40, kept the best-looking one, wrote “Sharpe 2.4, since 2021,” and raised. It worked — because the honest quant and the liar hand over identical-looking PDFs, and nothing could tell them apart. The liar wasn't punished. The liar was rewarded, and the honest one subsidized the search.

Same paper. Opposite truth. One check.
The deepest cut

The cruelty was never that fraud exists. It was that honesty and fraud became indistinguishable — so the market stopped paying for honesty at all.

Everything downstream — the $80B a year, the diligence tax, the pedigree game, the warm-intro economy — is a whole industry improvising around one question nobody could answer. Here is that question, stated exactly, and the forty-six years it stayed open.

§ 01
The deadlock

The deadlock

— Grossman & Stiglitz, 1980

To verify a trading record, the examiner must inspect the research behind it. To survive, the quant can never reveal that research — a disclosed edge stops being an edge. These two requirements are jointly unsatisfiable by any disclosure-based mechanism. That is why the problem stayed open for forty-six years: every prior verifier had to pick a side. Auditors demanded disclosure and got refused by anyone with real IP. Platforms demanded custody of the strategy and became gatekeepers nobody portable trusted. Allocators, unable to check, substituted proxies — pedigree, references, the warm introduction — none of which measure the only thing that matters: whether the record is real.

And the deadlock is not abstract — it is a toll booth every quant stands at today. To be believed, the strategy must be shown to someone. The allocator's diligence team asks for the methodology and the factor exposures. The verification firm asks for the research file. The platform asks for the code itself, held in custody on its servers. Each road ends in the same place: the edge, in someone else's hands — hands that belong, as often as not, to the party best equipped to redeploy it.

The standard defence — the NDA — fails exactly where it matters most. An edge is destroyed by being known, not by being published: damages are unprovable, knowledge is irreversible, and no contract un-teaches a reader what your research taught them. Every quant with something real knows this, which is why they refuse.

Disclosure-based verification therefore adversely selects. The strategies offered up for inspection are, disproportionately, the ones not worth protecting — while the real edges stay unverified by rational choice. The examiner sees everything except what matters. This is not a flaw in any particular verifier; it is a proof that the disclosure paradigm cannot work even in principle.
§ 02
The self-reported N

The number nobody could check

— Bailey & López de Prado, 2014; Harvey & Liu, 2015

Every statistical correction for luck — the Deflated Sharpe Ratio, White's Reality Check, Hansen's SPA, Harvey–Liu's multiple-testing haircuts — depends on a single input: N, the number of strategies tried before the one being shown. Understate N and every correction silently evaporates. And N has always been typed in by the person being judged. The Notices of the American Mathematical Society (2014) called failure to report the trial count a form of scientific fraud. It remained unreported — because no mechanism existed to make it reportable.

Because most financial analysts and academics rarely report the number of configurations tried for a given backtest, investors cannot evaluate the degree of overfitting in most investment proposals.

Bailey · Borwein · López de Prado · Zhu — Notices of the American Mathematical Society, 2014 · verbatim
The screening statistic the industry uses instead — backtest Sharpe — predicts out-of-sample performance with R² under 0.025 across 888 real trading algorithms. The primary number on which trillions are allocated carries almost no information about what happens next, and the field has known this in print for a decade.

Most claimed research findings in financial economics are likely false.

Harvey · Liu · Zhu — The Review of Financial Studies, 2016 · verbatim
§ 03
The AI acceleration · 2023–2026

The flood

Generating a compelling backtest once took months of skilled work; the cost was a natural rate-limiter on fabrication. Large language models reduced it to minutes and dollars. One 2026 study generated 898 AI strategies for $40 — zero survived an honest backtest. Published literature now clocks models at 100% accuracy in-sample and a coin flip out of sample. Manual diligence — reading, references, judgment — cannot filter output produced faster than it can be read. The trust problem became an arms race, and trust lost.

§ 04
The price

The price of the hole

CostDerivation
~$80B per year, recurringIf systematic capital ($8T) underdelivers its backtested expectation by just 1% — conservative, given R² < 0.025 — the annual shortfall is $80B, paid by allocators who had no way to check. The assumption is stated so a sceptic can move it; halving it still leaves $40B.
The diligence tax$50,000 manual audits per manager that structurally cannot check the trial count at all — the one number the corrections need.
The substitute billODD reviews at $25–75K per fund · GIPS verification at $10–50K a year · performance examinations at $50K · diligence-questionnaire platforms at up to $100K a year — billions a year already paid for verification that stops exactly where the number begins. The pain is not hypothetical; it is a standing budget line, spent today on strictly weaker instruments.
The invisible lossReal edge that goes unfunded because its owner cannot prove it without surrendering it: talent priced at zero for lack of a verification mechanism. Unmeasurable, and every allocator knows it exists. The stake for the excluded is not abstract: a single $100M allocation carries ~$2M a year in fees — careers are priced against a wall that measures pedigree, not truth.
The trust ceilingThe industry's answer — "invest in people you know" — caps the market at the size of everyone's rolodex. $23T allocated on handshake-era mechanisms.
§ 05
Prior attempts

Why every previous fix failed

AttemptWhat it verifiesWhat it structurally cannot do
Manual audit ($50K)That reported returns match statementsCannot see the trial count, the search, or hindsight; annual, not continuous
GIPS (1999)That the reporting process follows standardsVerifies process, not math; silent on overfitting, trial counts and backtests entirely
Platform track recordsLive performance inside one venue's wallsRecords die at the platform boundary; the platform is a player, not a referee; backtests unverified. The custody experiment ran at scale — hundreds of thousands of quants uploaded strategies to platform servers; when the largest venue closed in 2020, every record died with it
Disclosure under NDAThe strategy itself, shown to the examiner under legal promiseCannot un-reveal; damages unprovable; the examiner keeps the knowledge. Real edges rationally refuse — so the process adversely selects for strategies not worth protecting
Track-record databasesSelf-reported numbers, collectedGarbage in, garbage archived — the input is the thing that needed verifying
Trust + pedigreeWhere you worked, who vouchesMeasures social position, not statistical truth; excludes everyone outside the rolodex
The pattern: every mechanism either demands disclosure (and gets refused), takes custody (and becomes a conflicted gatekeeper), or accepts self-reports (and verifies nothing). The deadlock of § 01 is why — the constraint was mathematical, not commercial. No amount of diligence, regulation or reputation dissolves a jointly-unsatisfiable requirement. Only changing the mathematics does.

The number of small gains far exceeds the number of small losses… The discontinuity is absent in the 3 months culminating in an audit.

Bollen & Pool — The Journal of Finance, 2009 · verbatim · reported returns bend, and straighten only when an audit approaches
§ 06
The literature · 1980–2026

The record

— forty-six years, fifteen published works, every one openable

The testimony above is not selective. Here is the full record — a half-century of the field documenting the same open problem, in chronological order. Not one of these is a MIZAN claim. Each work is quoted and annotated in full on the evidence page.

YearThe workWhat it established
1980Grossman & Stiglitz — American Economic ReviewDisclosure destroys the edge: markets cannot pay for information once it is given away. The deadlock's first half.
2000White — EconometricaThe correction machinery is born — and takes the size of the search as its required input.
2000–03Fung & Hsieh · Barry — database-bias literatureThe databases themselves overstate: backfill and survivorship bias exceed 2% per year.
2002Lo — Financial Analysts JournalThe Sharpe ratio is an estimate — gameable, and overstated by up to 65% through serial correlation alone.
2004Getmansky, Lo & Makarov — Journal of Financial EconomicsReported returns are systematically smoother than economic reality — smoothing hides the risk.
2009Bollen & Pool — The Journal of FinanceThe kink at zero: small losses reported as small gains — and the kink vanishes when an audit approaches.
2011Agarwal, Daniel & Naik — Review of Financial StudiesReturns spike in December — exactly where the incentive to manage year-end numbers is strongest.
2012Brown, Goetzmann, Liang & Schwarz — Journal of Financial EconomicsOperational red flags predict fund failure — and investors, working only from what managers hand them, chase returns anyway.
2014Bailey, Borwein, López de Prado & Zhu — Notices of the AMSUnreported trial counts named a form of scientific fraud; investors cannot evaluate overfitting without N.
2014Bailey & López de Prado — Journal of Portfolio ManagementThe Deflated Sharpe Ratio: the correction for luck exists — but its decisive input, N, stayed self-reported.
2015Novy-Marx — NBER“Highly significant” backtests are easy to manufacture from combinations of random signals.
2016Harvey, Liu & Zhu — Review of Financial Studies“Most claimed research findings in financial economics are likely false.” Verbatim.
2016Wiecki, Campbell, Lent & Stauth — Journal of Investing888 real algorithms: backtest Sharpe predicts live performance with R² under 0.025.
2023Jensen, Kelly & Pedersen — The Journal of FinanceIncluded against interest — the strongest defence of the literature; both sides still deflate against a trial count nobody can check.
2025–26The LLM-strategy literature — arXivModels clock 100% in-sample and a coin flip out of sample; fabricating a compelling record now costs ~$40.
Forty-six years. Fifteen works. Five journals of record and the Notices of the AMS. One conclusion, never once closed — until the mathematics of verification changed. The full quotations, with verbatim and characterisation explicitly distinguished, are on the evidence page.
§ 07
The closure

What changed — and what “solved” means

Zero-knowledge proofs became practical — at this cost and speed, only in the last two years. A computation can now be performed inside a cryptographic circuit that proves the result correct without revealing the inputs. Applied to the deadlock: the record can be judged while the strategy stays sealed. The two jointly-unsatisfiable requirements stopped being joint.

MIZAN's construction closes each element of the problem in turn:

The problem elementThe mechanism that closes it
The deadlock (§ 01)The full evaluation — returns, costs, out-of-sample behaviour, the gate verdict — is computed inside a zero-knowledge proof on committed data. Anyone verifies the result offline in ~81 ms; nobody sees the strategy. Not the allocator. Not MIZAN.
The self-reported N (§ 02)The committed trial ledger: the candidate set is Merkle-committed before evaluation; N becomes the leaf count; the winner is proven maximal in-circuit; the Deflated Sharpe is recomputed against that N. The anti-luck correction can no longer be understated by the party it constrains. To our knowledge, the first mechanism of its kind (SSRN 7187038).
The AI flood (§ 03)A locked gate, identical for everyone, enforced in-circuit — fabricated backtests don't argue with diligence, they fail cryptography. And for AI strategies themselves: the sealed path proves a model's record with weights never revealed.
Hindsight (the last refuge)The live chain: strategy hashed and anchored (Bitcoin + RFC-3161) before the data it will be judged on exists, then re-proven on the growing window, every link chained to its parent. “Was this built in hindsight?” becomes a millisecond check. First live chain public — anchored in Bitcoin block 962,013.
§ 08
The proof of closure

The evidence that it is actually solved

Not a claim — a public record, re-verifiable by anyone, forever

The gate refused Apple buy-and-hold at +272,548% (Sharpe 0.89 vs the 1.20 bar) — REFUSED. It refused the S&P 500 itself, proving in-circuit that the benchmark is 99% market beta — REFUSED. It refused its maker's own +895% AI model, and shipped its maker's own flagship as NOT SIGNIFICANT at the 95% bar. 78 credentials on an append-only public wall, 54 of them refusals — because a gate that would pass everything certifies nothing. The standard behind it (VTR-1: 41 normative requirements, conformance vectors, public governance and errata) is published on SSRN under CC BY 4.0 — free for anyone to implement, including competitors. That is what makes it a standard rather than a product.

§ 09
Restated

The problem, restated as its solution

For forty-six years: “trust me — I can't show you.”
Now: “check me — I don't have to show you.”

Every era of finance installed a verification standard once the market outgrew handshakes — ratings in 1909, audit in 1923, GIPS in 1999. Systematic finance, the last $23T holdout, gets its standard now. The document half is published. What remains is the requirement — and the wedge producing it is already running.

MIZAN · the problem document · shareable · verify everything: studio.mizan.market/app/#registry · the standard: ssrn.com/abstract=7215664 · the mechanism: ssrn.com/abstract=7187038 · reproduce a credential: github.com/muaviamohammed/mizan-verify
Literature (all openable): Grossman & Stiglitz 1980 · White 2000 · Lo 2002 · Getmansky, Lo & Makarov 2004 · Hansen SPA 2005 · Bollen & Pool 2009 · Agarwal, Daniel & Naik 2011 · Bailey, Borwein, López de Prado & Zhu 2014 · DSR 2014 · Novy-Marx 2015 · Harvey, Liu & Zhu 2016 · Wiecki et al. 2016 · Jensen, Kelly & Pedersen 2023 · the full curated list with excerpts and verbatim quotations: mizan.market/evidence.
ENGINE v11 · 3ac3b10b… ● LIVE REGISTRY 78 CREDENTIALS · APPEND-ONLY VERIFY ~81 MS · OFFLINE · TRUSTING NO ONE ANCHOR BITCOIN #962,013 SPEC VTR-1 · FROZEN · CC BY ROOT 88298a2e…c6825a · MERKLE-COMMITTED