← MIZAN Research
№ 05 · ContributionsJuly 2026

What's new here — and what isn't.

MIZAN built the first engine we're aware of to prove, in production, a trading strategy's honesty in zero-knowledge, the complete anti-overfitting framework, live and re-verifiable. These are the contributions, written as claims you can check rather than adjectives you have to trust.

A verification company should be the last one to overstate. So this note states what MIZAN contributed as a short list of falsifiable claims, each one you can hold us to, each hedged to what we can actually defend. We did not invent the statistics we prove, and we did not invent zero-knowledge. Our contribution is the layer between them: turning corrections that were advisory into corrections that are enforced.

01The five contributions

Stated plainly, and in the order they matter:

01The anti-overfitting framework, enforced in zero-knowledge
✓ live · verifiable

The Deflated Sharpe Ratio, the Probability of Backtest Overfitting, and purged/embargoed combinatorial cross-validation are computed inside a STARK, on committed data, so the result is a proof anyone can re-check, not a number a manager reports. MIZAN is, to our knowledge, the first engine to prove the complete anti-overfitting framework in zero-knowledge in production: a live system minting credentials you can re-verify today, not a paper and not a prototype. Others have theorized proof-of-computation for finance; we shipped the one that enforces this framework.

Not ours: the statistics themselves are Bailey & López de Prado's published work; zero-knowledge proving is RISC Zero's. Ours: the enforcement — binding the two so the correction can't be softened by the person it judges.

02The committed trial ledger — closing “who counts the trials?”
✓ live · verifiable

Every anti-overfitting correction depends on N, the number of strategies searched, and in every ordinary workflow the researcher reports N themselves. MIZAN commits the full trial ledger in a Merkle tree; N is the leaf count, not a number anyone types, and the winner is bound in-circuit to be the maximum of the committed trials, so a decoy ledger cannot collapse the correction. This is the contribution we consider most original: it removes the one input the person being judged used to control.

The limit, stated here: a credential certifies the search you commit — pre-ledger experiments are unrecoverable, so a credential-grade N requires a pre-registered ledger, not the same request that names the winner.

03Both schools of backtest honesty, proven in-circuit
✓ live · v11

As of the v11 engine, MIZAN proves not only López de Prado's suite but the second major school: Harvey & Liu on multiple testing, Hansen's Superior Predictive Ability, the Probabilistic Sharpe Ratio, and minimum backtest length. A credential can now be required to clear both schools, which is materially harder to fake than clearing either. We believe proving both in zero-knowledge is, as of today, a first.

Boundary: minimum backtest length is currently consistency-checked verifier-side rather than fully re-derived in-circuit; the rest is proven in the circuit. We mark the line rather than blur it.

04Era-pinned credentials — the judge is bound to the verdict
✓ live · permanent

Every credential names the exact engine that judged it, by cryptographic fingerprint. Change one byte of the engine and the fingerprint changes and every verifier rejects it. There is no silent upgrade: eras are explicit and public, and a credential minted under one era verifies against that era's engine forever. v10 credentials remain checkable unchanged now that v11 is live.

Not ours: image-id binding is a RISC Zero primitive. Ours: operating it as a public era registry for financial credentials, so a verdict can never be quietly re-scored.

05The self-failing credential, as a discipline
✓ live · published

MIZAN mints and publishes credentials that fail: a real STARK proving a strategy did not clear the gate, anchored beside the ones that passed. Our own flagship reads “not significant” on the Deflated Sharpe, and the v11 engine computes the Probabilistic Sharpe in-circuit, where the flagship lands below the 0.95 bar as well. The contribution here is not technical; it is a standard of conduct: publishing losers as readily as winners — made concrete and permanent.

We did not invent the mathematics or the cryptography. We built the layer that makes an honest number unfakeable: and the discipline that publishes the dishonest ones too.

02What we are careful not to claim

⚠ The primacy we don't assert

“First in the world … in production” is a claim about a shipped, live, re-verifiable engine, not a claim that no one ever proposed these ideas before. Academics and cryptographers have sketched proof-of-computation for finance for years; where our work overlaps theirs, the credit is theirs. What we claim is precise and, we believe, unarguable: the first engine to prove this framework in zero-knowledge in production, with credentials anyone can re-verify today. That precision is deliberate, a claim you can check beats a claim you have to swallow.

And the deepest contribution is still ahead of us, not behind: a backtest proof certifies faithful history, not the future. The credential that will matter most, a pre-registered, anchored forward track: is on the roadmap, not yet shipped, and we say so plainly.

03Why we publish this at all

Two reasons, both honest. First, a contribution stated as a checkable claim is stronger than one dressed in adjectives — you can verify every item above against a live credential, which is the only kind of “innovative” a verification company should be willing to call itself. Second, being the clearly documented origin of this work is a better protection than secrecy: when the method is cited to us, a copy competes with the standard we defined. We would rather be the reference than the rumor.

References. Bailey & López de Prado (2014), The Deflated Sharpe Ratio. · Bailey, Borwein, López de Prado & Zhu (2014), The Probability of Backtest Overfitting. · Harvey & Liu (2015), Backtesting. · Hansen (2005), A Test for Superior Predictive Ability. · Every capability above corresponds to a real, re-verifiable credential in the MIZAN registry; the current engine is v11 (image id 3ac3b10b…).