A manifesto in six parts

An Honest Machine

The making of MIZAN.
I
The number that graded itself

Every backtest you have ever been shown was graded by the person selling it.

Sit with that, because the whole industry rests on it and almost no one says it out loud. A track record is a claim about the past made by the one party with every incentive to shade it. The tools that check that claim — the auditors, the administrators, the databases — verify that a number was reported, never that the number is true. Trust fills the gap. And trust is precisely what a liar exploits.

I did not arrive at this as a critic. I arrived at it as the liar.

II
The first graveyard

For years I hunted alone for an edge in the markets, and I found them — beautiful ones. Strategies with Sharpe ratios that quickened the pulse. Backtests that looked like printed money. For a season I believed I was brilliant, and belief is the most expensive thing a trader can own.

Then I began to audit my own work — not to improve it, but to break it. One by one, they died. A stop-loss model that booked profits at prices the market may never have traded. A signal that was quietly reading tomorrow's answer into today's decision. A universe that looked like genius only because history had already deleted its losers before I ever ran the test. Forty and more, killed by their own author. I did not bury them. I wrote the autopsy for each, and I published it.

Somewhere in that graveyard the question turned around and faced me. I had been asking how do I find an edge? The real question, the one that became a company, was quieter and worse: why should anyone believe a number — including mine?

III
The wound in the mathematics

I was not the first to see the disease. A generation of statisticians had spent twenty years diagnosing it. Bailey and López de Prado built the Deflated Sharpe Ratio — a way to correct a result for how many strategies were tried before it worked, because the best of a thousand coin-flips looks like skill. Hansen built a test for superior predictive ability on the foundation White had laid. Harvey showed that most of the "discoveries" in the published cross-section of returns are false. Gelman named the garden of forking paths, where even an honest researcher cannot count the choices that led to a result.

Airtight mathematics, all of it. And underneath, one open wound the entire twenty years: every one of these corrections is parameterized by the size of the search — how many trials, how wide the net — and that number is supplied by the very party being judged. Understate it, and the correction for luck evaporates. A thousand honest backtests with only the best one published is a lie that contains no lie anywhere inside it. No property of a single backtest can reveal how many siblings it had.

It was never a statistics problem. It was an enforcement problem, and statistics has no way to enforce anything. The tools would have to come from a different universe.

IV
The bolt

The cryptographers had built that universe without knowing what it was for. Goldwasser, Micali, and Rackoff, in 1985, invented a way to prove a statement true while revealing nothing else. Merkle gave us a way to freeze a body of data into a single value that cannot be quietly rewritten. Ben-Sasson and his collaborators made those proofs practical at scale, with no trusted setup, and a small team turned them into a machine you can actually run.

Two literatures, on the same campuses, across the same decades, that never once met in a single system. So the corrections stayed voluntary and the proofs stayed pointed at payments, and trillions of dollars kept moving on numbers no one could check.

MIZAN is the bolt between them. The set of trials is committed to a Merkle tree before anything is evaluated, so the trial count is the number of leaves on the tree, not a figure the researcher types in. The strategy presented as the winner is forced, inside a zero-knowledge proof, to be the true maximum of that committed set. The deflation is recomputed in the circuit, on prices pinned to their own root, after a cost floor that cannot be softened. What comes out is a credential: proof that a strategy is honest — after costs, no lookahead, corrected for the luck of the search — with the strategy itself never revealed. Proof, at last, instead of trust.

V
The machine that refused its maker

A verifier is only worth what it will refuse. So the first thing I did with it was turn it on myself.

My flagship strategy, the one I was proudest of, went in first. It failed. Deflated for the breadth of the search that found it, it did not reach significance — not deployable, not real enough to stand behind. I shipped that failure to the public wall with its proof attached, next to everyone else's. Then I pointed the engine at the S&P 500 with a strategy of my own making, on survivorship-free data with the dead companies left in, and it refused me again. I ran thirty-four American stocks through it, including every bar Nvidia has ever printed, the single best-performing name of the era. Not one passed. That is not a flaw in the machine. That is what an honest gate says about stock-picking edge in daily data, and anyone who tells you otherwise is selling you their backtest, not their edge.

A machine that will not lie for its own creator is the only kind fit to be trusted with anyone else's numbers. Every service, every fund, every track record you have been shown displays only its passes; the refusals, the thing that gives a pass its meaning, die in private. Mine live on a public wall, beside the passes, judged by the same locked engine, each one re-checkable on your own laptop in seconds, trusting no one, including me.

VI
The standard, and what it will not claim

I built all of it alone, over three years, before I raised a dollar — because I already knew, in a way most people are lucky enough never to learn, exactly what it costs to trust a number nobody proved.

So let me be as honest about the limits as about the machine. A credential is a seal on history made honest, not a forecast; it says nothing about tomorrow. It does not model the fills or the capacity that live beyond a committed cost floor. High-frequency and fill-dependent strategies I refuse permanently, because bar data cannot honestly prove a fill that never happened, and a proof of a fiction is worse than no proof at all. A verification standard is defined as much by what it will not sign as by what it will.

If this becomes the standard that systematic finance never had, I want the history to read correctly. I built the bolt. The people named in these pages built everything the bolt holds together. Twenty authors, one machine, and a founder it refused first.

The gate that refused me is open to anyone who believes their edge is real.

Prove it.

Prove the edge. Never reveal the strategy.

Read the receipts →The paper →Verify it yourself →