MIZAN Research · the forwardable

Seven questions an allocator should ask any manager

Written to be forwarded — to any manager, pitching any track record, in any asset class. None of the seven depend on our system existing.

This note is written to be forwarded. Not to managers who use MIZAN — to any manager, pitching any track record, in any asset class. The seven questions below do not depend on our system existing. They are the questions I would want asked of me. For each one I will say why it matters and what an honest answer looks like, so that the person across the table cannot retreat into vocabulary.

A disclosure first, because this note asks managers to disclose: I run a verification company, so I have an interest in these questions being asked. I have tried to write them so they are sound whether or not my company is in the room; judge that claim by the questions, not by me.

§ 01Question 1

1. How many strategies did you try before this one, and can you prove the count?

Why it matters. A Sharpe of 1.5 from the first strategy tried and a Sharpe of 1.5 from the best of two hundred are different objects. Selection across many trials manufactures apparent skill from noise — the multiple-testing problem, the largest silent inflator of backtested performance. The literature on backtest overfitting exists because the industry's default answer here is a shrug.

What an honest answer looks like. A number, first of all — even approximate, even embarrassing. Then some artifact that makes the number hard to revise later: a research log, a pre-registration, timestamped configuration records, a deflated performance figure that conditions on the trial count. The honest manager may not be able to prove the count — most research processes were never built to — but they will state the number without flinching, say it is self-reported, and say which direction the uncertainty runs. The tell is a manager who insists the question doesn't apply to them.

§ 02Question 2

2. Are your reported numbers net of committed, disclosed costs — committed before evaluation?

Why it matters. Costs chosen after seeing the results are a tuning parameter. A strategy that dies at 10bps per side and lives at 3bps can always be reported at 3bps, and the choice will be defended as "our realistic execution assumption." The word doing the work in this question is committed: fixed before the evaluation ran, so the cost model could not be fitted to rescue the numbers.

What an honest answer looks like. The per-side figures, stated. When they were fixed, stated. And — just as important — what the cost model does not cover: market impact, capacity, partial fills, fill quality under stress. An honest manager volunteers that a flat per-side cost is a floor, not a simulation of execution, and will not let net-of-costs be read as net-of-reality.

§ 03Question 3

3. What does your out-of-sample period look like, and who chose where it starts?

Why it matters. An out-of-sample window whose boundary was chosen after the fact is in-sample with better branding. If the split point was placed where it flatters the result — or moved when it didn't — the OOS number is a second in-sample number. The question is not "do you have OOS?" Everyone has OOS. The question is who drew the line, and when.

What an honest answer looks like. A rule, not a date: "final 30% of the window," "everything after the strategy was frozen," "walk-forward with these fixed folds" — stated before the evaluation, applied mechanically. And an honest OOS number is often worse than the headline. A manager whose out-of-sample matches in-sample too neatly, across many strategies, is describing a miracle or a methodology problem. Degradation, disclosed plainly, is what real looks like.

§ 04Question 4

4. If I re-run your backtest on my own copy of the data, will I get your numbers to the digit?

Why it matters. This question tests two things at once. Reproducibility: is the result an artifact of one environment, one library version, one vendor's quirks? And data provenance: were the numbers computed on the market's actual history, or on a copy with survivorship gaps, back-adjustments, or errors that flatter the strategy? A result that only exists on the manager's machine, on the manager's data, is a story.

What an honest answer looks like. Ideally: yes, to the digit, and here is what you need — the data specification, the exact window, the cost parameters, enough of the methodology that your rerun is a genuine replication rather than a guided tour. Failing that, honesty about why not: proprietary signal logic they won't disclose (legitimate), vendor data they can't redistribute (legitimate), or "it's complicated" (listen carefully to what follows). The unacceptable answer is treating the question as strange. Reproduction to the digit is what computation means; anything less is a request for trust, and should be named as one.

§ 05Question 5

5. What did you refuse to show me — where are your failures?

Why it matters. A track record with no failures attached is a numerator without a denominator. Every serious research program kills most of what it tries. If you are only shown survivors, you are evaluating a filter's output while being denied its rejection rate. The failures are not the embarrassing part of the record; they are the part that makes the successes interpretable.

What an honest answer looks like. Named dead strategies. Retractions, if there were errors — with the mechanism explained, not euphemized. Windows where the flagship lost money, stated before you find them yourself. A manager who can walk you through what they abandoned, and why, is showing you the research process. A manager who cannot name a single failure is showing you marketing.

§ 06Question 6

6. Who judged your track record, and can that judge's rules change retroactively?

Why it matters. Every performance claim was measured under some regime of rules — a methodology, an administrator, an auditor, a standard. If those rules can be revised after the fact, the claim is not stable: it means whatever the current rules say it means, and the current rules can change. This is why "audited" track records from defunct administrators are so hard to evaluate a decade later. A claim's judge is part of the claim.

What an honest answer looks like. The judge named — a methodology version, a named standard, an administrator, an auditor — and a straight answer on retroactivity. The strong answer is: the rules under which this number was produced are frozen and identifiable, and if we upgrade our methodology, old numbers remain labeled with the old methodology rather than silently restated. Silent restatement of history under new rules is the tell. History should accumulate; it should not be editable.

§ 07Question 7

7. What does your claim NOT cover — capacity, fills, regime transfer?

Why it matters. The most dangerous inflation in this business is not fabricated numbers — it is true numbers read as covering more than they do. A backtest, however honest, is silent on capacity: at what AUM the edge degrades into its own market impact. Silent on fills: what execution looks like against a live order book. Silent on regime transfer: whether an edge measured over one span of history survives the next one. These silences are not flaws; they are the boundary of what historical evaluation can say. The question is whether the manager knows the boundary and says it out loud.

What an honest answer looks like. An unprompted list of exclusions, stated with the same confidence as the performance figures. "This number does not model capacity. It assumes fills at these committed costs and nothing better or worse. It is a statement about this window, this instrument, this regime — not a forecast." A manager who bounds their own claim before you ask is showing you how they think. A manager who lets the claim stay unbounded is hoping you'll do the over-reading for them.

§ 08Question 8

Where MIZAN fits, and where it doesn't need to

Every one of these questions can be answered without us — with documents, with logs, with an administrator's letter, with the manager's word, and with your trust that none of it was edited after the fact. That has been the industry's answer for decades, and for many relationships it is enough.

MIZAN exists for the cases where trust is the expensive part. Our system is built so that several of these answers can be proven rather than attested: costs committed before evaluation and charged inside the proof; data pinned by cryptographic commitment, so a rerun on your own copy either matches to the digit or visibly doesn't; a judge — the verifier — whose rules are fixed per era and never retroactively revised; failures published alongside passes. Not all seven are provable — capacity and fill quality beyond committed costs sit outside what we model, and we say so on the credential itself — and a manager who answers the rest with documents and a straight face deserves a hearing.

But ask the questions either way. They are sound with or without us. If this note is useful, forward it — the questions are the product here, and they are free.

Prove the edge. Never reveal the strategy.

If your own screen carries a question missing here, we would rather hear it than guess. Tell us what we left out · the credential these questions describe: ALLOCATOR-V1 · back to for allocators.
ENGINE v11 · 3ac3b10b… ● LIVE REGISTRY 77 CREDENTIALS · APPEND-ONLY VERIFY ~81 MS · OFFLINE · TRUSTING NO ONE ANCHOR BITCOIN #962,013 SPEC VTR-1 · FROZEN · CC BY ROOT 88298a2e…c6825a · MERKLE-COMMITTED