The adjective, argued

The biggest solvable problem in quantitative finance.

We say this sentence in every letter we send. A sentence like that is either an argument or a boast. This page is the argument. It defines both words, states the tests, shows the evidence, and lists what we do not claim, so the adjective is yours to check rather than ours to assert.

§ 00Forty six years, in one line
1980Grossman & StiglitzA disclosed edge stops being an edge. Verify me, but I cannot show you.
2009Bollen & Pool · J. FinanceThe kink at zero: returns massaged after the fact, visible in the distribution.
2014Bailey & López de PradoThe deflated Sharpe ratio. Every correction divides by N. Not reporting N called a form of fraud, Notices of the AMS.
2017Harvey · AFA addressMost claimed research findings in financial economics are likely false.
2024STARKs go productionThe primitive that recomputes statistics over a committed set without revealing it, at laptop-verifiable cost.
2026VTR-1 frozen · 77 credentialsN is a Merkle leaf count. The gate refused its own maker first. The proof is public.
"Most claimed research findings in financial economics are likely false."Campbell Harvey · AFA Presidential Address, 2017
Backtest Sharpe predicts out-of-sample performance with R² under 0.025, measured across 888 real trading algorithms.Wiecki et al. · All That Glitters Is Not Gold
Failing to report the number of trials was called a form of fraud.Bailey, Borwein, López de Prado, Zhu · Notices of the AMS, 2014
§ 01Two words, two tests
Solvable

Structural, not natural.

A natural problem is a law of markets: non-stationarity, capacity, crowding, alpha decay, the efficient-markets limit. No mechanism repeals them. A structural problem is created by an institutional arrangement and can be removed by a different one. Solvable means structural. The test: does the problem survive a change of mechanism? If yes, it is natural and outside this claim.

Biggest

Upstream of the most, for the longest, at the highest price.

Among structural problems, size is measured on three axes a reader can check: dependency, how much of the field's machinery takes this problem's output as input; duration, how long the field's own literature has documented it as open; price, the capital screened on the unverified number. The biggest is the one that dominates on all three.

§ 02The candidates · every structural problem we could name

We tried to find a structural problem that sits upstream of the trial count. We could not. Every other one either already had a partial fix, or inherits from this one.

Structural problemPrior fix existed?Upstream of the trial count?What it inherits
The unverifiable trial count. Every overfitting correction divides by N, how many strategies were tried, and N was declared by the party being judged.NO · 46 yearsIT IS THE ROOTDeflated Sharpe, the Harvey–Liu haircuts, PBO, SPA, minimum track length: all take N or the trial set as input.
Survivorship bias in dataPartial · point-in-time datasetsNoA clean dataset with a fabricated N still fails the correction.
Look-ahead bias in backtestsPartial · tooling, purged CVNoDetectable only if the trials are committed; otherwise it is one more undeclared trial.
Cost and slippage misstatementPartial · cost modelsNoA cost model is one parameter of the search; it enters N.
Reported-return manipulationYes · audit, since 1933NoAudit verifies the statement, not the search that produced the strategy.
Process opacity in reportingYes · GIPS, 1999NoGIPS verifies process, and is silent on backtests and trial counts by design.
Disclosure destroys the edgeNoSame problem, other faceThis is the deadlock that kept the trial count unverifiable: verify me, but I cannot show you.
Allocation by pedigreeNoNo · a symptomThe substitute the industry adopted because the number could not be checked.
§ 03The three axes · evidence
Dependency

Every multiple-testing correction in the literature takes the trial count or the trial set as input. Bailey and López de Prado's deflated Sharpe ratio, Harvey and Liu's haircuts, the probability of backtest overfitting, Hansen's SPA test, the minimum track-record length. None of them can be computed honestly without N. No other structural problem is an input to all of them. The mechanism, in full →

// the deflated Sharpe ratio · Bailey & López de Prado, 2014 · the denominator in red
DSR = Z[ (SR̂ − SR₀*) · √(T − 1) / √(1 − γ₃·SR̂ + (γ₄ − 1)/4 · SR̂²) ]
SR₀* = √V[SR̂] · [ (1−γ)·Z⁻¹(1 − 1/N) + γ·Z⁻¹(1 − 1/(N·e)) ]
everywhere else   N ← declared by the manager   ·   under VTR-1   N = |leaves(ledger)| ✓   read by the circuit, never by a field
Duration

Grossman and Stiglitz stated the disclosure deadlock in 1980. Bailey and López de Prado formalised the correction in 2014 and, in the Notices of the American Mathematical Society, called failure to report the trial count a form of fraud. Harvey's 2017 AFA presidential address said most claimed findings are likely false. Bollen and Pool found the fingerprint of massaged returns in the Journal of Finance. Forty six years, five journals of record, never closed. Every citation, sourced →

Price

Roughly $8 trillion of systematic capital is screened on backtested numbers whose trial count cannot be checked. If those numbers overstate by one percent, that is $80 billion a year. The assumption is stated so a sceptic can turn it down; at a quarter of a percent it is still $20 billion. The industry already pays for weaker substitutes: operational due diligence at $25K to $75K a fund, GIPS verification at $10K to $50K a year, none of which can check N. The price, in full →

§ 04Why it stayed open, and why it is closed now

It stayed open because the two halves were jointly unsatisfiable by any disclosure-based mechanism: to verify the search you had to see it, and a seen strategy stops being an edge. Every prior fix picked a side. Manual audit, GIPS, platform custody, NDA disclosure, self-reported databases, pedigree: what each verifies and what each structurally cannot →

It is closed because the primitive that dissolves the pair, a proof that recomputes the statistics over a committed trial set without revealing it, became production-grade about two years ago. The trial ledger is committed as a Merkle tree before evaluation. The circuit reads its length. N stops being a declaration and becomes a count. VTR-1, the frozen specification →

The honest boundary

Closed means the impossibility is gone, not that the industry has switched. The leaf count binds only under register-before-run: the circuit proves the shown result is the maximum of the committed set. It cannot prove that no trials were run off-ledger. That boundary is stated in VTR-1 §3.5, on the credential, and here. A claim that hid it would be the thing this standard exists to abolish.

§ 05Check the adjective yourself
Test 1 · find something upstream

Name a structural problem in quantitative finance whose output does not enter the trial count, and which the trial count does not enter. If it exists, the word "biggest" is wrong and we will change it on this page, dated. We could not find one. Send it to [email protected].

Test 2 · find a prior closure

Name a mechanism, between 1980 and 2024, that verified the trial count without disclosure and without custody. Audit, GIPS, platforms, NDAs and databases each fail one of the two. If you find one, "solvable but unsolved for 46 years" is wrong and we will say so.

Test 3 · run it

Download the verifier and a credential. Recompute the data root from your own CSV. Watch the trial count come from the ledger, never from a field. About 81 milliseconds. Verify it yourself →

§ 06What we do not claim
Not the hardest problem. Finding alpha is harder, and it is not solvable: it is the natural problem the whole field exists to attempt.
Not the largest by capital. Returns are larger than verification. This is the largest problem that verification can close.
Not solved for everyone. The mechanism exists and runs. Adoption is a mandate the industry has not yet written. That is what the company is for.

The sentence, restated with its argument attached: the unverifiable trial count was the one deadlock in the field's own literature that was structural rather than natural, upstream of every overfitting correction, documented for forty six years, and priced in the tens of billions a year. It is closed. The proof is public.

The problem, in fullVerify it yourself