What we have tested, and what we have not
Every number this site shows, and whether it has been measured against what actually happened. The untested ones are listed first, with what it would take to test each and what is currently in the way.
A page carrying twenty numbers of which four have been measured is not research. The only way to fix that is to know which sixteen, so this is the list. It is published for the same reason the failures are: you should be able to see what we can and cannot stand behind without taking our word for either.
Of 34 things that make a claim about the future, 18 are tested, 4 partly, and 4 have not been measured yet. 2 more are plain facts from filings and need no test. · generated 2026-09-22 · what we tested and removed
Not tested yet — the work queue
Forty cards carry a 0-100 'right now' score per sector and theme. The market-level version of this blend lost to a moving average; the per-sector version has never been tested.
What would settle it: Reconstruct sector scores point-in-time and measure against that sector's forward drawdown.
In the way: Needs a point-in-time sector-score series; only 69 days are currently recorded.
The archive holds 549,587 rows across 28,842 tickers back to January 2025. SI-01 is written and runs from the REPORTED date (FINRA publishes about eight business days after settlement), testing the level, the change, and whether the fifths order cleanly. It needs more history before its answer means anything.
What would settle it: Rank companies by short interest and by change in short interest; forward returns and failure.
In the way: The backfill is 40 of 120 settlement dates. The manual endpoint now resumes from the cursor instead of restarting at today (v2.74), so the remaining 80 dates are one command rather than a fortnight of nightly slices.
Replaced the retired long-term score on the cards. SV-01 found sector valuation does not time sectors (-0.09 once a look-ahead bug was fixed), so this is presented as description and should be labelled that way.
What would settle it: Already effectively answered. The fix is labelling, not testing.
In the way: None.
Peers are chosen by size and industry distance. Whether that selection is useful has never been checked.
What would settle it: Do the peers actually behave alike? Correlation of forward returns within a peer set against random same-sector pairs.
In the way: None, but low value.
Partly tested
Each warning was scored on twenty years of tops, including how far the market rose after each one. Whether the COUNT of lit warnings is itself predictive has not been tested the way MR-02 tested the risk read.
What would settle it: Score the danger count with the same discrimination test used on every other market signal.
In the way: None.
Tested 2011–2025 as a return measure and it held. DCF-01 then split it on 33,742 filings: being profitable at all is worth +13.7 points against the median company, being the cheap fifth of the profitable rather than the dear fifth only +1.9 — so its power was solvency, not valuation. CP-01 then tested it AS a solvency measure against the four already on the page and it came last: 2.8x, against drawdown 21x, operating cash flow over assets 10.4x, Altman Z 10x and volatility (no failures at all in its safest fifth). The cause is the denominator — the same owner earnings divided by ASSETS scores 8.9x, so dividing by the share price drags the market into a question about the business. Inside each volatility band it separated in only 2 of 5, so it adds little on top of what is shown. It stays as the equity-bond lens it was built for; it is not promoted to a solvency read.
Useless at three months (18/100), real at ten years across 15 independent decades. Published only as a ten-year return expectation, with the caveat that today's level has one comparable episode in 150 years.
Carries whatever the board and the stock pages carry, so its coverage is theirs. The long-term dial was removed from the email in 2.62.
Tested
Ranked on companies that actually went bankrupt, calibrated into percentile bands. The riskiest 1% failed at about 27x the average rate.
131,232 company-quarters. Weak scorers failed at 5.4% against 1.4% for strong. Held in 3 of 3 eras and 12 of 12 sectors. Separates bad from less bad; does not find market-beaters.
206 combinations tested; only findings where the cheap fifth BEAT the index can appear.
Built because insider buying does not predict survival.
Used as the benchmark the bankruptcy models had to beat, so its own skill is measured.
Tested as price over owner cash flow: a +1.6 point gap and a failure ratio of 0.4x, which is the wrong way. It describes what the price assumes and predicts nothing.
Each valuation method tested inside each business type. 3 of 10 types have a method worth highlighting (commodity producer EV/Sales +7.6%, pre-profit earnings yield +5.7%, insurance cash-flow yield +2.3%); the other 7 now say plainly that no method wins.
Runway separates at 3.1x on failure and +7.8 points on return, consistent in every era. Dilution at 1.6x and +11.3 points. Both hold; neither beats the index.
Tangible book backfilled from the SEC data sets (goodwill was in 971,677 filings, just never extracted). P/TBV: cheapest fifth +2.3% against the median bank, a +3.0 point gap, holding in every era. Plain price-to-book manages +2.3 points, so the 0.7 point difference is inside the noise - but valuing a bank on book WORKS in either form, which is more than most measures here manage. ROTCE is still untested.
Chosen on 2006-16, checked on 2017-21, holdout read once. Each band carries the measured frequency of a 10% fall.
Generated only from findings where the cheap fifth beat the index. If nothing passes, the page says so.
518,320 rule combinations searched on filings before 2017, then checked on 2017 onward. 11 held up; 39 did not, and both numbers are published.
A 7.7 point spread, and MV-02 confirmed the direction on a much larger universe. Still does not beat the index.
Each piece renders its own study's numbers and states its method.
The two best failure separators on the whole site: 23.2x and 6.7x, both consistent across eras. But the most volatile fifth is ALREADY far below its own high when measured, so these largely describe trouble that has started rather than trouble to come. Presented as where a company already is.
Confirmed on the ledger: +13.7 point gap and a 4.2x failure ratio, the best of the fundamental scores, slightly ahead of the F-Score.
Low volatility plus low dilution: +6.4% against the median company when ranked quarterly, +4.4% at the fixed cut-offs the live screen uses. Survives a size control (+9.9 small, +3.5 mid, +3.7 large). Failure 0.52% against 2.66%. It does NOT beat an index fund (-2.1% against the S&P) and the screen says so.
10,838 tickers computed from the full daily history rather than waiting for each stock to be re-analysed. Same definition the study measured. Bands carry the measured bankruptcy frequency and its range across eras.
Tested and failed
Rebuilt and run as of every past filing date across 10,932 filings and 2,344 companies: owner earnings grown at the company’s own trailing five-year rate, faded to 2.5%, exit multiple 15x, solved for the implied return against the price on the day. The fifth it called cheapest beat the fifth it called dearest by only 0.8 points a year, and it was right in 47% of quarters — below a coin toss. A 4×3 sweep of exit multiple and terminal rate produced +0.6 to +0.8 points everywhere, so this is not one unlucky setting.
Rebuilt point-in-time across 213,282 company-quarters: the safest fifth returned 1.5 points WORSE than the riskiest, and the riskiest fifth failed slightly less often. 18% of it was inert by construction. Removed from the page.
Head to head against the F-Score: gap -2.1 points the wrong way, failure ratio 1.0x against the F-Score's 3.5x. Removed.
Book-value-per-share growth tested and failed: the fastest fifth returned -0.5% against the median insurer, the slowest +0.1%. The combined ratio is still untestable - it is not in the point-in-time panel. Insurance now leads with cash-flow yield (+2.3%, holds in every era).
For returns: +1.3 points, not consistent across eras, and a 0.8x failure ratio (the wrong way). For survival it had already failed.
3.8 point gap, reversed in 2017-2021, strong fifth still lost 10% to the index. Removed from the verdict block.
78 against a moving average's 81. Demoted to context.
Tested across every business type now that tangible book exists. It points the WRONG way for pre-profit companies (-6.2 point gap) and stable cash generators (-5.3), consistently. Cheap on tangible assets outside a bank is mostly a value trap.
Facts, not claims
Read from filings and prices. They say nothing about the future, so there is nothing to measure — they only have to be accurate.
Read from filings and public sources. No claim about the future, so no test applies. They must stay factual, which the page audit covers.
Filings, restated into plain language. No forward claim.
What "tested" has to mean here
- Measured against what actually happened next, not against a story about why it should work.
- On figures as they stood the day the filing reached the SEC — never a restated number, never one published after the date it is used on.
- On a history that keeps the companies that were later delisted, carrying whatever happened to them. A study run only on today's survivors flatters every result in it.
- Holding in periods that were not used to choose it. This is where most things die.
- For anything presented as a reason to buy: the favoured group has to beat the index, not merely beat the unfavoured group.
Nothing on this page is a promise about when the untested items will be tested. It is a list of what we know we do not know, which is the part most sites leave out.
