What we have tested, and what we have not

Every number this site shows, and whether it has been measured against what actually happened. The untested ones are listed first, with what it would take to test each and what is currently in the way.

A page carrying twenty numbers of which four have been measured is not research. The only way to fix that is to know which sixteen, so this is the list. It is published for the same reason the failures are: you should be able to see what we can and cannot stand behind without taking our word for either.

18
tested
4
partly tested
4
not tested yet
8
tested, failed
2
a fact, not a claim

Of 34 things that make a claim about the future, 18 are tested, 4 partly, and 4 have not been measured yet. 2 more are plain facts from filings and need no test. · generated 2026-09-22 · what we tested and removed

Not tested yet — the work queue

Market risk
Sector and theme risk scores
not tested yet

Forty cards carry a 0-100 'right now' score per sector and theme. The market-level version of this blend lost to a moving average; the per-sector version has never been tested.

What would settle it: Reconstruct sector scores point-in-time and measure against that sector's forward drawdown.

In the way: Needs a point-in-time sector-score series; only 69 days are currently recorded.

Stock page
Short interest on a single stock
not tested yet

The archive holds 549,587 rows across 28,842 tickers back to January 2025. SI-01 is written and runs from the REPORTED date (FINRA publishes about eight business days after settlement), testing the level, the change, and whether the fifths order cleanly. It needs more history before its answer means anything.

What would settle it: Rank companies by short interest and by change in short interest; forward returns and failure.

In the way: The backfill is 40 of 120 settlement dates. The manual endpoint now resumes from the cursor instead of restarting at today (v2.74), so the remaining 80 dates are one command rather than a fortnight of nightly slices.

Study: Market-wide short is in the risk blend
Market risk
Valuation stretch per sector
not tested yet

Replaced the retired long-term score on the cards. SV-01 found sector valuation does not time sectors (-0.09 once a look-ahead bug was fixed), so this is presented as description and should be labelled that way.

What would settle it: Already effectively answered. The fix is labelling, not testing.

In the way: None.

Study: SV-01 tested sector valuation as a timing signal and it failed
Stock page
Similar companies worth a look (peers)
not tested yet

Peers are chosen by size and industry distance. Whether that selection is useful has never been checked.

What would settle it: Do the peers actually behave alike? Correlation of forward returns within a peer set against random same-sector pairs.

In the way: None, but low value.

Partly tested

Market risk
Market-top monitor (6 warning signals)
partly tested

Each warning was scored on twenty years of tops, including how far the market rose after each one. Whether the COUNT of lit warnings is itself predictive has not been tested the way MR-02 tested the risk read.

What would settle it: Score the danger count with the same discrimination test used on every other market signal.

In the way: None.

Study: Market-top study
Stock page
What an owner keeps (owner-earnings yield)
partly tested

Tested 2011–2025 as a return measure and it held. DCF-01 then split it on 33,742 filings: being profitable at all is worth +13.7 points against the median company, being the cheap fifth of the profitable rather than the dear fifth only +1.9 — so its power was solvency, not valuation. CP-01 then tested it AS a solvency measure against the four already on the page and it came last: 2.8x, against drawdown 21x, operating cash flow over assets 10.4x, Altman Z 10x and volatility (no failures at all in its safest fifth). The cause is the denominator — the same owner earnings divided by ASSETS scores 8.9x, so dividing by the share price drags the market into a question about the business. Inside each volatility band it separated in only 2 of 5, so it adds little on top of what is shown. It stays as the equity-bond lens it was built for; it is not promoted to a solvency read.

Study: CL-01, DCF-01, CP-01
Market risk
Long-term (Shiller CAPE)
partly tested

Useless at three months (18/100), real at ten years across 15 independent decades. Published only as a ten-year return expectation, with the caveat that today's level has one comparable episode in 150 years.

Study: MR-06
Email and social
Weekly digest figures
partly tested

Carries whatever the board and the stock pages carry, so its coverage is theirs. The long-term dial was removed from the email in 2.62.

Tested

Stock page
Can it survive (failure rank)
tested

Ranked on companies that actually went bankrupt, calibrated into percentile bands. The riskiest 1% failed at about 27x the average rate.

Study: BK-01 / BK-05
Stock page
Quality of the business (F-Score)
tested

131,232 company-quarters. Weak scorers failed at 5.4% against 1.4% for strong. Held in 3 of 3 eras and 12 of 12 sectors. Separates bad from less bad; does not find market-beaters.

Study: PF-01
Stock page
Evidence strip for this business type
tested

206 combinations tested; only findings where the cheap fifth BEAT the index can appear.

Study: AV-01
Stock page
Distress warning checks
tested

Built because insider buying does not predict survival.

Study: Distress signals
Stock page
Altman Z-score
tested

Used as the benchmark the bankruptcy models had to beat, so its own skill is measured.

Study: BK-01
Stock page
Reverse DCF (what growth the price implies)
tested

Tested as price over owner cash flow: a +1.6 point gap and a failure ratio of 0.4x, which is the wrong way. It describes what the price assumes and predicts nothing.

Study: CL-01
Stock page
Football field (valuation range)
tested

Each valuation method tested inside each business type. 3 of 10 types have a method worth highlighting (commodity producer EV/Sales +7.6%, pre-profit earnings yield +5.7%, insurance cash-flow yield +2.3%); the other 7 now say plainly that no method wins.

Study: CL-03
Stock page
Cash runway and dilution (pre-profit types)
tested

Runway separates at 3.1x on failure and +7.8 points on return, consistent in every era. Dilution at 1.6x and +11.3 points. Both hold; neither beats the index.

Study: CL-01
Stock page
Bank lens: P/TBV and ROTCE
tested

Tangible book backfilled from the SEC data sets (goodwill was in 971,677 filings, just never extracted). P/TBV: cheapest fifth +2.3% against the median bank, a +3.0 point gap, holding in every era. Plain price-to-book manages +2.3 points, so the 0.7 point difference is inside the noise - but valuing a bank on book WORKS in either form, which is more than most measures here manage. ROTCE is still untested.

Study: TB-01
Market risk
Right now (150-day + macro)
tested

Chosen on 2006-16, checked on 2017-21, holdout read once. Each band carries the measured frequency of a 10% fall.

Study: MR-02 / MR-05
Screeners
Evidence-backed screens
tested

Generated only from findings where the cheap fifth beat the index. If nothing passes, the page says so.

Study: AV-01
Screeners
Screen Lab rules
tested

518,320 rule combinations searched on filings before 2017, then checked on 2017 onward. 11 held up; 39 did not, and both numbers are published.

Study: Discovery run
Screeners
Owner-yield / equity-bond table
tested

A 7.7 point spread, and MV-02 confirmed the direction on a much larger universe. Still does not beat the index.

Study: Owner-yield 2011-25
Research
Published research pieces
tested

Each piece renders its own study's numbers and states its method.

Study: various
Stock page
Volatility and drawdown from the 1-year high
tested

The two best failure separators on the whole site: 23.2x and 6.7x, both consistent across eras. But the most volatile fifth is ALREADY far below its own high when measured, so these largely describe trouble that has started rather than trouble to come. Presented as where a company already is.

Study: CL-01
Stock page
Altman Z-score
tested

Confirmed on the ledger: +13.7 point gap and a 4.2x failure ratio, the best of the fundamental scores, slightly ahead of the F-Score.

Study: BK-01 / CL-01
Screeners
Steady and not diluting you
tested

Low volatility plus low dilution: +6.4% against the median company when ranked quarterly, +4.4% at the fixed cut-offs the live screen uses. Survives a size control (+9.9 small, +3.5 mid, +3.7 large). Failure 0.52% against 2.66%. It does NOT beat an index fund (-2.1% against the S&P) and the screen says so.

Study: CL-03
Stock page
Volatility, computed for the whole universe
tested

10,838 tickers computed from the full daily history rather than waiting for each stock to be re-analysed. Same definition the study measured. Bands carry the measured bankruptcy frequency and its range across eras.

Study: CL-02 / VOL-01

Tested and failed

Stock page
Price against our model (DCF verdict)
tested, failed

Rebuilt and run as of every past filing date across 10,932 filings and 2,344 companies: owner earnings grown at the company’s own trailing five-year rate, faded to 2.5%, exit multiple 15x, solved for the implied return against the price on the day. The fifth it called cheapest beat the fifth it called dearest by only 0.8 points a year, and it was right in 47% of quarters — below a coin toss. A 4×3 sweep of exit multiple and terminal rate produced +0.6 to +0.8 points everywhere, so this is not one unlucky setting.

Study: DCF-01
Stock page
Per-stock risk score (0-100)
tested, failed

Rebuilt point-in-time across 213,282 company-quarters: the safest fifth returned 1.5 points WORSE than the riskiest, and the riskiest fifth failed slightly less often. 18% of it was inert by construction. Removed from the page.

Study: CL-01
Stock page
Fundamental health score
tested, failed

Head to head against the F-Score: gap -2.1 points the wrong way, failure ratio 1.0x against the F-Score's 3.5x. Removed.

Study: CL-01
Stock page
Insurer lens: combined ratio and book-value growth
tested, failed

Book-value-per-share growth tested and failed: the fastest fifth returned -0.5% against the median insurer, the slowest +0.1%. The combined ratio is still untestable - it is not in the point-in-time panel. Insurance now leads with cash-flow yield (+2.3%, holds in every era).

Study: CL-03
Stock page
Insider buying on a single stock
tested, failed

For returns: +1.3 points, not consistent across eras, and a 0.8x failure ratio (the wrong way). For survival it had already failed.

Study: CL-01
Stock page
Trend vs the 200-day
tested, failed

3.8 point gap, reversed in 2017-2021, strong fifth still lost 10% to the index. Removed from the verdict block.

Study: ST-01
Market risk
Our composite
tested, failed

78 against a moving average's 81. Demoted to context.

Study: MR-01
Stock page
Price-to-tangible-book outside banks
tested, failed

Tested across every business type now that tangible book exists. It points the WRONG way for pre-profit companies (-6.2 point gap) and stable cash generators (-5.3), consistently. Cheap on tangible assets outside a bank is mostly a value trap.

Study: TB-01

Facts, not claims

Read from filings and prices. They say nothing about the future, so there is nothing to measure — they only have to be accurate.

Stock page
Narrative sections (what they make, management, geography)
a fact, not a claim

Read from filings and public sources. No claim about the future, so no test applies. They must stay factual, which the page audit covers.

Stock page
Financial statements, anatomy of a share
a fact, not a claim

Filings, restated into plain language. No forward claim.

What "tested" has to mean here

  • Measured against what actually happened next, not against a story about why it should work.
  • On figures as they stood the day the filing reached the SEC — never a restated number, never one published after the date it is used on.
  • On a history that keeps the companies that were later delisted, carrying whatever happened to them. A study run only on today's survivors flatters every result in it.
  • Holding in periods that were not used to choose it. This is where most things die.
  • For anything presented as a reason to buy: the favoured group has to beat the index, not merely beat the unfavoured group.

Nothing on this page is a promise about when the untested items will be tested. It is a list of what we know we do not know, which is the part most sites leave out.