What we stopped showing, and why

Every measure below was once on this site, or was proposed for it, and is not there now. Each one says what it was, what we used it for, how it was tested, how it failed, and what happened to it.

A site that publishes only what worked is not publishing research. Most of these were our own ideas, several were on the page for years, and one of them turned out to be a bug in our own code rather than a signal. They are all here.

18 entries · register generated 2026-09-22 · the full list of what was tested, including the 156 that failed outright · what we have not tested yet

Shiller CAPE as a risk score

demoted · 2026-09-20
What it is
Price divided by ten years of inflation-adjusted earnings. We showed it as a 0-100 'long-term risk' dial on the market board, beside the short-term read.
What we used it for
Implying something about the months ahead.
How it was tested
MR-01 and MR-02. Scored on whether it ranked the days that preceded a 10% or 20% fall above the days that did not, at three and twelve months, over 2006-2026.
How it failed
18 out of 100, where 50 is a coin toss. Far worse than chance at the job we were using it for.
Why we retired it
It was the wrong tool for the question. A ten-year valuation measure says what you are paying, not what happens next quarter.
What it is still used for
A ten-year return expectation, in a dropdown. MR-06 tested CAPE on its OWN claim using Shiller's history back to 1871 and it holds: the direction is negative across 15 independent decades and strongest since 1990. It is real at ten years and useless at three months, and those are different sentences.
Study: MR-01 / MR-02 / MR-06

Long-term risk score, per sector and per theme

demoted · 2026-09-22
What it is
The same CAPE-driven 0-100 score, printed on each of forty sector and theme cards, and sortable.
What we used it for
Ranking which sectors carried the most structural risk.
How it was tested
It never was, at sector level. CAPE was only ever measured for the market as a whole.
How it failed
Two ways. The measure underneath scored 18 out of 100 at predicting falls, and applying a whole-market number per sector was never tested at all. MR-06 then showed that above a CAPE of about 14 it stops separating expensive from very expensive, which is the only range anything is in today.
Why we retired it
The board's headline had already disowned the number while forty cards below the fold still displayed it. A page arguing with itself is worse than a page with one fewer number.
What it is still used for
Valuation stretch, which is a measure those scopes actually have.
Study: MR-06

Our own composite as the headline risk read

demoted · 2026-09-20
What it is
A blend of valuation, smart money and macro conditions, shown as the near-term market risk score.
What we used it for
The main number on the market-risk board, in the weekly email, and in social posts.
How it was tested
MR-01, against the dullest alternative available: how far the S&P sits below a long moving average.
How it failed
It lost. 78 out of 100 against the moving average's 81 at ranking days that preceded falls.
Why we retired it
We were publishing the weaker number while holding the stronger one.
What it is still used for
Context beside the headline, labelled as our wider composite. Its valuation and insider parts describe things a price trend cannot see, and MR-02 found its macro component genuinely improves the moving average.
Study: MR-01 / MR-02

Plain trailing P/E as a market risk measure

demoted · 2026-09-22
What it is
The S&P 500's price divided by the last twelve months of earnings.
What we used it for
Proposed as a simpler replacement for CAPE on the risk board.
How it was tested
MR-07, the same discrimination test, on 2006-2026 with the ratio lagged one month so nothing is used before it could be known.
How it failed
82 out of 100 on the period that chose it, then 35 and 41 on the two that did not. Both worse than a coin toss.
Why we retired it
The 82 is an artefact with a mechanical cause. In a crash, earnings fall faster than price, so the P/E spikes DURING the collapse rather than before it. On a window containing 2008 that looks like prediction and is coincidence. This is exactly why Shiller averages ten years of earnings.
What it is still used for
Nothing on the risk board. Over ten years it ties CAPE across independent decades but loses badly since 1990.
Study: MR-07

Dividend yield as a market risk measure

removed · 2026-09-22
What it is
The S&P 500's dividend yield, with a low yield read as expensive.
What we used it for
Proposed alongside the P/E as a simple alternative.
How it was tested
MR-07, same test.
How it failed
23 out of 100 on the discovery period. Worse than a coin toss where it mattered most.
Why we retired it
Same reason as the P/E: a valuation level is not a forecast.
Study: MR-07

Price momentum as a stock-level signal

demoted · 2026-09-22
What it is
How far a share sits above or below its 200-day average, shown as a chip at the top of every stock page.
What we used it for
One of the reads in the page's opening verdict.
How it was tested
ST-01, on 193,000 point-in-time company-quarters: the strongest fifth on twelve-month momentum against the weakest, over the following year.
How it failed
A gap of only 3.8 points, it REVERSED in 2017-2021, and the strong fifth still lost 10% to the index.
Why we retired it
A signal that reverses between periods is not a signal. It never counted toward the verdict and now it does not appear in the verdict block at all.
What it is still used for
A plain description of the share price further down the page, where no claim is attached to it.
Study: ST-01

Sector-specific valuation ratios as a buy signal

demoted · 2026-09-22
What it is
Ranking companies on P/E, price-to-book, EV/EBIT, EV/sales, cash-flow yield or earnings yield against others in their own sector.
What we used it for
Proposed as panels on stock pages.
How it was tested
SV-01: 70 sector-and-measure combinations, cheapest fifth against dearest, within sector and within quarter, across three eras.
How it failed
48 of 70 had the cheap group ahead of the dear one. ZERO had the cheap group ahead of the index.
Why we retired it
Never shipped. 'Cheap within its sector and still lost to the market' is not a reason to buy anything, and a panel implying otherwise would be the thing this whole project exists to prevent.
What it is still used for
Nothing on stock pages.
Study: SV-01

Sector valuation as a sector-timing signal

removed · 2026-09-22
What it is
A sector's own median valuation, ranked against that sector's history, used to say which sectors were heading the wrong way.
What we used it for
Proposed as a sector risk read.
How it was tested
SV-01, correlation between a sector's valuation percentile and its companies' forward returns at 12 and 36 months.
How it failed
Nothing, once a bug was fixed. The first run ranked each quarter against the sector's ENTIRE history, so a 2012 quarter was ranked against 2024 data. That produced a publishable-looking -0.38. With an expanding window it collapsed to -0.09 and several measures flipped positive.
Why we retired it
The signal was the leak. This entry exists mostly as a warning: any time a value is ranked against 'its own history', check the window is expanding and not the full sample.
Study: SV-01

Net current asset value (NCAV) as a screen

demoted · 2026-09-22
What it is
Benjamin Graham's net-net test: current assets minus all liabilities, against the share price.
What we used it for
A screen linked from the homepage.
How it was tested
MV-01, top fifth against bottom fifth on point-in-time observations with forward returns.
How it failed
The gap was NEGATIVE, minus 3.4 points, and the direction did not hold across eras. The favoured group did worse.
Why we retired it
It is the only measure we have tested whose favoured group did WORSE than the unfavoured one, and it failed the era test as well. It was linked from the homepage as 'NCAV bargains' until 2026-09-22.
What it is still used for
Nothing. The homepage link now points here.
Study: MV-01

Our own valuation model as a market-beating signal

demoted · 2026-09-22
What it is
Owner earnings - operating cash flow minus capital spending - against the market value. The coupon the DCF is built on.
What we used it for
The 'price against our model' read at the top of every stock page.
How it was tested
MV-02 rebuilt it from the SEC data sets as of every past filing date: 33,742 filings across 5,264 companies, 2009-2025. MV-01 had asked the same question of 628.
How it failed
It separates strongly and consistently - 18.9 points at twelve months, 43.9 at three years, the direction holding in all three eras - but the highest-yielding fifth still lost to the index at EVERY horizon.
Why we retired it
Not retired; demoted and labelled. The model sorts companies. It does not find ones that beat the market, and a valuation model is usually sold as doing exactly that.
What it is still used for
Shown, and never counted toward the verdict. The stock-page label stays 'not tested' for the DCF verdict itself, which adds growth and terminal assumptions this test does not reach.
Study: MV-01 / MV-02

Insider buying as a survival signal

demoted · 2026-09-13
What it is
Corporate insiders buying their own stock on the open market.
What we used it for
Proposed as evidence a company would survive.
How it was tested
The bankruptcy study, on companies that actually went bankrupt.
How it failed
93.4% of the companies that failed had a qualifying insider buy beforehand.
Why we retired it
Insiders are not a distress signal. The finding is why the four failure-warning checks were built instead.
What it is still used for
An opportunistic-buy signal, which is a different claim and is tested separately.
Study: Distress signals

48 valuation findings that were real but lost to the index

demoted · 2026-09-20
What it is
Archetype-and-metric combinations where the cheap fifth genuinely and repeatably beat the dear fifth.
What we used it for
They would have qualified as evidence under a weaker bar.
How it was tested
AV-01, 213,000 point-in-time company-quarters, 206 combinations.
How it failed
In 48 of the 50 that survived the statistics, BOTH groups lost to the index. The cheap ones merely lost less.
Why we retired it
best_for() requires cheap_excess above zero, so none of them can ever back a screen or lead a stock page.
What it is still used for
Published in full on the evidence page under 'real, but not a reason to buy', next to the 156 that failed outright.
Study: AV-01

Routing valuation measures by business type

demoted · 2026-09-22
What it is
Our archetype router shows a different primary measure for each kind of company - EV/EBIT for a stable cash generator, EV/sales for a pre-profit company, price-to-tangible-book for a bank - instead of one measure for everyone.
What we used it for
The architecture behind every stock page: which numbers are shown prominently and which are suppressed.
How it was tested
MV-03. For each business type, its own policy measure, ranked inside that type and inside the quarter, against the same companies ranked on one measure for everyone. 213,282 company-quarters, 16 types attempted, 10 with enough data.
How it failed
Routing beat one-measure-for-everyone in 0 of 10 types. Only one type (pipeline/MLP) had its cheap fifth ahead of the index, and there the policy measure IS the one-for-all measure, so routing added nothing. In three types the policy measure pointed the WRONG way: cheap did worse than dear for pre-profit companies on EV/sales (-2.7 pts), pre-profit medtech on EV/sales (-15.7 pts), and insurers on price-to-book (-4.3 pts).
Why we retired it
Not retired. Demoted from claim to convention. Showing a bank price-to-tangible-book instead of a P/E is still the right way to present a bank, and we will keep doing it - but it is a presentation choice, not a tested edge, and the pages should not imply otherwise. The three reversed measures are the actionable part.
What it is still used for
How pages are laid out, labelled as a convention rather than evidence. Six of the ten rows used stand-in measures because the panel lacks the exact one our policy names, so this is a partial test and says so.
Study: MV-03

Per-stock risk score (0-100)

demoted · 2026-09-22
What it is
A blend of valuation 30, smart money 22, macro 18 and fundamental health 30, shown as a chip at the top of every stock page.
What we used it for
The headline risk read on an individual company.
How it was tested
CL-01 rebuilt it point-in-time across 213,282 company-quarters and put it through the same bar as everything else.
How it failed
It is faintly BACKWARDS. The fifth it called safest returned 1.5 points worse over the next year than the fifth it called riskiest, and the 'riskiest' fifth went bankrupt slightly LESS often (0.9x). Separately, 18% of it was inert by construction: the macro leg is market-wide, the same number for every company on a given day, so it cannot change the order of one stock against another.
Why we retired it
It was the largest untested number on the page, sitting beside four reads that do have evidence. A score that points the wrong way is worse than no score.
What it is still used for
The component breakdown is still reachable in a dropdown, labelled as a measure that failed, so anyone who saw the old number gets an explanation rather than a silent deletion.
Study: CL-01

Fundamental health score

demoted · 2026-09-22
What it is
A 0-100 score blending leverage, the three-year free-cash-flow trend and DCF applicability.
What we used it for
The Financial Health section, and 30% of the per-stock risk score.
How it was tested
CL-01, head to head against the Piotroski F-Score, which claims the same thing on the same page.
How it failed
Gap of -2.1 points, the wrong way, and a failure ratio of 1.0x - it separated bankrupt companies from survivors not at all. The F-Score managed +12.2 points and 3.5x on the same data, and Altman Z managed +13.7 and 4.2x.
Why we retired it
Two scores claiming the same thing on one page is one too many when only one has evidence. This was the one without.
What it is still used for
Nothing. The F-Score and Altman Z cover the ground it was meant to cover, and both are tested.
Study: CL-01

Insider buying as a reason to expect returns

demoted · 2026-09-22
What it is
Open-market insider purchases, shown per stock.
What we used it for
Implying a company with insider buying is a better bet.
How it was tested
CL-01, companies with any insider buying against companies with none, over the following year.
How it failed
A gap of 1.3 points that did NOT hold across eras, and companies with insider buying went bankrupt slightly MORE often (0.8x ratio, the wrong way). Its survival claim had already failed separately.
Why we retired it
Not retired. The opportunistic-buy signal - irregular, open-market, own-cash purchases with routine trades stripped out - is a narrower claim tested separately. Plain 'insiders bought' is not evidence of anything and is presented as an observation.
What it is still used for
An observation on the smart-money panel, and the separately tested opportunistic-buy signal.
Study: CL-01

Book-value-per-share growth as an insurance lens

demoted · 2026-09-22
What it is
Year-on-year growth in an insurer's book value per share, promoted to lead the insurance lens in v2.67.
What we used it for
The first thing a reader saw on an insurer's page, after price-to-book was found to point the wrong way.
How it was tested
CL-03, inside the insurance archetype, fastest fifth against slowest, against the median insurer that quarter.
How it failed
The FASTEST-growing fifth returned -0.5% against the median insurer and the slowest returned +0.1%. Slightly backwards, and certainly nothing.
Why we retired it
Demoted. It was promoted on reasoning alone in v2.67 to replace a measure that had just failed a test - which is the same mistake in a nicer suit. Of the measures we can test inside insurance only the cash-flow yield (+2.3%) and the P/E (+1.0%) separated consistently, so those lead now.
What it is still used for
Visible on the page, below the measures that were tested.
Study: CL-03

Our DCF verdict, as a read in the answer block

demoted · 2026-09-22
What it is
The headline “trading N% above/below what our model makes it worth” line, shown at the top of every stock page among the reads that answer “is this a good stock”. The model grows owner earnings at a trailing five-year rate for five years, fades to a 2.5% terminal rate over the next five, applies a 15x exit multiple and solves for the implied return against the price.
What we used it for
One of the four reads in the answer block, carrying a “not tested” badge beside three tested ones.
How it was tested
DCF-01. The model was rebuilt and run as of every past filing date — 10,932 filings across 2,344 companies, 2014–2026 — using only data public on the day. The growth rate is the company’s own trailing five-year owner-earnings growth, capped −5% to +15% exactly as the live model caps it; using realised future growth would have produced a beautiful and entirely fake result. Companies were sorted each quarter by the implied return and measured against the median listed company over the next year.
How it failed
The fifth the model called cheapest beat the fifth it called dearest by 0.8 points a year, and it was right in 47% of the 45 quarters — below a coin toss. By period: −0.3 points in 2017–2021, +3.5 in 2022–now, no reading before that. A 4×3 sweep of exit multiple (10–25x) and terminal rate (1.5–3.5%) returned +0.6 to +0.8 points in every cell, so the failure is the model and not one unlucky setting. The same study then split the cash coupon the model is built on, across the full 33,742 filings: being profitable at all is worth +13.7 points against the median company, while being the cheap fifth of the profitable rather than the dear fifth is worth +1.9. The ingredient was a solvency measure wearing a valuation measure’s clothes, and the growth-and-terminal machinery built on top of it adds nothing.
Why we retired it
The same decision ST-01 forced on the price trend, for the same reason. A measure that has been tested and failed does not belong in the block that answers whether a company is worth owning: sitting beside three tested reads, it invites a reader to weigh it. It was honest to show it while it was genuinely untested; it is not honest to keep it there now that it has been measured.
What it is still used for
Its own section further down the page — the verdict, the football field and the reverse DCF — where it describes what the price is assuming and makes no claim about what follows. That is what a discounted cash flow is genuinely good for, and it is not gated.
Study: DCF-01

How something ends up here

  • It is measured against what actually happened next, on a history that keeps the companies later delisted, using only figures filed on the day they are used.
  • It has to hold in periods that were not used to choose it. Several entries above scored well on one period and reversed on the next; that is the most common way something fails.
  • For anything presented as a reason to buy, the favoured group has to beat the index — not merely beat the unfavoured group. Both groups losing, with one losing less, is not a buy signal.
  • If it fails, it comes off the page and the entry above is written. It is not quietly deleted.

Retiring a measure is not a claim that it is worthless everywhere, only that we could not support it here, on this data, for the job we were using it for. Where something failed one job and holds at another, the entry says so.