How we ran the backtest
We went back and marked our own homework — and threw the results away three times before publishing them. One half of the system passed. The other half failed the moment we made the test strict enough. Both halves are on this page. The headline results are here; what follows is how they were produced, what we corrected, and what we decided we could not claim.
First, three words explained
A backtest is a simple question asked of the past: if we had run our system on a date years ago, using only what was knowable that day, would its picks have done well afterwards?
Fair value is what we think a business is worth based on the cash it actually produces — not what its share price happens to be today.
Margin of safety is the gap between the two. If we think a company is worth $100 a share and it trades at $70, the margin of safety is 30%. Buying with a margin of safety means paying meaningfully less than a business is worth, so that you can be a bit wrong and still be fine. It is the whole idea Numora is built on, and it is what these two studies test.
Part 1 — the math engine alone
Numora has two halves: a calculator that estimates what a business is worth from its financial filings, and an AI analyst that reads the company's own words and judges the things numbers can't see. Part 1 tests the calculator on its own.
What we did
- Twenty-eight starting dates. The last trading day of June in each year from 1998 to 2025 — every June we have prices for. Our price data begins in December 1997, so this is as far back as we can honestly go.
- The 1,500 biggest US companies on each of those days — including companies that later collapsed, got bought, or went bankrupt. If a stock stopped trading, we count whatever it was worth at the end. Disasters count against us; nothing gets quietly dropped from the record.
- Only what was public at the time. The calculator saw the financial filings that existed on that date and nothing else. No later news, no later numbers.
- No AI at all in Part 1. This is pure arithmetic on old filings. Because no language model is involved, there is no way for knowledge of what happened later to sneak in.
- 25,050 valuations came out of that process. We then measured what each stock actually did over the following 1, 3 and 5 years, counting dividends.
What we found: cheap beat expensive, and patience paid
We sorted every stock by how cheap the engine thought it was, then checked whether that ranking lined up with what happened next. A score of 0 would mean the ranking was worthless.
| If you held for | Ranking score | Strength of evidence | Verdict |
|---|---|---|---|
| 1 year | +0.051 | 2.04 | real (95% confidence) |
| 3 years | +0.061 | 2.58 | real (95% confidence) |
| 5 years | +0.064 | 2.66 | real (95% confidence) |
The scores are small in absolute terms — that is normal and expected for stock picking, where being right slightly more often than not is what compounds. What matters is that they are positive, and that the evidence clears the bar at every holding period, including one year. This is the finding that survived a bug that changed almost every other number on this page (more on that below), which is part of why we lead with it.
What we found: the buy line sorted outcomes correctly
Splitting every stock into three buckets by how the engine priced it, here is the total return that followed. Each cell shows the average first, then the typical (middle) result:
| What the engine said | After 1 year | After 3 years | After 5 years |
|---|---|---|---|
| Cheap (15%+ below our fair value) | +13.0% / +11.3% | +38.4% / +29.4% | +66.0% / +44.7% |
| About right (0–15% below) | +10.8% / +10.5% | +34.2% / +28.5% | +58.9% / +42.7% |
| Expensive (above our fair value) | +9.4% / +6.7% | +27.5% / +20.1% | +48.3% / +31.4% |
Over five years the cheap bucket beat the expensive bucket by 13.3 points for the typical stock, and 17.7 points on averages. The order never flips: cheap beat about-right, and about-right beat expensive, at every single holding period, on both measures. That is the engine's buy-versus-avoid line doing the job it claims to do.
“Typical” here means the median — the middle result, with half the stocks doing better and half worse. We show it next to the average throughout, because in stock returns the two say different things: a few enormous winners can lift an average well above what any ordinary pick did. Where they disagree, the median is usually the more useful number.
How it compares to the S&P 500
Beating the other stocks in a study is one thing. The question most people actually care about is simpler: did it beat just buying the index?
Here we use the real thing — the official S&P 500 Total Return index, which counts dividends, pulled from public market data. (An earlier version of this page compared against a stand-in we built ourselves. Replacing it with the official index is what uncovered the bug described below, so it earned its keep twice over.)
This is the part of our own results we like least, and it is the part we most want you to read. Percentage points better or worse than the index — average first, then typical:
| Group | 3-year vs index (avg / typical) | Beat it over 3 years | 5-year vs index (avg / typical) | Beat it over 5 years |
|---|---|---|---|---|
| Cheap | +3.2pp / −4.8pp | 46% | +1.3pp / −13.1pp | 44% |
| Expensive | −3.4pp / −9.9pp | 42% | −9.2pp / −21.3pp | 39% |
Read the cheap row honestly. Over five years those picks beat the index only 44% of the time, and the typical one finished −13.1pp behind it. The average was slightly ahead (+1.3pp) — but that is the gap between average and typical doing its work again.
Why most stocks lose to the index — even good ones
This sounds like a contradiction and isn't. It is one of the better-established facts in finance, and it is worth understanding before you read any stock-picking claim, ours included.
A stock index is not the average stock. It is a weighted blend in which the biggest companies count most, and its returns turn out to be carried by a small handful of enormous winners — the Apples and Amazons — while the majority of individual stocks underperform it. Researchers have found that most single stocks fail to beat even a savings account over their lifetime; the whole market's gain comes from a slim minority. So “the typical stock lags the index” is the normal state of the world, not a mark against a particular method.
That is exactly the shape in our cheap row: a positive average with a negative typical result means a minority of big winners carried the group. It is also why we keep putting both numbers in front of you.
So here is the claim we make, stated as narrowly as the evidence supports it. The engine ranks stocks against each other, and that ranking works — cheap consistently beat expensive, by double digits over five years, over 28 years of history, and the expensive group was worse than the cheap group on every single measure against the index too. What we are notclaiming is that any given pick will beat the S&P 500. Over this era, the typical one didn't.
Expect long droughts — they are in this data
Twenty-eight years is long enough to show something a shorter study would hide, and we would rather you hear it from us. This discipline does not work steadily. It works in bursts, and then it goes quiet for years at a stretch.
Its finest hour was the early stretch, 1998–2006, spanning the dot-com collapse: cheap names beat the index by +33.6pp on average, with 60% of them coming out ahead. Through the financial-crisis era (2007–2013) they roughly matched it (−5.6pp on average).
Then came the growth decade, and it was ugly. Across the 2014–2020 starting dates, cheap names trailed the index by −24.4pp on average, and only 29% of them beat it. That is the better part of a decade in which a handful of giant technology companies pulled the index away from almost everything else, and this approach simply did not pay. Anyone using it should plan on sitting through a stretch like that, because there is no version of value investing that skips them.
The numbers on this page are the corrected ones
Everything above was published in a different form a day earlier, with different figures. Here is why they changed.
To compare against the index we went and fetched the official S&P 500 — and the moment we lined our results up against it, they looked wrong. Chasing that led to a column in our database whose name promised adjusted prices and in fact held raw ones. Two things follow from that, and both are bad. Stock splits weren't accounted for, so any company that split during a holding period had its return mangled: Apple from 2009 to 2014 was recorded as losing 35% when it had actually gained 360%. And dividends were missing entirely, for every company.
So we rebuilt every return in the study from verified split-adjusted prices plus the cash dividends actually paid, spot-checked several companies against public market data, and re-ran the whole thing. The tables above are the rebuilt version. The rankings held — cheap still beat expensive at every horizon, and the spread got wider, not narrower — but nearly every absolute number on this page moved.
To be clear about the blast radius: the bug lived in the backtest only. Nothing on the live site reads that column, so no valuation or verdict we have published was affected.
This is the third time an outside check has made us delete results and start again — the other two are in Part 2 below, and each one took away a number we liked. We keep telling you about them because the alternative is asking you to trust figures whose failures you can't see. A study that has survived three deliberate attempts to break it is worth more than one that was never tested.
Where the engine alone hits its limit
The engine is arithmetic on filings, and arithmetic cannot tell the difference between a good business trading cheaply and a dying one. At the very cheap end that distinction is everything: those bins mix a few spectacular comebacks with a lot of companies that were cheap for an extremely good reason. The market calls the second kind value traps.
We tried the obvious mechanical fix — screening out low-quality businesses by a simple numeric rule — and it added nothing. Telling a bargain from a trap needs judgment about the business itself: is the moat real, is the industry being disrupted, is management trustworthy. That is exactly what the AI analyst layer is for, and exactly what Part 2 goes on to test — with a result we did not expect and did not want.
Part 2 — the full analyst pipeline
Part 1 showed the calculator works. Part 2 asked the harder question: does the AI analyst sitting on top of it add anything a backtest can see? We ran that test three times, each version stricter than the last. The strictest one took our best result away from us. Here it is anyway.
What we did
- Three starting dates — June 2016, June 2019 and June 2022 — deliberately picked to cover a calm market, a bad stretch for value investing, and a good one.
- We rebuilt the past as a sealed room. For each date we built a separate database containing only information that existed on that day: filings, prices, annual reports. We then audited every table to confirm nothing dated later had leaked in.
- We ran the real, unmodified production system — the same AI analyst, the same valuation engine, the same gates that run on the site today — against each sealed room.
- 438 companies were sampled, 338 could be analyzed. The rest were mostly delisted companies whose old annual reports we couldn't retrieve.
- We deliberately over-sampled the dangerous zone — the very cheap, low-quality names where Part 1 ran into trouble.
Run one: we threw it away
Our first attempt was contaminated, and we want you to know about it. The sealed rooms held the right filings and the right prices — but one channel had been left open by mistake: the live news feed. Present-day headlines were reaching the analyst while it was supposed to be sitting in 2016, and an audit found that 44% of the analyses quoted them.
That run produced our best-looking numbers. We deleted it, switched the news feed off, and ran the whole study again.
Run two: clean — and it looked like the AI was earning its keep
With today's headlines removed, the analyst's BUY picks still returned a typical +38.3% over three years, against +29.0% for the math engine on its own. A nine-point edge over the arithmetic, with the leak closed. We were fairly pleased with that.
Run three: we blindfolded it
One suspicion remained, and run two could not settle it. A language model has read enormous amounts of text written after2016. Even with no news feed, when it reads a filing and sees a famous company's name, it may simply remember how that company's story turned out.
So we ran the study a third time with every company name and ticker stripped out of the AI's reading material. Every business became “the Company.” The analyst saw the financials, the risk factors and the strategy — but had no idea who it was looking at.
| Version of the test | Status | BUY picks, 3-year typical | Deep-value names: kept vs rejected, 5-year typical |
|---|---|---|---|
| 1. First attempt — leaked today's news | discarded | withdrawn | withdrawn |
| 2. Clean — news channel switched off | counted | +38.3% | +86.2% vs +6.0% |
| 3. Blindfolded — company names hidden too | counted | +31.4% | +70.3% vs +38.6% |
The first run's figures are withdrawn rather than shown. Its returns were never rebuilt after the data correction described in Part 1, so any number for it would come from the broken column and would not be comparable with the other two.
Blindfolded, the BUY picks returned +31.4% — on 28 picks, that is the engine's own +29.0% again, near enough that we cannot honestly tell them apart. The nine-point edge from run two did not survive taking the company names away.
The last column is the one that stings, because it was the result we were proudest of. Inside the deep-value danger zone, look at the gap between the names the analyst kept and the ones it rejected. With names visible: +86.2% versus +6.0% — an eighty-point chasm, which looks like genuine skill at spotting a trap. Blindfolded, the same comparison narrows to +70.3% versus +38.6%. Less than half the gap. Most of what looked like judgment was the model knowing which companies these were.
One more measurement makes the point bluntly. Comparing the blindfolded verdicts against the sighted ones, they agreed only 70%of the time. Knowing a company's namechanged close to a third of our AI's judgments. Among the deep-value names it rejected, only 4 were rejected by both versions.
So we do not claim it
The plain reading is this: in a backtest, a good chunk of what looked like our AI's investing judgment was recognition. It knew these companies. It knew how their stories ended. Strip that away and the edge over plain arithmetic shrinks to something we cannot, at this sample size, honestly tell apart from zero.
So we are retiring the claim. What our backtest proves is the math engine — Part 1, hindsight-free by construction, 25,050 valuations, 28years. That result stands. The AI layer's contribution is not proven by backtesting, and we would rather say so than keep quoting a number that died the moment we tested it properly.
Why we still think the AI layer earns its place
Not as a consolation prize — because of what the test had to remove in order to be fair.
The AI analyst's actual job in production is reading currentinformation: this quarter's guidance, filings as they land, news the day it breaks, management changing its story. A fair backtest must take all of that away, because in the past it becomes knowledge of the future. And that is the awkward truth of run one: it discriminated well precisely because it had current information. That channel is illegitimate in a backtest and entirely legitimate in production — it is the whole point of the product.
Which leaves us with an honest position rather than a comfortable one. We are not claiming the backtest validates the AI layer. We are saying a backtest structurally cannot, and that anyone who tells you their AI stock-picker is backtest-proven should be asked whether they hid the company names.
The real test of the analyst is forward, in the open. Every verdict Numora publishes is stamped with the date it was made and stays on the record — including the ones that go badly. That track record is building in public, one dated call at a time, and it is the only evidence on this subject we would ask you to weigh.
The remaining caveats
- The samples are small. Around 28 BUY calls and roughly a dozen rejected traps per run. These are directional findings, not precise measurements — which cuts against the flattering numbers and the unflattering ones alike. Note also that a gap does remain in the blindfolded run; it is simply far smaller, and far too small a sample to build a claim on.
- The blindfold is imperfect. Hiding the name does not always hide the company: a large enough business is identifiable from its own products and segments inside the filing text. Some recognition almost certainly survived.
- The sample skews toward survivors. Of 438 companies drawn, 100 couldn't be analyzed because their old annual reports were no longer retrievable — and those skew toward companies that later disappeared. Separately, 7–10% of runs failed outright on a technical error and were excluded.
- The historical runs saw less than today's system. Management guidance and news feeds were thin or absent in the frozen worlds. That is the correct choice for a backtest, and it is another reason the backtest understates a live pipeline.
What all this does — and does not — mean
It does mean the core idea holds up on 28 years of real market history: paying less than a business is worth was rewarded, and stocks our engine called cheap consistently beat the ones it called expensive. That is Part 1, and it is the claim we stand behind — it survived a bug that rewrote almost every other figure here.
It does not meanany of the following. That these returns will repeat — they are what happened, not a forecast. That any individual pick will beat the S&P 500; over this period the typical one didn't, and we have put those numbers above rather than in a footnote. That any individual pick will work at all; these are patterns across thousands of names, and plenty of them lost money. That our AI analyst is proven to pick better stocks — Part 2 says the opposite of what we hoped, and we have left it on this page instead of quietly deleting it. Or that Numora is investment advice — it is research you should check, argue with, and make your own decision about.
If you take one thing from this page, take the shape of it: the part of our system we can test rigorously, we tested rigorously, and it works. The part we cannot test rigorously, we are not going to pretend we did. And when a check told us our own numbers were wrong, we republished them rather than the version that flattered us.
Fine print
Hypothetical backtested results, shown for research and education — not investment advice and not a promise of future performance. Past results, real or simulated, do not predict future results. Figures are gross of any fees, taxes, and trading costs; groups are equal-weighted; universe membership is limited to the covered data cross-section at each date; delisted names are carried at their final adjusted close. Coverage starts at the data floor: prices begin December 1997, making June 1998 the earliest vintage with forward returns, and the 1998 vintage is ranked by fundamentals-derived market capitalization because the valuation panel had not yet begun.
All figures on this page are the corrected ones. An earlier publication of this study computed forward returns from a database column that, contrary to its name, held unadjusted prices: returns for stocks that split were understated, sometimes severely, and dividends were omitted. Returns have been rebuilt from split-adjusted prices plus cash dividends per share accumulated over the holding window, without reinvestment compounding — a small understatement we accept rather than model. Superseded figures are not reported. The defect was confined to this study; no production valuation, verdict, or published analysis read that column. The benchmark is the official S&P 500 Total Return index (^SP500TR) sourced from public market data; excess returns are stated in percentage points against it and are not investable results, since they exclude fees, taxes, trading costs, and the cost of tracking an index.
Part 1 evaluates the deterministic valuation engine only, with no language model involved and therefore no hindsight; it is the result we treat as our validated claim. Part 2 evaluates the full production pipeline against date-clamped historical databases across three starting dates, and was run in three arms: one in which a live news channel was inadvertently left enabled (discarded in full, reported here only to describe how it was found and removed), one with that channel disabled, and one that additionally masked every company name and ticker in the model's inputs. Verdict agreement was 73% between the first two arms and 70% between the last two. Sample sizes (438 drawn, 338 analyzed, roughly 28BUY verdicts and a dozen rejected deep-value names per arm) are small enough that all Part-2 findings are directional rather than precise, and identity masking is imperfect for large, easily recognized businesses. Part-2 returns carried the same return-data defect and have been rebuilt on the same corrected basis; the discarded arm's returns were not recomputed and no figures for it are shown. Numora makes no claim that the analyst layer's selection value is established by backtest. Full methodology, per-date statistics, and the reproduction scripts are documented in the project repository.
← Back to the results · How the full six-stage pipeline works →