Third episode of The Ratio Wars. The first two pitted valuation ratios against each other: FCF Yield, whose edge came down to one decisive detail (dividing by enterprise value rather than market cap alone), then Greenblatt's Magic Formula, the series' first composite, which married cheapness and quality only to disappoint in its pure form. This time, a complete change of nature. The Piotroski F-Score isn't a valuation ratio: it's a quality score. Nine criteria, a rating from 0 to 9, and a logic concerned as much with a company's trajectory as with its present state.

The question differs from the previous two episodes. We're no longer asking "at what price are we buying?" but "is the company financially sound, and is it becoming more so?" Can a pure quality score, with no notion of valuation whatsoever, beat the European market on its own? And what happens once we finally add a guardrail on price? Twenty-two years of data will answer both, and the answer holds a lesson I wasn't quite expecting in this shape.
Table of Contents
The Piotroski F-Score explained
Where it comes from
In 2000, Joseph Piotroski, then an accounting professor at the University of Chicago, published a study with a blunt title: Value Investing: The Use of Historical Financial Statement Information to Separate Winners from Losers. The agenda is right there in the subtitle: separate winners from losers using nothing but historical financial statements.
The context is worth pausing on, because it frames the whole positioning of this series. Piotroski wasn't trying to build a strategy from scratch. He started from a known observation: stocks with a low price-to-book ratio (value in its most classic sense) outperform on average, but by construction many of them are also companies in genuine trouble, deservedly cheap. Within that value basket, a minority of solid companies drives most of the performance, while a majority of lame ducks drags it down. His idea: a simple accounting filter to distinguish, among cheap stocks, the ones whose fundamentals are improving from the ones that are deteriorating.
In other words, the F-Score was designed from the outset as a value-strategy improver. That's exactly this series' throughline, and we'll come back to it. But let's start by testing it for what it also is: a standalone quality signal.
The 9 criteria
The F-Score adds up nine binary tests: each criterion met is worth 1 point, each one missed is worth 0. The final score therefore ranges from 0 (poor financial health, and deteriorating) to 9 (a solid company, improving on every front). The nine criteria fall into three families.
| # | Family | Criterion | Point awarded if… |
|---|---|---|---|
| 1 | Profitability | Positive ROA | net income is positive |
| 2 | Profitability | Positive operating cash flow | CFO is positive |
| 3 | Profitability | Improving ROA | ROA exceeds last year's |
| 4 | Profitability | Earnings quality | CFO exceeds net income (low accruals) |
| 5 | Financial structure | Deleveraging | the long-term debt ratio declines |
| 6 | Financial structure | Liquidity | the current ratio improves |
| 7 | Financial structure | No dilution | no net new share issuance |
| 8 | Efficiency | Rising gross margin | gross margin improves |
| 9 | Efficiency | Rising asset turnover | assets turn over faster |
The originality stands out when you compare this score to the Greenblatt-style quality measure from the previous episode. The Magic Formula's ROIC measures a level: is the company profitable, right now? Five of Piotroski's nine criteria instead measure a change: is ROA rising, is the margin widening, is leverage falling, is dilution being avoided? It's a different philosophy of quality: not the persistence of a high level, but the improvement of a trajectory. Where the previous article corrected the Magic Formula with a five-year average ROIC (quality as duration), Piotroski captures quality as direction.
Reading the score
The reading convention is simple: a score of 0 to 2 signals a fragile company, 3 to 6 a neutral zone, 7 to 9 solid financial health on an upward trend. The 7-9 cutoff is the classic one for selection, and 0-2 for avoidance. The whole point of the backtest is to check whether that intuition translates into returns.
Backtest methodology
Universe and period
The backtest covers January 1, 2004 to June 20, 2026, 22 years including the 2008-2009 crisis and several major corrections. Benchmark: MSCI Europe. Reference currency: the euro, consistent with a European universe. I'm reusing the universe from the earlier episodes (stocks on main markets, penny stocks excluded, minimum daily liquidity, special situations cleaned out), aligned geographically with the composition of MSCI Europe.
One methodological point deserves an honest mention, since it sets this article apart from the previous one. The "academic" F-Score applies to non-financial companies: several of its criteria (current ratio, gross margin, asset turnover) barely make sense for a bank or an insurer, whose balance sheet structure is of a different nature. Greenblatt, for the same reason, excluded financials and utilities from the Magic Formula. For consistency, I could have excluded them here too.
I didn't, and that's a small lesson in itself. I tested both versions: including financials and utilities doesn't hurt the result, it improves it very slightly. The score stays discriminating even where theory said it shouldn't apply, likely because the four profitability and financing criteria still make sense for these sectors. So I kept the full universe. The trade-off: this episode's universe differs marginally from the Magic Formula's (which excluded those sectors). As already flagged for the gap between FCF Yield and the Magic Formula, this calls for some caution when comparing figures directly across articles, a gap the series' final episode will settle by re-running every strategy on a strictly identical base.
Calculating the score
For ranking, I use Portfolio123's standard F-Score implementation (PiotFScore factor documentation), which faithfully reproduces Piotroski's nine criteria. Note: Portfolio123 also offers an "All-Stars Piotroski" ranking model that adds a valuation factor (price-to-book) to the score. I did not use it: partly because it distorts the original score by injecting value into it, partly because, tested on this universe, it delivers worse results, price-to-book has been through a rough decade, more on that below. Here, the F-Score is taken for what it is: a pure quality score, with no price dimension at all.
Simulation parameters
So that the only variable changing from one episode to the next is the criterion being tested, I'm reusing the simulation parameters from the earlier articles unchanged.
| Parameter | Value used |
|---|---|
| Number of stocks | 25 |
| Rebalancing | Annual (52 weeks) |
| Commissions | 0.15% of amount |
| Slippage | Variable by liquidity |
| Transaction price | (High + Low + 2×Close) / 4 |
| Weighting | Equal-weighted |
| Ranking | Across the whole universe |
| Benchmark | MSCI Europe (EUR) |
One quirk comes from the nature of the F-Score itself. Being a discrete score from 0 to 9 with many ties, it doesn't split cleanly into exact-value buckets (very uneven group sizes, unstable handling of missing values). So it's tested the standard way a factor is: rank the universe by score, then split it into five quintiles, from worst to best rated. The split isn't perfectly clean, since stocks with identical scores can fall on either side of a boundary, but each quintile is still dominated by a coherent score range, which is all that matters here. That's exactly suited to the central question: do the top-rated stocks actually return more than the bottom-rated ones? Annual rebalancing, finally, is consistent with the pace at which the underlying financial statements are published.
The score's signal
Before building an investable portfolio, one question comes first: does the F-Score genuinely discriminate future returns across the whole score range, not just at the extremes? The universe is ranked into five quintiles by score, from worst to best rated, and we look at each one's annualized return over the 22 years.

The result is as clean as one could hope for. The best-rated quintile returns 8.76% a year, versus 1.42% for the worst-rated, nearly seven annualized points apart over two decades. And it's not merely an extremes effect: the progression is strictly increasing from one quintile to the next, without a single hiccup (Spearman rank correlation of 1). The top quintile clearly beats the universe (6.02%), while the bottom quintile falls well below it. Each quintile rests on a solid sample, on the order of 300 to 600 stocks, which rules out any small-sample artifact.
One important caveat before going further: these quintiles are broad and equal-weighted (several hundred stocks each). They measure the factor's signal, not a strategy you'd actually hold in a portfolio. For that, you need to concentrate.
The investable strategy: the top 25
So let's narrow down to the universe's 25 best scores, equal-weighted, rebalanced once a year. Here's what 22 years of European markets deliver.
| Metric | Piotroski (top 25) | Europe benchmark |
|---|---|---|
| Annualized CAGR | 12.73% | 7.52% |
| Total return | 1,377.06% | 410.55% |
| Sharpe ratio | 0.71 | 0.49 |
| Sortino ratio | 0.94 | 0.64 |
| Max drawdown | -62.06% | -58.42% |
| Annualized std. dev. | 17.06% | 14.04% |
| Beta | 1.03 | n/a |
| Annualized alpha | 5.19% | n/a |
The verdict, this time, is unambiguous: the Piotroski F-Score beats the market as a standalone strategy. 12.73% annualized return versus 7.52%, a Sharpe of 0.71 versus 0.49, an annualized alpha of over five points. One euro invested in the strategy in 2004 is worth nearly 15 today; the same euro in the index is worth a little over 5.
The contrast with the previous episode is striking. The pure Magic Formula (a value+quality composite, more sophisticated on paper) failed to beat the market (6.18% CAGR, Sharpe of 0.35, negative alpha). The F-Score, a pure quality score with no notion of price at all, does twice as well. Quality measured as trajectory is therefore far more self-sufficient than quality measured as level. All due caution kept in mind (one universe, one period, 25 stocks, this is a sample), it's a strong result, and it confirms what you'll find in Les Déterminants de la Richesse (available in French): the F-Score isn't just an improver, it holds up on its own too.
Two caveats, in fairness. First, the result isn't a concentration artifact: widened to 45 stocks, the strategy still returns 12.95% a year for a Sharpe of 0.76, the performance doesn't hinge on a handful of lucky names. Second, and this cuts across the whole series, the last three years are disappointing: over that window the strategy underperforms the index (rolling alpha of -1.27, Sharpe of 0.67 versus 0.99 for the benchmark). Value and cheap quality have been going through a rough patch, and the F-Score is no exception.
The most likely explanation lies in the market regime: these years rewarded expensive growth, with the AI-megacap rally as its most spectacular expression, and in that environment anything tilted toward cheap lags behind, in Europe as in the US. That's precisely the setup discussed in what 2000-2003 taught me to see (FR): a growth premium stretched to extreme valuations that, historically, has eventually turned back in value's favor. Nothing guarantees history repeats, but whoever holds cheap quality today has optionality on their side.
The limits
Before getting to what I think is this article's most instructive point, three limits worth keeping in mind.
The F-Score needs a full history. Five of its nine criteria compare one fiscal year to the previous one. A recently listed company, or one with incomplete accounts, is ineligible or poorly scored by construction. The score is a tool for mature companies.
It's not a valuation detector. And that's its cardinal weakness: a stock can post a F-Score of 9 and trade at a fortune. The score says the company is sound and improving; it says nothing about the price you're paying for it. Hence the need, which we're about to see in action, to pair it with a valuation ratio.
The small-cap edge didn't show up here. The literature (and my own book, on other universes and periods) credits the F-Score with stronger performance among small caps, on the order of several extra points. On this European universe and this specific period, my test restricted to small and micro caps gives a result nearly identical to the full universe, with no significant boost in return, Sharpe, or alpha. Flagging it for accuracy's sake: the edge may exist elsewhere, it doesn't show up here.
Quality has a price
Let's go back to the performance table and stop on the line that's uncomfortable: a max drawdown of -62.06%, worse than the market's (-58.42%), and worse still than the pure Magic Formula's (-57.24%). That runs against intuition. You'd expect a "quality score" to cushion the downside; it does the opposite. In 2008, the strategy lost 46% for the year. A portfolio of financially solid companies should have held up better. Why does it dig deeper?
Here's the hypothesis, and it checks out remarkably well: hiding among the best F-Scores are quality names trading at expensive prices. The score doesn't look at price; nothing stops part of the top 25 from being made up of excellent but overvalued companies. And those are precisely the ones that suffer most in a correction, when the market brutally disqualifies high multiples. Quality with no price guardrail unknowingly buys vulnerability.
To check this, I add a simple valuation filter to the F-Score, the cheapest 20% of stocks in their sector on the price-to-sales ratio (P/S), before picking the top 25 scores. The result speaks for itself.
| Metric | Piotroski alone | P/S + Piotroski | Benchmark |
|---|---|---|---|
| Annualized CAGR | 12.73% | 16.32% | 7.52% |
| Sharpe ratio | 0.71 | 0.87 | 0.49 |
| Sortino ratio | 0.94 | 1.16 | 0.64 |
| Max drawdown | -62.06% | -58.07% | -58.42% |
| Annualized alpha | 5.19% | 8.56% | n/a |
The value filter does everything at once: it lifts returns (from 12.73% to 16.32%), it improves the Sharpe ratio, and above all it brings the drawdown back down to market level. The mechanical detail is subtle but telling: the filtered version is actually slightly more volatile day to day (17.72% std. dev. versus 17.06%), which makes sense since cheap stocks are jumpier, but it digs less in the worst-case scenario. Volatility measures everyday agitation; max drawdown measures the depth of a crash. The value filter doesn't calm the agitation, it tames the extreme losses.
The head-to-head on the years that matter makes the mechanism visible:
| Year | Piotroski alone | P/S + Piotroski | Benchmark |
|---|---|---|---|
| 2008 | -46.40% | -43.32% | -42.71% |
| 2011 | -20.19% | -19.24% | -7.96% |
| 2020 | +23.73% | +2.59% | -3.22% |
| 2022 | -19.14% | -17.39% | -9.15% |
There are two things to read here, and the second matters as much as the first. In 2008, the year of all dangers, the value filter cushions the fall by three points and, over the full 2008-2009 stretch, brings the max drawdown back from -62% to -58%, index level. The margin of safety caps the damage: this is Benjamin Graham's intuition, borne out by experience. Piotroski measures soundness, a good business; Graham demands you don't overpay for it, the margin of safety.
But 2020 is a reminder that there's no free lunch. That year, the value filter cost twenty-one points (+2.59% versus +23.73% for the raw Piotroski): the post-Covid rebound was led by expensive quality, precisely what the filter screens out. The honest takeaway isn't "value protects", it's: the value filter tames extreme losses in crashes, at the cost of missed upside in growth-led rebounds. And make no mistake: even filtered, the strategy is still notably more downside-prone than the index in 2011 and 2022 (-19% versus -8%). The value+quality marriage improves the risk profile; it doesn't turn this into a defensive strategy.
Conclusion
The Piotroski F-Score comes out of the European test looking strong. As a pure signal, it cleanly discriminates returns across the whole score range. As a standalone strategy, it beats the market over 22 years where the pure Magic Formula failed: quality measured as trajectory is, on its own, more robust than commonly assumed. And paired with a simple valuation ratio, it gains further in return and in risk control. At this stage of the series, it's the most convincing building block yet.
But this episode has also produced, from the inside, the problem the next ones will need to solve. We now have both ingredients: the best valuation ratio (FCF Yield) and the best quality measure (Piotroski). What remains is marrying them intelligently, and we've just had a striking preview of why that's not trivial: the choice of value ratio isn't neutral, price-to-book has been disappointing for years, while a simple P/S transforms the strategy. So should you bet on one valuation ratio, hoping to pick the right one, or combine them all so you never have to choose?
That's exactly the idea behind O'Shaughnessy's Value Composite, and the subject of the next episode.
En savoir plus sur dividendes
Subscribe to get the latest posts sent to your email.