Outgrowing the Safety Stock Formula
The classical safety stock formula, Ss = zσ√(L + R), which most software vendors and consultants use, is based on multiple assumptions and shortcuts that make it unfit for real supply chains. This article proposes step-by-step improvements by highlighting its limitations and how they can be addressed (to an extent).
This article is 100% human-written thanks to numerous reviewers. You can find their names in the acknowledgment section at the end of the article.
Throughout this article, we’ll define 4 levels of maturity in using the safety stock formula, each building on the previous one — with level 4 being a paradigm shift away from the classical equation. Here’s already a quick summary. We’ll go through these steps and explain them in the following sections.
Why do we Need Safety Stocks?
We could define safety stock in a broad sense as any extra inventory you would want to have beyond what you would need if demand and supply were 100% predictable (you could use the term safety stock interchangeably for buffer, extra, or caution stock). To put it differently, we need safety stocks to meet specific service-level targets despite inaccurate forecasts and unreliable suppliers.
Safety stock lies between these three inputs, ensuring you meet your requested service level despite disruptions.
This article won’t discuss how to set service-level targets, as we already tackled the subject here (in short, you want to balance risks, costs, and rewards).
Level 1 — Basic Safety Stock Model
In its most basic form, the safety stock equation links safety stocks to demand variability (measured per period), lead time L, and review period R. We have,
Where z is the service level target.[1]
Level 0. The first common error is forgetting to include the review period R in the square root (Ss=zσ√L). When it comes to safety stocks, the review period has the same impact as the lead times: it expands the risk-horizon against which you need protection.[2]
If we represent this model visually, we obtain the following,
As we will discuss throughout this article, this equation contains many limitations due to its underlying assumptions. Let’s fix the worst one first: demand variability.
[1] Despite all my research, I never could identify the origin of this formula or the person who first used or recommended it.
[2] I originally coined the term risk-horizon in my books, Inventory Optimization: Models and Simulations, and Demand Forecasting Best Practices, to denote the lead time plus the review period, as both are equally important when it comes to (safety) stock setting. Unfortunately, most practitioners (software vendors and consultants alike) often forget to account for the review period. I think it is because originally, most academics analyzed continuous inventory policies without review periods. In practice, true continuous policies are extremely rare.
Level 2 — Forecast-based Safety Stocks
Forecasting Accuracy vs. Demand Variability
We need safety stocks not because demand varies, but because forecasts are wrong. Unfortunately, Ss=zσ√(L+R) doesn’t account for forecasting accuracy: It looks at σ, which is the demand variability (some also use its scaled version, the demand coefficient of variation, or COV).
Let’s summarize the differences between demand variability and forecasting accuracy,
- Demand Variability: measures how demand values deviate from the ex-post average demand.
Ex-post means that we measure this after the fact. In this case, we compare demand values to the average demand over the same period. That is not the same as computing the forecast error of a 12-month moving average against future observations.
- Forecasting Accuracy: measures how forecasts deviate from demand values.
These are drastically different concepts. Especially when looking at seasonal, promotion-driven, or mid/long-term forecasts. For the interested reader, see a live explanation of the differences between forecasting accuracy and demand variability from one of my previous webinars, here.
We can easily fix the formula by replacing σ with RMSE, which would be a theoretically correct replacement. But feel free to try other metrics as well — as long as it’s not MAPE (we’ll discuss how to determine which one is best through simulations in the following section).
We now have a much better formula tied to forecasting accuracy rather than demand variability. Let’s take a look at two examples to illustrate the difference between RMSE and COV (which is the scale version of demand variability).
Level 2.5 — Cumulative Error vs. Period Error
Historically, the safety stock formula Ss=zσ√(L+R) measured demand variability per period, and assumed that all periods were identically and independently distributed. In simple words, it assumed that all periods behaved the same, and they were all independent from each other — that is, if you sold more in January, it would not have any impact on February sales. This (wrong) assumption led to the classical formula in which demand variability is computed per period and multiplied by the square root of the risk-horizon to obtain the demand variability over the risk-horizon.
Demand Variability over the Risk-Horizon = Period Demand Variability x √(risk-horizon)
Similarly, our formula zRMSE√(L+R) computes the RMSE per period and multiplies it by the square root of the risk-horizon — following the same assumption of independence.
But there is no need to assume this independence. Instead, we can simply compute the cumulative forecast error over the risk horizon directly. We would get,
Level 3 — Getting Rid of z
In our formula, we haven’t discussed the service level factor z yet. It connects the expected cycle service level you want to achieve to a safety stock level by assuming normally distributed forecast errors. Let’s unpack both this connection and the underlying assumptions. Because both are plain wrong.
Cycle Service Level
Let’s take a minute to define service levels properly. As for forecasting accuracy, many metrics can be used — and choosing the wrong one could hurt your supply chain.
This is one of the first subjects in my supply training course, and when the planners I train read these definitions, they usually recognize the fill rate, or the order fill rate, as the KPI they use to measure service level (whereas retailers typically use on-shelf availability). On the other hand, the cycle service level usually confuses everyone. This is quite normal: the cycle service level is the most counter-intuitive and complicated metric you could possibly use to measure service. It also doesn’t make sense in a business context for a supply chain: different product-market combinations with different lead times would be measured differently (because their order cycles would differ). And if you change suppliers, it could also impact how service is measured, as lead times might change.
Unfortunately, in the safety stock formula, the service level factor z, is mathematically connecting safety stocks to a cycle service level target.[1] Not to the fill rate. Not to the order fill rate. And, in practice, the cycle service level isn’t aligned to these two: You could achieve a 60% cycle service level while delivering 90%+ fill rate. In general, the cycle service level is much more conservative than the fill rates, but the difference depends on many factors (risk-horizon, MOQ, magnitude of the forecast errors).
The only reason planners use z is that it’s convenient, simple, and reassuring. You simply have to put a number in an Excel formula, z=norm.inv(service level, 0,1), to get a safety stock number that holds the promise of service level. Unfortunately, it won’t materialize. I have once witnessed a planner entering 99.99% as a service-level target in this equation — hoping for the best.
As explained in my book, you could upgrade the safety stock formula to connect to the fill rate instead of the cycle service level. It would, unfortunately, result in an intractable equation (that is, you can’t solve it without the help of a solver) and be extremely sensitive to extreme values. In my experience, because of this sensitivity, the fill rate safety stock formula variant results in the worst results in practice (and it’s not just me: no one used it for VN2!).
[1] See my book Inventory Optimization: Models and Simulations for more details.
Assumption of Normality
There is a second reason why z doesn’t work in practice: it assumes that demand (and forecast errors) are normally distributed. Unfortunately, this assumption doesn’t stand up to reality.
In practice, most demand distributions are right-skewed with many low-value observations and a few big values here and there. This usually fits more gamma distributions, as shown in the following graphs.
That said, if you analyze forecast error rather than demand — especially if you look at cumulative forecast error over the risk-horizon — you will find distributions that look more like normal distributions. While still not being normal.
See below an example taken from one of our 2023 case studies. The error over the risk-horizon is more normally distributed, though not normal: the distribution is still left-skewed (there is a small probability of a large negative error, corresponding to cases with extremely high demand, as error = forecast — demand).
Solution: Optimize k using Simulations
We can change and simplify the formula to,
Where RMSE is ideally computed based on the cumulative error over the risk-horizon (or per period, but that’s often less robust as discussed in level 2.5), and k is again a service-level factor — but unrelated to the cycle service level.
Unfortunately, we can’t rely on any formula to optimize k. To tune it, you’ll need to run historical simulations (using historical values, not distribution sampling) to see how different values of k affect costs, inventory, and service, and pick the value that yields the right trade-off for your supply chain. Running a simulation (to optimize k) is more difficult than using a simple formula (to optimize z), but it is much more robust because it uses actual historical demand and forecast values. Simulations can also include most (if not all) business requirements, such as MOQs, random lead times, production calendars, and use more advanced service metrics such as average time in backlog.
RMSE or MAE?
As a last piece of advice, in my experience, using Ss = k×MAE rather than k×RMSE usually yields a better trade-off between inventory and service levels. In practice, as MAE doesn’t overreact as much as RMSE to extreme forecast errors, it avoids safety stock “surges” after big orders (which often result in big forecast errors). In any case, feel free to experiment with both approaches to see which one delivers the best results for your supply chain. If you want to learn how to do that yourself, that’s exactly what I teach in my latest Udemy course.
What About Lead Times?
If you face unreliable suppliers and lead times, it is likely to cause more harm to service levels than inaccurate forecasts. Properly accounting for these in your inventory policies (or safety stock targets) should be one of your priorities.
The safety stock formula often combines a protection against demand variability and lead time reliability.
Where z stands for the service level factor, L the lead time, R the review period (don’t forget about it!), d the period demand, σL the lead time variability, and σd the demand variability.
This equation combines all the shortcomings and poor assumptions we have discussed so far and mixes them with even poorer assumptions about lead times. As with demand, lead times are not normally distributed, which is the core assumption of this formula. I would argue that lead times are even less normal than demand.
Moreover, estimating lead time variability (even without assuming normally distributed lead times) is very difficult, if not irrelevant: with lead times, past performance is a poor indicator of future performance. Let’s illustrate with two examples,
- Your supplier could have faced one-off logistical issues (e.g., an earthquake or flood), resulting in exceptional delays.
- One of your suppliers might have been late with many orders in the past because of limited production capacity on a specific product line. But they just installed a new one.
In my experience with SupChains’ clients, it’s easier to ask planners to set Supplier Reliability Ratings (based on their judgment) and use these in your simulations.
Before we move on to Level 4, let’s recap on the possible evolutions of the classical safety stock formula,
Level 4 — Connect Safety Stocks to the Future
The issue with k×RMSE or k×MAE is that it connects future safety stocks to past performance. But past and future performances do not always perfectly correlate. Moreover, expressing safety stocks in units doesn’t scale well with higher/lower forecasts. To put it simply, such models will recommend the same safety stock target at the start and the end of your high season.
To solve this last shortcoming, we’ll need to outgrow the safety stock formula. I see two main ways to do this,
- Use forecast-coverage periods as safety stocks. Forecast coverages are optimized per product- location using the same simulations as per k×RMSE (see my Udemy course for an example of how to run these simulations). One of the drawbacks is that, even though stock targets are now dynamic (higher at the start of the high season and lower at the end), the coverage still relies on historical performance to set future targets.
In VN2, I proposed a benchmark using non-optimized forecast-coverage-based policies. It beat 86% of the competitors. - Simulate the future using distribution forecasts where every possibility is evaluated. This technique is best explained by VN2’s top competitors. This technique usually no longer relies on historical performance to set stock targets — Great! But it makes the opposite assumption: the distribution forecasts are assumed to be 100% correct when assessing the best order to make. So, using probabilistic forecasts is no guarantee that you will get good policies: Probabilities can be wrong, too. And even totally off. Especially if they rely on… normality assumptions!
Conclusion
Many planners still rely on poor formulas to set safety stock targets, implicitly relying on erroneous assumptions. At times, they also drown in segmentation techniques, resulting in ever-growing complexity, with little to no results. In the inventory competition VN2, of the 25 participants who beat the simple benchmark, a single one relied on segmentation, and no one relied on classical safety stock formulas (levels 0 to 2.5 in this document).
Yet these practices are still at the core of many (if not most) supply planning processes.
It’s time, as a community, that we move on.
Acknowledgments
Koen Cobbaert, Richard Plat, Philip Stubbs, Justin Matthews, Jeff Baker, Franco Gaston Espinaco, John Cao, Bayu Anggoro, Jürgen Defroyère, Prince Boadu, Nural Efe, Gabor Farkas, Thamin Rashid, Dmytro Suzyi, Prasanna Srinivasan
