Outliers in Inventory Forecasting
One wholesale order can inflate every average built from it. How to tell an outlier from a real signal, and whether to exclude, cap, or simply flag it.
One corporate gift order for 200 candles lands in week eight. It is a real sale, it is in your Shopify data, and it is now sitting inside every average built from that quarter. What comes out the other side does not describe your store. Deciding what to do with that week is a judgment call, and it is the step most forecasting advice skips on its way to the formula.
What counts as an outlier
An outlier is not a big number. It is a data point produced by a cause other than the demand process you are forecasting. That makes the test causal rather than numerical: if you cannot name what produced the number, you do not have an outlier yet, only an unexplained one.
Three families cover almost all of them. Single non-repeating transactions: a wholesale or corporate order, a reseller clearing your shelf, one customer buying twelve of something people buy one of. Data artifacts: a test order, a mis-keyed quantity, a duplicated import, bot traffic. And the one everybody forgets, the low-side outlier: a week of zero sales because the product was out of stock. That week is not evidence of no demand. It is evidence that your sales report could not see the demand.
Here is "Cedar & Fig, 250g" over twelve weeks, units sold per week, oldest first:
34, 33, 36, 37, 62, 35, 36, 238, 34, 45, 47, 46
Average all twelve and you get 683 ÷ 12 = 56.9 units a week, a figure that store has never had. Average the seven quiet weeks alone, weeks 1, 2, 3, 4, 6, 7 and 9: (34 + 33 + 36 + 37 + 35 + 36 + 34) ÷ 7 = 245 ÷ 7 = 35 units. That 35 is the real baseline. The whole gap between the two numbers is three weeks, and those three weeks want three different treatments.
Outlier or real signal
Four questions separate them, in this order.
- Can you name the cause? A wholesale buyer, a discount, a press mention, a warehouse error. No nameable cause means no decision yet, only a number to investigate.
- Will the cause recur, and can you anticipate it? A one-off corporate order may recur, but not on a schedule you control. December recurs every year. A promotion recurs whenever you run one.
- Did it persist? One high period is an event. Three consecutive periods at a new level is a level.
- Was it one order or many? Two hundred extra units on one order line, and two hundred spread across sixty orders, are opposite findings. The first is a transaction; the second is demand.
Run the three weeks through that. Week 8's 238 was one 200-unit gift order on top of an ordinary 38-unit retail week: nameable, non-repeating, one order. Exclude the 200. Week 5's 62 was a discount you ran: nameable, recurring, spread across many orders. Flag it, keep it. Weeks 10 to 12, at 45, 47 and 46, persisted across three periods and are made of ordinary-sized orders from different customers. That is not an outlier. That is the number your next forecast should be built on.
The costliest mistake in this area is that last one in reverse. A seasonal peak and a bulk order look identical to any rule that only sees magnitude, and deleting December as an anomaly teaches next year's forecast that December is ordinary. The peak is the only evidence the pattern exists.
Spotting them in your Shopify data
Assume the export is in front of you already; pulling it, and the first-pass cleanup of cancellations and refunds, is covered in using your own Shopify sales data.
Sort the SKU's order lines by quantity, largest first, before aggregating anything. A weekly total hides whether 238 units came from one buyer or ninety, and that distinction decides the treatment. Then check the customer column on the top few lines and the dates against your own calendar of promotions, press and paid campaigns. At twelve data points your eyes beat any rule.
Statistical rules are for when you have too many SKUs to eyeball. They are conventions, not laws, and each carries an assumption worth stating before you trust it.
The common one flags anything more than three standard deviations above the mean. On these twelve weeks the mean is 56.9 and the standard deviation is 55.2, so the cut-off sits at 56.9 + 3 × 55.2 = 222.5. The 238-unit week clears it by about fifteen units. Notice what did the work: the 238 is most of the reason the standard deviation is 55, so the point had to clear a threshold it inflated itself. Replace it with a 100-unit week, a 62-unit wholesale order on top of an ordinary 38, and the same rule on the same series gives a mean of 45.4, a standard deviation of 18.3, and a cut-off of 100.4. The spike sits under its own threshold and passes as normal.
That rule also assumes the values are spread roughly symmetrically around the mean. Weekly sales for one SKU are bounded below at zero and usually skewed right, so the percentage the convention borrows from the normal distribution does not describe the data. The interquartile-range convention, flagging anything more than one and a half interquartile ranges above the upper quartile, is harder for one big point to distort, because quartiles are positions in a sorted list rather than averages. It fails differently. Take a slow mover that sold on seven days out of eighty-four: three quarters of the days are zero, so the upper quartile is zero, the interquartile range is zero, and the threshold is zero. The rule flags every sale the product ever made.
Exclude, cap, or flag
Three treatments, and the wrong one is worse than none.
Exclude the portion, not the period. Week 8's 238 becomes 38 once the wholesale line is removed, and the twelve-week average drops from 56.9 to 483 ÷ 12 = 40.25 units. Deleting the whole week instead would have thrown away 38 units of genuine retail demand and shortened your window by a period.
Cap the value when the demand was real but not repeatable at that size: keep the period, limit it to a ceiling. Useful on erratic-but-genuine SKUs, and the ceiling is your own judgment, which is worth recording as such rather than dressing up as a calculation.
Flag the period and exclude it only from the calculation it would distort. This is the right treatment for promotions and seasonal peaks, because you need those periods intact to plan the next one. Forecasting around promotions covers the uplift estimate and baseline reset in full.
Now finish the arithmetic. Take the wholesale order out and substitute the 35-unit baseline for the promo week's 62, and the twelve-week average lands at 456 ÷ 12 = 38 units. Cleaner, and still wrong, because the last three weeks average (45 + 47 + 46) ÷ 3 = 46. Cleaning removed the contamination; it did not fix a twelve-week window being the wrong window for a series whose level moved three weeks ago. That is a window-length problem, covered in moving average forecasting.
By hand this means holding a promotion calendar, a wholesale customer list and a window-length decision in your head, per SKU, every month. StockCue excludes outlier sales from the velocity behind its suggestions on every plan including Free, and its seasonal indices are damped and capped so one unusually strong month does not triple the next order. Which is exactly why the next section matters: an automatic rule is still a rule, and it has not read your calendar.
The risk of over-cleaning
Every exclusion does two things at once. It lowers the level, and it narrows the measured variability. Narrower variability feeds a smaller safety buffer, so an over-cleaned history produces a forecast that is both lower and less protected than the store it describes. The tidiest series in your spreadsheet is often the least defensible.
Three losses are worth naming. Deleting a seasonal peak removes the only evidence the pattern exists, and no seasonal index can be derived from history that no longer contains the season. Treating a stockout week's zero as real demand drags the baseline down and guarantees you under-order a product that was already selling out. And deleting the start of a step change costs most: weeks 10 to 12 above sit ten units clear of the old baseline and look like three candidate anomalies to anyone measuring against 35. Remove them and you forecast 35 into a store now selling 46.
The guard is a log rather than a better rule. Every exclusion gets a row: period, SKU, what was removed, why, who decided. Without it, nobody can tell six months later whether a thin history means the product sold badly or means somebody cleaned it. Then check whether the cleaning helped, using the forecast-versus-actual review in improving forecast accuracy, bearing in mind that post's warning about percentage error metrics on low-volume SKUs, which are exactly the SKUs where one excluded order changes everything. Skipping the outlier step is already on the site's list of common forecasting mistakes. Doing it too enthusiastically belongs on the same list.
STOCKCUE
StockCue keeps outlier sales out of the velocity behind every reorder suggestion, on every plan including Free, and damps and caps its seasonal indices so one strong month does not multiply the next order. The judgment calls in this post stay yours; the bookkeeping does not.
Install StockCue on Shopify →Frequently Asked Questions
What counts as an outlier in sales data?
An outlier is a data point produced by something other than the demand you are trying to forecast: a single wholesale or corporate order, a test order or mis-keyed quantity, bot traffic, or a week of zero sales caused by an empty shelf rather than by nobody wanting the product. Size alone does not make a number an outlier. A cause you can name does.
Should you remove promotional sales from a forecast?
Not remove: flag. A promotion has a known cause that recurs the next time you run one, so those weeks are the evidence you need in order to plan the next promotion. Exclude them from the baseline that drives everyday replenishment, keep them labelled in the history, and do not delete them outright.
How do you tell a trend from an outlier?
Ask whether it persisted and whether it was spread across many orders. A single high week driven by one large order line is a transaction. Three or more consecutive periods at a new level, made up of ordinary-sized orders from different customers, is a level change, and forecasting from the old baseline after that point will systematically under-order.
What happens if you clean too much out of your sales history?
Every exclusion lowers the level and narrows the measured variability, so an over-cleaned history produces a lower forecast and a smaller safety buffer at the same time. Deleting a seasonal peak is the most expensive version: the peak is the only evidence the pattern exists, so next year's forecast treats that period as an ordinary one.
