All guides

By September 6, 20268 min read

How to Measure Supplier Performance

A supplier scorecard built from data you already hold: on-time rate, fill rate, lead-time drift and quality, all of it out of your own purchase orders.

Ask a merchant which of their suppliers is the unreliable one and the answer comes out of memory: the shipment that landed three weeks late during a promotion, or the carton that turned up half empty. Two vivid orders end up deciding the reputation of a supplier you have ordered from fifteen times, and the supplier that quietly slipped from twelve days to nineteen over the same period never gets mentioned.

Why a scorecard beats a feeling

The data to settle it already exists and you generated all of it. Every order you place has a date you sent it and a date you were told to expect the goods. Every receipt has a date the goods landed and a quantity that was counted in. Four facts per delivery, all known at the time. What is missing is not measurement, it is the habit of writing them down somewhere they can later be read in sequence.

Shopify holds the raw material and does none of the arithmetic. A purchase order records products, quantities, costs, payment terms and supplier details, and once it is marked as Ordered the linked inventory transfer is where shipments get tracked and stock gets received, in full or in part. The supplier record holds contact information, address details and recent purchase order line items, per Shopify's suppliers documentation, checked in September 2026. No on-time rate, no fill rate, no lead-time history. Shopify's own supplier guidance lists performance tracking as good practice without supplying a tool for it, so anyone running a scorecard is running it in a spreadsheet or an app.

There is no standard to compare against

Before the formulas, the thing most articles on this subject get wrong. There is no published on-time or fill-rate standard for small ecommerce suppliers. The percentages that circulate come from two places: the supplier-compliance programmes of very large retailers, and companies selling supplier-management software to merchants who have just been told their suppliers are underperforming. The most-quoted retailer programme is reported inconsistently across the trade sources that quote it. The one figure with a named academic attribution behind it describes leading UK retail chains measured at SKU level, not a candle supplier shipping you eight pallets a year.

So this post gives you no target. What it gives you is a way of producing numbers that are comparable in the only two directions that matter: a supplier against its own previous period, and your suppliers against each other on identically defined measures.

A supplier score is not a grade. It is a comparison against the same supplier six months ago.

On-time delivery rate

The standard shape of the formula, stated as it is normally given:

On-time rate = deliveries received on time ÷ total deliveries × 100

Two decisions sit underneath it and both change the answer. The first is the unit you count. Deliveries, purchase orders, order lines and cases all get used, and a supplier's rate can move several points on that choice alone. Pick one, write it down, and never compare two suppliers counted differently.

The second is which date you measure against. A supplier that quotes generously and then hits its own quote scores perfectly against the promised date, and can look poor against the date you actually asked for. Both are legitimate. The promise tells you whether the supplier keeps its word; the request tells you whether it fits your calendar. Track the promise if you track only one, because it is already written on the order.

Here are eight deliveries of Cedar & Fig, 250g from one supplier, Harbourside, over eight months. Each order was placed with a twelve-day promise.

OrderedPromisedReceivedDays late
Jan 6Jan 18Jan 191
Feb 10Feb 22Feb 220
Mar 9Mar 21Mar 265
Apr 13Apr 25Apr 241 early
May 11May 23May 230
Jun 8Jun 20Jun 288
Jul 13Jul 25Jul 316
Aug 10Aug 22Aug 297

With zero tolerance, three of the eight arrived on or before the promised date: 37.5%. Allow a two-day grace period and January joins them, giving four of eight, or 50%. Same supplier, same eight deliveries, thirteen points of difference decided by a rule you set. Which is the argument for writing the rule down before you calculate anything.

Eight deliveries plotted against their promised datesA scatter of eight deliveries from one supplier over eight months, plotted vertically by how many days each one arrived after the date promised on the purchase order. A horizontal line marks the promised date. The January delivery sits one day above the line, February sits on it, March sits five days above, April sits one day below the line because it arrived a day early, and May sits on the line. The last three deliveries, in June, July and August, sit six, seven and eight days above the line. The first five points scatter around the promise with no direction; the final three sit clearly above it, which is the difference between a supplier that is occasionally late and a supplier that is getting slower.Eight deliveries against the promised dateCedar & Fig, 250g from Harbourside, January to August+8 days+4 dayson timethe last three land 6 to 8 days latearrived a day earlyJanFebMarAprMayJunJulAugthe date promised on the purchase order
Collapse these eight points into one average and the shape disappears. Here the order the deliveries came in carries more information than the mean does.

Fill rate

On time says nothing about whether the right amount turned up. Fill rate is the second half:

Supplier fill rate = units received ÷ units ordered × 100

Say which fill rate you mean, every time. The one above is supplier-side: did the supplier ship what you bought. It is not the customer-side fill rate, a service-level measure of how much of your own demand you met immediately, which inventory textbooks split into five related but numerically different definitions. Putting the two next to each other as though they were one number is a common and expensive confusion.

Across those eight Harbourside orders you bought 150 units each time, so 1,200 units in total. Three deliveries came up short: 12 units in March, 30 in June, 8 in August. That is 50 units short, 1,150 received, and a 95.8% fill rate by units. Count on a whole-order basis instead, where an order either arrived complete or it did not, and five of eight were complete: 62.5%. Both describe the same eight deliveries, and the 33-point gap between them is what a supplier that is accurate but not reliably complete looks like.

Combining the two tests gives on time in full, usually written OTIF: the share of deliveries that were both on the promised date and complete. Three of the eight qualify, or 37.5%. Some practitioners multiply the two rates instead, giving 0.375 × 0.625, or 23.4%. Two conventions, one supplier, a fourteen-point spread. Pick one and keep it. For the mechanics of recording a short shipment, see managing a partially received purchase order.

Lead-time drift

On-time rate measures the supplier against its own promise. Lead-time drift measures what the promise cannot hide: how long an order actually takes, and whether that number is moving.

Order to receipt across the same eight deliveries, in sequence: 13, 12, 17, 11, 12, 20, 18, 19 days. The average is 15.25 days. No delivery took 15 days. The first four average 13.25; the last four average 17.25. The supplier has added four days over eight months and never mentioned it, and the average is the one summary that would let you miss it.

That matters because lead time is a direct input to your reorder point. Cedar & Fig sells about 5 units a day, so four extra days is 20 units of cover you were counting on and no longer have. Plan at 12 days while the supplier runs at 17 and the buffer absorbs the difference until the month it does not.

Measuring lead time properly, including the standard deviation that feeds a safety-stock calculation, is worked through in how supplier lead times affect your forecast. One caution before you go: that maths needs observations. With three deliveries behind you, a standard deviation is decoration. Below roughly ten, plot the dates and look at them.

One more convention to settle: which two events you clock between. Order issued to receipt, supplier acknowledgement to receipt, shipped to received, and received to available-to-sell are four different measurements, and the same supplier looks days faster or slower depending on which you pick. Choose the one you can capture every time.

Quality and accuracy

The third group is the one merchants notice most and record least, because the problems surface at unboxing rather than in a system. Four are worth a column each, filled in while you are counting rather than remembered later:

  • Short shipments. Already counted in fill rate, but note whether the supplier flagged the shortage or let you discover it.
  • Wrong items. A different variant, a different size, a substitution nobody agreed to.
  • Damage. Units unsellable on arrival, counted separately from units missing.
  • Paperwork. An invoice that does not match the purchase order, missing barcodes, labelling that fails your own receiving process.

Practitioners bundle these into the idea of a perfect order: accurate, on time, undamaged, complete, correctly documented. It is useful as a way to remember the four columns, but it is a concept rather than a defined standard, so do not treat it as something to benchmark.

Building the scorecard

The minimum useful version is one row per delivery and five fields: supplier, date promised, date received, units ordered, units received. A note column if you want the quality half. That is it. Everything above is derived from those five fields with arithmetic a spreadsheet does in one pass.

Harbourside's eight months, summarised:

37.5%

delivered on the promised date

95.8%

of ordered units received

62.5%

of orders arrived complete

15.3

days average, trending to 19

None of those four is a pass mark, and none means anything alone. They mean something in two comparisons. First against Harbourside's own first half of the year, when the average lead time was 13.25 days and two of four deliveries were on time. Second against your other suppliers scored identically: if Milden delivered five of six on the promised date over the same period on the same definition, you have a reason to prefer Milden where timing matters, and a specific thing to raise with Harbourside.

Sample size decides how hard you lean on any of it. At eight deliveries, one bad one is worth 12.5 percentage points on the on-time rate. At three it is worth 33. A percentage built on four data points should be read as four dates.

Most stores never build this because the promised date and the received date live in different places and neither gets written down at the moment it is known. StockCue does not grade your suppliers, and no honest description of it would say otherwise. What it keeps is the record you grade them from: purchase orders and receiving, full or partial, syncing back to Shopify, on Starter and up, with CSV export on Growth if you want the history in a sheet. For the day-to-day mechanics, see tracking purchase orders in Shopify.

Acting on a bad score

A scorecard that changes nothing is a hobby. Three responses, roughly in order of cost.

Plan around the real number. The cheapest response and often the right one: if the supplier now runs 19 days, plan on 19 instead of the 12 they quote. It costs cash, because the extra cover has to be bought and held, and it buys certainty without a conversation or a change of supplier.

Have the conversation with dates. "You have been late a lot lately" gets a polite reply. "Your last three deliveries landed six, seven and eight days after the promised date, and order to receipt has gone from 13 days to 19 since January" gets a reason, and sometimes the reason is fixable at their end.

Change the exposure. Second-source the SKUs where timing hurts most, or move the supplier to backup. This is the expensive option, because a second supplier means a second minimum order quantity and a second relationship, and it is worth it only where a stockout costs you a customer rather than a sale.

One boundary. Chasing a specific late shipment, expediting it and covering the gap it leaves is a different job, covered in how to manage late purchase orders. The scorecard is for the pattern, not the incident. The wider relationship it sits inside is set out in the Shopify supplier management guide.

Frequently Asked Questions

How many purchase orders do I need before a supplier score means anything?

There is no threshold that switches a score from meaningless to reliable, but the arithmetic tells you how much weight to give it. With eight deliveries, one bad one moves the on-time rate by 12.5 points; with three deliveries, one bad one moves it by 33. Below roughly eight to ten deliveries, read the individual dates rather than the percentage, and do not try to compute a standard deviation of lead time at all.

What counts as an on-time delivery?

Whatever you decide, as long as you decide once and apply it to every supplier the same way. The two choices that matter are which date you measure against, the supplier's promised date or the date you actually asked for, and how many days of tolerance you allow. A supplier that promises late and then hits its own promise scores perfectly against the promised date, which is why the requested date is often the more honest comparison.

Should a partial delivery count as late?

Not as late, but not as complete either. On time and in full are two separate tests, and a shipment that arrives on the promised date with 90 of 100 units is on time and short. Counting it once under each measure keeps both numbers meaningful; folding it into a single rate hides which of the two problems the supplier actually has.

How often should I update a supplier scorecard?

Record the numbers whenever a delivery is received, because that is the one moment when the promised date, the actual date and the received quantity are all in front of you. Read the scorecard on a fixed cadence instead, quarterly for most stores. Reviewing more often than you have deliveries is the common mistake: a supplier you order from four times a year produces four data points a year.

STOCKCUE

A supplier scorecard needs dated purchase orders and dated receipts, which is exactly the pair most stores never capture in one place. StockCue records both: purchase orders, full and partial receiving that syncs back to Shopify, on Starter and up, with forecasting on every plan including Free.

Install StockCue on Shopify →
Rahat Khan, Ecommerce Operations Analyst at Devmerx

Rahat Khan

Ecommerce Operations Analyst

Rahat Khan writes about Shopify inventory operations for Devmerx, the studio behind StockCue: Inventory Forecast.

Need help with your Shopify store?

Devmerx builds and optimises Shopify stores for DTC brands. Book a free 20-minute consultation.