🎧 Listen to this article (15 min)
Most mid-market manufacturers can tell you their on-time-in-full rate to the decimal. Far fewer can tell you which supplier caused last month's miss, when the first warning sign appeared, or whether anyone acted on it. The number exists. The accountability behind it does not.
That gap has a specific cause. OTIF is calculated from ERP data, but the events that determine OTIF happen in email. A supplier replies to a purchase order confirming a ship date two weeks later than requested. A planner reads it, makes a mental note, and moves on. The ERP still shows the original date. Six weeks later the line goes down and the postmortem blames the supplier, when the truth is the signal arrived on time and nobody captured it.
This article covers how to build supplier delivery scorecards that reflect what actually happened, what data you need beyond the ERP, and how to connect supplier performance to operational consequences instead of leaving it as a number on a slide.
Your ERP records two dates reliably: the date you requested and the date you received. Everything between those two points, the confirmation, the revision, the partial shipment notice, the carrier delay, lives in an inbox.
According to Gartner, 50% of purchase order lines undergo changes after issuance, making real-time supplier visibility a procurement priority. If half your PO lines change and those changes arrive as unstructured email, then ERP-derived OTIF is measuring outcomes while staying blind to causes. You get a score with no diagnostic value.
The practical failure modes look like this. A supplier acknowledges a PO but proposes a later date, and because nobody updates the ERP the variance shows up as a surprise at receipt. A partial shipment closes the line in the system while the balance is still open in reality. A supplier warns of a delay three weeks out and the note sits unread in a shared mailbox. In all three cases the supplier communicated. The system just did not listen.
A Deloitte supply chain study found that 70% of supply chain disruptions originate before materials leave the supplier's facility. That is the window where email holds the only record, and it is exactly the window most scorecards ignore.
Plenty of scorecards track OTIF and stop. That single metric compresses too much into one number to be actionable. If a supplier is at 82%, you cannot tell whether they quote unrealistic dates, quote well and ship late, or ship on time to a date they revised three times.
The metrics worth separating out:
Tracked separately, these tell you what to do. A supplier weak on promise-date accuracy needs a lead-time conversation. One weak on advance notice needs a communication expectation set. One weak on acknowledgement rate needs a process fix, possibly on your side.
Most scorecard disputes are definition disputes. Fix these before you publish a number, because a supplier who successfully challenges your methodology once will challenge every number after that.
Decide explicitly: Is early delivery on time? Receiving three weeks early creates carrying cost and space problems, so many manufacturers define a tolerance window rather than treating early as automatically good. What is your tolerance? A one-day grace period is common, and whether it is calendar or business days matters. Which date is the baseline, your original request date or the supplier's confirmed date? Track both. Requested-date performance tells you about capacity planning; confirmed-date performance tells you about supplier reliability. What counts as "in full," and does a 98% quantity fill pass? Which date stamps the delivery, carrier pickup, dock arrival, or receipt posting? The gap between dock arrival and posting is your own process, not the supplier's.
Write the definitions down and share them with suppliers before the first scorecard goes out. A supplier who knows the rules can improve against them. One who gets surprised by a score will spend the meeting arguing about methodology.
Scorecard quality is a data-capture problem before it is a reporting problem. There are a few ways teams close the email gap.
Manual logging. A planner updates the ERP or a spreadsheet whenever a supplier email arrives. It works at low volume and degrades fast. At 50-plus active suppliers it becomes the first thing dropped in a busy week, and the gaps are invisible until you try to run a report.
Supplier portals. Ask suppliers to enter updates directly. Clean data when it works, but adoption is the constraint. Smaller suppliers with thin admin staff will keep emailing, and you end up maintaining a portal and an inbox.
EDI. Reliable and structured for suppliers who support it. Setup cost and per-connection maintenance mean it rarely extends past your top tier, leaving the long tail uncovered, which is usually where the delivery problems concentrate.
Automated email parsing. Read supplier replies where they already arrive, extract confirmations, dates, quantities, and exceptions, and write them back to the ERP. No supplier behavior change required, which is why it covers the long tail the other three approaches miss.
Aberdeen Group research shows that automated PO tracking reduces operational costs by up to 30% for mid-market manufacturers. The savings come from eliminating manual status-chasing, but the more durable benefit is that the data becomes complete enough to trust.
Whether your procurement team runs on SAP, Oracle NetSuite, Microsoft Dynamics 365, Epicor, or Infor, the constraint is the same: the ERP is a good system of record and a poor system of capture for unstructured supplier communication. For teams running Microsoft Dynamics 365, whether Business Central, Finance and Supply Chain, or Navision, Leverage AI integrates directly with your existing ERP environment to automate supplier PO confirmations, flag exceptions in real time, and surface OTIF data without custom development or ERP modification.
This distinction causes real evaluation mistakes. Most tools marketed as procurement analytics are built for sourcing questions: spend by category, savings realized, contract compliance, supplier consolidation. Those matter to a CPO building a category strategy.
They do not answer the operations question, which is whether the parts arrive when the schedule needs them. Delivery performance requires PO-line-level granularity, promise-date history including revisions, exception events with timestamps, and communication records. Spend analytics runs on invoice and contract data, which is downstream of everything that determines whether a line runs on Tuesday.
When you evaluate, ask a concrete question: can this tool show me every date revision on PO 48213 with who communicated what and when? Spend analytics platforms cannot. Delivery-focused tools can, and that transcript is what makes a supplier review productive.
The two are complementary. Just do not expect a spend tool to tell you why line 12 is late, and be skeptical of any vendor claiming to do both equally well.
The hardest part is showing that a supplier's delivery performance caused a specific operational consequence. Without that link, scorecards stay informational and nothing changes.
The connection runs through the bill of materials and the production schedule. A late component becomes a late work order, which becomes a missed customer commitment. Making that traceable requires PO lines mapped to the work orders and customer orders that consume them, so a supplier miss can be traced to a specific downstream effect.
Once you have it, supplier reviews change character. Instead of "your OTIF is 84%," the conversation becomes "three late shipments in Q2 pushed six customer orders, two of which triggered penalty clauses." That is a business conversation, and it usually produces a different response.
It also changes internal prioritization. A supplier at 78% on non-critical consumables is a nuisance. A supplier at 91% on a single-sourced component with a 12-week lead time is an existential risk. OTIF alone ranks the first as worse. Weighting by operational exposure ranks them correctly.
According to McKinsey, companies with mature supply chain visibility capabilities outperform peers by 15-20% on OTIF metrics. Maturity here means the feedback loop closes: signals get captured, exceptions get routed, performance gets measured against the operations that depend on it.
Scorecards can improve supplier relationships or damage them, and the difference is mostly rollout.
Share the methodology before the first score. Send definitions, the calculation, and the data source. Let suppliers dispute the method before it is applied to them. Start with your top 20 by spend or risk exposure rather than all 200 at once, so you can absorb the disputes that will surface real data-quality problems on your side. Expect that and fix it quietly.
Send scores monthly or quarterly with enough detail to act on. A single number invites argument; a PO-level breakdown invites correction. And separate the metrics in the review conversation, because "you were late 14 times" produces defensiveness while "your confirmed dates were accurate 94% of the time but you only acknowledged 61% of POs" produces a specific fix.
IDC projects that 60% of enterprise procurement teams will transition to AI-powered automation by 2025. The teams getting value are the ones using automation to capture and route supplier signals, not just to generate more reports.
What tools generate supplier scorecards using both ERP data and supplier communications?
You need a tool that reads unstructured supplier email and writes structured events back to the ERP. Most procurement analytics platforms only consume ERP and invoice data, which means the confirmation and revision history never enters the scorecard. Look specifically for automated parsing of supplier replies, promise-date revision tracking, and exception capture with timestamps. Leverage AI does this ERP-agnostically, so the same capture layer works across SAP, NetSuite, Microsoft Dynamics 365, Epicor, and Infor without ERP modification.
How do mid-market manufacturers track OTIF when updates come through email and PDFs?
The three viable approaches are manual logging, requiring suppliers to use a portal, or automatically parsing the emails and attachments as they arrive. Manual logging breaks down past roughly 50 active suppliers. Portals work only where suppliers adopt them, which usually excludes the smaller suppliers where delivery problems concentrate. Automated parsing requires no supplier behavior change, which is why it tends to be the only option that produces complete data.
What is the difference between a supplier delivery dashboard and procurement spend analytics?
Spend analytics answers sourcing questions using invoice and contract data: category spend, savings, contract compliance. Delivery dashboards answer operational questions using PO-line-level data: promise-date accuracy, revision history, exception events. A spend tool cannot tell you why line 12 is late because it does not carry promise-date history or supplier communication records.
Should early delivery count as on time?
Usually not without a tolerance window. Receiving three weeks early creates carrying cost and warehouse space problems, so most manufacturers define an acceptable window, for example on time meaning within one business day early to zero days late, and treat anything outside it as a variance. The important part is defining it explicitly and telling suppliers before you score them.
How do we connect supplier delivery performance to customer impact?
Map PO lines to the work orders and customer orders that consume them. That lets you trace a late component to a specific late work order and a specific missed customer commitment. It also lets you weight suppliers by operational exposure, so a single-sourced long-lead-time component with 91% OTIF gets prioritized above a consumable supplier at 78%.
How many suppliers should we score at once?
Start with your top 20 by spend or risk exposure. The first cycle surfaces data-quality problems, many of them on your side, and a smaller cohort makes those manageable. Expand once definitions are stable and the data holds up to supplier scrutiny.