Batch Traceability Without an MES: The 4-Hour Test
Ask a plant manager whether their batch traceability is any good and you get a confident yes. Ask when they last traced a lot end to end with a stopwatch running, and the answer is usually the annual audit, performed with advance notice, on a batch chosen because it would go smoothly.
Batch traceability is not a property of your software. It is a property of your last honest mock recall. You can pass on paper and spreadsheets, and plenty of certified sites do. What fails is not the paper, it is the reconstruction: three people, a filing cabinet, a WhatsApp thread and an afternoon spent proving something that was true all along but was never written down in a connected way.
This is the traceability companion to automating production reporting without an MES. Same capture problem, different output.
What is actually required
Three layers, and they are frequently confused with each other.
- The legal floor. In the EU and UK, Regulation (EC) 178/2002 requires one step back and one step forward: who supplied you, and who you supplied. It does not require you to know which specific lot went into which specific pallet.
- The certification layer. This is where the real demands live. BRCGS Food Safety Issue 9 expects full traceability within four hours, tested at least annually, including a quantity check. That is a considerably higher bar than one-up-one-down, and it is the one most sites are actually being measured against.
- The customer layer. Frequently the strictest of the three. A large retailer or a pharmaceutical customer will impose lot-level requirements through the supply agreement that no regulation asked for, and they audit against their own document.
If you ship into the US, check FSMA 204 obligations separately with someone who tracks its timeline; the requirements and dates have moved more than once and are outside the scope of this article.
One distinction worth holding onto: completing the trace in four hours is not the same as being able to run a recall in four hours. The trace is a data exercise. The recall involves customers, transport, regulators and a decision about how much product to pull. Sites pass the first and discover the second is a different animal.
The chain breaks at transformations, not in storage
Receiving and dispatch are the best-recorded events in most plants, because somebody signs for them and somebody else invoices against them. Nobody loses a lot number at goods-in.
The chain snaps in the middle, wherever one lot identity turns into another:
- Blending and batching. Four raw lots go in, one output batch comes out. If the record says the batch used "material X" rather than four specific lot numbers, backward traceability is finished at that point.
- Repacking and splitting. One inbound lot becomes six output pallets with new codes. The link between old and new exists in one person's memory of what they were doing that morning.
- Rework. Yesterday's out-of-spec output gets fed back into today's batch. This is the one that turns a two-pallet recall into a two-day one, because rework quietly connects batches that were otherwise unrelated.
- Unrecorded substitution. The most common of all. The first pallet ran out mid-run, an operator opened the next one, and the batch record still names only the first lot. The document is complete, tidy, signed and wrong.
- Cleaning boundaries and carry-over. Where the allergen or the changeover matters, the trace has to know what ran before, not just what ran.
Every one of these is a data-capture problem at a specific moment on the floor. None of them is solved by buying a system that stores lot numbers, because the lot number was never captured in the first place.
Run an honest mock recall
The exercise is cheap and it is the only measurement that means anything. Do it unannounced, on a batch somebody else picks.
- Pick backwards and forwards. Choose one finished lot shipped in the last quarter and one raw material lot received in the last quarter. Trace each in both directions. Most sites are decent one way and weak the other.
- Start a clock and write down the timestamps. Not a rough sense of how long it took. Actual times, per step, because the audit asks for them and because the timings are where you learn what to fix.
- Do it without the person who knows. If the exercise only works when Marta is in the building, you do not have a traceability system, you have Marta. This is the same single-point-of-failure that shows up in a shared inbox in August, and it is the finding that most changes behaviour.
- Include the quantity check. Reconcile received against used, produced, shipped and remaining. The documents can look complete while the numbers disagree, and the numbers are the part that is hard to fake.
- Record what you could not prove, rather than what you eventually found. The gap list is the deliverable.
Judge the result against this rather than against a feeling:
| Check | Pass | Warning sign |
|---|---|---|
| Time to full trace | Under 4 hours | Anything requiring "we'd have to dig out the folders" |
| Direction | Both ways, equally fast | Forward fine, backward vague |
| Quantity reconciliation | Balances within your tolerance | Nobody has ever tried it |
| Dependence on individuals | Any supervisor can run it | One named person or nothing |
| Rework and substitutions | Visible in the record | Known to have happened, not written down |
| Recall width | Bounded to specific lots | "We'd pull the whole week to be safe" |
That last row is the commercial one, and it is covered below.
Capture at the moment, not at the audit
The instinct after a bad mock recall is to buy an MES. For a site with a handful of lines and no appetite for a two-year project, the cheaper fix is to stop reconstructing.
Keep the paper batch record if it works. What changes is that the record becomes machine-readable at the moment it is created rather than at the moment it is needed:
- Photograph or scan the batch sheet at the end of the run, from a phone on the floor. An agent reads the handwriting, extracts the lot numbers, quantities, times and operator, and writes them into a structured record linked to the work order. This is the same document extraction problem as any inbound paperwork, with worse handwriting.
- Let people report substitutions in their own words. "Ran out of the 20kg flour lot 4471, opened 4488 at about half two" is a perfectly good input. Mapping that free text onto structured lot events is exactly the pattern used for downtime reasons in OEE, and it works for the same reason: people will tell you what happened if you do not make them fill in a form to do it.
- Reconcile against the ERP daily. Lot consumption in the system versus lot consumption on the floor, checked while the shift is still recent enough for someone to remember. Where the ERP is old enough that this sounds impossible, the routes are in connecting a legacy ERP.
- Flag the breaks as they happen. A batch record that names one lot for a quantity larger than the lot contained is arithmetically impossible, and a system that notices on the day is worth more than one that stores it neatly for eighteen months.
- Keep the human decision human. Deciding to recall is not an automation candidate, and it never becomes one. The system's job is to put the affected lot list and the customer list in front of a person in minutes, the standard human-in-the-loop boundary.
Recall width is what this is worth
The traceability business case is usually written as audit compliance, which is why it competes badly for budget against things that make money.
The honest framing is width. In an incident, you recall what you cannot exclude. A site that can prove which four pallets contained the suspect lot recalls four pallets. A site that can only prove the week recalls the week, plus the customer conversations, plus the credit notes, plus the retailer's view of you as a supplier afterwards.
The difference between those two outcomes is a handful of unrecorded substitutions and one blending step where the lot numbers were not written down. That is the entire project, and it is worth quantifying before the incident rather than during it, using the same payback method as any other automation.
Rollout order that finishes: one product family, the one with the most transformations rather than the most volume. Capture its batch records for a month. Run an unannounced mock recall on it. Fix what the gap list says. Then take the next family, with a working method instead of a proposal.
Start here
Pick a finished lot you shipped six weeks ago. Hand it to a supervisor who did not make it. Start a timer.
Whatever number comes back is your real traceability figure, and it is almost certainly not the one in your quality manual. If it is over four hours, the fix is nearly always three or four specific moments on the floor where a lot number exists in someone's head and nowhere else.
Oido captures shop-floor records as they are written: batch sheets photographed at the end of a run and read into structured lot events, substitutions reported in plain language and mapped to the right lot, daily reconciliation against the ERP, and a lot genealogy that answers backwards and forwards in minutes rather than afternoons. See the pattern applied to manufacturing and food distribution, or send us one batch record and we will show you what comes out of it.
Frequently asked questions
What is batch traceability?
The ability to say which raw material lots went into a finished batch, and which customers received that batch, for any batch you made. Backwards to your suppliers and forwards to your customers. In the EU and UK the legal floor is one step back and one step forward under Regulation (EC) 178/2002; retailer and certification schemes then ask for considerably more.
Do I need an MES for batch traceability?
No. Plenty of certified sites trace on paper batch records and a spreadsheet. What matters is whether the links between lots survive, and whether you can reassemble them under time pressure. An MES makes that automatic and expensive; the cheaper route is to keep the paper and capture it as it is written rather than reconstructing it during an audit.
How long should a traceability exercise take?
BRCGS Food Safety Issue 9 expects full traceability within four hours, tested at least annually and including a quantity check. Other GFSI schemes are less specific about the clock. Treat four hours as the working target regardless of which scheme you hold, because it is roughly the point at which a real incident starts making the decision for you.
What is a mass balance check?
Reconciling the quantity of a lot you received against what you used, what you made, what you shipped and what remains. It is the part of a traceability test that catches fiction, because a chain of documents can look complete while the numbers underneath it do not add up. Auditors ask for it precisely because it is hard to fake.
Where do traceability chains usually break?
At transformations, not in storage. Receiving and shipping are recorded well because someone signs for them. Blending, repacking, rework and unrecorded substitutions are where lot identity is destroyed, and the substitution an operator made at 3am because the first pallet ran out is the single most common reason a recall widens.
How does this connect to production reporting?
They are the same capture problem with two outputs. The shift record that tells you how many units ran is the record that says which lots they were made from. Sites that collect one and not the other are usually doing the work twice, which is why traceability tends to arrive free once shift capture is in place.