Reason codes rot. Not all at once — a code here gets duplicated during a MES version bump, a new machine gets bolted onto the line with a generic “Equipment Fault” bucket nobody bothered to break out, a product changeover introduces a failure mode that doesn’t map to anything in the picklist. Six months later your OEE dashboard is technically producing a number, but the loss breakdown underneath it is mostly noise. “Other” has crept from a rounding error to the second-largest category on the Pareto chart, and nobody trusts the availability trend enough to act on it.
This happens most predictably around two events: an MES upgrade (even a minor version bump can reset picklists, remap code IDs, or introduce new default categories) and a line change (new equipment, a retooled cell, a product family with different failure modes). Mid-year is also when a lot of plants are doing OEE program health checks ahead of Q3/Q4 planning, which makes it a natural point to stop patching around the problem and actually rebuild the taxonomy. Here’s how to do that without shutting the line down or blowing up your historical trend data.
Step 1: Pull the code-usage frequency report before you touch anything
Before you redesign a single code, get a hard count of how every existing code has actually been used over a meaningful window — a full quarter minimum, a rolling twelve months if your MES retains it. You want three things out of this report:
- Frequency and duration by code, sorted both ways. A code that fires often but for seconds at a time is a different problem than one that fires rarely but eats hours.
- The “Other”/”Miscellaneous”/”Uncategorized” share of total downtime, by line and by shift. If it’s above the single digits as a percentage of total downtime hours, your taxonomy has already failed operationally, whatever the dashboard says.
- Codes with near-zero usage. These are candidates for retirement, but don’t cut them yet — a code that only fires twice a year might be the one time it matters (a rare changeover fault, a utility interruption). Flag, don’t delete.
Do this segmented by line and by shift before you aggregate. Reason-code drift is rarely plant-wide — it’s usually concentrated on the newest equipment, the newest product, or the crew least invested in the picklist because nobody retrained them after the last change.
Step 2: Find the duplicates and the ambiguous overlaps
MES migrations are notorious for creating duplicate codes with slightly different labels — “Changeover,” “Product Changeover,” “CO – Setup” all coexisting because a config migration copied old codes forward instead of merging them into the new schema. Operators pick whichever one is closest to the top of the list or whichever one a shift lead told them to use once, and now your changeover loss is split three ways and none of the three looks significant on its own.
Go through the full picklist code by code and ask: could two operators reasonably disagree about which of these to pick for the same event? If yes, they’re either duplicates to merge or you have a genuine ambiguity that needs a decision rule (a written one, not tribal knowledge), not just a label fix. This is also where you catch codes that have silently become “wastebasket” categories — vague labels like “Equipment Issue” or “Process Issue” that get used because the real failure mode isn’t in the list.
Step 3: Map every surviving code to the Six Big Losses
This is the step plants skip, and it’s the one that actually connects reason codes to OEE math. Every reason code should trace to one of the six big losses — breakdowns, setup/adjustment, minor stops, reduced speed, startup rejects, production rejects — and by extension to availability, performance, or quality loss. If you can’t map a code to one of the six with confidence, that’s a sign the code is either too granular (fine detail that belongs as a sub-code, not a top-level category) or too vague (a wastebasket code masquerading as diagnostic data).
Build this mapping as an explicit table, not an assumption baked into dashboard configuration that only the MES admin understands. Reason code, six-loss category, OEE factor affected. When engineering or ops leadership asks why performance loss jumped in a given month, you want to be able to trace it to specific codes, not argue about definitions.
Watch for the new-equipment blind spot
When you’ve added a machine or retooled a cell, resist the urge to bolt its downtime onto existing generic codes just to get it live faster. New equipment has new failure modes — different sensors, different fault codes coming off the PLC, sometimes an OPC UA tag structure that doesn’t map cleanly to your existing reason-code hierarchy. Give it real, specific codes from day one, even if that means a slightly longer configuration cycle before go-live. The alternative is baking ambiguity into your data from the first shift the equipment runs.
Step 4: Rebuild the picklist and retrain — don’t just push a config change
Once you’ve merged duplicates, retired dead codes, and mapped everything to the six big losses, the temptation is to push the new picklist through the MES config and move on. That’s how you get a second round of drift six months later. Reason codes are a human-interface problem as much as a data-model problem: operators pick from a list under time pressure, often mid-fault, and the list has to make sense to someone who didn’t sit in on the taxonomy redesign meeting.
- Keep the picklist short and ordered by actual frequency for that specific line, not alphabetically and not identical across every line in the plant.
- Put a one-line definition and a photo or example on the HMI or job aid for any code that’s remotely ambiguous — “Minor Stop – Jam” should show what a jam looks like on that machine, not assume shared understanding.
- Retrain in person, on shift, at the machine — not via an email announcing the new codes. Fifteen minutes with each crew showing the before/after picklist and explaining why codes moved goes a long way toward adoption.
- Assign an owner (usually the shift lead or a designated MES super-user) to answer “which code do I pick” questions in the first couple of weeks, since that’s when bad habits either form or get corrected.
Step 5: Run two weeks in parallel before you cut over for good
Don’t flip the switch cold on a taxonomy that feeds a KPI leadership actually looks at. Run the new reason-code set in parallel with the old logic for roughly two weeks — long enough to cover both shift patterns and at least one full changeover cycle for your dominant product mix. During that window:
- Compare “Other” usage rate week-over-week. It should be dropping, not just moving around.
- Spot-check a sample of logged events against what actually happened on the floor (pull maintenance work orders, cross-reference with operator logs) to confirm codes are being applied correctly, not just consistently.
- Check that your OEE trend line doesn’t show an artificial step-change from the taxonomy switch alone. If availability or performance jumps or drops sharply the week you cut over, that’s a mapping issue, not a real production shift — chase it down before it gets read into a business review as if it were.
Document the old-code-to-new-code crosswalk permanently. Anyone pulling year-over-year trend data eighteen months from now needs to know that “Changeover” in this year’s data and “CO – Setup” in last year’s data are the same bucket, or your historical comparisons quietly become apples-to-oranges without anyone noticing.
What “done” actually looks like
A healthy reason-code taxonomy has a few tells: “Other” sits in the low single digits of total downtime, every code on the active picklist gets used often enough to justify its existence, the six-big-losses mapping is written down somewhere other than one engineer’s head, and operators can explain, without hesitating, the difference between any two adjacent codes on their picklist. If you can’t say yes to all four after a line change or MES upgrade, the taxonomy isn’t done — it’s just live, and it’ll drift again by the next changeover.
This article was written with the assistance of artificial intelligence. While we aim for accuracy, the information may be incomplete, out of date, or incorrect, and should be independently verified before you rely on it for any decision. It is provided for general information only and does not constitute professional advice.
