Every major MES and manufacturing cloud vendor has spent the last release cycle doing the same rebrand: the “copilot” that used to suggest a schedule change is now an “agent” that can make the schedule change. Same underlying model in a lot of cases, different verb. Suggest became execute. And that shift is showing up in your upgrade notes as a feature flag, often defaulted to something that sounds reasonable — “supervised autonomy” or “assisted mode” — with a vendor rep telling you it’s safe because a human can always override it after the fact.
That’s not a governance model. That’s a liability shape.
The honest problem is that “autonomy level” as vendors define it is a single dial applied across a system that makes wildly different kinds of decisions. Adjusting a sequence-dependent changeover schedule and voiding a quality hold are not the same risk category, even though both might live under one agent’s “production optimization” umbrella in the platform’s marketing. If you accept the vendor’s autonomy tiers as-is, you’re letting a product manager who has never seen your floor decide where your human-in-the-loop gates sit. You need your own map, built around your decisions, not theirs.
Stop asking “how autonomous” and start asking “which decisions”
The useful framework isn’t a slider from manual to full-auto. It’s a grid. On one axis: reversibility — can this action be cleanly undone before it causes downstream harm, or does it propagate the moment it fires? On the other axis: blast radius — does this affect one work order, one line, or does it touch genealogy, a customer shipment, or a regulatory record?
Plot your actual shop-floor decision types on that grid and the picture gets uncomfortable fast, because the categories vendors love to demo as agent-friendly are not uniformly low-risk.
Scheduling adjustments
Sequencing and dispatch changes are usually the easiest sell for autonomy, and often rightly so. Reordering jobs within a line to reduce changeover time is generally reversible — you can re-sequence again — and the blast radius is typically contained to throughput and labor, not product identity or compliance. This is the category where letting an agent act, with a fast human review loop rather than a pre-approval gate, is defensible for most discrete and batch environments.
Material substitution
This is where things get sharper. A substitution that swaps an approved-equivalent raw material or component is fine for an agent to propose and even execute if your engineering change process has already pre-qualified the substitution and the ERP/MES bill of materials reflects it cleanly. But an agent making a substitution call based on inventory availability alone, without a tie back to an approved substitutes list, is a genealogy problem waiting to surface at your next audit or, worse, at a customer complaint. Blast radius here isn’t the line — it’s every unit built with that lot, potentially after it’s shipped.
Quality holds
Placing a hold should be cheap to grant an agent — false positives cost you a little throughput and nothing else, and the action is reversible by a quality engineer’s release. Releasing a hold is the opposite. That’s low reversibility (product may have already moved), high blast radius (a nonconforming lot in the field), and it is exactly the kind of decision that regulators, customers, and your own quality system expect to see a named, accountable human sign against. An agent should never be the entity that clears a hold. Full stop. It can recommend, assemble the evidence, even draft the disposition — but the release needs a person’s name on it.
Maintenance dispatch
Auto-generating a work order from a condition-based trigger is generally safe to automate; worst case you get an unnecessary PM and a technician’s mildly annoyed. But letting an agent take a machine out of the production schedule to force maintenance, or reroute an order to another line/asset to accommodate that decision, moves you into blast-radius territory that touches delivery commitments. That one deserves a gate, at minimum a notify-and-delay-by-default pattern where the action executes only if nobody stops it within a defined window.
Set the gates in writing, per category — not per vendor default
Once you’ve mapped your categories, the governance model is just this: for each decision type, define whether the agent can act autonomously, act with a delayed human veto window, or only recommend pending explicit approval. Write it down. Put it in your validation documentation if you’re in a regulated industry, and put it in your standard operating procedures if you’re not. Do not let “the platform’s default agent policy” stand in for that document. Vendor defaults are tuned for demo-ability and adoption metrics, not for your specific exposure.
What the audit trail owes you when the agent acted, not a person
This is the part that gets glossed over in the sales cycle, and it’s the part that matters most when something goes wrong six months later and you’re trying to explain a nonconformance to a customer or an auditor. When an agent — not a person — makes a change that touches genealogy, inventory disposition, or a shipment, your audit trail needs to capture more than “system changed status.” Demand, contractually if you have to:
- The specific model or ruleset version that produced the decision, not just “AI agent,” so you can reproduce or challenge it later.
- The full input context at decision time — sensor readings, inventory state, open orders — not just the output action.
- A distinct actor identity in the record. “Agent-3, policy v2.4” logged separately from any human who later reviewed it, with no blending of the two into a single ambiguous log line.
- A confidence or rationale field, even if it’s imperfect, so a quality engineer isn’t reverse-engineering why the system did what it did from scratch.
- An immutable timestamp on both the decision and any human review or override, matched to your ISA-95 event model so it ties cleanly to the batch or work order record rather than sitting in a separate AI subsystem log that nobody thinks to pull during an investigation.
If your vendor can’t produce that level of traceability today, treat “agentic” as a beta feature regardless of what the release notes call it. Plenty of platforms have genuinely useful recommendation engines wrapped in agent branding; the branding got ahead of the traceability. That’s not a reason to avoid the technology — it’s a reason to insist your MES admin and quality team define the gates before the feature gets flipped on in a routine upgrade, rather than after an agent quietly clears something it shouldn’t have.
The vendors are going to keep shipping more autonomy, faster than most plants can build governance around it. The plants that come out ahead in 2026 won’t be the ones with the most agentic features enabled. They’ll be the ones that can say, specifically, which decisions their AI is allowed to make alone, and prove it in a log the next time someone asks.
This article was written with the assistance of artificial intelligence. While we aim for accuracy, the information may be incomplete, out of date, or incorrect, and should be independently verified before you rely on it for any decision. It is provided for general information only and does not constitute professional advice.
