The Harness Is the Risk: Why Enterprise AI Governance Starts With Permissions, Not Models
There is a version of the AI safety debate that lives entirely in the realm of existential philosophy — Dario Amodei warning about frontier models outrunning humanity’s ability to control them, Anthropic researchers assigning double-digit probability estimates to civilizational catastrophe. That debate is real and worth having. But for the CIO deploying contact center agents this quarter, or the CISO trying to scope an AI audit, or the COO wondering why an automated workflow touched customer data it had no business touching, that debate is also almost entirely beside the point.
The more actionable argument — and the one that deserves serious enterprise attention — comes from NiCE Chief AI Officer Philipp Heltewig, who has made the same point in two different registers over the past several months: the model is not where the risk lives. The harness is. And most enterprises are still building the harness as an afterthought.
The Harness Argument, Precisely Stated
Heltewig’s framing is deceptively simple but analytically precise. A language model, on its own, generates text and proposes actions. It cannot do anything to a customer record, a billing system, or a downstream workflow unless something connects it to those systems and grants it the permissions to act. That connecting infrastructure — the tools, the API integrations, the permission scopes, the rules governing what the model can invoke and under what conditions — is what Heltewig calls the harness.
His analogy to human behavior is instructive: the model is the brain; alignment research is the attempt to shape its values; the harness is the equivalent of laws and physical constraints that determine which tools a person can legally and practically use. A person with poor judgment is dangerous. A person with poor judgment holding a scalpel in an operating theater connected to a patient is categorically more dangerous. The scalpel, the operating theater, and the connection to the patient are the harness.
This reframing has a direct operational implication. Most enterprise AI risk discussions focus on model behavior — hallucination rates, prompt injection vulnerabilities, bias in outputs. Those are real concerns. But they are upstream of the actual damage vector in a production deployment. An agent that misidentifies a customer’s intent and then proposes the wrong resolution is a nuisance. An agent that misidentifies a customer’s intent and then executes a refund, updates a case record, cancels a service, or sends a confirmation email is a liability event. The difference between those two outcomes is not the model. It is what the harness permitted the model to do with its wrong answer.
Why This Is Not Just a Security Question
The natural instinct for CISOs and compliance teams is to file the harness argument under “security and access controls” and treat it as a solved problem domain. It is not, and the reason it is not comes from the architecture of agentic AI systems specifically.
Traditional access control frameworks are built around human users with predictable, bounded task sets. A customer service representative gets access to the CRM, the ticketing system, and the knowledge base. Their access is scoped to their role. The number of actions they can take per hour is naturally constrained by human cognitive speed. Auditing their activity is relatively straightforward because humans leave comprehensible trails.
AI agents break every one of those assumptions simultaneously. An agent operates at machine speed, can be invoked concurrently at massive scale, can chain tool calls in ways that no individual tool call would flag as suspicious, and can be repurposed or redeployed in ways that outrun the access review cycle that initially approved it. Heltewig’s comment in his June CMSWire interview is worth quoting precisely for its operational specificity: “You can give me JavaScript and Claude, and I’m going to build a cool voice demo without Cognigy, but does that demo hold up when I’m getting bombarded with 20,000 concurrent calls? Does it hold up under the security and compliance reviews? Does it have things like auditing, logging and alerting?”
That is not a rhetorical question. It is a checklist that most enterprise AI deployments currently fail. The demo-to-production gap in agentic AI is not primarily a model capability gap. It is a governance infrastructure gap. And the organizations closing that gap fastest are the ones treating harness design as a first-class engineering concern rather than a compliance checkbox.
The McKinsey and Microsoft Corroboration
Heltewig’s argument does not exist in isolation. McKinsey’s 2026 AI Trust Maturity Survey, drawing on approximately 500 organizations, identifies governance and permission controls as the critical differentiator as enterprises move from generative AI toward agentic systems capable of executing tasks. The distinction McKinsey draws — between a system that generates content and a system that takes actions — maps precisely onto Heltewig’s harness framework. The risk profile changes categorically when the model can act, not just answer.
Microsoft’s 2026 Responsible AI Transparency Report covers the same terrain from a platform perspective, emphasizing access controls, permission scoping, and behavioral guardrails for agents connected to external tools and services. Both documents point toward the same enterprise reality: the responsible AI conversation has moved past model evaluation into infrastructure design. Organizations that are still primarily asking “is this model safe?” are asking the wrong question. The right question is “what have we allowed this model to do, and can we reconstruct every action it took?”
Weis Sharpens the Operational Edge
Michael Weis, CRO at Synergy Group AI, adds the operational visibility dimension that Heltewig’s harness argument implies but does not fully develop. Weis’ framing is blunter and more immediately actionable: do you actually know which agent accessed which data, which LLM it called, which user triggered the interaction, what it cost, and whether you can reconstruct the sequence of events when something goes wrong?
For most enterprises currently running AI agents in production, the honest answer to at least some of those questions is no. That is not a frontier AI problem. It is a logging and observability problem. It is the kind of problem that seems manageable when an agent handles a hundred interactions a day and becomes an audit catastrophe when it handles a hundred thousand.
Weis’s summary position — that the urgent question is not whether to slow AI down but whether organizations have put controls in place to safely keep moving forward — is the correct operational framing for enterprise leaders who are not waiting for Anthropic to resolve the alignment problem before deploying agents in the contact center. The frontier safety debate is a parallel track. The enterprise governance question is live and requires answers now.
The Customer Trust Data Makes the Stakes Concrete
Five9’s 2026 CX research, covering 3,000 consumers and 600 decision-makers across three markets, found that 92% of organizations surveyed had implemented or piloted AI in customer service. It also found that 80% of consumers were willing to use AI-powered customer service — but two-thirds still preferred speaking with a human. That gap is not irrational sentiment. It reflects a reasonable consumer assessment of where AI agents currently fail: Qualtrics research across more than 7,000 consumers found higher failure rates for AI-powered customer service than for other AI applications and identified understanding as a specific weakness.
The implication for customer experience leaders is that the harness is not just a governance concern. It is a trust architecture concern. An AI agent that fails to understand a customer’s intent and then takes a consequential action on that misunderstanding does not just create a bad interaction. It creates a trust withdrawal that is harder to reverse than a refund. Customers who experience AI-driven errors in high-stakes service contexts — billing disputes, account changes, service cancellations — do not distinguish between “the model misunderstood me” and “the company used AI to do something harmful to my account.” They attribute the failure to the company. The harness is where the company’s culpability lives.
The Workforce Dimension: Augmentation Beats Substitution
Christian Macaro’s contribution to this cluster of arguments addresses a different strategic question but connects to the governance theme in an underappreciated way. His argument — that using AI primarily to eliminate skilled employees produces a competitive advantage that compresses as AI becomes commoditized — is correct as a strategic matter. The IKEA example he cites is genuinely instructive: routing routine customer service calls to AI while redeploying 8,500 workers into remote interior design consulting turns a cost center into a revenue generator. That is a strategic choice about what AI is for, not just how to implement it.
The connection to the harness argument is this: organizations that deploy AI primarily as a labor cost reduction mechanism have a structural incentive to minimize governance overhead, because governance overhead looks like cost on a spreadsheet. Organizations that deploy AI as an augmentation tool — to make skilled employees more effective and to create more valuable customer interactions — have a structural incentive to invest in the harness, because the harness is what makes the augmentation trustworthy and sustainable. Governance is not just risk mitigation. It is the foundation of the value proposition that makes augmentation work at scale.
The Position Worth Taking
Here is the analytical conclusion that the evidence supports, stated plainly: enterprise AI risk is predominantly a harness problem, not a model problem, and most enterprises are underinvesting in harness design relative to their investment in model capability and deployment speed.
The practical implications separate out cleanly by function. For CIOs and CTOs, the harness question is an infrastructure priority: permission scoping, tool access governance, and agent observability need to be treated as first-class engineering requirements, not add-ons. For CISOs, the threat model needs to shift from “what can a compromised model do?” to “what can a legitimately operating model do with the permissions we have already granted it?” — which is often the larger attack surface. For COOs and CX leaders, the harness question is operational: if an agent goes off course, can your team reconstruct what happened, and do you have intervention mechanisms that work at the speed agents operate? For CFOs, the governance investment question has a clear ROI frame: the cost of a harness failure — regulatory, reputational, and operational — is almost always larger than the cost of building the harness correctly in the first place.
Heltewig’s closing observation in the CMSWire piece deserves to be taken seriously by every executive currently deploying agents in customer-facing contexts: “Keeping a model away from a gun means more than withholding the gun. It means controlling who can hand it one — and securing the seemingly harmless systems it could use to build one.” The organizations that internalize that principle before they learn it from an incident will be in a materially better position than the ones that learn it after.
The model is not the risk. The harness is. Build the harness accordingly.
Based on reporting from Is the AI Model the Risk, or Is It the Harness Around It?, originally published 2026-09-14 17:00:00.
