All essays

AI and the enterpriseGovernance

The Leadership Chasm Between Assistance and Transformation: Why AI Must “Grow Up” to Create Real Value

Why the limiting factor is rarely the model, and far more often the organization around it

8 min read

A copper bridge reaching across a ravine from a small navy screen on one cliff to an automated production line on the other.

The most useful way to read the Copilot debate is not as a contest of model quality, but as a question of value creation. Copilot-style assistants tend to fit the organization as it exists today; agentic systems force the organization to evolve. That is why adoption often hits a ceiling: the limiting factor is rarely the model, and far more often the organization's willingness to redesign workflows, reassign decision rights, and absorb the governance and accountability implications that come with machine action. The real limit of assistance is structural, not qualitative. Copilots sit outside the execution loop: they help people sense, search, draft, and prepare decisions, but they do not own decision rights, they do not execute across systems, and they do not carry accountability when something goes wrong. That is why the value often plateaus even when the outputs look impressive — because the organization's core loop remains unchanged: sense > decide > execute > measure > learn. Assistance improves the "sense" and "decide" phases, but it rarely reaches "execute," and therefore it rarely closes the loop into measurement and learning. Without delegated authority, enforceable policies, and auditable actions, AI cannot compound value; it can only accelerate human work inside the existing system. Transformation begins when AI is moved inside the loop under governance — where it can act, be constrained, be observed, and be improved against outcomes rather than activity. The quickest way to make this cultural point tangible is to classify copilots by "thickness," not by vendor. Thin, medium, and thick copilots map remarkably well to leadership courage. Thin copilots are the least confrontational: generic "assistants everywhere" that rewrite text, summarize, and polish documents. They spread fast because they are lowdisruption and psychologically safe, but they also plateau because they improve activity more than they move outcomes. Medium copilots go further, accelerating knowledge work through drafting and retrieval, yet their return remains structurally uneven because "time saved" does not automatically convert into P&L impact unless leaders change how work is staffed, routed, and measured.

Thick copilots are where the organization accepts real change: they are fused to operational workflows and KPIs and can reliably trigger the next step in a system of record, which is why they are the ones that tend to survive budget scrutiny.

In service operations, that difference shows up clearly in the data: Gartner's framing of the most valuable customer service and support use cases clusters around assisted agents, customer self-service, operational automation, and agentic AI across the stack, rather than generic chat, and Metrigy's published numbers show why queue economics matter, with agent assist associated with a 27.2% reduction in average handle time. The thicker the copilot, the more it forces leaders to confront operating model questions; the thinner the copilot, the easier it is to adopt without changing anything important.

Another clear illustration of this "thickness" is found in modern supply chain orchestration, a domain where 2026 is seeing a strategic shift toward gaining an "Agentic Advantage" to power business growth. While a thin assistant might summarize a shipping delay notification and a medium one answers which orders are affected by searching a database, a thick agentic system moves from information to execution. It identifies a delay, calculates the specific "Line Down" impact, automatically sources alternative suppliers with available stock, and drafts the necessary purchase orders for human approval. By shifting from "polishing an update" to mitigating a million-dollar production stoppage, the AI is no longer just a digital assistant; it is fused to Revenue Protection and Working Capital Optimization. Once you see thickness as a proxy for courage, the "permission wall" becomes easier to explain without slipping into security naivety. In a typical enterprise copilot deployment the wall is not an accident; it is the point. Microsoft is explicit that Copilot only accesses data a user is authorized to access and only surfaces organizational data where the user has permissions, which is foundational safety, not a bug. The real issue is that many valuecreating workflows are end-to-end across tools, teams, and sometimes partners, anda tenant-bounded assistant will naturally feel constrained when the outcome lives in the seams. The moment leaders ask AI to do more than draft — when they ask it to coordinate work across systems — they run into the true enterprise constraint: authorization and accountability, not "better prompts."

That is why leadership is a better metaphor for this transition than technology. Leaders don't grow by getting smarter; they grow by taking responsibility for larger surfaces under clearer accountability. The move from individual contributor to manager is not a boost in output; it is a shift from tasks to outcomes, from personal excellence to orchestration, and from improvisation to doctrine. AI maturity follows the same pattern, but with one crucial twist: AI does not feel consequences. It has no intrinsic sense of responsibility. "Grown-up AI" therefore cannot mean "AI becomes wise." It can only mean that the organization becomes disciplined enough to grant bounded authority under enforceable rules.

This is where the parenting analogy works — but only if it is read as governance, not sentiment. A good parent designs staged autonomy because staged autonomy is the safest path to competence. First you supervise, then you let them try, then you step back while keeping boundaries: what you can do, what you cannot do, and what happens when something goes wrong. In AI terms that is mandates, decision rights, guardrails, escalation paths, and consequences explicit enough to withstand real pressure. If we keep AI permanently in the passenger seat, we get comfortable productivity; if we throw it the keys without doctrine, we get liability.

A structured transition therefore needs an agentic architecture, otherwise "agentic" remains a slogan. The right mental model is to separate intelligence from authority: the model is the reasoning engine, the execution stack is what makes action safe. Practically, this means an intent layer that translates outcomes into plans, an orchestration layer that sequences steps across approved tools, a policy layer that enforces least-privilege decision rights before anything executes, and an observability layer that logs prompts, tool calls, data access, approvals, and outcomes so humans can audit and improve the system.

This is the only credible way to cross the agentic security paradox — how to give an AI enough access to be useful without giving it enough access to be catastrophic — because you stop relying on the model to behave and start relying on external controls to constrain behavior. The OpenID Foundation's work on identity management for agentic AI makes the same point in security language: agents should not impersonate users; they should operate with delegated authority and remain identifiable as agents to reduce risk and preserve accountability.

Traditional operational systems rely on deterministic controls: clear thresholds, explicit checklists, sensors, and procedures that can reliably detect abnormal conditions and stop error propagation early. With probabilistic AI, we cannot pretend the model will pull its own Andon cord — an agent can be confidently wrong, and confidence is not correctness. That means the "stop cord" must live outside the model: boundary checks on inputs and outputs, verification of tool results, reconciliations against systems of record, policy constraints that block high-risk actions, and escalation triggers when uncertainty or blast radius is high. This is exactly why OWASP's LLM guidance is useful: it names failure modes such as prompt injection and insecure output handling that cannot be fixed by better prompting and must be mitigated through external controls, governance, and engineering discipline. Between shallow copilot wins and deep agentic transformation there is a very real valley of death, and acknowledging it makes the argument stronger, not weaker. Shallow wins are cheap and psychologically safe; deep agency demands integration, controls, ownership, and measurement, which raises cost and friction before payback is visible. Gartner predicts at least 30% of GenAI projects will be abandoned after proof of concept by the end of 2025, and Reuters reports Gartner's view that over 40% of agentic AI projects will be cancelled by 2027 due to high costs and unclear business outcomes. This is not a lack of courage in the motivational-poster sense; it is a recognition that governance is expensive when bespoke and ROI collapses when every use case builds its own control plane. The way through is precisely where leadership courage matters: treat autonomy like a promotion — earned, scoped, reviewable, and reversible — and build reusable governance once, then deploy it into a small number of high-value workflows where outcomes are measurable and blast radius is bounded. The era of investing 'on faith’ is closing, replaced by an urgent demand for disciplined capital. Investors are now sharply discriminating between organizations delivering genuine structural value and those merely 'opportunistically surfing the AI hype wave’. According to the Financial Times, the industry has entered a phase of 'hard-headed evaluation, where practical reliability and commercial viability must justify an extraordinary investment surge — with AI capital expenditure projected to top $500 billion in 2026. In this high-stakes environment, assistive copilots often appear as mere incremental UX polish against market expectations that have already priced in total transformation. Durable value will only accrue where AI fundamentally rewires end-to-end workflows, shifts unit economics, and creates a governance-backed moat that survives the accelerating commoditization of underlying models. That is why I'll still say it provocatively: copilot-style assistance must die as the center of gravity — not as a category, but as the end state. If the default mindset remains "keep AI shallow so it does not threaten structures," AI will never be transformative, because transformation lives on the other side of redesigned workflows and adult governance. The destination is not magic; it is governed execution and leadership courage.

References

  • Financial Times, How the AI 'bubble' compares to history (Dec 30, 2025).
  • Financial Times, Three questions AI needs to answer (Jan 2, 2026).
  • Microsoft Learn, Microsoft 365 Copilot architecture and Data, Privacy, and Security for Microsoft 365 Copilot (permission-bound access).
  • OpenID Foundation, Identity Management for Agentic AI (delegated authority; agent identity).
  • OWASP, Top 10 for Large Language Model Applications and related guidance (prompt injection, insecure output handling, and other controls outside the model).
  • Gallup, AI Use at Work Rises (Q3 2025 adoption and usage frequency; tool-category usage patterns).
  • Gartner, Most valuable AI use cases for customer service and support fall into four areas (Oct 8, 2025).
  • Metrigy/ICMI, agent assist impact (27.2% AHT reduction).
  • Gartner, 30% of GenAI projects abandoned after PoC by end of 2025 (press release).
  • Reuters on Gartner, Over 40% of agentic AI projects will be scrapped by 2027 (June 25, 2025).

Written by

Matteo Gatta

Commercial leadership, corporate development and infrastructure

Chief executive of a global communications carrier through its turnaround, and the strategy director behind a national fibre and spectrum position before that.

Full backgroundLinkedIn

Contact

If this describes a decision you are holding, write and say so.

A short note on the situation, and what has to be decided, is enough to establish whether either principal is the right person to be holding it.

Start a conversationMore writing