LinkedInXEmail
regular-banner-bg
Why Most Enterprise AI Agents Never Make It Past the Pilot.
October 7, 26

Why Most Enterprise AI Agents Never Make It Past the Pilot.

Gartner expects more than 40 percent of agentic AI projects to be canceled before the end of 2027, citing escalating costs, unclear business value, or inadequate risk controls. (Gartner, June 2025) That is a strange number for a technology that, in every demo, looks like it already works.

Watch one of these demos and the agent reads an inbox, checks a policy, drafts a response, and takes an action, in seconds, without a human touching the keyboard. Watch the same agent six months later, deep inside a real enterprise environment, and it is usually doing a fraction of that work, with a person checking most of its output before anything happens.

And yet the direction of travel is clear. The idea of the Autonomous Enterprise, a business where AI agents increasingly run everyday work while people set the direction and step in when something falls outside the rules, has become one of the defining visions for enterprise software. Gartner predicts that by 2028, at least 15 percent of day-to-day work decisions will be made autonomously by agentic AI, up from virtually none in 2024.

So the ambition is not the problem.

The gap between the demo and that autonomous future is not primarily a model problem either. It is a systems problem, and it shows up in the same place almost every time.

An agent is only as good as what it is allowed to touch

A demo agent usually works against a clean, cooperative, well-labeled slice of data, built for the demo.

A production agent has to work against the enterprise systems that already run the business: the ones with decades of process exceptions, inconsistent records, custom logic, and approval chains that exist for real financial, operational, and compliance reasons.

And that distinction matters.

The way a business actually operates rarely fits neatly into a standard process. Pricing rules have evolved over years. Plants have their own procedures. Approval paths reflect organizational reality. Supply chain teams have exceptions that everyone working in the business understands, even if they are documented nowhere.

In SAP environments in particular, much of that institutional knowledge has been translated over decades into custom code and business logic. It is often precisely this customization that makes one company’s SAP landscape different from another’s.

An AI agent that can see the standard process but not that context does not fully understand the business it is being asked to act on.

Most agent projects are built by teams who are experts in the model and newcomers to the system of record it needs to act inside. So the agent gets read access to a database export, or a narrow API someone built in a sprint, and the moment it needs to do something that touches a real transaction, a real approval, or a real piece of master data, someone has to step in and do it by hand.

The project quietly becomes an expensive way to draft things a human still has to execute.

The execution gap is bigger than AI

There is another problem hiding underneath all of this.

Enterprise systems were built to manage the business. They were not always built around the way people actually do the work.

A warehouse operator might move through multiple screens to complete one relatively simple task. A field technician might work somewhere without a reliable connection and record information to enter later. An office worker might move between SAP and several other systems just to finish a single process.

Over time, businesses build around these gaps. They create custom applications, spreadsheets, manual processes, integrations and workarounds. Individually, each workaround can seem small. Across thousands of employees and millions of transactions, they become the actual operating reality of the enterprise.

Then AI arrives and inherits exactly the same problem.

Giving an agent access to a language model does not suddenly give it access to the context, processes and business logic required to get work done.

This is the execution gap: the distance between what the core system manages and how the business actually runs.

And for AI agents, closing that gap is the difference between producing an intelligent answer and completing useful work.

“Agentic” and “generative” are not the same claim

Generative AI produces something: a draft, a summary, a suggestion.

Agentic AI is supposed to go further and act on what it produces, inside a real workflow, with real consequences if it gets something wrong.

That second part, acting inside the workflow rather than describing what the workflow should be, is the harder engineering problem by a wide margin. It is also the part many “AI agent” pilots quietly give up on once they hit a system they cannot safely write to.

A useful question for any agent initiative under consideration is this:

When this agent finishes its reasoning, does it change something in a system of record directly, or does it hand a suggestion to a person who then goes and makes the change themselves?

If the honest answer is the second one, the project is a generative AI assistant wearing agentic AI’s name, and it will be measured against agentic AI’s expectations regardless.

That does not make the assistant useless. Generating a response, surfacing information or helping someone make a decision can create enormous value.

But it is not autonomy.

Autonomy is not a switch you turn on

This is also where some of the conversation around AI agents becomes misleading.

There is a tendency to imagine a jump from today’s copilots directly to fully autonomous agents running entire business processes. In reality, enterprise autonomy is much more likely to develop in stages.

First, software made work faster and simpler, while people still held the knowledge, reasoning and decisions.

Then agents began to share some of that load. They could read what was happening, retrieve information, check a stock level, prepare a transaction or take on a defined routine task. But the person still owned the decision.

The next step is more significant: agents that can take a task from start to finish across systems, tools and people, within a clearly defined boundary.

And only then do you reach the version of autonomy everyone is talking about: agents that do not always wait to be asked. They can recognize that something needs doing and act proactively, while escalating anything that falls outside the rules.

The difference between those stages is not simply a better model.

It is how much context the agent has, which systems it can reach, what actions it can take, and how clearly its authority has been defined.

Governance is not the thing slowing you down

It is the thing that was missing from day one.

The projects Gartner expects to survive past 2027 share a pattern: someone decided, before the agent went live, exactly what it was allowed to decide on its own, what it had to escalate, and who was accountable when it got something wrong.

The projects that get quietly shelved usually skipped that step, treating governance as a blocker to work around rather than the thing that makes it safe to give an agent more authority over time.

Because a production agent needs more than access.

It needs boundaries.

Can it read this customer record? Can it update it? Can it create a purchase order up to a certain value? Can it approve one? What happens if the data is incomplete? At what point does a person have to take over? Which authorization determines whether an action is allowed?

These are not questions to answer once the pilot has proved itself. They are part of the architecture of the pilot.

And that is also why failures tend to be expensive rather than small. Scrapping a project after it is already wired into a live workflow costs far more than scoping the authority boundary correctly at the start.

The real challenge is getting AI into the work

The most useful way to think about enterprise agents may therefore be less about the AI itself and more about where the AI sits.

If it sits alongside the business, reading exported data and producing suggestions, it can assist the work.

If it is connected to the applications, processes, business logic and governed data where the work actually happens, it can begin to participate in it.

That requires a few things to come together.

The experience needs to be built around the work rather than around individual systems. It needs to reach people wherever that work happens, including mobile and offline environments. And AI needs to be embedded into those workflows so that it can guide, automate and eventually coordinate work while people remain in control where they need to be.

That is a much harder problem than adding a chat interface to enterprise data.

It is also where the real value of agentic AI begins.

What this means before you fund the next pilot

Before adding another agent pilot to the list, ask a few uncomfortable questions.

Does this agent act directly inside the system that holds the real data, or does it work from a copy, an export, or a narrow read-only view built for the pilot?

Can it understand the business logic and exceptions that make your processes different from the standard?

Can it actually complete an action, or does a person still have to take its output and execute the work somewhere else?

If it needs to escalate, is there an actual person and process on the other end today, or does that get designed later?

And if this pilot is quietly discontinued in six months, would it be because the model was not good enough, or because nobody decided in advance what it was allowed to do?

None of this is an argument against agentic AI.

Quite the opposite.

The Autonomous Enterprise becomes much more credible once we stop treating autonomy as something a model magically creates and start treating it as an operating model that has to be engineered.

People set the direction. Agents take on more of the execution. And when something falls outside the rules, people step back in.

But for that to happen, agents need a governed place to execute. They need access to the real data, business logic and processes that already run the company. And they need the authority to act inside clearly defined boundaries.

For enterprises running on SAP specifically, that question becomes even sharper. The systems holding the real data, approvals, custom logic and process history are also the systems an agent has to understand and act within safely.

That is the gap between an AI agent that works in a demo and one that works in the business.
Closing that gap is what turns agentic AI from an interesting experiment into real operational work. But building that bridge requires a clear roadmap, one that aligns system authority, custom SAP logic, and user workflows into a single governed architecture.

Download The Autonomous Enterprise Playbook to see how you can safely deploy, govern, and scale fully integrated AI agents across your SAP landscape.
Ready to bridge the AI execution gap?