Editor's note: The following is a guest post from Jason Gu, managing director in BRG’s AI and decision intelligence practice. The views are the author's own and do not represent those of their employer.
AI promises to improve business processes, making them more efficient and less resource-intensive. This explains why AI spending is projected to grow by 47% this year to more than $2.5 trillion.
But the hard truth is that integrating the technology at scale often makes those processes more complicated — and costly — especially without effective pilot programs.
For example, healthcare providers are spending millions on new AI tools to file medical records into existing digital systems and redeploy staff to higher-value activities. Yet when they rush towards adoption, a variety of process errors — such as inconsistent nomenclature — can lead to mistakes that require more staff, not less.
Unfortunately, not all AI pilots can work out the necessary kinks and most end up going nowhere. Nearly one-third of organizations have run more than 100 AI pilots, according to a recent Zapier survey, but just 13% have broadly deployed these projects across their business. About 40% also said their longest-running AI pilot has been locked in testing for over a year.
I see this regularly in my day-to-day consulting work. At a high level, failures typically stem from overreach (“AI can do everything!") or underreach (“The process/data problems we’re experiencing are model limitations.”)
Both indicate that leadership control issues are actually at the heart of stalled pilots, as organizations focus on the technology or AI model itself rather than the system surrounding it.
In my experience, the leadership control issues that inhibit AI pilots from reaching their full potential can be placed into four buckets:
1. Getting started: identify the business problem and who owns it
It sounds simple: start with a measurable, repeatable business problem and an accountable owner. In reality, not so much.
Some of this derives from an incomplete framing of the problem itself. For example, it might be intuitive for an executive to say, “Let’s use AI to review invoices.” That’s great, but it won’t take you very far.
Instead, leaders should identify the desired business outcome that will result if AI helps solve the problem. A better framing might therefore be, “Let’s use AI to reduce processing time by 95%, achieve 99% accuracy and route the remaining 1% of exceptions for manual review and validation.”
That’s step one. Yet even if everyone supports the pilot, a single accountable business leader needs to establish a measurable baseline, develop necessary process redesigns and evaluate the exception queue to put the plan into action.
This isn’t always going to be the responsibility of an AI-specific leader. While a head of AI can (and should) centralize governance and best practices, outcome ownership should fall to someone with relevant business experience.
2. Designing the pilot right: Match the technique to the problem and ensure outputs can be verified economically
Consider again the invoice example. Ideally, the right AI technology would extract and normalize data from incoming invoices, after which deterministic logic would compare the invoice against the purchase order and goods receipt within defined tolerance rules.
Invoices that pass could be approved according to company policy, while exceptions would be routed to a person with the failure reason automatically attached.
Selecting the right document extraction technology, however, is key. Should you go with optical character recognition? A multimodal model? A large language model? It all depends on document variability.
Verifying potential failures is equally important: teams often assume that, for instance, every invoice that “passes” is automatically correct. Trusting the model at scale may require auditing a meaningful slice of auto-approved outputs to confirm real error rates.
Two high-level design errors are also important to keep in mind. One is using generative AI where conventional analytics would be more reliable; another is building workflows in which every output still requires human review. A robust production design typically combines several techniques — and less generative AI than one might think.
3). Assessing production economics: Evaluate fully loaded cost for each successful task
A critical aim of any AI pilot is to measure how much it will cost to integrate the new technology at scale.
Management might be tempted to evaluate this using cost per token or initial response. That would be a mistake. The metric should be the fully loaded cost per successful task, inclusive of the model, tool and infrastructure costs across all attempts, retries and rework, human review and exception resolution, technology integration, evaluation and ongoing operations, governance and expected losses from incorrect outputs.
Mapping this onto the invoice example, leadership should ask: What share could be processed automatically? What share might enter the exception queue? How much human time does each exception require? What does an incorrect or duplicate payment cost?
4). Choosing the right operating model: redesign workflows around clear controls, human decision boundaries, auditability and continuous learning
The value of a given AI integration comes from slotting the technology into a defined workflow, measuring the outcome and redesigning the surrounding process as needed. That’s where selecting the right operating model comes in.
These three practices should be top of mind for decision-makers:
- Implement a control layer
Connect the workflow to authoritative systems and machine-checkable rules, with defined evaluation, authorization and escalation criteria. Additionally, maintain enough activity history to reconstruct inputs, outputs, tool actions, approvals, exceptions and model and prompt versions.
JPMorgan did this with its Contract Intelligence (COiN) platform: by linking the platform to defined default terms, renewal conditions and regulatory requirements — and building human workflows and validation processes around it — the AI model could process thousands of commercial credit agreements in seconds, with near-zero error rates.
- Assign decision boundaries:
Automate routine cases that meet explicit criteria and route exceptions or high-consequence decisions to people. For instance, Klarna’s customer service AI assistant helped drive a $40 million profit improvement; as the tool scaled, however, the company recognized that it needed to add prominent human escalation paths to handle the most complex issues.
- Create a learning loop:
Convert corrections, exceptions and business outcomes into evaluation cases so the organization can determine whether performance, exception rates and economics are improving.
The best practices noted above align with several of the factors the Zapier survey says leads to successful AI pilots: executive sponsorship from the start, identifying a high-value and easy-to-justify use case and being able to clearly measure ROI.
What tends to be tricky is thinking holistically enough to bring a pilot to life. That’s where effective operating model redesign comes center stage, encompassing not just the technology or model itself but the systems surrounding it.
Today, most AI pilots tend to stall before ever reaching lift-off. With the right leadership control and operational structures in place, tomorrow’s will be better primed to scale.