Building an AI agent that works in a demo takes days. Getting one into production, where it touches real customers, real data and real money, is where most stall. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, "due to escalating costs, unclear business value or inadequate risk controls." Each of those can be checked before launch. That's what gates are for.
Why agents stall after the pilot
LangChain's State of Agent Engineering survey of 1,340 practitioners found that 57% of organizations already have agents in production. Asked what holds them back, the top answer was quality, cited by 32%, ahead of latency at 20%. Yet only 52% run offline evaluations before release, and 37% evaluate agents once they're live. Teams know quality is the problem, but many aren't measuring it.
Gartner adds a warning about scope: "Many use cases positioned as agentic today don't require agentic implementations." A simple automation or a search tool can do some jobs with less cost and risk. Check what you're buying, too: Gartner estimates that only about 130 of the thousands of vendors selling agentic AI offer the real thing, a practice it calls agent washing.
The five gates

1. Value and risk
Name the task, the metric that proves success and the cost per completed task you can afford. Then rate the risk: what's the worst outcome if the agent is wrong? An agent that drafts internal summaries and one that issues refunds shouldn't pass through the same gate. If a rule-based workflow would do the job, this gate should say so.
2. Evaluation
Build a test set from real cases, including the awkward ones, and set the pass mark before you look at results. Add adversarial cases that try to push the agent off course. Combine automated scoring with people: in LangChain's survey, 60% of teams use human review and 53% use a model as judge. Re-run the full set whenever the model, prompt or tools change, so a fix in one place doesn't break another.
3. Permissions and guardrails
Give the agent only the tools and data the task needs. Cap the steps, spend and retries per task. Require human approval for anything irreversible, such as payments, deletions and messages to customers. Give the agent its own identity and access, so every action can be traced and revoked.
4. Staged rollout
Run in shadow mode first, where the agent suggests and people decide, then widen access in steps. Compare its results with the people doing the same work before it handles anything alone, and keep a switch that turns it off in seconds.
5. Monitoring and ownership
Trace every run, keep evaluating in production and track cost per outcome. Name an owner who reviews the numbers and handles incidents. LangChain found that 89% of organizations have some observability for their agents, but only 62% can trace every step an agent takes.
Governance that doesn't slow you down
The NIST AI Risk Management Framework gives these gates a backbone in four functions: govern, map, measure and manage. Its Generative AI Profile, released in July 2024, adds actions for the risks specific to generative AI. The practical lesson is to scale the gates to the risk. A low-risk internal agent can pass in days, while one that acts on customers' accounts earns a longer review. Gates applied the same way to everything become bureaucracy. Gates matched to risk let safe projects move fast.
Before you sign off
Ask four questions of any agent heading to production. What does one completed task cost, and what is it worth? What pass rate did it reach on real cases, and who set the bar? What can it do without a person approving it? Who gets the call when it goes wrong? If the answers are clear, it's ready. If they aren't, it's still a pilot.
Our data and AI team designs and reviews agents with these gates in place, from the first use case to production monitoring.
