On October 5, 2026, MIT Technology Review Insights published a survey of 300 data, AI, and other technology executives. Its headline finding is that, on average, 34% of organizations' agentic AI projects make it into production.
Anthropic's 2026 State of AI Agents report seems to point the other way. Eight in ten organizations in that survey say agents have already delivered measurable ROI. Anthropic, which sells agents, ran it with the research firm Material, surveying more than 500 US technical leaders in late 2025.
The two figures do not contradict each other, because the MIT Technology Review Insights figure of 34% is a share of projects and Anthropic's 80% is a share of organizations. Take a company with one working agent and three stalled pilots. In Anthropic's survey, that company counts toward the 80%, because it has an agent delivering ROI. In MIT Technology Review Insights' survey, the same company has a 25% production rate, because one of its four projects reached production.
Read together, the reports describe companies that have shipped one agent and are struggling to ship the next. In our experience, the next agent takes on more complex work, which crosses teams and pulls in rules from each of them. Those rules often conflict, so an agent needs retrieval to get the rules in front of it and precedence to decide which one governs.
Retrieval gets the rules to the agent
The MIT Technology Review Insights report makes a strong case for retrieval, getting the right rules in front of the agent. It defines knowledge as understanding "what the data means in the context of individual organizations." Data fragmentation was the most common top challenge to expanding agents' access to knowledge, cited by 55% of respondents.
For the year ahead, respondents plan to prioritize ingestion pipelines, AI-ready APIs, and retrieval-augmented generation (RAG), along with AI evaluation agents and knowledge graphs. The industry is right to invest there, because an agent cannot apply a rule it cannot reach.
In our view, this investment solves retrieval, so the agent sees every rule that applies, but the choice between conflicting rules is still left to the model.
The problem that survives perfect retrieval
Consider an accounts payable team running a three-way match.
- Company policy sets a 0.5% price tolerance between purchase order and invoice, which is a strict policy.
- One strategic vendor was granted 2%, but only until its contract renewal, and that renewal happened last quarter.
- The exception still sits on a wiki page, with nothing on it marking the exception as expired.
- An invoice from that vendor arrives 1.8% over the PO price.
- Retrieval works perfectly and returns both the 0.5% policy and the 2% exception.
- The exception names this vendor, so it looks more relevant to the case than the general policy does.
- The agent applies 2%, clears the invoice, and writes a confident justification.
- The variance surfaces months later, in an audit.
Every retrieval step did its job here. A retriever ranks candidates by similarity to the question, and the exception is the closer match because it names the vendor. Nothing in a similarity score records that the exception lapsed, so in our view retrieval ranks by similarity, not by authority.
An ERP can handle part of this case. In Microsoft Dynamics 365, a vendor-specific tolerance stored as a structured record beats a group or default tolerance through a fixed lookup order. The setup documentation describes no end date for that tolerance, and the failure above lives in a document the agent retrieves.
We call the missing piece precedence. Retrieval tells an agent which rules relate to a case, while precedence records which one governs, under what conditions, and why. Our expectation, which no study has tested, is that a stronger retriever surfaces more relevant exceptions and so hands the model more conflicts to settle alone.
Teams often patch the prompt with a line such as "vendor exceptions expire at renewal." That works for a few rules. A process with hundreds or thousands of rules accumulates patches nobody can audit, and a missing patch fails without any signal.
Our position: an agent that cannot say which rule governs, and why, should not act alone on a case where two rules apply.
Why the first agent ships and the next ones stall
This section is our argument, and neither survey tests it. A first pilot is usually scoped to one team whose people know which rule wins when two disagree. When the agent picks the wrong rule, someone on that team notices.
Cross-team workflows such as order-to-cash and procure-to-pay inherit rules from procurement, finance, sales, and legal that nobody has reconciled in writing. Someone usually did reconcile them, case by case, using judgment that was rarely documented. An agent built on retrieval receives every relevant rule from every team, with no record of how that person chose between them.
Anthropic's data fits this shape, although it does not explain it. In that survey, 57% of organizations deploy agents for multi-stage workflows, including 29% within a single department and 16% across functions. Another 29% plan cross-functional deployments in 2026.
For a CIO, this matters because cross-team workflows carry the largest budgets and the most exposure to unreconciled rules. It matters for audit too, since a wrong rule applied with a clean justification raises no error for monitoring to catch. Deterministic checks still belong on money and safety limits, but a fixed check does not record why an exception was granted or when it ends.
How to tell which problem your stalled project has
Not every stall is a knowledge problem. Gartner, for example, attributes forecast cancellations to escalating costs, unclear business value, and inadequate risk controls. These three checks show whether precedence is part of yours.
- Sort your failures. Take the last 20 to 30 agent errors and sort them into two buckets. In the first, the agent lacked a rule. In the second, it had the right rules and applied the wrong one. No reliable public figure exists for that split, so your own log is the best evidence available, and only the second bucket needs recorded precedence.
- Ask the conflict question. Put this to your team and to every vendor: "Two of your rules both apply to this invoice and they disagree. What does your system return?" An answer of "both, with sources" describes retrieval, which leaves the decision to the model.
- Find the person who resolves conflicts today. Every process has someone who knows that the vendor exception ended at renewal. If that knowledge lives only with them, an agent cannot use it, and capturing it becomes part of the project.
A knowledge graph alone does not pass the second check. Graph database vendors and knowledge-graph platforms give teams a solid way to store rules and the relationships between them. In a case like the vendor exception, though, they tend to leave four gaps:
- No rule authoring for the business. Someone has to build the graph and fill it with rules before an agent can use it, and the people who own those rules rarely can.
- No precedence. The graph can link the policy and the exception, but it does not say which one governs, under what conditions, or why.
- No approvals. Nothing routes a new rule or exception to its owner before it goes live.
- No conflict detection. Two rules that disagree sit side by side until an agent retrieves both.
A team taking the graph route builds those parts itself, so plan for that work.
Where Trail fits
Trail is the context layer your agents run on, and precedence is the problem we built it around. Rules and exceptions are nodes, dependencies and overrides are edges, and the reason each rule exists is stored on its node. A graph database gives a team the tools to build a context graph, while Trail is the context graph, already populated with rules, explicit precedence, and built-in conflict detection.
In the AP example, the 2% exception is a node with an override edge to the 0.5% policy, valid until the vendor's contract renewal. Because the renewal has happened, Trail returns the 0.5% policy as the governing rule. The agent holds the 1.8% invoice and cites the policy, the lapsed exception, and the reason the exception no longer applies.
When two rules conflict and no precedence exists, Trail asks the owner to set precedence before either rule goes live. When a case falls outside what it knows, it asks instead of guessing. Approvals, overrides, and exceptions from live runs feed back so the rules stay current. Trail sits alongside your retrieval over MCP, native retrievers, or REST and GraphQL, or you can build the agent on Trail directly.
What this means for next year's AI budget
Security and privacy will take a large share of next year's agent budget, and they should. The companies furthest ahead on agents are the most likely to call them a major concern, with 72% doing so. Locking down what an agent can see does not tell it which rule to follow. That means context will be the other critical play.
Retrieval spend is reasonable, because an agent needs access to rules before anything else works. If many of your agent's errors come from applying the wrong rule, invest in a way to record which rule wins, with auditable conditions and reasons. That record is what your next cross-team agent needs to reach production.

