A context graph is a graph-structured memory for AI agents that stores three things about a business process: the entities involved, the rules that govern them, and the past decisions made about them, including why each was made. Agents read a small part of it before acting and write each new decision back.
This guide covers how to create a context graph and how to use one. For why agents need one in the first place, read our earlier explainer, What Are Context Graphs?
TL;DR
- A context graph has three layers: entities and their relationships, rules and their dependencies, and decision traces with their reasoning.
- You create a context graph by modeling one exception-heavy workflow, turning its policies into rule nodes, and capturing a decision trace at the moment each decision is made.
- AI agents use a context graph by pulling a small subgraph, checking the rules, retrieving eligible precedents, deciding, and writing the new decision back.
- A past decision in a context graph should count as precedent only if a human approved it, its outcome is known, and its policy version is still current.
- A context graph is worth building when agents make repeated judgment calls with exceptions. For simple lookups, agentic search is enough.
What is a context graph?
A context graph is a memory structure for AI agents in which nodes hold business entities, rules and past decisions, and edges hold the relationships between them. A context graph tells an agent what exists, which rules apply, and how similar cases were decided before. The agent then does not rebuild that picture on every task.
The term spread after Foundation Capital's December 2025 essay, AI's trillion-dollar opportunity: Context graphs. That essay framed a context graph as a record of decision traces: the exceptions, overrides and approvals that systems of record never store.
Our definition is broader. It adds a rule layer, because an agent that cannot follow interlocking rules produces decisions nobody should reuse as precedent.
How is a context graph different from a knowledge graph, GraphRAG, or agent memory?
A knowledge graph stores facts about entities. GraphRAG uses a graph to retrieve document passages. Agent memory stores what an agent learned in past sessions. A context graph differs from all three because it also stores rules and decision traces. Each trace records which rule applied, what exception was granted, who approved it, and why.
| Structure | What it stores | Question it answers | What it misses |
|---|---|---|---|
| Knowledge graph | Entities and facts | What is true? | Why a decision was made |
| GraphRAG | Document chunks linked by a graph | Which passages are relevant? | Decisions that were never written down |
| Agent memory | Facts and preferences from past sessions | What did this agent learn before? | Rules and approvals shared across the organization |
| Context graph | Entities, rules and decision traces | What applies here, and how was this decided before? | Anything not captured at decision time |
The term "context graph" is itself used in at least six ways. When a vendor says "context graph", ask which one they mean.
| Meaning of "context graph" | Source |
|---|---|
| A record of decision traces across entities and time | Foundation Capital, December 2025 |
| Agent memory that links knowledge, conversation history and reasoning | Neo4j |
| A metadata foundation that decision traces attach to | DataHub |
| A live graph of entity state changes for proactive agents | Kumar, arXiv, July 2026 |
| A structure for compacting an agent's long interaction history | Yang et al., arXiv, September 2026 |
| Rules as nodes and rule dependencies as edges | Trail, Nanonets |
This guide uses "context graph" to mean entities, rules and decisions held in one graph.
What goes into a context graph?
A context graph contains three layers. The entity layer holds business objects such as invoices, purchase orders, vendors and approvers. The rule layer holds policies as nodes, with edges showing which rules activate, override or re-check each other. The decision layer holds decision traces: what was decided, by whom, and why.
| Layer | Nodes | Example edges | Failure it prevents |
|---|---|---|---|
| Entity | Invoice, purchase order, budget, vendor, contract, approver | references, draws on, governed by | Re-deriving links from text chunks on every task |
| Rule | Thresholds, limits, approval gates, payment terms | activates, overrides, narrows, re-checks | Dropping or misordering rules that depend on each other |
| Decision | Decision traces | applied rule, granted exception, cites precedent, led to outcome | Re-deciding what the company has already decided once |
The rule layer is the one most definitions leave out, and it is where frontier models fail. On the ComplexConstraints benchmark, which tests densely interlocking rules, our Trail context graph engine fully solved 45.0% of prompts. Gemini 3.1 Pro solved 40.4% and GPT 5.5 solved 38.7%. Trail satisfied 90.0% of individual rules across 1,559 evaluation items.
Here is the failure in one purchase order. The order is 600 units at contract pricing with a 10% volume discount, a $50,000 credit limit, and delivery in 3 days.
- A model with flat context checks the credit limit first and passes it.
- It applies contract pricing and the volume discount.
- The 3-day delivery triggers a $2,000 rush-ship surcharge.
- The final total is $51,200. The credit limit is exceeded and nothing flags it.
In a context graph, the surcharge node has a "re-checks" edge to the credit limit node. The limit is tested again after the surcharge fires.
Our position: build the rule layer before the decision layer. Precedents from an agent that misapplies rules are not worth retrieving.
How do you create a context graph?
You create a context graph in seven steps. Pick one exception-heavy workflow, model its entities, and convert its policies into rule nodes. Define a decision trace schema and capture a trace whenever a decision is made. Record the reason for every human override, then link each decision to its outcome.
- Pick one exception-heavy workflow. Choose a process where the answer is often "it depends", such as invoice approval, deal desk or claims. One workflow gives the context graph a boundary and a way to measure results.
- Model the entities. List the objects the workflow touches and connect them with typed edges. For accounts payable, that means invoice, purchase order, budget, vendor, contract and approver.
- Convert policies into rule nodes. Break each policy document into single rules. Add an edge wherever one rule activates, overrides, narrows or re-checks another.
- Define the decision trace schema. Fix the fields every trace must carry before the first trace is written. The next section lists eleven.
- Capture traces at decision time. Write the trace as a side effect of the agent doing the work. Rebuilding the reasoning later from old chat threads and call notes is lossy.
- Ask why on every human override. When a person changes the agent's proposal, ask one question: why? Store the answer in the trace and mark the trace human-approved.
- Link outcomes later. When the result is known, attach it to the trace. A decision with no recorded outcome cannot show whether it was a good one.
Step 7 is the easiest to skip, because nothing breaks on the day you skip it. The cost arrives later, when the agent cannot tell a good precedent from a bad one.
What should a decision trace record?
A decision trace should record eleven fields: the trigger, the entities involved, the rules applied and their version, the options rejected, the decision, who decided, links to the evidence, the precedents cited, the outcome, a review date, and a trust level. The trace must show what the decision rested on, not only what was decided.
| Field | What it holds | Why it matters |
|---|---|---|
| Trigger | The event that forced a decision | Makes similar cases findable |
| Entities | Links to the invoice, purchase order, vendor or account | Connects the trace to the entity layer |
| Rules applied | Rule IDs and the policy version | Shows whether the policy has changed since |
| Options rejected | The alternatives and why they lost | Stops the agent proposing them again |
| Decision | The action taken | The one field a flat record already has |
| Decided by | The agent or a named human approver | Establishes authority |
| Evidence links | Pointers to the message, call or document | Lets anyone verify the rationale |
| Precedents cited | IDs of earlier traces the decision relied on | Builds the chain of precedent |
| Outcome | What happened afterwards, added later | Separates good precedents from bad ones |
| Review date | When the trace stops being reusable | Prevents stale precedent |
| Trust level | Human-approved or agent-inferred | Gates what can be reused |
The written rationale is the least reliable field in a decision trace, for two reasons.
- Models rationalize after the fact. A study of frontier models, Chain-of-Thought Reasoning In The Wild Is Not Always Faithful (Arcuschin et al.), found post-hoc rationalization in up to 13% of question pairs for production models. No model tested was fully faithful.
- People write shorthand. One practitioner reviewed the "reason for discount" field in a B2B commerce system he had built. The entries were short labels such as "CEO approved", which name an approver but explain nothing.
Our position: store pointers to evidence, and treat the written rationale as a summary of that evidence. Label an agent-written rationale "agent-inferred" until a human confirms it.
What does a context graph look like on a real task?
On an invoice, a context graph gives the agent three things: the linked records, the rules that apply, and a comparable past decision. In the illustrative example below, an agent asked to pay a $12,400 invoice finds the purchase order is $8,800 short, requests an amendment, cites a precedent, and records a new decision trace.
This example is illustrative. The figures are invented to show the mechanics and do not come from a customer.
The question: can the agent pay invoice #842?
| Record | Value |
|---|---|
| Invoice #842 | Acme Corporation, $12,400 |
| Purchase order 7731 | $3,600 of budget remaining |
| Goods receipt | Posted 12 March |
| Vendor status | Not on payment hold |
| Approval policy v4 | Invoices of $10,000 or more need manager sign-off |
| Contract | Payment terms Net-60 |
Entity layer. The agent pulls a six-node subgraph: the invoice, the purchase order it references, the budget that order draws on, the budget owner, the vendor, and the vendor's contract. Every other record stays out of the context window.
Rule layer. The agent walks the rules attached to those nodes.
| Rule | Result |
|---|---|
| Goods must be received before payment | Pass: receipt posted 12 March |
| Vendor must not be on payment hold | Pass |
| Invoice must fit the remaining purchase order budget | Fail: $12,400 against $3,600, short by $8,800 |
| A budget shortfall activates a purchase order amendment | Activated: budget owner must approve |
| Invoices of $10,000 or more need manager sign-off | Required |
| Payment follows contract terms | Net-60 |
An agent with flat context can pass the first two checks, collect the manager sign-off, and pay. The budget rule lives in a different document, and no edge connects it to the payment step.
Decision layer. The agent retrieves one eligible precedent, trace DEC-2026-017 from February. An Acme invoice ran $5,200 over its purchase order. The budget owner approved an amendment because the requester had agreed a scope change by email, and the invoice was paid on terms with no dispute.
The decision. The agent does not pay. It sends the budget owner an $8,800 amendment request with the precedent attached. The budget owner approves and gives the reason: extra units arrived on 12 March. The manager signs off, and payment is scheduled on Net-60 terms.
The new decision trace:
| Field | Value |
|---|---|
| ID | DEC-2026-058 |
| Trigger | Invoice #842 exceeds the remaining budget on purchase order 7731 by $8,800 |
| Rules applied | Budget fit (failed), amendment (activated), manager sign-off at $10,000 (required), policy v4 |
| Options rejected | Pay now (breaks the budget rule); reject the invoice (goods were received) |
| Decision | Amend purchase order 7731 by $8,800, then pay on Net-60 terms |
| Decided by | Budget owner and AP manager |
| Evidence links | Amendment approval, goods receipt of 12 March |
| Precedents cited | DEC-2026-017 |
| Outcome | Pending until payment clears |
| Review date | When approval policy v4 is replaced |
| Trust level | Human-approved |
What a context graph saves here. The saving is in the hops the agent no longer repeats. In the AgenticRAG paper, agentic search on FinanceBench averaged 114.8K tokens per query, 7.8 times the single-shot cost. At 1,000 exception invoices a month, that is about 115 million tokens against about 15 million.
A context graph aims for agentic-search accuracy at a cost closer to the lower figure, because each link is stored once. FinanceBench questions cover financial filings, so treat this as an order-of-magnitude estimate for accounts payable.
How do AI agents use a context graph?
AI agents use a context graph in a five-step loop. The agent pulls the subgraph around the task, checks the rules attached to it, retrieves eligible precedents, makes or proposes a decision, and writes a new decision trace back. Teams also use a context graph to audit decisions and find policies that are often overridden.
- Pull the subgraph. Find the task's main entity, then follow its edges to the connected records. The agent reads a handful of nodes instead of thousands of documents.
- Check the rules. Walk the rule nodes attached to those entities in dependency order. Re-check any rule that a later rule points back to.
- Retrieve eligible precedents. Search past decision traces by similarity, then filter by entity properties and by the eligibility gates in the next section.
- Decide or propose. Act alone when the rules pass and precedent agrees. Send the case to a person when a rule fails or no precedent exists.
- Write the trace back. Store the new decision trace and link it to the precedents it cited.
People use the same context graph for two other jobs.
- Audit. Each decision carries its rules, approver and evidence, so a reviewer can see why a payment was released without asking anyone.
- Policy repair. A rule that is overridden again and again shows up as a cluster of exception traces. That cluster is a prompt to review the rule itself.
When should a past decision not be used as precedent?
A past decision should not be used as precedent when its outcome is unknown or bad, when the policy it was made under has changed, or when its rationale was inferred by an agent and never confirmed by a person. A decision that fails any of these checks is a record of what happened, not guidance.
| Gate | Test | If the trace fails |
|---|---|---|
| Outcome | Is the outcome recorded, and was it acceptable? | Keep it for audit; exclude it from precedent search |
| Currency | Is the policy version the same as today's? | Exclude it until someone reviews it |
| Trust | Did a person approve the decision or confirm the rationale? | Label it agent-inferred; do not reuse it |
Critics of context graphs have a fair point here. FlexRule argues that guiding agents by decision traces treats precedent as policy and repeats the organization's past mistakes. The risk is real, and the answer is to gate precedent, not to discard it.
Twenty identical exceptions can mean two opposite things. The policy may be wrong, or a bad habit may have become normal. Only the outcome field tells them apart. If the exceptions ended well, change the policy. If they ended badly, stop granting them.
Record denials and failures as well as approvals. Jason Stanley argues that failure traces need different primitives from ordinary decision traces if the goal is prevention. A context graph that holds only approvals teaches the agent to approve.
Our position: a precedent without a recorded outcome is an anecdote. Do not let an agent act on it.
Can a context graph be poisoned?
Yes. A context graph can be poisoned when false or malicious content is stored as a decision trace and later retrieved as precedent. Research on agent memory shows such attacks succeed through ordinary interaction. The defense is provenance: only traces approved by an authorized human should be eligible as precedent.
The published evidence comes from agent memory in general.
- MINJA. This attack makes an agent store poisoned reasoning records using queries alone. A follow-up study cites over 95% injection success and 70% attack success under idealized conditions.
- MemoryGraft. Ten poisoned entries in a 110-entry experience store produced 48% poisoned retrieval, as summarized in a 2026 defense paper.
- OWASP. Memory and context poisoning is listed as ASI06 in the OWASP Top 10 for Agentic Applications 2026, as WorkOS describes.
The same pattern applies to decision traces, because both are stored records that agents retrieve and imitate. Consider an illustrative accounts payable case. A vendor email says invoices up to 10% over the purchase order were agreed last quarter. If the agent stores that claim as a decision, every later invoice from that vendor retrieves it as precedent.
Four controls close the gap.
- Record the source of every trace. Store who or what wrote it.
- Restrict precedent to approved traces. Only traces approved by an authorized person are eligible.
- Treat outside content as evidence. A vendor email can support a decision. It can never be one.
- Audit new traces. Review a sample of newly written traces every month.
How do you measure whether a context graph is working?
Measure a context graph with five workflow metrics: precedent hit rate, human override rate, exception recurrence, tokens per decision, and trace quality on a sampled audit. Public agent-memory benchmarks do not measure these. They test recall of conversation facts, and vendor-reported scores on them conflict widely.
| Metric | Definition | Healthy direction |
|---|---|---|
| Precedent hit rate | Share of decisions where an eligible precedent was found | Rises as traces accumulate |
| Human override rate | Share of agent proposals a person changed | Falls over time |
| Exception recurrence | How often the same exception type is granted | Falls after a policy fix |
| Tokens per decision | Total tokens the agent used to reach a decision | Falls as stored links are reused |
| Trace quality | Share of sampled traces with evidence links and a specific rationale | Stays high |
Public benchmarks are a weak guide for two reasons.
- A graph does not always win. On the LoCoMo benchmark, a filesystem-based Letta agent scored 74.0%. The full-context baseline scored 72.9% and Mem0's graph variant scored 68.4%, according to published results compiled under one protocol.
- Vendor scores do not replicate cleanly. An independent test of Mem0's open-source edition scored 32.4% on LongMemEval. The vendor reports 93.4% for its managed platform.
Our position: until a public benchmark tests decision precedent, measure a context graph on your own workflow.
Do you need a context graph, or is agentic search enough?
You need a context graph when agents make repeated judgment calls that involve exceptions, approvals and several systems. Agentic search is enough when the answer already exists in written records and each question is asked once. Agentic search finds what was written down; a context graph also keeps what was decided and why.
Agentic search is strong, and the evidence says so. In the AgenticRAG paper, an agentic loop reached 49.6% recall@1 on BRIGHT against 8.41% for single-shot search. On FinanceBench it reached 92% answer correctness, close to the 94% ceiling with perfect evidence.
It has two limits. It pays the search cost again on every query, and it cannot find reasoning that nobody wrote down.
| Build a context graph when | Stay with agentic search when |
|---|---|
| The same type of decision recurs daily or weekly | Questions are one-off |
| Decisions depend on exceptions and approvals | The answer is stated in a document |
| Rules interlock across several policies | One rule applies at a time |
| Auditors ask why a decision was made | Nobody needs the reasoning later |
| Context spans several systems | One system holds everything |
The two approaches combine well. Let the agent search when the context graph has no answer, then write what it finds back to the graph.
Frequently asked questions about context graphs
Is a context graph the same as a knowledge graph?
No. A knowledge graph stores entities and facts. A context graph also stores rules and decision traces, so it records why a state came to be as well as what the state is.
Which database should you use for a context graph?
A context graph does not need a special database. A graph database works, and so do relational tables with an embedding index for similarity search. The trace schema and the capture process matter more than the storage engine.
How long does it take to build a context graph?
It depends on the workflow. The entity and rule layers of a context graph can be built from existing records and policy documents. The decision layer cannot be backfilled reliably, so it grows only as new decisions are made.
Can you move a context graph from one vendor to another?
There is no open standard for business decision traces yet. Software has an early example in Agent Trace, an open specification from Cursor for recording AI code attribution. Foundation Capital's follow-up essay notes that a context graph captured by an orchestration layer is tied to that tool. Before buying, ask whether traces export in a documented format.
Who should be allowed to read decision traces in a context graph?
Decision traces can hold sensitive reasoning about customers, employees and risk. Apply the access controls of the source systems to each trace. An agent should retrieve only the precedents its user is allowed to see.
Does the EU AI Act require decision traces?
Article 12 of the EU AI Act requires high-risk AI systems to record events automatically over their lifetime. Whether a workflow is high-risk depends on its use case, so confirm the classification with counsel. A provisional agreement would move the high-risk deadline from 2 August 2026 to 2 December 2027.

