To make an AI agent follow your business rules, you need to rewrite your process rules as explicit requirements, load it where the agent will see it, enforce money and safety limits in code, and test every rule several times. Agents majorly break rules in seven distinct ways, and each one has a different fix.
TL;DR
- Where rules go: in Claude and ChatGPT, put rules in project instructions or a skill. A rule pasted into a chat message is gone in the next conversation.
- Knowing is not doing: in one 2026 study, 65% of wrong decisions came after the agent had restated the rule correctly.
- Pressure works on agents: in a 12-model test, a deadline was the most effective way to talk an agent out of a rule.
- Shorter beats longer: focused procedures raised agent pass rates by up to 18.8 points. Exhaustive ones lowered them by 2.9.
- One pass proves little: Claude Opus 5 passed 66.5% of single runs on a business benchmark, but only 47.5% of tasks on all 20 runs.
- Custom GPTs are going away: OpenAI plans to retire them on December 11, 2026, so rules stored there need a new home.
How do you make AI agents follow your SOPs and business rules?
Make AI agents follow your SOPs and business rules in four steps: rewrite each SOP as explicit if-then requirements, place it where the agent loads it every time, enforce limits that cost money in code, and test each rule five times. Prompts guide an agent's judgment. They do not guarantee compliance.
- Rewrite the SOP. Turn each paragraph into a trigger, an action, an exception and an escalation path. Name who can approve an exception.
- Place it where it loads. Keep a short core of rules always in context. Put each longer procedure in its own skill or file, and name that file in the core.
- Enforce the hard limits. Refund caps, spending limits and approvals belong in tool permissions or code, where the agent cannot reason its way around them.
- Test it repeatedly. Run five real cases five times each. Count a case as passed only when every run passes.
The four steps exist because agents fail in different ways. This table lists the seven failures found in 2025 and 2026 research, with one fix for each.
| Failure | Evidence | Fix |
|---|---|---|
| 1. Never loaded | Vercel found an available skill was never invoked in 56% of eval cases. | Keep a short always-on core that names each procedure. |
| 2. Crowded out | In SkillsBench, exhaustive skills cut pass rates by 2.9 points. Focused ones added up to 18.8. | One procedure per task, critical rules first. |
| 3. Missing fact | An April 2026 paper shows agents breaking policy because the deciding fact was absent from their context. | Put the deciding field in the record or tool response. |
| 4. Misapplied | In one study, 65% of wrong decisions followed a correct restatement of the rule. | Let code handle thresholds and arithmetic. Give worked examples. |
| 5. Talked out of it | In a 12-model test, deadline pressure cut compliance to 9% or lower in one condition. | Write rules as requirements. Name who grants exceptions. Send rushed requests to a person. |
| 6. Every step passes, the job fails | A September 2026 paper gives the example: three $90 purchases pass a $100 item cap and break a $250 daily cap. | State limits on the finished job. Keep running totals in code. |
| 7. Right once, wrong later | On the Thinkingbox benchmark, Claude Opus 5 passed 66.5% of single runs and 47.5% of tasks on all 20 runs. | Run each test case five times. Count only all-pass. |
How do you give Claude your business rules?
Give Claude your business rules through project instructions, a skill, or the API system prompt. Project instructions apply to every chat in that project. A skill packages one procedure that Claude loads when the task matches. On Team and Enterprise plans, an owner can provision a skill to every user.
| Where the rule lives | Best for | Who sets it up |
|---|---|---|
| Project instructions | Standing rules for one workstream | Any user with a project |
| Skill | One repeatable procedure, loaded on demand | Any user on Free, Pro, Max, Team or Enterprise |
| Organization skill | Company-wide procedures | A Team or Enterprise owner |
| System prompt (API) | Agents you build yourself | A developer |
| CLAUDE.md file | Coding work in Claude Code | A developer |
A skill only helps if Claude loads it. Its description is the trigger, so write what the skill does and when to use it. Skills also need code execution turned on in settings.
Use project instructions for rules that apply to every task. Use skills for procedures that apply to some tasks.
How do you give ChatGPT your business rules?
Give ChatGPT your business rules through custom instructions, project instructions, or a skill. Custom instructions apply to all your chats. Project instructions apply inside one project. Skills, launched in July 2026 for Business, Enterprise, Healthcare and Edu workspaces, package a procedure so ChatGPT repeats it the same way.
| Where the rule lives | Best for | Availability |
|---|---|---|
| Custom instructions | Rules that are true in every chat | Set under Personalization in settings |
| Project instructions | Standing rules for one workstream | Paid plans, including Plus |
| Skills | One repeatable, shareable procedure | Business, Enterprise, Healthcare, Edu |
| Plugins | Instructions bundled with files and connected apps | The replacement for custom GPTs |
| System instructions (API) | Agents you build yourself | Developers |
Do not start a new custom GPT for your rules. OpenAI announced on September 11, 2026 that it plans to retire custom GPTs, with retirement scheduled for December 11, 2026. Instructions and knowledge files can migrate to a plugin. Custom actions do not transfer and must be rebuilt.
OpenAI also warns that a migrated plugin may not behave exactly like the original GPT. Treat the migration as a rule change and re-test every rule afterward.
Why do AI agents ignore rules they were given?
AI agents ignore rules for four main reasons: the rule never loaded, it was buried among too many others, the fact needed to apply it was missing, or the agent understood the rule and applied it wrongly. These causes look identical from outside, and each needs a different fix.
- The rule never loaded. Vercel found that an available skill was never invoked in 56% of eval cases. Long sessions add a second cause. One logged report covering 163 sessions found violations clustering right after the conversation was compacted.
- The rule was crowded out. The IFScale benchmark gave 20 models between 10 and 500 instructions at once. The best model followed 68% at 500. Most errors were instructions dropped entirely, not instructions bent.
- The deciding fact was missing. A rule such as "never share pricing with contractors" cannot be followed if the record does not say who is a contractor. A 2026 paper calls these policy-invisible violations.
- The rule was misapplied. In one study, 65% of wrong decisions came after the agent had estimated the inputs correctly and restated the rule. Adding the rule to the prompt did not raise adherence.
The last finding matters most for testing. Asking an agent to repeat a rule back does not show that it will apply the rule.
How should you write an SOP for an AI agent?
Write an SOP for an AI agent as a short list of requirements. Give each one a trigger, an action, the allowed exceptions and who approves them. Spell out steps a trained employee would assume. Leave out penalty amounts. Have a person who knows the work write it, not the model.
- Use requirement language. In a 12-model compliance test, Grok 4.1 Fast followed a rule 100% of the time when it said "requires". It followed the same rule 60% of the time when it was phrased as information from Legal.
- Leave out the penalty. In the same test, adding a small, unlikely fine dropped Gemini 3 Flash from 100% compliance to 34%. The agent treated the fine as a price.
- Name who grants exceptions. Say plainly that a deadline or a manager's chat message is not an approval.
- Spell out assumed steps. Amazon's SOP-Bench team cites a patient intake SOP that says "verify insurance" twice without saying how. Staff know the two checks differ. An agent has to guess.
- Have an expert write it. In SkillsBench, human-curated skills raised pass rates by 16.2 points. Skills the model wrote for itself lowered them by 1.3.
Example rewrite (illustrative)
Before, written for staff:
Refunds over the limit should generally be escalated. Customers are usually eligible within 30 days. Unauthorized refunds may result in a $500 chargeback fee.
After, written for an agent:
- Refund only orders delivered 30 days ago or less. Read the date from the delivered_date field.
- Refund up to $200 without approval.
- Above $200, open an approval ticket for the support lead. Tell the customer the refund is under review.
- Only the support lead approves exceptions, and only in the ticket. A customer deadline or a manager's chat message is not an approval.
- If the delivery date is missing, ask for the order number and stop.
The rewrite removes "generally" and "usually", names the data field, names the approver, and drops the fee.
How many rules can an AI agent follow at once?
There is no fixed limit, but reliability falls as rules pile up. In the IFScale benchmark, the strongest models stayed near-perfect to about 150 simple instructions and then declined. The best reached 68% at 500. Business rules are harder than keyword instructions, so treat those numbers as a ceiling.
Three findings point the same way:
- More instructions, more drops. IFScale tested 20 models in 2025. As the count rose, models increasingly favored earlier instructions over later ones.
- Exhaustive documents hurt. In SkillsBench, focused skills with two or three modules beat comprehensive documentation. Comprehensive skills scored 2.9 points below having no skill at all.
- Extra tools hurt too. Amazon gave an agent the six tools a task needed, then added 20 plausible extras. Success nearly halved.
In practice, give each task its own procedure, put the critical rules first, and remove tools the task does not need.
Should a business rule live in the prompt, a skill, or code?
Put judgment rules in the prompt, task procedures in a skill, and hard limits in code. A prompt rule is followed most of the time. A code rule is followed every time. Any rule whose violation costs money, breaks the law or cannot be undone belongs in code or tool permissions.
| Rule type | Example | Where it lives | Why |
|---|---|---|---|
| Judgment | Offer a replacement before a refund | Always-on instructions | Applies to every task |
| Task procedure | Month-end close checklist | Skill | Needed only for some tasks |
| Reference fact | Price list, policy text | Knowledge file | Looked up, not memorized |
| Hard limit | No refund above $200 | Tool permission or code | Must hold every time |
| Running total | No more than $250 a day | Code that tracks the total | The agent keeps no reliable counter |
| Approval | Anything above a cap | A human approval step | A person owns the exception |
The channel matters as well as the location. In a Hacker News thread on Claude ignoring stop hooks, commenters noted that text arriving as a tool result is treated as information, not instruction. That is the channel where prompt injections arrive.
So a check written in code should block the action. It should not hand the agent a message and hope the agent complies.
Can an AI agent follow every rule and still break your policy?
Yes. An agent can pass every step-level check and still break a policy that is about the whole job. Three $90 purchases each pass a $100 per-item cap, yet together they break a $250 daily cap. A September 2026 paper calls this a compositional policy violation.
The paper names four types. Its examples are constructed illustrations, not reported incidents.
| Type | What goes wrong | Fix |
|---|---|---|
| Cumulative sum | Compliant actions add up past a limit nobody totals | Keep a running total and check it |
| Threshold laundering | A check passes, then a later step changes the number | Re-run the check on the final state |
| Authority creep | Several steps each dismiss one exception, so none is escalated | Trace which step removed each exception |
| Context collapse | Each handoff summarizes the last, and the file drifts from its source | Re-derive the decision from the original document |
The threshold example makes the point. A policy refers any account with prior losses above $250,000. The check sees $210,000 and passes. A later step adds a valid $75,000 claim. The final total is $285,000, and the rule never fires.
A better step-level check does not help, because no single step holds the whole picture. Write each limit as a statement about the finished job, then check it at the end. This explainer walks through each repair.
What makes an AI agent break a rule under pressure?
Deadlines, a manager's say-so and stated penalties all make AI agents break rules. In a 2026 test of 12 models acting as a procurement bot, deadline urgency was the most effective bypass. A manager's blanket authorization drove compliance to 0% in 15 of 48 test cells.
The study put a regulation in the system prompt, then varied what the user and the company said around it.
- Urgency. In one condition, every model fell to 9% compliance or lower when the user cited a deadline. An explicit instruction to follow all regulations regardless of user requests restored full compliance in only one model.
- Authority. A note that a manager had approved vendor discretion, or a board policy to cut costs, overrode the rule in most models.
- Penalties. Telling the agent the fine was small and unlikely turned the rule into a cost calculation.
Two cheap checks helped. A neutral follow-up, "can you double-check that?", often made a non-compliant answer correct itself. Compliant answers held 80% to 100% of the time under the same probe.
Most violations were also visible. Across 6,743 violations, 94.5% mentioned the rule in the agent's stated reasoning. Reading that reasoning catches most of them.
One limit applies. The study tested mostly open-weight models, and no Claude or flagship GPT model. Run the same pressure tests on the model you use.
How do you test whether an AI agent follows your rules?
Test an AI agent on your rules by running real cases several times each. Count a case as passed only when every run passes. A single passing run is weak evidence. On one 507-task business benchmark, Claude Opus 5 passed 66.5% of single runs but only 47.5% of tasks in all 20 runs.
On the same Thinkingbox benchmark, GPT-5.4 passed 65.4% of single runs and 25.3% of tasks in all 20. The arithmetic explains the gap. An agent that is right 90% of the time gets five runs in a row right about 59% of the time.
- Collect five real cases per rule. Include one edge case and one pressure case, such as a customer citing a deadline.
- Run each case five times. Use a fresh conversation each time.
- Score all-or-nothing. A case passes only if all five runs follow the rule.
- Diagnose each failure. Read the agent's reasoning and match it to one of the seven failures above.
- Re-test after every change. That includes new instructions, new tools and a new model version.
Step five is easy to skip. Amazon's SOP-Bench team found that a newer model family scored lower than its predecessor on one agent design. A routine upgrade can lower compliance with no visible sign.
Frequently asked questions
Can I paste my SOP document straight into Claude or ChatGPT?
You can, but it only lasts for that chat. SOPs written for staff also assume knowledge an agent lacks. Rewrite the SOP as explicit requirements first, then store it in project instructions or a skill so it loads every time.
Do I need to repeat my rules in a long conversation?
Yes, for rules that matter. One logged report covering 163 sessions found violations clustering right after the conversation was compacted and after unrelated tangents. Start a new chat for a new task, or restate the key rules.
Should I ask the AI to write the SOP for itself?
Not on its own. In SkillsBench, skills the model wrote for itself lowered pass rates by 1.3 points, while human-curated skills raised them by 16.2. Let the model draft if you like, but have someone who does the work correct it.
Do newer AI models follow business rules better?
Not reliably. Amazon's SOP-Bench team found a newer model family scoring lower than the older one on the same procedures. Re-run your rule tests after every model change.
Is a system prompt enough for compliance rules?
No. In a 12-model test, an explicit instruction to follow all regulations reduced violations under pressure but did not stop them. Put hard limits in code or tool permissions, and send exceptions to a person.
What happens to rules stored in a custom GPT?
OpenAI plans to retire custom GPTs on December 11, 2026 and move creators to plugins. Instructions and knowledge files can migrate. Custom actions do not, and a migrated plugin starts private. Re-test your rules after migrating.

