Rules that say when they apply
Two preprints on rules that carry their own conditions and expiry, OpenAI's contracting evaluation with Ironclad, and a practitioner case for agreed business definitions.
-
SAGA: agent evolution via knowledge abstraction
- What changed
- SAGA turns an agent's experience into a hierarchy of examples, procedures and reusable principles. Each principle says when it applies and keeps a link to the evidence behind it. The authors report better performance in ScienceWorld and ALFWorld.
- Limits
- Both are simulated environments, not customer operations.
- Why it matters
- It's a concrete research model for turning staff experience into reusable context without losing the circumstances it came from.
Idea to test
Take ten customer-service exceptions and write a card for each: the example, the rule, when it applies, its exceptions and the evidence. Compare AI answers that use the cards with answers that use the original notes.
-
OpenAI's contracting research with Ironclad
- What changed
- Domain experts from Ironclad helped define 11 contracting tasks, covering approval routing, document configuration and reusable clauses, each graded against 8 to 50 criteria. GPT-6 Astra averaged 55.0% against 41.6% for GPT-5.6 Sol.
- Limits
- A small internal evaluation. The two models ran at different reasoning settings, and the reported time savings are simulated estimates, not measured customer savings.
- Why it matters
- The evaluation checks whether business requirements survive a whole workflow, exceptions included. For context design, that suggests a written context brief is worth more when it comes with explicit acceptance criteria.
Idea to test
Build a 12-case evaluation for one business process: normal cases, cases right at a threshold, and exceptions. Score whether each required rule was followed.
-
From memory to guide: procedural coding memory
- What changed
- The proposed Composer adapts stored procedures to the current workspace and the current stage of the work, including when a piece of guidance should expire. It targets the risk of applying a useful procedure where local conditions don't fit. On 13 long-horizon coding tasks, the authors report a pass rate 7.2 points higher on the hardest tasks, with the main agent using 32% fewer tokens.
- Limits
- The work is about coding agents. Whether it transfers to other business workflows is unproven.
- Why it matters
- Reusable knowledge needs a defined scope. A campaign rule may depend on stock, channel, audience or season, and retrieving it doesn't establish that it still applies.
Idea to test
Add three fields to one campaign's context brief: applies when, overridden by, and expires when. Test it against a change in stock and a different channel.
-
The next enterprise AI advantage is better business context
- What changed
- A practical account of governed definitions, machine-readable rules and keeping context maintained. It describes how answers go wrong when what a term means in the business changes, even though the data pipelines still work.
- Limits
- The examples are the author's own, with no independent evaluation.
- Why it matters
- It gives a specific place to start a diagnosis: the questions whose answers depend on disputed definitions or unwritten conventions.
Idea to test
Ask three people on the team to define five terms, such as "active customer", "urgent order" and "available stock". Turn the disagreements into written conditions, each with a named owner.
---
title: Rules that say when they apply
date: 2026-10-09
summary: Two preprints on rules that carry their own conditions and expiry, OpenAI's
contracting evaluation with Ironclad, and a practitioner case for agreed business
definitions.
pick: 'Items 1 and 2 together, as one small demonstration: ten exception cards plus
a 12-case evaluation. That could produce evidence of whether structured context
actually improves decisions.'
sources:
- title: 'SAGA: agent evolution via knowledge abstraction'
url: https://arxiv.org/abs/2610.06964
source: arXiv
kind: Research preprint
published: 2026-10-03
what_changed: SAGA turns an agent's experience into a hierarchy of examples, procedures
and reusable principles. Each principle says when it applies and keeps a link
to the evidence behind it. The authors report better performance in ScienceWorld
and ALFWorld.
limits: Both are simulated environments, not customer operations.
why_it_matters: It's a concrete research model for turning staff experience into
reusable context without losing the circumstances it came from.
idea_to_test: 'Take ten customer-service exceptions and write a card for each: the
example, the rule, when it applies, its exceptions and the evidence. Compare AI
answers that use the cards with answers that use the original notes.'
- title: OpenAI's contracting research with Ironclad
url: https://openai.com/index/advancing-computer-use-with-ironclad/
source: OpenAI
kind: First-party research report
published: 2026-10-06
what_changed: Domain experts from Ironclad helped define 11 contracting tasks, covering
approval routing, document configuration and reusable clauses, each graded against
8 to 50 criteria. GPT-6 Astra averaged 55.0% against 41.6% for GPT-5.6 Sol.
limits: A small internal evaluation. The two models ran at different reasoning settings,
and the reported time savings are simulated estimates, not measured customer savings.
why_it_matters: The evaluation checks whether business requirements survive a whole
workflow, exceptions included. For context design, that suggests a written context
brief is worth more when it comes with explicit acceptance criteria.
idea_to_test: 'Build a 12-case evaluation for one business process: normal cases,
cases right at a threshold, and exceptions. Score whether each required rule was
followed.'
- title: 'From memory to guide: procedural coding memory'
url: https://arxiv.org/abs/2610.04868
source: arXiv
kind: Research preprint
published: 2026-10-04
what_changed: The proposed Composer adapts stored procedures to the current workspace
and the current stage of the work, including when a piece of guidance should expire.
It targets the risk of applying a useful procedure where local conditions don't
fit. On 13 long-horizon coding tasks, the authors report a pass rate 7.2 points
higher on the hardest tasks, with the main agent using 32% fewer tokens.
limits: The work is about coding agents. Whether it transfers to other business
workflows is unproven.
why_it_matters: Reusable knowledge needs a defined scope. A campaign rule may depend
on stock, channel, audience or season, and retrieving it doesn't establish that
it still applies.
idea_to_test: 'Add three fields to one campaign''s context brief: applies when,
overridden by, and expires when. Test it against a change in stock and a different
channel.'
- title: The next enterprise AI advantage is better business context
source: Express Computer
kind: Practitioner analysis
published: 2026-10-08
what_changed: A practical account of governed definitions, machine-readable rules
and keeping context maintained. It describes how answers go wrong when what a
term means in the business changes, even though the data pipelines still work.
limits: The examples are the author's own, with no independent evaluation.
why_it_matters: 'It gives a specific place to start a diagnosis: the questions whose
answers depend on disputed definitions or unwritten conventions.'
idea_to_test: Ask three people on the team to define five terms, such as "active
customer", "urgent order" and "available stock". Turn the disagreements into written
conditions, each with a named owner.
---
Items 1 and 2 together, as one small demonstration: ten exception cards plus a 12-case evaluation. That could produce evidence of whether structured context actually improves decisions.
If you're wondering which of these matters for your own business, the audit is the quickest way to find out what your AI is missing: one workflow, one 90-minute session, and a written brief on what to fix first.
Book the audit Book a free 20-minute call
Not sure an audit is what you need? Bring one AI answer your team had to fix to the call, and I'll tell you whether missing context is the problem.