← All briefings
Briefing · 9 October 2026

Rules that say when they apply

Two preprints on rules that carry their own conditions and expiry, OpenAI's contracting evaluation with Ironclad, and a practitioner case for agreed business definitions.

  1. 01 Research preprint · arXiv · 3 October 2026

    SAGA: agent evolution via knowledge abstraction

    What changed
    SAGA turns an agent's experience into a hierarchy of examples, procedures and reusable principles. Each principle says when it applies and keeps a link to the evidence behind it. The authors report better performance in ScienceWorld and ALFWorld.
    Limits
    Both are simulated environments, not customer operations.
    Why it matters
    It's a concrete research model for turning staff experience into reusable context without losing the circumstances it came from.

    Idea to test

    Take ten customer-service exceptions and write a card for each: the example, the rule, when it applies, its exceptions and the evidence. Compare AI answers that use the cards with answers that use the original notes.

  2. 02 First-party research report · OpenAI · 6 October 2026

    OpenAI's contracting research with Ironclad

    What changed
    Domain experts from Ironclad helped define 11 contracting tasks, covering approval routing, document configuration and reusable clauses, each graded against 8 to 50 criteria. GPT-6 Astra averaged 55.0% against 41.6% for GPT-5.6 Sol.
    Limits
    A small internal evaluation. The two models ran at different reasoning settings, and the reported time savings are simulated estimates, not measured customer savings.
    Why it matters
    The evaluation checks whether business requirements survive a whole workflow, exceptions included. For context design, that suggests a written context brief is worth more when it comes with explicit acceptance criteria.

    Idea to test

    Build a 12-case evaluation for one business process: normal cases, cases right at a threshold, and exceptions. Score whether each required rule was followed.

  3. 03 Research preprint · arXiv · 4 October 2026

    From memory to guide: procedural coding memory

    What changed
    The proposed Composer adapts stored procedures to the current workspace and the current stage of the work, including when a piece of guidance should expire. It targets the risk of applying a useful procedure where local conditions don't fit. On 13 long-horizon coding tasks, the authors report a pass rate 7.2 points higher on the hardest tasks, with the main agent using 32% fewer tokens.
    Limits
    The work is about coding agents. Whether it transfers to other business workflows is unproven.
    Why it matters
    Reusable knowledge needs a defined scope. A campaign rule may depend on stock, channel, audience or season, and retrieving it doesn't establish that it still applies.

    Idea to test

    Add three fields to one campaign's context brief: applies when, overridden by, and expires when. Test it against a change in stock and a different channel.

  4. 04 Practitioner analysis · Express Computer · 8 October 2026

    The next enterprise AI advantage is better business context

    What changed
    A practical account of governed definitions, machine-readable rules and keeping context maintained. It describes how answers go wrong when what a term means in the business changes, even though the data pipelines still work.
    Limits
    The examples are the author's own, with no independent evaluation.
    Why it matters
    It gives a specific place to start a diagnosis: the questions whose answers depend on disputed definitions or unwritten conventions.

    Idea to test

    Ask three people on the team to define five terms, such as "active customer", "urgent order" and "available stock". Turn the disagreements into written conditions, each with a named owner.

---
title: Rules that say when they apply
date: 2026-10-09
summary: Two preprints on rules that carry their own conditions and expiry, OpenAI's
  contracting evaluation with Ironclad, and a practitioner case for agreed business
  definitions.
pick: 'Items 1 and 2 together, as one small demonstration: ten exception cards plus
  a 12-case evaluation. That could produce evidence of whether structured context
  actually improves decisions.'
sources:
- title: 'SAGA: agent evolution via knowledge abstraction'
  url: https://arxiv.org/abs/2610.06964
  source: arXiv
  kind: Research preprint
  published: 2026-10-03
  what_changed: SAGA turns an agent's experience into a hierarchy of examples, procedures
    and reusable principles. Each principle says when it applies and keeps a link
    to the evidence behind it. The authors report better performance in ScienceWorld
    and ALFWorld.
  limits: Both are simulated environments, not customer operations.
  why_it_matters: It's a concrete research model for turning staff experience into
    reusable context without losing the circumstances it came from.
  idea_to_test: 'Take ten customer-service exceptions and write a card for each: the
    example, the rule, when it applies, its exceptions and the evidence. Compare AI
    answers that use the cards with answers that use the original notes.'
- title: OpenAI's contracting research with Ironclad
  url: https://openai.com/index/advancing-computer-use-with-ironclad/
  source: OpenAI
  kind: First-party research report
  published: 2026-10-06
  what_changed: Domain experts from Ironclad helped define 11 contracting tasks, covering
    approval routing, document configuration and reusable clauses, each graded against
    8 to 50 criteria. GPT-6 Astra averaged 55.0% against 41.6% for GPT-5.6 Sol.
  limits: A small internal evaluation. The two models ran at different reasoning settings,
    and the reported time savings are simulated estimates, not measured customer savings.
  why_it_matters: The evaluation checks whether business requirements survive a whole
    workflow, exceptions included. For context design, that suggests a written context
    brief is worth more when it comes with explicit acceptance criteria.
  idea_to_test: 'Build a 12-case evaluation for one business process: normal cases,
    cases right at a threshold, and exceptions. Score whether each required rule was
    followed.'
- title: 'From memory to guide: procedural coding memory'
  url: https://arxiv.org/abs/2610.04868
  source: arXiv
  kind: Research preprint
  published: 2026-10-04
  what_changed: The proposed Composer adapts stored procedures to the current workspace
    and the current stage of the work, including when a piece of guidance should expire.
    It targets the risk of applying a useful procedure where local conditions don't
    fit. On 13 long-horizon coding tasks, the authors report a pass rate 7.2 points
    higher on the hardest tasks, with the main agent using 32% fewer tokens.
  limits: The work is about coding agents. Whether it transfers to other business
    workflows is unproven.
  why_it_matters: Reusable knowledge needs a defined scope. A campaign rule may depend
    on stock, channel, audience or season, and retrieving it doesn't establish that
    it still applies.
  idea_to_test: 'Add three fields to one campaign''s context brief: applies when,
    overridden by, and expires when. Test it against a change in stock and a different
    channel.'
- title: The next enterprise AI advantage is better business context
  source: Express Computer
  kind: Practitioner analysis
  published: 2026-10-08
  what_changed: A practical account of governed definitions, machine-readable rules
    and keeping context maintained. It describes how answers go wrong when what a
    term means in the business changes, even though the data pipelines still work.
  limits: The examples are the author's own, with no independent evaluation.
  why_it_matters: 'It gives a specific place to start a diagnosis: the questions whose
    answers depend on disputed definitions or unwritten conventions.'
  idea_to_test: Ask three people on the team to define five terms, such as "active
    customer", "urgent order" and "available stock". Turn the disagreements into written
    conditions, each with a named owner.
---
What I'd test first

Items 1 and 2 together, as one small demonstration: ten exception cards plus a 12-case evaluation. That could produce evidence of whether structured context actually improves decisions.


If you're wondering which of these matters for your own business, the audit is the quickest way to find out what your AI is missing: one workflow, one 90-minute session, and a written brief on what to fix first.

Book the audit Book a free 20-minute call

Not sure an audit is what you need? Bring one AI answer your team had to fix to the call, and I'll tell you whether missing context is the problem.