InoGen

Lessons that stick.

Our controlled benchmarks · Updated 10 September 2026

A correction you give an AI agent should help the next person’s agent too. Making that happen takes more than an instruction file. Teams need a way to submit corrections, record them, update the guidance and make it available to agents on other machines.

OMS provides those mechanics. It also uses AI to turn corrections into useful shared guidance.

Our benchmarks tested both parts. In a handoff experiment, we corrected one agent and let OMS publish guidance for another agent in a new session. The second agent followed the rule 76% of the time.

We then compared how different approaches maintain guidance. The main alternative also uses AI: a model chooses an instruction file, reads it in full and returns an updated version. This is a sensible way to maintain ordinary files. OMS instead works with individual, connected rules, finds relevant existing guidance and reconciles changes before publishing instructions.

After 640 corrections covering 160 rules, we gave agents the resulting documents and tested them on the same tasks. Agents using OMS guidance followed the tested rule first time in 59% of tasks, compared with 40% for the file-editing approach. They needed 24% fewer corrective messages in the automated test.

These results support value beyond delivering guidance: the documents OMS produced helped agents follow it more reliably. The larger test gave OMS the correct document hint while the alternative selected its own, so the gain includes differences in locating guidance as well as maintaining it. Smaller tests covering 33 rules found no clear advantage.

Same tasks. Different guidance.

We made 640 corrections across 160 rules, then gave agents the finished instructions.

View results and comparisons
Followed the rule on the first try

Both task groups combined · Higher is better

OMS59%

189 of 319 tasks

AI-edited files40%

128 of 320 tasks

Less correcting in the test

24%

fewer reminders

About 55 reminders per 100 tasks with OMS, down from 73 with AI-edited files.

Sent by the test, not people. Staff time was not measured.

Our controlled test: 160 rules, 640 corrections, 40 instruction files. 10 September 2026.

These scores measure following the tested rule, not the whole task.

Both approaches use AI.

The difference is how they update the instructions.

OMS

  1. 1Find related rules
  2. 2Update the guidance
  3. 3Publish instructions

Finds related rules, checks what needs to change, then publishes updated instructions.

AI-edited files

  1. 1Choose a file
  2. 2Read the whole file
  3. 3Return the edited file

Chooses a file, reads it in full, then edits it. It is told to remove old advice, merge repeats and keep unrelated guidance.

Help the next agent learn it too.

In a separate test, we corrected one agent. OMS updated the guidance, and the test gave it to another agent in a fresh session.

View the handoff results
  1. Teach one agent

    Submit a correction
  2. OMS updates the guidance

    Process and publish
  3. Another agent reads it

    Start a fresh session
33 cases, each repeated three times. The test used OMS publication with scripted review handling. It did not test installation across employees’ devices or different agent products.

76%

Next agent followed the rule

The lesson helped in a new session. Receiving an instruction does not mean an agent will always follow it.

99/99

Checks found the rule’s marker

A marker is a word or phrase we looked for. In 6 checks, it was already there. Finding it does not prove the whole rule arrived correctly.

How did shared files compare?

Shared files worked well when kept up to date. OMS handled the tested handoff automatically. We tried 11 recording rates for files. These were test settings, not measured habits of teams.

Separate handoff test. Selected rates; rounded figures.

OMS

Marker found
100%
Rule followed
76%

Files: every correction recorded

Marker found
98%
Rule followed
81%

Files: 90% recording rate

Marker found
88%
Rule followed
70%

Files: 50% recording rate

Marker found
48%
Rule followed
41%

Files: none recorded

Marker found
0%
Rule followed
1%

Look inside the tests

In the larger test, agents using OMS guidance followed more rules on the first try and needed fewer reminders. Explore the comparisons, method and limits below.

View the evidence and method

Did the agent follow the rule on the first try?

After 640 corrections across 160 rules, agents used the finished instructions to answer the same tasks. Each first answer was checked for the tested rule.

“Familiar wording” uses the rule’s original terms. “Reworded tasks” asks agents to apply the rule in a differently worded task or situation.

First answers that followed the tested rule · Higher is better · Rounded percentages
How corrections were handledFamiliar wordingReworded tasksBoth combined
OMSFinds related rules, reconciles changes and publishes updated instructions.64%54%59%
AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction.44%36%40%

First answers that followed the tested rule · Higher is better · Rounded percentages

OMS

Finds related rules, reconciles changes and publishes updated instructions.

Familiar wording
64%
Reworded tasks
54%
Both combined
59%

AI-updated instruction files

An AI chooses a file, reads it and rewrites it with the correction.

Familiar wording
44%
Reworded tasks
36%
Both combined
40%

“Both combined” adds the results from both groups. These are the 59% and 40% scores in the chart above.

Agents using OMS guidance followed the rule more often in both groups. These scores measure the tested rule, not the quality of the whole answer.

Additional test comparisons

These checks help explain the result. They are not three equivalent products a team would buy.

Supporting task tests · First answers that followed the tested rule · Counts shown below percentages
How corrections were handledFamiliar wordingReworded tasks
Unedited correction historyKeeps every correction as written, including conflicting instructions.Tests whether keeping everything is enough.50.6%81 of 16046.3%74 of 160
OMS with existing-rule lookup disabledProcesses corrections without finding related guidance already stored in OMS.Tests the contribution of finding earlier rules. This is a test setting, not a separate product.42.1%67 of 15939.4%63 of 160
No supplied guidanceGives the agent the task without the team's instructions.Checks what the agent can do without being taught these rules.11.3%18 of 16011.9%19 of 160

Supporting task tests · First answers that followed the tested rule · Counts shown below percentages

Unedited correction history

Keeps every correction as written, including conflicting instructions.

Tests whether keeping everything is enough.

Familiar wording
50.6%81 of 160
Reworded tasks
46.3%74 of 160

OMS with existing-rule lookup disabled

Processes corrections without finding related guidance already stored in OMS.

Tests the contribution of finding earlier rules. This is a test setting, not a separate product.

Familiar wording
42.1%67 of 159
Reworded tasks
39.4%63 of 160

No supplied guidance

Gives the agent the task without the team's instructions.

Checks what the agent can do without being taught these rules.

Familiar wording
11.3%18 of 160
Reworded tasks
11.9%19 of 160
Separate test: questions about the guidance

Before the task-and-retry experiment, a separate run asked questions about the same 160 rules using each approach’s guidance. The same questions were later used in the task test; these are separate results, not extra first-try task results.

Separate question run · Answers that passed the rule check · Not included in the first-try chart
How corrections were handledFamiliar wordingReworded questions
OMSFinds related rules, reconciles changes and publishes updated instructions.63.8%102 of 16059.1%94 of 159
AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction.40.9%65 of 15933.8%54 of 160
Unedited correction historyKeeps every correction as written, including conflicting instructions.54.1%86 of 15945.0%72 of 160
OMS with existing-rule lookup disabledProcesses corrections without finding related guidance already stored in OMS.43.8%70 of 16039.4%63 of 160
No supplied guidanceGives the agent the task without the team's instructions.15.6%25 of 16010.6%17 of 160

Separate question run · Answers that passed the rule check · Not included in the first-try chart

OMS

Finds related rules, reconciles changes and publishes updated instructions.

Familiar wording
63.8%102 of 160
Reworded questions
59.1%94 of 159

AI-updated instruction files

An AI chooses a file, reads it and rewrites it with the correction.

Familiar wording
40.9%65 of 159
Reworded questions
33.8%54 of 160

Unedited correction history

Keeps every correction as written, including conflicting instructions.

Familiar wording
54.1%86 of 159
Reworded questions
45.0%72 of 160

OMS with existing-rule lookup disabled

Processes corrections without finding related guidance already stored in OMS.

Familiar wording
43.8%70 of 160
Reworded questions
39.4%63 of 160

No supplied guidance

Gives the agent the task without the team's instructions.

Familiar wording
15.6%25 of 160
Reworded questions
10.6%17 of 160
Exact counts and statistical checks
Main task comparison · Passing first answers out of measured first answers
How corrections were handledFamiliar wordingReworded tasksBoth combined
OMSFinds related rules, reconciles changes and publishes updated instructions.64.2%102 of 15954.4%87 of 16059.2%189 of 319
AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction.43.8%70 of 16036.3%58 of 16040.0%128 of 320

Main task comparison · Passing first answers out of measured first answers

OMS

Finds related rules, reconciles changes and publishes updated instructions.

Familiar wording
64.2%102 of 159
Reworded tasks
54.4%87 of 160
Both combined
59.2%189 of 319

AI-updated instruction files

An AI chooses a file, reads it and rewrites it with the correction.

Familiar wording
43.8%70 of 160
Reworded tasks
36.3%58 of 160
Both combined
40.0%128 of 320

OMS’s lead over AI-updated instruction files was statistically significant in both groups (paired McNemar’s exact test). Familiar wording: p = 0.000014. Reworded tasks: p = 0.000204.

Each group had 160 planned tasks per approach. Missing first answers are excluded, which is why some counts use 159.

How we ran and scored the tests
  1. 1Teach the rules

    Start with 40 instruction files. Add 640 corrections about 160 rules, including changes to earlier advice.

  2. 2Give each approach the same tasks

    Some tasks look like the original correction. Others put the rule in a new situation. Each task starts in a fresh session.

  3. 3Check the answer

    Check for set words and phrases. If the agent misses the rule, send a reminder and allow up to two more tries.

Five approaches, 320 planned task sessions each. The chart combines familiar wording and reworded tasks. Average reminders per measured task (rounded): 0.552 for OMS and 0.728 for AI-edited files.

Read the test method
What these results do and do not show
What made the difference?

OMS finds earlier rules and updates them when instructions change. Agents using full OMS guidance followed the rule first time in 59% of tasks, versus 41% with that lookup switched off. This supports the value of finding existing guidance.

OMS uses skill context supplied with a correction to direct the update, avoiding a separate search for the right instruction file. The benchmark supplied OMS with the correct skill context, while the file-editing agent selected its own file. The result therefore reflects both locating guidance and maintaining it.

First tries and retries differ

Agents using OMS guidance needed 24% fewer corrective messages per task. Its first-try advantage over AI-edited files was statistically significant for both familiar tasks and new situations.

Size and scope matter

OMS had the highest first-try scores of all five approaches in the 160-rule test, after 640 corrections including changes to earlier instructions.

Smaller tests with 33 rules found no clear advantage. This does not establish a rule-count threshold. Keeping every correction as written also worked well; OMS’s first-try lead over it was inconclusive on new situations.

Our tests, not a customer trial

We checked for set words and phrases, not overall answer quality. The recorded model label was ‘sonnet’, not a fixed model version. These controlled results are not independent validation or a promise for every setup.

See exactly what we asked.

Download the rules, questions, scoring checks and test method. The pack contains test inputs, not model answers or full result logs.

Download the questions and method

ZIP · 7 CSV files and a method guide · 109 KB

Make your own benefit estimate

Separate from the test results: choose how many repeat corrections you think your team could avoid. This is an estimate of staff time, not measured savings, net return or an OMS price.

Estimate the value for your teamIllustrative, not measured
250
2
5 min
GBP 450

You are estimating how many repeat corrections OMS would avoid. It is your assumption, not a measured OMS result. Set it to zero to see a no-benefit case.

Potential hours freed per year

2,000

Value of that staff time

GBP 120,000

Assumptions

  • 48 working weeks a year and a 7.5-hour working day.
  • Staff time is valued at the loaded day rate you set above.
  • Software, implementation and operating costs are excluded.
  • This estimates the value of time released. It is not OMS pricing, a cash saving or a net return.

Discuss the results against your own guidance

See what OMS would manage for your team and agree how to evaluate its value.

Book an OMS demo