OMS
- 1Find related rules
- 2Update the guidance
- 3Publish instructions
Finds related rules, checks what needs to change, then publishes updated instructions.
Measurement
Our controlled benchmarks · Updated 10 September 2026
A correction you give an AI agent should help the next person’s agent too. Making that happen takes more than an instruction file. Teams need a way to submit corrections, record them, update the guidance and make it available to agents on other machines.
OMS provides those mechanics. It also uses AI to turn corrections into useful shared guidance.
Our benchmarks tested both parts. In a handoff experiment, we corrected one agent and let OMS publish guidance for another agent in a new session. The second agent followed the rule 76% of the time.
We then compared how different approaches maintain guidance. The main alternative also uses AI: a model chooses an instruction file, reads it in full and returns an updated version. This is a sensible way to maintain ordinary files. OMS instead works with individual, connected rules, finds relevant existing guidance and reconciles changes before publishing instructions.
After 640 corrections covering 160 rules, we gave agents the resulting documents and tested them on the same tasks. Agents using OMS guidance followed the tested rule first time in 59% of tasks, compared with 40% for the file-editing approach. They needed 24% fewer corrective messages in the automated test.
These results support value beyond delivering guidance: the documents OMS produced helped agents follow it more reliably. The larger test gave OMS the correct document hint while the alternative selected its own, so the gain includes differences in locating guidance as well as maintaining it. Smaller tests covering 33 rules found no clear advantage.
01 / Following the rule
We made 640 corrections across 160 rules, then gave agents the finished instructions.
Both task groups combined · Higher is better
Less correcting in the test
24%
About 55 reminders per 100 tasks with OMS, down from 73 with AI-edited files.
Sent by the test, not people. Staff time was not measured.
Our controlled test: 160 rules, 640 corrections, 40 instruction files. 10 September 2026.
These scores measure following the tested rule, not the whole task.
The difference is how they update the instructions.
Finds related rules, checks what needs to change, then publishes updated instructions.
Chooses a file, reads it in full, then edits it. It is told to remove old advice, merge repeats and keep unrelated guidance.
02 / Sharing the lesson
In a separate test, we corrected one agent. OMS updated the guidance, and the test gave it to another agent in a fresh session.
Teach one agent
Submit a correctionOMS updates the guidance
Process and publishAnother agent reads it
Start a fresh session76%
The lesson helped in a new session. Receiving an instruction does not mean an agent will always follow it.
99/99
A marker is a word or phrase we looked for. In 6 checks, it was already there. Finding it does not prove the whole rule arrived correctly.
Shared files worked well when kept up to date. OMS handled the tested handoff automatically. We tried 11 recording rates for files. These were test settings, not measured habits of teams.
Separate handoff test. Selected rates; rounded figures.
03 / The evidence
In the larger test, agents using OMS guidance followed more rules on the first try and needed fewer reminders. Explore the comparisons, method and limits below.
After 640 corrections across 160 rules, agents used the finished instructions to answer the same tasks. Each first answer was checked for the tested rule.
“Familiar wording” uses the rule’s original terms. “Reworded tasks” asks agents to apply the rule in a differently worded task or situation.
| How corrections were handled | Familiar wording | Reworded tasks | Both combined |
|---|---|---|---|
| OMSFinds related rules, reconciles changes and publishes updated instructions. | 64% | 54% | 59% |
| AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction. | 44% | 36% | 40% |
First answers that followed the tested rule · Higher is better · Rounded percentages
Finds related rules, reconciles changes and publishes updated instructions.
An AI chooses a file, reads it and rewrites it with the correction.
“Both combined” adds the results from both groups. These are the 59% and 40% scores in the chart above.
Agents using OMS guidance followed the rule more often in both groups. These scores measure the tested rule, not the quality of the whole answer.
These checks help explain the result. They are not three equivalent products a team would buy.
| How corrections were handled | Familiar wording | Reworded tasks |
|---|---|---|
| Unedited correction historyKeeps every correction as written, including conflicting instructions.Tests whether keeping everything is enough. | 50.6%81 of 160 | 46.3%74 of 160 |
| OMS with existing-rule lookup disabledProcesses corrections without finding related guidance already stored in OMS.Tests the contribution of finding earlier rules. This is a test setting, not a separate product. | 42.1%67 of 159 | 39.4%63 of 160 |
| No supplied guidanceGives the agent the task without the team's instructions.Checks what the agent can do without being taught these rules. | 11.3%18 of 160 | 11.9%19 of 160 |
Supporting task tests · First answers that followed the tested rule · Counts shown below percentages
Keeps every correction as written, including conflicting instructions.
Tests whether keeping everything is enough.
Processes corrections without finding related guidance already stored in OMS.
Tests the contribution of finding earlier rules. This is a test setting, not a separate product.
Gives the agent the task without the team's instructions.
Checks what the agent can do without being taught these rules.
Before the task-and-retry experiment, a separate run asked questions about the same 160 rules using each approach’s guidance. The same questions were later used in the task test; these are separate results, not extra first-try task results.
| How corrections were handled | Familiar wording | Reworded questions |
|---|---|---|
| OMSFinds related rules, reconciles changes and publishes updated instructions. | 63.8%102 of 160 | 59.1%94 of 159 |
| AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction. | 40.9%65 of 159 | 33.8%54 of 160 |
| Unedited correction historyKeeps every correction as written, including conflicting instructions. | 54.1%86 of 159 | 45.0%72 of 160 |
| OMS with existing-rule lookup disabledProcesses corrections without finding related guidance already stored in OMS. | 43.8%70 of 160 | 39.4%63 of 160 |
| No supplied guidanceGives the agent the task without the team's instructions. | 15.6%25 of 160 | 10.6%17 of 160 |
Separate question run · Answers that passed the rule check · Not included in the first-try chart
Finds related rules, reconciles changes and publishes updated instructions.
An AI chooses a file, reads it and rewrites it with the correction.
Keeps every correction as written, including conflicting instructions.
Processes corrections without finding related guidance already stored in OMS.
Gives the agent the task without the team's instructions.
| How corrections were handled | Familiar wording | Reworded tasks | Both combined |
|---|---|---|---|
| OMSFinds related rules, reconciles changes and publishes updated instructions. | 64.2%102 of 159 | 54.4%87 of 160 | 59.2%189 of 319 |
| AI-updated instruction filesAn AI chooses a file, reads it and rewrites it with the correction. | 43.8%70 of 160 | 36.3%58 of 160 | 40.0%128 of 320 |
Main task comparison · Passing first answers out of measured first answers
Finds related rules, reconciles changes and publishes updated instructions.
An AI chooses a file, reads it and rewrites it with the correction.
OMS’s lead over AI-updated instruction files was statistically significant in both groups (paired McNemar’s exact test). Familiar wording: p = 0.000014. Reworded tasks: p = 0.000204.
Each group had 160 planned tasks per approach. Missing first answers are excluded, which is why some counts use 159.
Start with 40 instruction files. Add 640 corrections about 160 rules, including changes to earlier advice.
Some tasks look like the original correction. Others put the rule in a new situation. Each task starts in a fresh session.
Check for set words and phrases. If the agent misses the rule, send a reminder and allow up to two more tries.
Five approaches, 320 planned task sessions each. The chart combines familiar wording and reworded tasks. Average reminders per measured task (rounded): 0.552 for OMS and 0.728 for AI-edited files.
Read the test methodOMS finds earlier rules and updates them when instructions change. Agents using full OMS guidance followed the rule first time in 59% of tasks, versus 41% with that lookup switched off. This supports the value of finding existing guidance.
OMS uses skill context supplied with a correction to direct the update, avoiding a separate search for the right instruction file. The benchmark supplied OMS with the correct skill context, while the file-editing agent selected its own file. The result therefore reflects both locating guidance and maintaining it.
Agents using OMS guidance needed 24% fewer corrective messages per task. Its first-try advantage over AI-edited files was statistically significant for both familiar tasks and new situations.
OMS had the highest first-try scores of all five approaches in the 160-rule test, after 640 corrections including changes to earlier instructions.
Smaller tests with 33 rules found no clear advantage. This does not establish a rule-count threshold. Keeping every correction as written also worked well; OMS’s first-try lead over it was inconclusive on new situations.
We checked for set words and phrases, not overall answer quality. The recorded model label was ‘sonnet’, not a fixed model version. These controlled results are not independent validation or a promise for every setup.
Download the rules, questions, scoring checks and test method. The pack contains test inputs, not model answers or full result logs.
Download the questions and methodZIP · 7 CSV files and a method guide · 109 KB
Separate from the test results: choose how many repeat corrections you think your team could avoid. This is an estimate of staff time, not measured savings, net return or an OMS price.
You are estimating how many repeat corrections OMS would avoid. It is your assumption, not a measured OMS result. Set it to zero to see a no-benefit case.
Potential hours freed per year
2,000
Value of that staff time
GBP 120,000
Assumptions
See what OMS would manage for your team and agree how to evaluate its value.
Book an OMS demo