Deployed Intelligence

Why AI agent skills fail silently — and how we gate them

The Forge Operating System treats agent skills the way serious teams treat model weights: versioned, scored, and only promoted when held-out work improves.

The edit that sounds better

Teams write a skill, watch one demo go well, then tighten the wording. The next week the agent skips a check it used to do. Nothing in the prose looks wrong. The failure shows up in the work.

Skill-edit research puts a number on that. About a quarter of edits transfer negatively: the skill is worse on tasks it was not tuned on. Asking a model to judge the skill, with no held-out tasks, lands near chance. The checks that hold up are specific: name the failure, say what to do, and blacklist the moves that cause real damage.

What we do instead

  1. Seed the skill as it is actually used.
  2. Score it on a pack of hard pass/fail tasks, split into train and held-out.
  3. Try a bounded edit. Do not rewrite the operating canon.
  4. Adopt only when the held-out score goes up and no passed held-out item flips to a fail.
  5. A person promotes the file. The log keeps the rejected edit so it is not retried as new.

On an internal OpenClaw ops skill, the seed scored 5 of 6 held-out checks. An edit that told the agent to skip a remote check fell to 4 of 6 and was rejected. An edit that named the missing checks rose to 6 of 6 and was staged. That is our pack, not a client result. Agent-heavy engagements get their own log.

What this is not

We did not invent the papers, and we do not sell the optimizer. Aligned with Microsoft Research SkillLens / SkillOpt; Forge owns the governance contract and the evidence. SkillOpt is not installed in client work, and transcripts are not sent out for optimization.

Questions

Why do AI agent skills fail after an edit?

A skill edit can sound clearer and still make the next task worse. Research on skill edits finds a large share of changes transfer negatively. Judging the skill by asking a model if it looks good is close to chance.

What is Forge Skill Governance?

It is Realee’s method for agent skills: version them, score them on a task pack, and promote a change only when held-out work improves. A person adopts the winner. We do not auto-rewrite the operating rules.

Do you ship Microsoft SkillOpt?

No. SkillLens and SkillOpt are research we align with. Forge owns the governance contract and the evidence log. SkillOpt is not a Realee product.

What does a client receive?

On agent, skill, or recurring AI-ops work, a skill evolution log: the starting score, a rejected edit, the edit we kept, and the delta. Brochure proposals do not carry this appendix.

Book a discovery call if the work includes agents that have to keep working after the demo.