feat(gate): P17 keeps the engine tool surface out of skill content

The first-party agent-tool layer is Experimental and its surface moves between
builds. A recipe that names a toolset does not fail loudly on the next build --
the tool is simply absent, the search finds nothing, and "nothing found" reads
as "no problem here". That is the exact confusion this bundle documents, so the
tokens are barred rather than discouraged.

Grounded in measurement against a live editor, not in reading:
- the aggregator plugin lists 21 dependencies; the server reported 53
  registered toolsets from 22 plugins, one dependency contributing none;
- the visible tool count flips between 3 meta-tools and every tool registered
  natively, on one project setting.

- gate.py: ENGINE_TOOL_TOKENS and check_engine_tool_surface, registered in
  CHECKS; rule floor raised to 17.
- test_gate.py: one poison naming a toolset, a meta-tool and a host:port.
- ADR-0004 records the decision, its confirmation criteria and what is
  deliberately deferred.
- README and CONTRIBUTING state the rule where a contributor meets it.

Verified: gate.py 0 violations over 17 rules; test_gate.py 17/17 redden on
their fixtures, baseline clean, surface coverage intact. The token list matches
nothing under skills/ today, checked before the rule landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-06 00:00:59 +07:00
parent cd041ec1bf
commit 433ee61131
5 changed files with 165 additions and 4 deletions
+4 -2
View File
@@ -63,8 +63,10 @@ claude plugin validate . --strict
The gate is not advisory. It enforces, among other things: no absolute paths, no source
citations, no donor project name outside a Provenance section, no Cyrillic, no vendor or
harness name anywhere under `skills/`, and the six required fields on every failure-mode
entry.
harness name anywhere under `skills/`, no engine toolset, meta-tool or endpoint name
either (P17, see
[ADR-0004](docs/architecture/decisions/0004-engine-tool-surface-out-of-skill-content.md)),
and the six required fields on every failure-mode entry.
If you add a rule to the gate, **add a poisoned fixture that makes it fire.** The
self-test fails if any declared rule has no fixture. This is deliberate: the validator
+1
View File
@@ -22,6 +22,7 @@
- [ADR-0001: Discovery over PID file](docs/architecture/decisions/0001-discovery-over-pid-file.md) — поиск запущенного редактора через UDP-broadcast, а не через файл-PID.
- [ADR-0002: Harness-neutral skill bundle](docs/architecture/decisions/0002-harness-neutral-skill-bundle.md) — репозиторий как маркетплейс, `catalog.json` как источник истины, манифест вендора как адаптер. Нейтральность проверяется гейтом, а не обещается.
- [ADR-0003: Split licensing](docs/architecture/decisions/0003-split-licensing-prose-and-code.md) — проза бандла под CC BY-ND 4.0, код и метаданные под Apache-2.0; правообладатель назван, атрибуция в титрах — просьба, а не условие.
- [ADR-0004: Engine tool surface out of skill content](docs/architecture/decisions/0004-engine-tool-surface-out-of-skill-content.md) — имена тулсетов, мета-тулов и эндпоинтов движка не попадают в `skills/`: слой экспериментальный, а его отказ молчалив. Проверяется правилом гейта P17.
## Лицензия
@@ -0,0 +1,97 @@
---
status: accepted
date: 2026-09-05
deciders: project owner
consulted: live measurement of the first-party agent-tool layer in UE 5.8
informed: future contributors
---
# The engine's own tool surface stays out of skill content
## Context and Problem Statement
The engine now ships an agent-tool layer of its own: an MCP server, a registry of
toolsets, and skill assets discovered from the project. The obvious move is to write
recipes against it — name the toolset, name the tool, and the reader runs it.
Two measurements taken against a live editor say why that is a trap.
**The surface is larger and differently shaped than its own manifest suggests.** The
aggregator plugin lists 21 dependencies. The running server reported **53 registered
toolsets** from 22 plugins, one aggregated dependency contributing none at all and two
non-aggregated plugins contributing two. Any number written into a recipe is a number
about one project's plugin set, not about the engine.
**The number of visible tools is a project setting, not a property of the engine.** With
tool search on — the default — `tools/list` returns exactly three meta-tools and every
real tool is reached indirectly. With it off, every toolset tool registers natively. A
recipe that says "call tool X" is correct or nonsense depending on one checkbox.
Both layers are marked Experimental by their author, with an explicit notice that APIs
and data formats may change at any time. That is the vendor telling us the surface is
not a stable address.
## Decision Drivers
- **This bundle's subject is what fails silently.** A recipe naming a tool that no longer
exists does not error: the tool is absent, nothing is found, and "nothing found" reads
as "no problem here" — the exact confusion the bundle exists to prevent.
- **ADR-0002 already forbids naming one harness.** An engine build is the same category
of dependency, and the argument for P13 is the argument here.
- **The first party asks for the same thing.** Its published skill-authoring norm lists
*Durable* and *Agnostic* among six properties; this rule is those two, enforced.
- **A promise that lives in prose rots.** P13 was made checkable the hour it was decided.
## Considered Options
1. **Write recipes against the native tools.** Cheapest to read, dates fastest.
2. **Name tools but mark them version-bound.** A warning next to a command is read as a
command; and the failure is silent, so the warning arrives too late.
3. **Keep the tool surface out of skill content, and enforce it.**
## Decision Outcome
Chosen: **option 3**, as gate rule **P17**.
Nothing under `skills/` may name a toolset, a meta-tool, an endpoint host:port, or the
Experimental class names of the engine's skill-asset API. Detect recipes may say *what*
to look for; they may not say *which first-party tool to look with*. Everything outside
`skills/` — the catalog, adapters, the gate, the research archive — is unaffected, and
that is where any tool-specific knowledge belongs.
### Consequences
- Good: a recipe written today still means something on the next engine build, or fails
loudly by being unrunnable rather than quietly by finding nothing.
- Good: the bundle's neutrality claim now covers both axes it can drift on — the harness
that runs the agent, and the engine that hosts the tools.
- Bad: recipes stay one step further from executable. The reader translates "find fields
with no writer" into whatever their build offers. This is the accepted price.
- Bad: the rule cannot distinguish a tool name from an ordinary word that ends in the
same token. Checked against the whole bundle before landing: zero matches, so the cost
today is zero and any future collision is visible at the moment it is introduced.
- Neutral: knowledge about the native layer is not lost, only relocated. It lives in the
research archive, where it can name builds, ports and classes freely.
## Confirmation
Checked at the time of writing, and required for every later change:
- `python _gate/gate.py .` reports zero violations. *(passing, 17 rules)*
- `python _gate/test_gate.py` reports every rule reddening on its own poisoned fixture,
with a clean baseline and full surface coverage. *(passing, 17/17)*
- The P17 fixture names a toolset, a meta-tool and a `host:port` in one line, and the
rule fires on it. *(verified)*
- The token list produces no match anywhere under `skills/` today. *(verified before the
rule landed; a match would be a new violation, not a false positive to be excused)*
## Deferred
- **Whether the bundle ever ships an engine-side adapter** that maps a recipe to whatever
tools the current build offers. That is the natural home for the knowledge this rule
keeps out, and it should be written when a build is worth targeting — not before.
- **Whether skill assets inside the engine become a second delivery channel.** Measured as
workable: an asset created in a live session appeared in the project's skill list with
no editor restart. Also measured as fragile: an asset created with an empty description
is created successfully and then never listed, with no error on either call. A channel
whose failure mode is silent deletion needs a gate of its own before it carries anything.
+55 -2
View File
@@ -49,6 +49,22 @@ HARNESS_TOKENS = (
r'(?<![A-Za-z0-9_./-])\.claude/',
)
# Engine tool surface. The first-party agent-tool layer shipped with the editor
# is Experimental and says so: "APIs and data formats are subject to change at
# any time". A recipe that names a toolset, a meta-tool or an endpoint dates
# itself to one engine build and fails silently on the next -- the tool is
# simply absent, which reads as "nothing found here" rather than as an error.
# Measured surface drift is why this is a rule and not advice: one editor
# session exposed 53 toolsets from 21 aggregated plugins, and the visible tool
# count flips between 3 and several hundred on a single project setting.
ENGINE_TOOL_TOKENS = (
r'\b(?:list_toolsets|describe_toolset|call_tool)\b',
r'\b[A-Za-z][A-Za-z0-9_]*Toolsets?\b',
r'\bAgentSkill[A-Za-z]*\b',
r'\bModelContextProtocol[A-Za-z]*\b',
r'\b(?:127\.0\.0\.1|localhost)\s*:\s*\d+',
)
TEXT_SUFFIXES = {'.md', '.json', '.txt', '.yaml', '.yml'}
JUNK_SUFFIXES = {'.pyc', '.pyo', '.html', '.htm', '.zip', '.exe', '.dll',
'.uasset', '.umap', '.pdb', '.log'}
@@ -115,12 +131,14 @@ RULES = [
Rule('P15', 'failure-mode entry ids are not sequential in document order'),
Rule('P16', 'two skills share an entry-id prefix, so the archive resolver '
'silently collapses them'),
Rule('P17', 'skill content names an engine tool surface (toolset, meta-tool '
'or endpoint) that is Experimental and version-bound'),
]
# A rule that never fires on its fixture does not exist. The fixture corpus in
# _gate/test_gate.py asserts one poison per rule; this guards against a rule
# being silently dropped from the list itself.
assert len(RULES) >= 16, 'rule set shrank; a rule was lost'
assert len(RULES) >= 17, 'rule set shrank; a rule was lost'
assert len({r.id for r in RULES}) == len(RULES), 'duplicate rule id'
@@ -300,6 +318,40 @@ def check_harness_neutral(root: Path) -> list[Violation]:
return out
def check_engine_tool_surface(root: Path) -> list[Violation]:
"""P17: nothing under skills/ may name the engine's own agent-tool surface.
P13 keeps the content free of one *harness*; this keeps it free of one
*engine build*. The distinction matters because the failure looks different:
a harness token tells the reader to run something they do not have, while a
toolset name tells them to run something that used to exist. The second is
quieter -- the tool is missing, the search returns nothing, and "nothing
found" is indistinguishable from "no problem here", which is the exact
failure this bundle exists to document.
Detect recipes may still say what to look for. They may not say which
first-party tool to look with.
"""
out: list[Violation] = []
base = root / 'skills'
if not base.is_dir():
return out
patterns = [re.compile(p) for p in ENGINE_TOOL_TOKENS]
for path in sorted(base.rglob('*')):
if not path.is_file() or path.suffix.lower() not in TEXT_SUFFIXES:
continue
rel = path.relative_to(root).as_posix()
for n, line in enumerate(path.read_text(
encoding='utf-8', errors='replace').splitlines(), 1):
for pat in patterns:
m = pat.search(line)
if m:
out.append(Violation('P17', rel, n,
f'engine tool surface: {m.group(0)}'))
break
return out
def check_catalog(root: Path) -> list[Violation]:
"""P14: catalog.json describes every skill without vendor vocabulary.
@@ -555,7 +607,8 @@ def check_manifest(root: Path) -> list[Violation]:
CHECKS = (check_text_rules, check_top_level, check_junk, check_skills,
check_manifest, check_harness_neutral, check_catalog,
check_manifest, check_harness_neutral, check_engine_tool_surface,
check_catalog,
check_id_prefixes)
@@ -58,6 +58,14 @@ POISONS: dict[str, tuple] = {
# shipped entries vanish from the archive resolver without any count
# disagreeing loudly enough to notice.
'P16': ('copyskill', 'skills/sample-skill', 'sample-twin'),
# The engine's own agent-tool layer is Experimental and its surface moves
# between builds: one measured session exposed 53 toolsets, and the count
# of visible tools flips between 3 and hundreds on a single project
# setting. A recipe naming a toolset does not error on the next build, it
# finds nothing -- which reads as "no problem here".
'P17': ('append', FM,
'\nRun this through the EditorAppToolset via call_tool at '
'localhost:8000.\n'),
}