From ecd87ac96d063f7122478c2dbd78fe0c788889f2 Mon Sep 17 00:00:00 2001 From: MagentaDolphin Date: Sat, 5 Sep 2026 23:48:55 +0700 Subject: [PATCH] feat(skills): ship ue-design-skills bundle, licensing and delivery gate Phase 0 of the handoff plan, as a marketplace rather than a flat skills/ directory. Content moved out of the LyraResearch archive and depersonalised: addresses stay in the archive, recipes ship. - plugins/ue-design-skills: 17 skills, 232 failure-mode entries, each with the six required fields; catalog.json as the harness-neutral source of truth and .claude-plugin/ as one adapter over it. - _gate: 16 rules, one poisoned fixture per rule, plus surface coverage so a declared file cannot silently miss the line rules. - ADR-0002 (harness-neutral bundle behind a marketplace) and ADR-0003 (split licensing: CC BY-ND 4.0 prose, Apache-2.0 code and metadata). - LICENSE files at both levels, CONTRIBUTING.md, docs/licensing-options.md as the material the licence decision grew from. Verified: gate.py 0 violations; test_gate.py 16/16 rules redden on their fixtures with a clean baseline and 2 root files reaching the line rules. Co-Authored-By: Claude Opus 5 --- .claude-plugin/marketplace.json | 29 + CONTRIBUTING.md | 110 +++ LICENSE | 239 ++++++ .../0002-harness-neutral-skill-bundle.md | 128 +++ .../0003-split-licensing-prose-and-code.md | 147 ++++ docs/licensing-options.md | 236 ++++++ .../.claude-plugin/plugin.json | 16 + plugins/ue-design-skills/LICENSE | 554 +++++++++++++ plugins/ue-design-skills/README.md | 122 +++ .../fixtures/clean/.claude-plugin/plugin.json | 4 + .../_gate/fixtures/clean/README.md | 5 + .../_gate/fixtures/clean/catalog.json | 14 + .../clean/skills/sample-skill/SKILL.md | 25 + .../sample-skill/references/failure-modes.md | 37 + plugins/ue-design-skills/_gate/gate.py | 592 ++++++++++++++ plugins/ue-design-skills/_gate/test_gate.py | 191 +++++ plugins/ue-design-skills/catalog.json | 199 +++++ .../ue-architecture-guardrails/SKILL.md | 319 ++++++++ .../references/failure-modes.md | 379 +++++++++ .../references/patterns.md | 330 ++++++++ .../ue-asset-loading-and-memory/SKILL.md | 298 +++++++ .../references/failure-modes.md | 389 +++++++++ .../references/patterns.md | 360 +++++++++ .../skills/ue-cosmetics-and-teams/SKILL.md | 315 ++++++++ .../references/failure-modes.md | 406 ++++++++++ .../references/patterns.md | 288 +++++++ .../ue-data-driven-architecture/SKILL.md | 413 ++++++++++ .../references/failure-modes.md | 401 +++++++++ .../references/patterns.md | 268 ++++++ .../skills/ue-evidence-discipline/SKILL.md | 242 ++++++ .../references/failure-modes.md | 327 ++++++++ .../ue-game-settings-architecture/SKILL.md | 321 ++++++++ .../references/failure-modes.md | 348 ++++++++ .../references/patterns.md | 229 ++++++ .../skills/ue-gameplay-cues/SKILL.md | 376 +++++++++ .../references/failure-modes.md | 665 +++++++++++++++ .../ue-gameplay-cues/references/patterns.md | 200 +++++ .../skills/ue-gameplay-messaging/SKILL.md | 435 ++++++++++ .../references/failure-modes.md | 646 +++++++++++++++ .../references/patterns.md | 300 +++++++ .../ue-gameplay-tag-governance/SKILL.md | 386 +++++++++ .../references/failure-modes.md | 670 +++++++++++++++ .../references/patterns.md | 205 +++++ .../skills/ue-gas-architecture/SKILL.md | 452 +++++++++++ .../references/failure-modes.md | 502 ++++++++++++ .../references/patterns.md | 340 ++++++++ .../skills/ue-input-architecture/SKILL.md | 387 +++++++++ .../references/failure-modes.md | 405 ++++++++++ .../references/patterns.md | 225 ++++++ .../skills/ue-modular-gameplay/SKILL.md | 425 ++++++++++ .../references/failure-modes.md | 701 ++++++++++++++++ .../references/patterns.md | 302 +++++++ .../skills/ue-multiplayer-authority/SKILL.md | 390 +++++++++ .../references/failure-modes.md | 764 ++++++++++++++++++ .../references/patterns.md | 178 ++++ .../ue-reference-project-adoption/SKILL.md | 343 ++++++++ .../references/failure-modes.md | 739 +++++++++++++++++ .../references/patterns.md | 294 +++++++ .../SKILL.md | 265 ++++++ .../references/failure-modes.md | 384 +++++++++ .../references/patterns.md | 260 ++++++ .../SKILL.md | 291 +++++++ .../references/failure-modes.md | 388 +++++++++ .../references/patterns.md | 202 +++++ .../skills/ue-ui-architecture/SKILL.md | 380 +++++++++ .../references/failure-modes.md | 518 ++++++++++++ .../ue-ui-architecture/references/patterns.md | 331 ++++++++ 67 files changed, 21630 insertions(+) create mode 100644 .claude-plugin/marketplace.json create mode 100644 CONTRIBUTING.md create mode 100644 LICENSE create mode 100644 docs/architecture/decisions/0002-harness-neutral-skill-bundle.md create mode 100644 docs/architecture/decisions/0003-split-licensing-prose-and-code.md create mode 100644 docs/licensing-options.md create mode 100644 plugins/ue-design-skills/.claude-plugin/plugin.json create mode 100644 plugins/ue-design-skills/LICENSE create mode 100644 plugins/ue-design-skills/README.md create mode 100644 plugins/ue-design-skills/_gate/fixtures/clean/.claude-plugin/plugin.json create mode 100644 plugins/ue-design-skills/_gate/fixtures/clean/README.md create mode 100644 plugins/ue-design-skills/_gate/fixtures/clean/catalog.json create mode 100644 plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/SKILL.md create mode 100644 plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/references/failure-modes.md create mode 100644 plugins/ue-design-skills/_gate/gate.py create mode 100644 plugins/ue-design-skills/_gate/test_gate.py create mode 100644 plugins/ue-design-skills/catalog.json create mode 100644 plugins/ue-design-skills/skills/ue-architecture-guardrails/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-architecture-guardrails/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-architecture-guardrails/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-asset-loading-and-memory/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-cosmetics-and-teams/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-data-driven-architecture/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-data-driven-architecture/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-data-driven-architecture/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-evidence-discipline/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-evidence-discipline/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-game-settings-architecture/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-game-settings-architecture/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-game-settings-architecture/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-cues/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-cues/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-cues/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-messaging/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-messaging/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-messaging/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-tag-governance/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-gas-architecture/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-gas-architecture/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-gas-architecture/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-input-architecture/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-input-architecture/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-input-architecture/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-modular-gameplay/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-modular-gameplay/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-modular-gameplay/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-multiplayer-authority/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-multiplayer-authority/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-multiplayer-authority/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-reference-project-adoption/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-reference-project-adoption/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-reference-project-adoption/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-streaming-and-platform-budgets/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-streaming-and-platform-budgets/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-streaming-and-platform-budgets/references/patterns.md create mode 100644 plugins/ue-design-skills/skills/ue-ui-architecture/SKILL.md create mode 100644 plugins/ue-design-skills/skills/ue-ui-architecture/references/failure-modes.md create mode 100644 plugins/ue-design-skills/skills/ue-ui-architecture/references/patterns.md diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json new file mode 100644 index 0000000..e8795e3 --- /dev/null +++ b/.claude-plugin/marketplace.json @@ -0,0 +1,29 @@ +{ + "name": "ue-toolchain", + "owner": { + "name": "MagentaDolphin" + }, + "metadata": { + "description": "Unreal Engine tooling for LLM agents: design-review skills, an editor bridge and an MCP server.", + "version": "0.1.0" + }, + "plugins": [ + { + "name": "ue-design-skills", + "source": "./plugins/ue-design-skills", + "description": "Design-review skills for Unreal Engine 5.6 projects: norms plus reproducible detection recipes for silent architectural failures.", + "version": "0.1.0", + "author": { + "name": "MagentaDolphin" + }, + "license": "CC-BY-ND-4.0", + "keywords": [ + "unreal-engine", + "ue5", + "gameplay-ability-system", + "game-features", + "architecture-review" + ] + } + ] +} diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..d6d263c --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,110 @@ +# Contributing + +Thanks for reading this far. Two things are unusual about this repository, and both +change how a contribution is judged. Read them before opening a pull request. + +## 1. Prose and code are licensed differently + +| Path | License | +|---|---| +| `plugins/ue-design-skills/skills/**` and the bundle README | CC BY-ND 4.0 | +| `catalog.json`, `.claude-plugin/**`, `_gate/**`, Detect recipes | Apache-2.0 | +| Everything else in this repository | Apache-2.0 | + +See [ADR-0003](docs/architecture/decisions/0003-split-licensing-prose-and-code.md) for +why, and `plugins/ue-design-skills/LICENSE` for the exact grants. + +**NoDerivatives does not block contribution.** Publishing a fork in order to propose a +change back is explicitly permitted (additional permission 7). What it blocks is +publishing a modified version as a separate product. + +### Inbound licensing + +By opening a pull request you agree that your contribution may be distributed under the +license that already applies to the file you changed — CC BY-ND 4.0 for prose, +Apache-2.0 for code and metadata — and you confirm that you have the right to grant +this. + +If your employer owns your work, get clearance before submitting. This matters more than +usual here: a bundle under NoDerivatives has to be able to state a single copyright +holder, and an unclear contribution is worse than a missing one. + +## 2. A claim without a source you can point at is not a contribution + +The whole bundle rests on one invariant, and it applies to changes as much as to the +original material: + +> A skill that is confident and wrong is worse than a missing one. + +Read `skills/ue-evidence-discipline/SKILL.md` first. It is short, and it is the standard +the review will use. In particular: + +- **Every Detect recipe must have been run**, against a real tree, in both directions: + it finds the known instance, and it stays quiet where there is none. A recipe that has + not been run is a hypothesis. Say which tree you ran it against. +- **An empty search result proves the pattern, not the absence.** A negative claim must + state the searches that produced it. +- **Measured, derived and open are never mixed in one sentence.** +- **Counts state their scope.** "94 definitions in 31 files across Source and Plugins" is + a fact; "94 definitions" is a number waiting to be misused. + +Corrections to existing entries are the most welcome kind of change, especially ones +that show a recipe failing. If you found a case where a recipe returns a false negative, +that is a finding, not a nuisance — the bundle documents three of its own. + +## 3. Before you open the pull request + +```bash +cd plugins/ue-design-skills +python _gate/gate.py . # must print 0 violations +python _gate/test_gate.py # must print all rules reddening on their fixtures +claude plugin validate . --strict +``` + +The gate is not advisory. It enforces, among other things: no absolute paths, no source +citations, no donor project name outside a Provenance section, no Cyrillic, no vendor or +harness name anywhere under `skills/`, and the six required fields on every failure-mode +entry. + +If you add a rule to the gate, **add a poisoned fixture that makes it fire.** The +self-test fails if any declared rule has no fixture. This is deliberate: the validator +this gate replaced ran green for months while checking a condition that could never +become true. + +### Adding a failure-mode entry + +Every entry carries six fields, in this order: + +``` +**Mechanism.** +**Why it is silent.** +**Why the obvious check misses it.** +**Symptom.** +**Detect.** +**Guardrail.** +``` + +Three of them — silence, why the obvious check misses it, and a Detect recipe that has +been run — cannot be written by someone who has not done the work. That is the point. +Entries missing any field are rejected by rule P09 before a human reads them. + +Entry identifiers are sequential in document order (P15) and their prefix is unique +across skills (P16). If you insert an entry in the middle, renumber. + +## 4. What is likely to be declined + +- A recipe that has not been run, or whose output is not shown. +- An entry whose "why it is silent" restates the mechanism in other words. +- A skill that names a specific agent runtime, path variable or command syntax. The + bundle must work with any harness; this is enforced by rule P13. +- Reformatting passes that touch many files without changing a claim. +- A new skill proposed before a discussion issue. Open the issue first — the bundle is + deliberately small, and a sixteenth subsystem needs a reason. + +## 5. Reporting something you cannot fix + +Open an issue. A well-reported false negative in a Detect recipe is worth more than a +patch, because it usually means the recipe is wrong in a way the author could not see. + +Include the tree you ran against (engine version is enough — no source needed), the +command, the output you got and the output you expected. diff --git a/LICENSE b/LICENSE new file mode 100644 index 0000000..9d45592 --- /dev/null +++ b/LICENSE @@ -0,0 +1,239 @@ +ue-toolchain + +Copyright (c) 2026 MagentaDolphin + +=============================================================================== +SCOPE +=============================================================================== + +The software in this repository -- the editor plugin, the MCP server, the +installer, the delivery gate, the build and packaging scripts, and every +machine-readable manifest -- is licensed under the Apache License, Version 2.0, +reproduced in full below. + +One component carries a different license for its prose. + + plugins/ue-design-skills/ + +That bundle is documentation: design norms and detection recipes written by +hand. Its prose is under Creative Commons Attribution-NoDerivatives 4.0 +International, and its own LICENSE file states the exact boundary and the +additional permissions that widen it -- including the parts of that bundle +which are Apache-2.0 anyway: the catalog metadata, the delivery gate under +_gate/, and the shell commands inside "Detect" fields. + + plugins/ue-design-skills/LICENSE + +If you are vendoring code from this repository, Apache-2.0 covers it. If you +are republishing the skills prose, read that file first. + +Nothing here is asserted over what you build. No claim of any kind is made +over your game, your source, your documentation or your revenue. + +=============================================================================== +APACHE LICENSE, VERSION 2.0 +Full license text follows. +=============================================================================== + + + Apache License + Version 2.0, January 2004 + http://www.apache.org/licenses/ + + TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION + + 1. Definitions. + + "License" shall mean the terms and conditions for use, reproduction, + and distribution as defined by Sections 1 through 9 of this document. + + "Licensor" shall mean the copyright owner or entity authorized by + the copyright owner that is granting the License. + + "Legal Entity" shall mean the union of the acting entity and all + other entities that control, are controlled by, or are under common + control with that entity. For the purposes of this definition, + "control" means (i) the power, direct or indirect, to cause the + direction or management of such entity, whether by contract or + otherwise, or (ii) ownership of fifty percent (50%) or more of the + outstanding shares, or (iii) beneficial ownership of such entity. + + "You" (or "Your") shall mean an individual or Legal Entity + exercising permissions granted by this License. + + "Source" form shall mean the preferred form for making modifications, + including but not limited to software source code, documentation + source, and configuration files. + + "Object" form shall mean any form resulting from mechanical + transformation or translation of a Source form, including but + not limited to compiled object code, generated documentation, + and conversions to other media types. + + "Work" shall mean the work of authorship, whether in Source or + Object form, made available under the License, as indicated by a + copyright notice that is included in or attached to the work + (an example is provided in the Appendix below). + + "Derivative Works" shall mean any work, whether in Source or Object + form, that is based on (or derived from) the Work and for which the + editorial revisions, annotations, elaborations, or other modifications + represent, as a whole, an original work of authorship. For the purposes + of this License, Derivative Works shall not include works that remain + separable from, or merely link (or bind by name) to the interfaces of, + the Work and Derivative Works thereof. + + "Contribution" shall mean any work of authorship, including + the original version of the Work and any modifications or additions + to that Work or Derivative Works thereof, that is intentionally + submitted to Licensor for inclusion in the Work by the copyright owner + or by an individual or Legal Entity authorized to submit on behalf of + the copyright owner. For the purposes of this definition, "submitted" + means any form of electronic, verbal, or written communication sent + to the Licensor or its representatives, including but not limited to + communication on electronic mailing lists, source code control systems, + and issue tracking systems that are managed by, or on behalf of, the + Licensor for the purpose of discussing and improving the Work, but + excluding communication that is conspicuously marked or otherwise + designated in writing by the copyright owner as "Not a Contribution." + + "Contributor" shall mean Licensor and any individual or Legal Entity + on behalf of whom a Contribution has been received by Licensor and + subsequently incorporated within the Work. + + 2. Grant of Copyright License. Subject to the terms and conditions of + this License, each Contributor hereby grants to You a perpetual, + worldwide, non-exclusive, no-charge, royalty-free, irrevocable + copyright license to reproduce, prepare Derivative Works of, + publicly display, publicly perform, sublicense, and distribute the + Work and such Derivative Works in Source or Object form. + + 3. Grant of Patent License. Subject to the terms and conditions of + this License, each Contributor hereby grants to You a perpetual, + worldwide, non-exclusive, no-charge, royalty-free, irrevocable + (except as stated in this section) patent license to make, have made, + use, offer to sell, sell, import, and otherwise transfer the Work, + where such license applies only to those patent claims licensable + by such Contributor that are necessarily infringed by their + Contribution(s) alone or by combination of their Contribution(s) + with the Work to which such Contribution(s) was submitted. If You + institute patent litigation against any entity (including a + cross-claim or counterclaim in a lawsuit) alleging that the Work + or a Contribution incorporated within the Work constitutes direct + or contributory patent infringement, then any patent licenses + granted to You under this License for that Work shall terminate + as of the date such litigation is filed. + + 4. Redistribution. You may reproduce and distribute copies of the + Work or Derivative Works thereof in any medium, with or without + modifications, and in Source or Object form, provided that You + meet the following conditions: + + (a) You must give any other recipients of the Work or + Derivative Works a copy of this License; and + + (b) You must cause any modified files to carry prominent notices + stating that You changed the files; and + + (c) You must retain, in the Source form of any Derivative Works + that You distribute, all copyright, patent, trademark, and + attribution notices from the Source form of the Work, + excluding those notices that do not pertain to any part of + the Derivative Works; and + + (d) If the Work includes a "NOTICE" text file as part of its + distribution, then any Derivative Works that You distribute must + include a readable copy of the attribution notices contained + within such NOTICE file, excluding those notices that do not + pertain to any part of the Derivative Works, in at least one + of the following places: within a NOTICE text file distributed + as part of the Derivative Works; within the Source form or + documentation, if provided along with the Derivative Works; or, + within a display generated by the Derivative Works, if and + wherever such third-party notices normally appear. The contents + of the NOTICE file are for informational purposes only and + do not modify the License. You may add Your own attribution + notices within Derivative Works that You distribute, alongside + or as an addendum to the NOTICE text from the Work, provided + that such additional attribution notices cannot be construed + as modifying the License. + + You may add Your own copyright statement to Your modifications and + may provide additional or different license terms and conditions + for use, reproduction, or distribution of Your modifications, or + for any such Derivative Works as a whole, provided Your use, + reproduction, and distribution of the Work otherwise complies with + the conditions stated in this License. + + 5. Submission of Contributions. Unless You explicitly state otherwise, + any Contribution intentionally submitted for inclusion in the Work + by You to the Licensor shall be under the terms and conditions of + this License, without any additional terms or conditions. + Notwithstanding the above, nothing herein shall supersede or modify + the terms of any separate license agreement you may have executed + with Licensor regarding such Contributions. + + 6. Trademarks. This License does not grant permission to use the trade + names, trademarks, service marks, or product names of the Licensor, + except as required for reasonable and customary use in describing the + origin of the Work and reproducing the content of the NOTICE file. + + 7. Disclaimer of Warranty. Unless required by applicable law or + agreed to in writing, Licensor provides the Work (and each + Contributor provides its Contributions) on an "AS IS" BASIS, + WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or + implied, including, without limitation, any warranties or conditions + of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A + PARTICULAR PURPOSE. You are solely responsible for determining the + appropriateness of using or redistributing the Work and assume any + risks associated with Your exercise of permissions under this License. + + 8. Limitation of Liability. In no event and under no legal theory, + whether in tort (including negligence), contract, or otherwise, + unless required by applicable law (such as deliberate and grossly + negligent acts) or agreed to in writing, shall any Contributor be + liable to You for damages, including any direct, indirect, special, + incidental, or consequential damages of any character arising as a + result of this License or out of the use or inability to use the + Work (including but not limited to damages for loss of goodwill, + work stoppage, computer failure or malfunction, or any and all + other commercial damages or losses), even if such Contributor + has been advised of the possibility of such damages. + + 9. Accepting Warranty or Additional Liability. While redistributing + the Work or Derivative Works thereof, You may choose to offer, + and charge a fee for, acceptance of support, warranty, indemnity, + or other liability obligations and/or rights consistent with this + License. However, in accepting such obligations, You may act only + on Your own behalf and on Your sole responsibility, not on behalf + of any other Contributor, and only if You agree to indemnify, + defend, and hold each Contributor harmless for any liability + incurred by, or claims asserted against, such Contributor by reason + of your accepting any such warranty or additional liability. + + END OF TERMS AND CONDITIONS + + APPENDIX: How to apply the Apache License to your work. + + To apply the Apache License to your work, attach the following + boilerplate notice, with the fields enclosed by brackets "[]" + replaced with your own identifying information. (Don't include + the brackets!) The text should be enclosed in the appropriate + comment syntax for the file format. We also recommend that a + file or class name and description of purpose be included on the + same "printed page" as the copyright notice for easier + identification within third-party archives. + + Copyright [yyyy] [name of copyright owner] + + Licensed under the Apache License, Version 2.0 (the "License"); + you may not use this file except in compliance with the License. + You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + + Unless required by applicable law or agreed to in writing, software + distributed under the License is distributed on an "AS IS" BASIS, + WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + See the License for the specific language governing permissions and + limitations under the License. diff --git a/docs/architecture/decisions/0002-harness-neutral-skill-bundle.md b/docs/architecture/decisions/0002-harness-neutral-skill-bundle.md new file mode 100644 index 0000000..b3d9bb1 --- /dev/null +++ b/docs/architecture/decisions/0002-harness-neutral-skill-bundle.md @@ -0,0 +1,128 @@ +--- +status: accepted +date: 2026-09-02 +deciders: project owner +consulted: LyraResearch depersonalization pilot +informed: future contributors +--- + +# Harness-neutral skill bundle behind a marketplace + +## Context and Problem Statement + +The `skills/` component of `ue-toolchain` is not speculative work: it already exists as +18 skill bundles in the `LyraResearch` archive, grown from a line-by-line audit of a +first-party UE 5.6 sample. What is new is turning them into something a stranger can +install. + +Two forces pull against each other. + +**Distribution wants a vendor format.** Agent runners discover plugins through their own +manifest, in their own directory, with their own rules — component directories must sit +at the plugin root, and paths outside it are rejected. Writing to one such format is the +only way to get a one-command install today. + +**The product must outlive that vendor.** The owner's constraint is explicit: *"не +привязывайся к клоду... инструмент должен нормально работать с любым харнесом или +агентом."* A skill that says "run this vendor command" is not portable content, it is a +script for one runner. + +There is also a physical question. An agent runner copies the entire plugin root into its +cache on install. If the plugin root is the repository root, then the UE C++ plugin, the +MCP server and `examples/` are copied too — for a bundle whose actual payload is +Markdown. The precedent is measured: the VibeUE plugin copied into the reference project +weighs 161 MB. + +## Decision Drivers + +- **Content portability is a property of the files, not of the installer.** It has to be + checkable, or it will rot on the first convenient exception. +- **One vendor adapter must not become the interface.** Adding a second harness must not + require editing a single file under `skills/`. +- **The install surface should carry only what it delivers.** Source trees for the editor + plugin and the MCP server are not part of a documentation bundle. +- **Discoverability without a runtime.** A harness that has never heard of any plugin + format still needs to know what is in here and when to reach for each piece. + +## Considered Options + +1. **Whole repository is one plugin.** Vendor manifest at the repo root, `skills/` beside + `Engine/` and `mcp-server/`. +2. **Separate repository for the skills.** Independent versioning and release cadence. +3. **Repository is a marketplace; the bundle is a subdirectory.** + +## Decision Outcome + +Chosen: **option 3**, with a neutral catalog inside the bundle. + +```text +ue-toolchain/ +├── .claude-plugin/marketplace.json one adapter, points at the bundle +└── plugins/ue-design-skills/ the bundle, and the whole install surface + ├── catalog.json source of truth: what is here, when to use it + ├── .claude-plugin/plugin.json the same adapter, one level down + ├── skills//SKILL.md plain Markdown + YAML frontmatter + └── _gate/ rules, poisoned fixtures, self-test +``` + +Three commitments follow: + +**The catalog is the source of truth; the vendor manifest is an adapter.** `catalog.json` +lists every skill with `id`, `path`, `entry`, `description`, `use_when` and its +references. Any harness can read it, or ignore it and load the directory as context. A +second adapter is a new file beside the first, never an edit under `skills/`. + +**Skill content names no harness.** Not the vendor, not its path variables, not its +command syntax. The rule holds for everything under `skills/`; the adapter directories are +exempt because being vendor-specific is their entire job. + +**Both commitments are enforced, not documented.** They became gate rules P13 and P14 the +same hour the decision was made. A portability promise that lives only in an ADR is a +promise nobody checks. + +### Consequences + +- Good: the install surface holds only Markdown, JSON and the gate. C++ and examples stay + out of the plugin cache. +- Good: skills release on their own cadence, before the editor plugin (phase 1) and the + MCP server (phase 2) exist. +- Good: the neutrality claim is falsifiable — a single grep decides it. +- Bad: the layout is one level deeper than the component table in the handoff implies + (`plugins/ue-design-skills/`, not `skills/`). The README is updated to match. +- Bad: two manifests describe the same bundle. They will drift unless the catalog stays + authoritative and adapters are generated or reviewed against it. +- Neutral: nothing here prevents later extraction into its own repository. The bundle is + self-contained by construction — the gate rejects any link that escapes a skill + directory, so it moves with a `mv`. That was verified when it moved out of the archive. + +## Confirmation + +Checked at the time of writing, and required for every later change: + +- `claude plugin validate . --strict` passes at the marketplace root and at the bundle + root. *(passing)* +- `python _gate/gate.py .` reports zero violations on the bundle. *(passing, 14 rules)* +- `python _gate/test_gate.py` reports every rule reddening on its own poisoned fixture + with a clean baseline. A rule that never reddens is treated as absent and fails the + build. *(passing, 14/14)* +- No file under `skills/` matches the harness token list — vendor names, vendor path + variables, vendor command syntax. *(P13)* +- Every skill directory has a catalog entry whose `path` holds a `SKILL.md`, and the + catalog names no skill that is absent. *(P14, both directions)* +- Adding a second harness adapter changes no file under `skills/`. *(structural; re-check + when the second adapter is written)* + +## Deferred + +- **Which second adapter comes first.** Not guessed. It will be written when a real + harness needs it, and it will be the test of whether the catalog was actually enough. +- **License.** ~~Still open (handoff question 2). Deliberately *removed* from + `catalog.json` and the vendor manifest rather than filled in with the planned value: a + license asserted in shipped metadata before it is chosen is a claim with no author.~~ + **Resolved by [ADR-0003](0003-split-licensing-prose-and-code.md)**: prose CC BY-ND 4.0, + code and metadata Apache-2.0, holder named. The original wording is kept struck through + rather than deleted — the reason the field was left empty is part of how it was filled. +- **Whether the addressed audits are ever published.** For now the delivery carries + detection recipes and stable entry identifiers; the source addresses stay in the + research archive, where `_tools/resolve_id.py` resolves any identifier back to the file + and line it came from. diff --git a/docs/architecture/decisions/0003-split-licensing-prose-and-code.md b/docs/architecture/decisions/0003-split-licensing-prose-and-code.md new file mode 100644 index 0000000..1b736ec --- /dev/null +++ b/docs/architecture/decisions/0003-split-licensing-prose-and-code.md @@ -0,0 +1,147 @@ +--- +status: accepted +date: 2026-09-03 +deciders: project owner +consulted: licence text review (CC BY-ND 4.0 legal code, Apache-2.0, PolyForm family) +informed: future contributors, studios adopting the bundle +--- + +# Split licensing: CC BY-ND 4.0 for prose, Apache-2.0 for code + +## Context and Problem Statement + +Handoff question 2 has been open since 2026-09-01. ADR-0002 deliberately *removed* +`license` from the shipped metadata rather than filling in the planned value, on the +grounds that a licence asserted before it is chosen is a claim with no author. That left +the bundle shippable but not releasable. + +The owner's requirements, stated directly: + +> Free use by studios · my name credited, including in game credits · no forking my work +> into someone else's product without permission, commercial ones especially · no claim +> over the games and code built with the tool. + +Translated into licensing terms: a grant on **use** including commercial (R1); attribution +triggered by **use** (R2); a restriction on **Share Adapted Material** (R3); an explicit +**output** carve-out (R4). + +R1 and R3 are compatible: they govern different acts. Using a work and republishing a +modified version of it are separate permissions, and granting the first while withholding +the second is an ordinary construction. + +**R2 as stated is not achievable through any licence.** Attribution obligations attach to +distribution. A studio that reads the skills, applies them and ships a game has +distributed nothing, so no licence term can reach that act. This was measured against the +candidate texts rather than assumed, and the owner withdrew the requirement once the +mechanism was clear, replacing it with a request in the README. + +## Decision Drivers + +- **The bundle is prose, not software.** ~19 000 lines of hand-written norms and recipes. + Licences written for code do not carry the obligations prose needs. +- **Phase 1 puts C++ in this repository.** A restriction that is tolerable on a document + is expensive on code a studio compiles into its editor. Licences cannot be changed + retroactively for those who already took the work. +- **Adoption is the point.** A restriction that sends every studio's legal department into + review defeats R1 more effectively than a permissive licence defeats R3. +- **The bundle's own doctrine constrains the licence.** It instructs the reader to verify + every recipe against their own tree. Forbidding publication of the correction would make + the delivery contradict its content. + +## Considered Options + +1. **MIT or Apache-2.0 throughout.** Fails R3 explicitly: both permit commercial forks. +2. **CC BY 4.0 throughout.** Real attribution on redistribution, but derivatives allowed. +3. **CC BY-NC.** Blocks commercial *use*, which is R1's audience. The owner's objection is + to commercial forks, not commercial users; NC is aimed at the wrong act. +4. **PolyForm Internal Use.** Fits R1 and R3 closely, but is little known; an unfamiliar + licence more often produces a refusal than a question. +5. **A bespoke four-clause licence.** Delivers R1–R4 including credits, at the cost of + mandatory legal review at every adopting studio — the cost falls precisely on R1. +6. **Split: CC BY-ND 4.0 for prose, Apache-2.0 for code and metadata.** + +## Decision Outcome + +Chosen: **option 6**, with seven explicit additional permissions. + +Copyright holder: **MagentaDolphin**, a pseudonym under which the owner publishes. +Section 3(a)(1)(A)(i) of the licence provides for attribution by pseudonym, so this is a +supported form rather than a compromise, and substituting a legal name later refines the +same holder rather than changing licence. Publisher: **Kodlo.art**, recorded in +`catalog.json`. Collective pseudonyms were rejected for the copyright line: naming a +holder that stands for several people asserts joint ownership, and every later act — a +relicence, a translation grant, an enforcement — would then require their agreement. + +| Surface | Licence | +|---|---| +| `plugins/ue-design-skills/` prose: `SKILL.md`, `references/`, `README.md` | CC BY-ND 4.0 | +| `catalog.json`, `.claude-plugin/plugin.json` | Apache-2.0 | +| `_gate/` | Apache-2.0 | +| Shell commands inside `Detect` fields | Apache-2.0 | +| Editor plugin, MCP server, installer (phases 1–3) | Apache-2.0 | + +Additional permissions in `plugins/ue-design-skills/LICENSE`, each widening the grant: +metadata, gate and Detect commands are Apache-2.0; format adapters may be published; +translations may be published under stated conditions; private adaptation is restated as +already permitted; contributions via fork-and-pull are permitted. + +**Attribution in game credits is a request in the README, not a licence condition.** It is +written as a request so that nobody mistakes it for an obligation they have breached. + +### Consequences + +- Good: R1, R3 and R4 hold. Commercial use is free; republishing modified prose is not; + nothing is claimed over what users build. +- Good: the carve-outs keep ADR-0002 intact. A second harness adapter is explicitly + permitted, so the neutrality commitment is not quietly revoked by the licence. +- Good: the code path stays permissive before any code exists, so phase 1 does not force + a retroactive change. +- Bad: two licences in one repository require explanation. Mitigated by stating the + boundary in three places — both `LICENSE` files and both READMEs. +- Bad: CC BY-ND is not an OSI-approved open-source licence. Some organisations restrict + non-OSI content by policy. This is the accepted price of R3. +- Bad: R2 as originally stated is not delivered. Recorded here so that the gap is visible + rather than assumed closed. +- Neutral: the licence is irrevocable. Anyone who takes the work under these terms keeps + them. True of any licence, more expensive to get wrong with a restrictive one. + +## Confirmation + +Checked at the time of writing, and required for every later change: + +- `claude plugin validate . --strict` passes at the marketplace root and the bundle root. + *(passing; `publisher` and `license` were rejected inside marketplace `metadata` and + moved to `catalog.json`, which is the source of truth per ADR-0002)* +- `python _gate/gate.py .` reports zero violations with both LICENSE files in place. + *(passing)* +- `python _gate/test_gate.py` reports every rule reddening on its own fixture, plus every + file in `SCAN_ROOT_FILES` reaching the line-level rules. *(passing, 16 rules, 2 root + files)* +- The donor name inside the LICENSE `PROVENANCE` banner passes; the same name one line + outside it fires P03. *(verified in both directions)* +- Licence texts are the canonical ones, fetched from `apache.org` and + `creativecommons.org` rather than transcribed. + +### A gate defect found by this work + +`LICENSE` was declared in the gate's `SCAN_ROOT_FILES` from the start, and the line-level +rules filtered it out one line later because they matched on file extension and `LICENSE` +has none. A `LICENSE` carrying an absolute path, a source citation, the donor name, +Cyrillic and a wikilink passed the gate green — verified by poisoning it before the fix. + +This is the gate's own P08 shape: a check whose condition could not become true, on the +one file a lawyer reads. The fix treats declared root files as text whatever their suffix, +and the self-test now asserts surface coverage separately from rule coverage — proving not +just that each rule *can* fire, but that each file the gate *claims* to scan actually +reaches the rules. Reverting the fix reddens the suite, which was checked. + +## Deferred + +- **Whether the repository is public.** Handoff question 1, still open. The licence choice + does not decide it. +- **A legal review of the UE EULA question.** The bundle ships no third-party source, no + file addresses and no verbatim excerpts, and the gate enforces the first two. That + lowers the risk and is not a legal determination. If the release is public, this needs + someone qualified. +- **Whether to register the pseudonym to a legal name.** Not required to publish; a + refinement of the same holder when it happens. diff --git a/docs/licensing-options.md b/docs/licensing-options.md new file mode 100644 index 0000000..2d5afe9 --- /dev/null +++ b/docs/licensing-options.md @@ -0,0 +1,236 @@ +# Выбор лицензии: варианты и что они означают + +**Статус:** ~~решение не принято~~ — **решено 2026-09-03**, +[ADR-0003](architecture/decisions/0003-split-licensing-prose-and-code.md): проза +бандла CC BY-ND 4.0, код и метаданные Apache-2.0, правообладатель — MagentaDolphin. +Документ сохраняется как материал, из которого решение выросло: §§1–6 ниже описывают +состояние **до** выбора, включая отсутствовавший тогда `LICENSE`. Читать как разбор +вариантов, не как состояние репозитория. +**Открытый вопрос:** №2 из `handoff.md` — закрыт. +**Это не юридическое заключение.** Автор документа не имеет квалификации его дать. + +--- + +## 0. Состояние на сейчас + +| Что | Значение | +|---|---| +| Файл `LICENSE` в репозитории | отсутствует **[изм]** | +| `README.md:25` | «TBD (планируется MIT)» **[изм]** | +| `plugin.json` → `author.name` | `"ue-design-skills"` — заглушка, автором значится сам бандл **[изм]** | +| Поле `license` в манифестах | отсутствует; удалено намеренно (ADR-0002) **[изм]** | +| Правообладатель | **не определён** | + +Лицензируемый объект сегодня — бандл `plugins/ue-design-skills`: ~19 000 строк +авторской прозы, 232 записи с рецептами обнаружения, 335 блоков кода (в +подавляющем большинстве — команды поиска и публичный API движка), плюс скрипт +гейта `_gate/` на Python. + +Через фазу 1 к этому добавится **C++-модуль UE и MCP-сервер**. Это меняет расчёт +и разобрано в §6. + +--- + +## 1. Требования владельца, переведённые в лицензионные термины + +Сформулировано владельцем: + +> Свободное использование студиями · упоминание имени, в том числе в титрах · +> никакого форка в чужие продукты без разрешения, тем более коммерческие · +> никаких претензий на игры и код, созданные с помощью инструмента. + +В терминах, которыми оперируют лицензии: + +| № | Требование | Лицензионный термин | +|---|---|---| +| **R1** | студии, включая коммерческие, пользуются свободно | грант на **use**, без ограничения по коммерческому характеру | +| **R2** | имя указано, в том числе в титрах игры | **attribution**, причём с триггером на *использование*, а не на распространение | +| **R3** | нельзя форкать в свой продукт без разрешения | запрет на **Share Adapted Material** / **redistribution** | +| **R4** | претензий на созданное с помощью инструмента нет | явная **output clause** | + +--- + +## 2. Главное различие, вокруг которого всё вращается + +**Использование — не распространение.** Почти все лицензионные обязанности +привязаны ко второму. + +- Студия читает скилы, проектирует по ним систему, выпускает игру. **Ничего + вашего наружу не ушло.** Ни одна лицензия здесь ничего не требует. +- Студия берёт бандл, переписывает под себя и публикует как свой продукт. **Вот + здесь** лицензия работает, и здесь живёт R3. + +Отсюда следствие, которое важно принять до выбора строки в таблице: + +> **R1 и R3 совместимы** — они про разные действия. Свободно пользоваться и не +> иметь права перевыпускать — нормальная, часто встречающаяся конструкция. +> +> **R2 в лицензионной форме недостижим.** Ни одна из существующих лицензий не +> обязывает упоминать автора в титрах игры, при создании которой инструмент +> просто использовали. Обязанность атрибуции срабатывает при распространении +> копии, а копия здесь не распространяется. **[выв]** + +Это не придирка к формулировке. Это разница между «в лицензии написано» и «есть +механизм, который заставит». + +--- + +## 3. Требование R2: чем его закрывают на практике + +Раз лицензия не даёт, перечислю то, что даёт — по убыванию силы. + +| Механизм | Как работает | Чего стоит | +|---|---|---| +| **Договор** | студия подписывает условия использования, где атрибуция в титрах — обязательство | требует переговоров с каждой студией; убивает «скачал и работай» | +| **Двойное лицензирование** | бесплатная лицензия с ограничением + платная/свободная по запросу, в обмен на атрибуцию | рабочая схема, но это уже бизнес-процесс, а не файл в репозитории | +| **Пункт-рекомендация в README** | «если это помогло — упомяните в титрах» | ноль принуждения, ненулевой эффект; студии этим пользуются добровольно чаще, чем кажется | +| **Историческая advertising clause** (BSD-4-Clause) | требовала упоминания «во всех рекламных материалах» | пункт признан вредным и отозван Беркли; вызывает несовместимость и отказ юротделов | +| **Атрибуция при распространении** (CC BY / Apache NOTICE) | обязательна, но только когда копию передают дальше | закрывает форки и маркетплейсы, **не** закрывает титры игры | + +Практический вывод: атрибуция в титрах достижима как **норма сообщества плюс +условие для тех, кто просит расширенных прав**, а не как строка лицензии. **[выв]** + +--- + +## 4. Кандидаты против четырёх требований + +Легенда: ✅ даёт · ⚠️ даёт частично · ❌ не даёт или противоречит. + +| Лицензия | R1 свободное использование | R2 титры | R3 запрет форка | R4 выход свободен | Что это на самом деле означает | +|---|---|---|---|---|---| +| **MIT** | ✅ | ❌ | ❌ | ✅ | Форк в коммерческий продукт разрешён явно. Прямо противоречит R3 | +| **Apache-2.0** | ✅ | ⚠️ `NOTICE` при распространении | ❌ | ✅ | То же, что MIT, плюс патентный грант и пометки об изменениях. R3 не закрывает | +| **CC BY 4.0** | ✅ | ⚠️ при распространении | ❌ | ✅ | Сделана для текста, атрибуция реальна — но производные разрешены, включая коммерческие | +| **CC BY-SA 4.0** | ✅ | ⚠️ | ⚠️ форк разрешён, но обязан остаться открытым под SA | ✅ | Не запрещает форк, а заражает его. Для внутренних вики студий — стоп-сигнал юротдела | +| **CC BY-ND 4.0** | ✅ | ⚠️ при распространении | ✅ | ✅ | **Ближайшее совпадение по смыслу.** В версии 4.0 адаптации создавать *можно*, нельзя **Share** их. Студия правит под себя внутри — законно; публикует свою версию — нет | +| **CC BY-NC 4.0** | ❌ | ⚠️ | ⚠️ | ✅ | Запрещает коммерческое использование — отсекает ровно ту аудиторию, которая нужна по R1 | +| **PolyForm Noncommercial** | ❌ | ❌ | ✅ | ✅ | Та же ошибка, что у NC: блокирует коммерческие студии | +| **PolyForm Internal Use** | ✅ | ❌ | ✅ | ✅ | «Пользуйтесь для своих нужд, но не поставляйте другим». Точно ложится на R1+R3 | +| **PolyForm Shield** | ✅ | ❌ | ⚠️ запрещает только конкуренцию с автором | ✅ | Широкое принятие, блокируются лишь прямые конкуренты. Мягче, чем R3 | +| **CC0 / public domain** | ✅ | ❌ | ❌ | ✅ | Отказ от всего, включая имя | +| **Проприетарная / своя** | по вашему усмотрению | можно вписать | можно вписать | можно вписать | Единственный способ получить ровно R1–R4. Цена — в §5 | + +Замечания к таблице: + +- **CC для кода использовать не следует** — это позиция самой Creative Commons. + Если выбор уходит в CC-семейство, гейт и будущий C++-модуль всё равно требуют + отдельной строки. +- **PolyForm-лицензии не одобрены OSI** — все накладывают ограничение по области + использования. То же верно для всех CC, кроме CC0/BY/BY-SA. +- **CC BY-ND 4.0 — не то же самое, что ND 3.0.** В 3.0 создание адаптации было + за пределами гранта; в 4.0 адаптации явно разрешены, пока их не публикуют. + Общая страница CC об этом до сих пор пишет по-старому, авторитетен здесь текст + legal code, а не сводка. + +--- + +## 5. Три реалистичные комбинации + +### Вариант A — CC BY-ND 4.0 на тексты, Apache-2.0 на код + +Что получает пользователь: читает, применяет, правит под себя внутри студии; +опубликовать свою версию бандла не может; код гейта и будущего плагина берёт +свободно. + +- **R1** ✅ · **R3** ✅ · **R4** ✅ · **R2** — только атрибуция при + распространении, титры остаются просьбой в README. +- Готовые тексты, ничего сочинять не нужно. +- Минус: нужно объяснять границу «текст/код» в каждом обсуждении. Минус второй: + ND-лицензия закрывает и *дружественный* вклад — исправление в вашем тексте + нельзя опубликовать даже как pull request-форк. + +### Вариант B — PolyForm Internal Use на бандл, Apache-2.0 на код + +Формулировка «пользуйтесь для внутренних операций своей компании, но не +поставляйте другим» ближе к вашей интонации, чем язык авторского права CC. + +- **R1** ✅ · **R3** ✅ · **R4** ✅ · **R2** ❌ полностью. +- Минус: лицензия малоизвестна, юротдел студии её читает целиком, а незнакомая + лицензия чаще приводит к отказу, чем к вопросу. **[гип]** + +### Вариант C — своя лицензия из четырёх пунктов + +Ровно R1–R4, включая пункт про титры и пункт «на созданное с помощью инструмента +претензий нет». + +- **R1–R4** ✅ на бумаге. +- Минус, который перевешивает: **каждая студия обязана провести юридическую + проверку**, а нестандартная лицензия — типовая причина автоматического отказа + во внутренних политиках. Свободное распространение по R1 упирается ровно в это. + Плюс текст должен написать юрист, иначе неоднозначность будет толковаться не в + вашу пользу. + +--- + +## 6. Что изменится в фазе 1 и почему это важно решить сейчас + +Сегодня продукт — текст. После фазы 1 в том же репозитории появится **C++-модуль +UE и MCP-сервер**, то есть код, который студия компилирует внутрь своего +редактора. + +Отсюда два следствия **[выв]**: + +1. **Ограничительная лицензия на исполняемый код бьёт по принятию сильнее, чем + на текст.** Юротдел, спокойно пропускающий документ с ограничениями, обычно + гораздо строже к коду, попадающему в сборку. +2. **Менять лицензию задним числом нельзя** — те, кто уже взял по прежним + условиям, остаются на них. Поэтому раздельная схема «код свободный, тексты + ограничены» выгоднее выбрать сразу, чем переигрывать после первого релиза. + +--- + +## 7. Правообладатель + +Любой из вариантов начинается со строки `Copyright (c) 2026 <ПРАВООБЛАДАТЕЛЬ>`. + +| Форма | Когда уместна | +|---|---| +| Настоящее имя физлица | автор один; самая простая и внятная форма | +| Псевдоним / handle | если публичность имени нежелательна; авторство труднее доказывать при споре | +| Юрлицо | если работа сделана в рамках компании — тогда выбор может быть и не ваш | +| ` contributors` | принято при нескольких авторах; сейчас это описание будущего, а не факт | + +Имя правообладателя не выводится из содержимого репозитория и не может быть +подставлено по умолчанию. Требование R2 — про **это** имя, поэтому пока оно не +названо, R2 не реализуемо ни одним из механизмов §3. + +--- + +## 8. Слой, который выбором из таблицы не закрывается + +Содержание бандла выведено из разбора Epic Lyra, распространяемого по UE EULA. + +Измерено: в поставке нет кода Epic, нет адресов вида `Файл.ext:строка`, нет имён +с префиксом донора вне блоков `Provenance` (правило гейта P03, ноль нарушений). +Присутствуют публичные имена API движка — `UPROPERTY`, `TArray`, `FGameplayTag` и +подобные — и собственные наблюдения об архитектуре. **[изм]** + +Это существенно снижает риск, но **не является юридическим заключением** и не +заменяет его. Если релиз публичный — вопрос адресуется человеку с квалификацией. +Обезличивание, уже сделанное, принималось ради качества продукта; то, что оно +попутно уменьшает этот риск, — побочный эффект, а не доказательство. **[гип]** + +--- + +## 9. Что требуется от владельца, чтобы закрыть вопрос + +1. **Правообладатель** — имя в строку copyright (§7). +2. **Комбинация** — A, B, C или иная (§5). +3. **Форма требования о титрах** — просьба в README, условие двойного + лицензирования, или снятие требования (§3). + +После этого механическая часть — `LICENSE`, поле `license` в `catalog.json` и +`plugin.json`, раздел в `README.md`, правило гейта на присутствие файла — делается +за один заход. + +--- + +## Источники + +- [PolyForm Project — исходные тексты лицензий](https://github.com/polyformproject/polyform-licenses) +- [PolyForm Noncommercial License 1.0.0](https://polyformproject.org/licenses/noncommercial/1.0.0) +- [License Round-Up — разбор семейства PolyForm](https://writing.kemitchell.com/2021/06/20/License-Round-Up) +- [CC BY-ND 4.0 — deed](https://creativecommons.org/licenses/by-nd/4.0/deed.en) +- [CC 4.0 — обращение с адаптациями (wiki)](https://wiki.creativecommons.org/wiki/4.0/Treatment_of_adaptations) +- [CC — о лицензиях](https://creativecommons.org/share-your-work/cclicenses/) +- [CC FAQ](https://creativecommons.org/faq/) diff --git a/plugins/ue-design-skills/.claude-plugin/plugin.json b/plugins/ue-design-skills/.claude-plugin/plugin.json new file mode 100644 index 0000000..8e38c32 --- /dev/null +++ b/plugins/ue-design-skills/.claude-plugin/plugin.json @@ -0,0 +1,16 @@ +{ + "name": "ue-design-skills", + "version": "0.1.0", + "description": "Design-review skills for Unreal Engine 5.6 projects: modular gameplay, GAS, input, UI, messaging, tags, network authority, cosmetics, settings, asset loading, streaming budgets, allocation. Norms plus reproducible failure-mode detection recipes.", + "author": { + "name": "MagentaDolphin" + }, + "license": "CC-BY-ND-4.0", + "keywords": [ + "unreal-engine", + "ue5", + "gameplay-ability-system", + "game-features", + "architecture-review" + ] +} diff --git a/plugins/ue-design-skills/LICENSE b/plugins/ue-design-skills/LICENSE new file mode 100644 index 0000000..71fe89b --- /dev/null +++ b/plugins/ue-design-skills/LICENSE @@ -0,0 +1,554 @@ +ue-design-skills + +Copyright (c) 2026 MagentaDolphin + +=============================================================================== +SCOPE +=============================================================================== + +The prose in this bundle -- every SKILL.md, every file under references/, and +README.md -- is licensed under the Creative Commons Attribution-NoDerivatives +4.0 International License, reproduced in full below. + +Three things are NOT under that license. They are Apache-2.0, and the exact +carve-outs are stated under ADDITIONAL PERMISSIONS: + + * machine-readable metadata: catalog.json and .claude-plugin/plugin.json + * the delivery gate: everything under _gate/ + * the shell commands inside "Detect" fields + +A copy of the Apache License, Version 2.0 is at: +https://www.apache.org/licenses/LICENSE-2.0 + +=============================================================================== +WHAT THIS MEANS, IN PLAIN LANGUAGE +=============================================================================== + +This summary is not the license. Where the two disagree, the license text +below governs. + + You MAY, at no cost, for any purpose including commercial work: + + - read this material and apply it to your own projects; + - run the Detect recipes against your own code; + - redistribute this bundle unchanged, with attribution; + - modify it privately -- inside your studio, for your own pipeline, + for your own conventions -- without asking anyone. Section 2(a)(1)(B) + of the license below grants exactly this: you may produce and + reproduce adapted material, you may not Share it. + + You MAY NOT: + + - publish a modified version of this prose as your own product. + + NOT COVERED AT ALL: + + - anything you build using this material. No claim of any kind is made + over your game, your source code, your documentation, your internal + standards or your revenue. This bundle describes how to look at a + codebase; what you find and what you build is yours. + +If you want to do something the license does not allow, ask. The answer is +often yes. + +=============================================================================== +ADDITIONAL PERMISSIONS +=============================================================================== + +These permissions are granted by the copyright holder in addition to the +license below. They only ever widen what you may do; nothing here takes +anything away. + +1. METADATA IS APACHE-2.0 + + catalog.json and .claude-plugin/plugin.json are configuration, not prose. + Read, copy, edit and redistribute them freely, including modified, under + the Apache License 2.0 -- rewriting paths at install time, generating an + adapter for another agent runtime, or emitting a different manifest format + entirely. + +2. THE DELIVERY GATE IS APACHE-2.0 + + Everything under _gate/ is software: rules, poisoned fixtures and the + self-test. Use it, fork it, adapt it to your own delivery, under the + Apache License 2.0. + +3. DETECT RECIPES ARE APACHE-2.0 + + The shell commands inside "Detect" fields -- the search invocations + themselves, not the prose around them -- are Apache-2.0. You may publish + corrected or improved versions. + + This one is deliberate. This bundle teaches that a recipe which has not + been run is a hypothesis, and asks you to verify every recipe against your + own tree. Forbidding you to publish the correction would contradict the + material. + +4. FORMAT ADAPTERS ARE PERMITTED + + You may write and publish an adapter that converts this bundle into the + format another agent runtime expects -- a different manifest, a different + index, a different packaging -- and distribute it, provided the prose + itself travels unmodified and attribution is retained. + + Harness neutrality is the point of the catalog. A license that blocked the + second adapter would defeat it. + +5. TRANSLATIONS ARE PERMITTED + + A translation is an adaptation, which the license below would otherwise + forbid you to publish. You may translate this material into any language + and publish the result, provided that: + + - attribution is retained, with a link to the original; + - it is marked clearly as an unofficial translation; + - the meaning is preserved -- a translation is not a place to add, + remove or soften claims; + - it carries this same license and these same permissions. + +6. PRIVATE ADAPTATION IS ALREADY ALLOWED + + Stated explicitly because NoDerivatives is widely read as stricter than it + is. Under version 4.0, adapting the material is inside the grant; only + Sharing the adaptation is not. Editing these skills to match your studio's + conventions, adding your own entries, trimming what does not apply, and + circulating the result inside your organisation requires no permission. + +7. CONTRIBUTIONS + + Publishing a fork for the purpose of proposing a change back -- a pull + request, a patch, an issue with a diff -- is permitted. See CONTRIBUTING.md. + +=============================================================================== +ATTRIBUTION +=============================================================================== + +When redistributing, attribution in this form is sufficient: + + ue-design-skills by MagentaDolphin, licensed under CC BY-ND 4.0. + +Attribution is required when you redistribute this material. It is NOT +required when you merely use it -- if these skills helped you ship something, +a credit is welcome and appreciated, but nothing obliges you to give one. + +=============================================================================== +PROVENANCE +=============================================================================== + +The material was produced by reading Epic's Lyra Starter Game on Unreal +Engine 5.6 as source rather than running it. What ships is the classification +and the detection recipes, not the audit: no third-party source code, no file +addresses, and no verbatim excerpts are included. Names of public engine API +appear as any technical documentation names them. + +That description is a statement of how this bundle was made. It is not legal +advice and it is not a determination about your use of it. + +=============================================================================== +NO WARRANTY +=============================================================================== + +Provided as-is, without warranty of any kind. The delivery gate checks the +form of an entry -- that fields exist, that links resolve, that no address +leaked. It cannot check whether an entry is true. Verify every claim against +your own tree before acting on it; the material says so throughout, and means +it. + +=============================================================================== +CREATIVE COMMONS ATTRIBUTION-NODERIVATIVES 4.0 INTERNATIONAL +Full license text follows. +=============================================================================== + +Attribution-NoDerivatives 4.0 International + +======================================================================= + +Creative Commons Corporation ("Creative Commons") is not a law firm and +does not provide legal services or legal advice. Distribution of +Creative Commons public licenses does not create a lawyer-client or +other relationship. Creative Commons makes its licenses and related +information available on an "as-is" basis. Creative Commons gives no +warranties regarding its licenses, any material licensed under their +terms and conditions, or any related information. Creative Commons +disclaims all liability for damages resulting from their use to the +fullest extent possible. + +Using Creative Commons Public Licenses + +Creative Commons public licenses provide a standard set of terms and +conditions that creators and other rights holders may use to share +original works of authorship and other material subject to copyright +and certain other rights specified in the public license below. The +following considerations are for informational purposes only, are not +exhaustive, and do not form part of our licenses. + + Considerations for licensors: Our public licenses are + intended for use by those authorized to give the public + permission to use material in ways otherwise restricted by + copyright and certain other rights. Our licenses are + irrevocable. Licensors should read and understand the terms + and conditions of the license they choose before applying it. + Licensors should also secure all rights necessary before + applying our licenses so that the public can reuse the + material as expected. Licensors should clearly mark any + material not subject to the license. This includes other CC- + licensed material, or material used under an exception or + limitation to copyright. More considerations for licensors: + wiki.creativecommons.org/Considerations_for_licensors + + Considerations for the public: By using one of our public + licenses, a licensor grants the public permission to use the + licensed material under specified terms and conditions. If + the licensor's permission is not necessary for any reason--for + example, because of any applicable exception or limitation to + copyright--then that use is not regulated by the license. Our + licenses grant only permissions under copyright and certain + other rights that a licensor has authority to grant. Use of + the licensed material may still be restricted for other + reasons, including because others have copyright or other + rights in the material. A licensor may make special requests, + such as asking that all changes be marked or described. + Although not required by our licenses, you are encouraged to + respect those requests where reasonable. More considerations + for the public: + wiki.creativecommons.org/Considerations_for_licensees + + +======================================================================= + +Creative Commons Attribution-NoDerivatives 4.0 International Public +License + +By exercising the Licensed Rights (defined below), You accept and agree +to be bound by the terms and conditions of this Creative Commons +Attribution-NoDerivatives 4.0 International Public License ("Public +License"). To the extent this Public License may be interpreted as a +contract, You are granted the Licensed Rights in consideration of Your +acceptance of these terms and conditions, and the Licensor grants You +such rights in consideration of benefits the Licensor receives from +making the Licensed Material available under these terms and +conditions. + + +Section 1 -- Definitions. + + a. Adapted Material means material subject to Copyright and Similar + Rights that is derived from or based upon the Licensed Material + and in which the Licensed Material is translated, altered, + arranged, transformed, or otherwise modified in a manner requiring + permission under the Copyright and Similar Rights held by the + Licensor. For purposes of this Public License, where the Licensed + Material is a musical work, performance, or sound recording, + Adapted Material is always produced where the Licensed Material is + synched in timed relation with a moving image. + + b. Copyright and Similar Rights means copyright and/or similar rights + closely related to copyright including, without limitation, + performance, broadcast, sound recording, and Sui Generis Database + Rights, without regard to how the rights are labeled or + categorized. For purposes of this Public License, the rights + specified in Section 2(b)(1)-(2) are not Copyright and Similar + Rights. + + c. Effective Technological Measures means those measures that, in the + absence of proper authority, may not be circumvented under laws + fulfilling obligations under Article 11 of the WIPO Copyright + Treaty adopted on December 20, 1996, and/or similar international + agreements. + + d. Exceptions and Limitations means fair use, fair dealing, and/or + any other exception or limitation to Copyright and Similar Rights + that applies to Your use of the Licensed Material. + + e. Licensed Material means the artistic or literary work, database, + or other material to which the Licensor applied this Public + License. + + f. Licensed Rights means the rights granted to You subject to the + terms and conditions of this Public License, which are limited to + all Copyright and Similar Rights that apply to Your use of the + Licensed Material and that the Licensor has authority to license. + + g. Licensor means the individual(s) or entity(ies) granting rights + under this Public License. + + h. Share means to provide material to the public by any means or + process that requires permission under the Licensed Rights, such + as reproduction, public display, public performance, distribution, + dissemination, communication, or importation, and to make material + available to the public including in ways that members of the + public may access the material from a place and at a time + individually chosen by them. + + i. Sui Generis Database Rights means rights other than copyright + resulting from Directive 96/9/EC of the European Parliament and of + the Council of 11 March 1996 on the legal protection of databases, + as amended and/or succeeded, as well as other essentially + equivalent rights anywhere in the world. + + j. You means the individual or entity exercising the Licensed Rights + under this Public License. Your has a corresponding meaning. + + +Section 2 -- Scope. + + a. License grant. + + 1. Subject to the terms and conditions of this Public License, + the Licensor hereby grants You a worldwide, royalty-free, + non-sublicensable, non-exclusive, irrevocable license to + exercise the Licensed Rights in the Licensed Material to: + + a. reproduce and Share the Licensed Material, in whole or + in part; and + + b. produce and reproduce, but not Share, Adapted Material. + + 2. Exceptions and Limitations. For the avoidance of doubt, where + Exceptions and Limitations apply to Your use, this Public + License does not apply, and You do not need to comply with + its terms and conditions. + + 3. Term. The term of this Public License is specified in Section + 6(a). + + 4. Media and formats; technical modifications allowed. The + Licensor authorizes You to exercise the Licensed Rights in + all media and formats whether now known or hereafter created, + and to make technical modifications necessary to do so. The + Licensor waives and/or agrees not to assert any right or + authority to forbid You from making technical modifications + necessary to exercise the Licensed Rights, including + technical modifications necessary to circumvent Effective + Technological Measures. For purposes of this Public License, + simply making modifications authorized by this Section 2(a) + (4) never produces Adapted Material. + + 5. Downstream recipients. + + a. Offer from the Licensor -- Licensed Material. Every + recipient of the Licensed Material automatically + receives an offer from the Licensor to exercise the + Licensed Rights under the terms and conditions of this + Public License. + + b. No downstream restrictions. You may not offer or impose + any additional or different terms or conditions on, or + apply any Effective Technological Measures to, the + Licensed Material if doing so restricts exercise of the + Licensed Rights by any recipient of the Licensed + Material. + + 6. No endorsement. Nothing in this Public License constitutes or + may be construed as permission to assert or imply that You + are, or that Your use of the Licensed Material is, connected + with, or sponsored, endorsed, or granted official status by, + the Licensor or others designated to receive attribution as + provided in Section 3(a)(1)(A)(i). + + b. Other rights. + + 1. Moral rights, such as the right of integrity, are not + licensed under this Public License, nor are publicity, + privacy, and/or other similar personality rights; however, to + the extent possible, the Licensor waives and/or agrees not to + assert any such rights held by the Licensor to the limited + extent necessary to allow You to exercise the Licensed + Rights, but not otherwise. + + 2. Patent and trademark rights are not licensed under this + Public License. + + 3. To the extent possible, the Licensor waives any right to + collect royalties from You for the exercise of the Licensed + Rights, whether directly or through a collecting society + under any voluntary or waivable statutory or compulsory + licensing scheme. In all other cases the Licensor expressly + reserves any right to collect such royalties. + + +Section 3 -- License Conditions. + +Your exercise of the Licensed Rights is expressly made subject to the +following conditions. + + a. Attribution. + + 1. If You Share the Licensed Material, You must: + + a. retain the following if it is supplied by the Licensor + with the Licensed Material: + + i. identification of the creator(s) of the Licensed + Material and any others designated to receive + attribution, in any reasonable manner requested by + the Licensor (including by pseudonym if + designated); + + ii. a copyright notice; + + iii. a notice that refers to this Public License; + + iv. a notice that refers to the disclaimer of + warranties; + + v. a URI or hyperlink to the Licensed Material to the + extent reasonably practicable; + + b. indicate if You modified the Licensed Material and + retain an indication of any previous modifications; and + + c. indicate the Licensed Material is licensed under this + Public License, and include the text of, or the URI or + hyperlink to, this Public License. + + For the avoidance of doubt, You do not have permission under + this Public License to Share Adapted Material. + + 2. You may satisfy the conditions in Section 3(a)(1) in any + reasonable manner based on the medium, means, and context in + which You Share the Licensed Material. For example, it may be + reasonable to satisfy the conditions by providing a URI or + hyperlink to a resource that includes the required + information. + + 3. If requested by the Licensor, You must remove any of the + information required by Section 3(a)(1)(A) to the extent + reasonably practicable. + + +Section 4 -- Sui Generis Database Rights. + +Where the Licensed Rights include Sui Generis Database Rights that +apply to Your use of the Licensed Material: + + a. for the avoidance of doubt, Section 2(a)(1) grants You the right + to extract, reuse, reproduce, and Share all or a substantial + portion of the contents of the database, provided You do not Share + Adapted Material; + + b. if You include all or a substantial portion of the database + contents in a database in which You have Sui Generis Database + Rights, then the database in which You have Sui Generis Database + Rights (but not its individual contents) is Adapted Material; and + + c. You must comply with the conditions in Section 3(a) if You Share + all or a substantial portion of the contents of the database. + +For the avoidance of doubt, this Section 4 supplements and does not +replace Your obligations under this Public License where the Licensed +Rights include other Copyright and Similar Rights. + + +Section 5 -- Disclaimer of Warranties and Limitation of Liability. + + a. UNLESS OTHERWISE SEPARATELY UNDERTAKEN BY THE LICENSOR, TO THE + EXTENT POSSIBLE, THE LICENSOR OFFERS THE LICENSED MATERIAL AS-IS + AND AS-AVAILABLE, AND MAKES NO REPRESENTATIONS OR WARRANTIES OF + ANY KIND CONCERNING THE LICENSED MATERIAL, WHETHER EXPRESS, + IMPLIED, STATUTORY, OR OTHER. THIS INCLUDES, WITHOUT LIMITATION, + WARRANTIES OF TITLE, MERCHANTABILITY, FITNESS FOR A PARTICULAR + PURPOSE, NON-INFRINGEMENT, ABSENCE OF LATENT OR OTHER DEFECTS, + ACCURACY, OR THE PRESENCE OR ABSENCE OF ERRORS, WHETHER OR NOT + KNOWN OR DISCOVERABLE. WHERE DISCLAIMERS OF WARRANTIES ARE NOT + ALLOWED IN FULL OR IN PART, THIS DISCLAIMER MAY NOT APPLY TO YOU. + + b. TO THE EXTENT POSSIBLE, IN NO EVENT WILL THE LICENSOR BE LIABLE + TO YOU ON ANY LEGAL THEORY (INCLUDING, WITHOUT LIMITATION, + NEGLIGENCE) OR OTHERWISE FOR ANY DIRECT, SPECIAL, INDIRECT, + INCIDENTAL, CONSEQUENTIAL, PUNITIVE, EXEMPLARY, OR OTHER LOSSES, + COSTS, EXPENSES, OR DAMAGES ARISING OUT OF THIS PUBLIC LICENSE OR + USE OF THE LICENSED MATERIAL, EVEN IF THE LICENSOR HAS BEEN + ADVISED OF THE POSSIBILITY OF SUCH LOSSES, COSTS, EXPENSES, OR + DAMAGES. WHERE A LIMITATION OF LIABILITY IS NOT ALLOWED IN FULL OR + IN PART, THIS LIMITATION MAY NOT APPLY TO YOU. + + c. The disclaimer of warranties and limitation of liability provided + above shall be interpreted in a manner that, to the extent + possible, most closely approximates an absolute disclaimer and + waiver of all liability. + + +Section 6 -- Term and Termination. + + a. This Public License applies for the term of the Copyright and + Similar Rights licensed here. However, if You fail to comply with + this Public License, then Your rights under this Public License + terminate automatically. + + b. Where Your right to use the Licensed Material has terminated under + Section 6(a), it reinstates: + + 1. automatically as of the date the violation is cured, provided + it is cured within 30 days of Your discovery of the + violation; or + + 2. upon express reinstatement by the Licensor. + + For the avoidance of doubt, this Section 6(b) does not affect any + right the Licensor may have to seek remedies for Your violations + of this Public License. + + c. For the avoidance of doubt, the Licensor may also offer the + Licensed Material under separate terms or conditions or stop + distributing the Licensed Material at any time; however, doing so + will not terminate this Public License. + + d. Sections 1, 5, 6, 7, and 8 survive termination of this Public + License. + + +Section 7 -- Other Terms and Conditions. + + a. The Licensor shall not be bound by any additional or different + terms or conditions communicated by You unless expressly agreed. + + b. Any arrangements, understandings, or agreements regarding the + Licensed Material not stated herein are separate from and + independent of the terms and conditions of this Public License. + + +Section 8 -- Interpretation. + + a. For the avoidance of doubt, this Public License does not, and + shall not be interpreted to, reduce, limit, restrict, or impose + conditions on any use of the Licensed Material that could lawfully + be made without permission under this Public License. + + b. To the extent possible, if any provision of this Public License is + deemed unenforceable, it shall be automatically reformed to the + minimum extent necessary to make it enforceable. If the provision + cannot be reformed, it shall be severed from this Public License + without affecting the enforceability of the remaining terms and + conditions. + + c. No term or condition of this Public License will be waived and no + failure to comply consented to unless expressly agreed to by the + Licensor. + + d. Nothing in this Public License constitutes or may be interpreted + as a limitation upon, or waiver of, any privileges and immunities + that apply to the Licensor or You, including from the legal + processes of any jurisdiction or authority. + +======================================================================= + +Creative Commons is not a party to its public +licenses. Notwithstanding, Creative Commons may elect to apply one of +its public licenses to material it publishes and in those instances +will be considered the “Licensor.” The text of the Creative Commons +public licenses is dedicated to the public domain under the CC0 Public +Domain Dedication. Except for the limited purpose of indicating that +material is shared under a Creative Commons public license or as +otherwise permitted by the Creative Commons policies published at +creativecommons.org/policies, Creative Commons does not authorize the +use of the trademark "Creative Commons" or any other trademark or logo +of Creative Commons without its prior written consent including, +without limitation, in connection with any unauthorized modifications +to any of its public licenses or any other arrangements, +understandings, or agreements concerning use of licensed material. For +the avoidance of doubt, this paragraph does not form part of the +public licenses. + +Creative Commons may be contacted at creativecommons.org. + diff --git a/plugins/ue-design-skills/README.md b/plugins/ue-design-skills/README.md new file mode 100644 index 0000000..79befae --- /dev/null +++ b/plugins/ue-design-skills/README.md @@ -0,0 +1,122 @@ +# ue-design-skills + +Design-review skills for Unreal Engine 5.6 projects: norms for how to build a +system, plus reproducible recipes for finding the failures that never announce +themselves. + +Every claim here came out of a line-by-line audit of a large first-party +reference project. What ships is not the audit — it is the recipe. See +[Provenance](#provenance). + +## What is in the box + +`catalog.json` is the index. It is plain JSON with no vendor vocabulary: + +```json +{ "skills": [ { "id": "...", "entry": "skills//SKILL.md", "use_when": "..." } ] } +``` + +Each skill is a directory: + +```text +skills// +├── SKILL.md YAML frontmatter + normative body +└── references/ + ├── patterns.md worked classification, measured scale, hygiene + └── failure-modes.md entries with a Detect recipe you can run +``` + +## Using it without any particular agent runtime + +Nothing here needs a specific harness. Three ways in, cheapest first: + +1. **Read it.** `SKILL.md` is Markdown. Open it. +2. **Load it as context.** Point your agent at `skills//` — the frontmatter + `description` says when the skill applies, `use_when` in `catalog.json` says + the same thing in one line for routing. +3. **Route on the catalog.** Parse `catalog.json`, match `use_when` against the + task, load the matching `entry`. + +Skill content is checked to contain no harness-specific tokens, so a recipe that +works for one agent works for the next. The manifest under `.claude-plugin/` is +one adapter over this catalog, not the source of truth; adding another adapter +does not touch `skills/`. + +## Reading a failure-mode entry + +Each entry carries six fields, in this order: + +| Field | What it answers | +|---|---| +| **Mechanism** | what the code actually does | +| **Why it is silent** | why nothing crashes, logs or fails a test | +| **Why the obvious check misses it** | which natural search returns the wrong answer | +| **Symptom** | what a human reports, in their words | +| **Detect** | a command you run against your own tree | +| **Guardrail** | the rule that stops it recurring | + +The three middle fields are the point. Anyone can write "check that fields have +writers"; only someone who did the work can say why `grep` answers "yes" when the +truth is "no". + +## Detect recipes are meant to be run, not admired + +Every recipe was executed against the reference project before shipping. That +caught two recipes which returned the wrong answer in exactly the case they were +written for — a member declared with an initializer matches every assignment +pattern you can write, and a character class containing `>` matches the `->` +operator. Both are documented in place, because the trap teaches more than the +command. + +If a recipe returns nothing in your tree, widen it before concluding the problem +is absent. **An empty search result proves the pattern, not the absence.** + +## What this bundle does not claim + +- It is not a defect list for any current engine version. It is a worked example + of a classification, measured once, on one version, in one workspace. +- Entries are evidence about *shapes*, not guarantees about your code. +- The delivery gate checks form — no absolute paths, no dangling links, no + missing fields. It cannot check whether an entry is true. + +## Provenance + +The material comes from a research project that read Epic's Lyra Starter Game on +Unreal Engine 5.6 as source rather than running it. Source addresses stay in that +archive: an address in a tree you do not have is not actionable, and a recipe is. + +Entry identifiers (`RA-01` and up) resolve to the audited locations in the +archive, so any specific claim can be produced on request. + +## License + +Copyright (c) 2026 MagentaDolphin. The prose is +[CC BY-ND 4.0](https://creativecommons.org/licenses/by-nd/4.0/); see +[LICENSE](LICENSE) for the exact boundary and the permissions that widen it. + +In short: + +| | | +|---|---| +| Read it, apply it, run the recipes on your code | yes, commercial work included | +| Edit it privately for your studio | yes, no permission needed | +| Redistribute it unchanged | yes, with attribution | +| Publish a modified version as your own product | no | +| Anything you build with it | yours, no claim made | + +Four things are deliberately outside the NoDerivatives restriction, and all four +are Apache-2.0 or explicitly granted: `catalog.json` and +`.claude-plugin/plugin.json`, everything under `_gate/`, the shell commands +inside `Detect` fields, and format adapters and translations. The reasoning is +in `LICENSE` next to each grant — a bundle that tells you to verify every recipe +against your own tree cannot also forbid you to publish the correction. + +### A request, not a condition + +If these skills helped you ship something, a credit is welcome: + +> ue-design-skills by MagentaDolphin + +Nothing obliges you to give one. Attribution is required only when you +redistribute the material itself, and that requirement lives in the license +rather than here. diff --git a/plugins/ue-design-skills/_gate/fixtures/clean/.claude-plugin/plugin.json b/plugins/ue-design-skills/_gate/fixtures/clean/.claude-plugin/plugin.json new file mode 100644 index 0000000..5eee594 --- /dev/null +++ b/plugins/ue-design-skills/_gate/fixtures/clean/.claude-plugin/plugin.json @@ -0,0 +1,4 @@ +{ + "name": "sample-plugin", + "description": "Gate fixture manifest. Clean baseline, not shipped." +} diff --git a/plugins/ue-design-skills/_gate/fixtures/clean/README.md b/plugins/ue-design-skills/_gate/fixtures/clean/README.md new file mode 100644 index 0000000..611d84f --- /dev/null +++ b/plugins/ue-design-skills/_gate/fixtures/clean/README.md @@ -0,0 +1,5 @@ +# Gate fixture: clean baseline + +A minimal plugin tree that satisfies every delivery rule. The gate must report +zero violations here. Each poisoned variant in `test_gate.py` is this tree plus +exactly one documented mutation. diff --git a/plugins/ue-design-skills/_gate/fixtures/clean/catalog.json b/plugins/ue-design-skills/_gate/fixtures/clean/catalog.json new file mode 100644 index 0000000..c02c06d --- /dev/null +++ b/plugins/ue-design-skills/_gate/fixtures/clean/catalog.json @@ -0,0 +1,14 @@ +{ + "schema": "ue-skills-catalog/1", + "name": "sample-plugin", + "description": "Gate fixture catalog. Clean baseline, not shipped.", + "skills": [ + { + "id": "sample-skill", + "path": "skills/sample-skill", + "entry": "skills/sample-skill/SKILL.md", + "description": "Gate fixture skill proving the catalog contract holds.", + "use_when": "Never in production; this entry exists so the baseline has one valid row." + } + ] +} diff --git a/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/SKILL.md b/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/SKILL.md new file mode 100644 index 0000000..9040f54 --- /dev/null +++ b/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/SKILL.md @@ -0,0 +1,25 @@ +--- +name: sample-skill +description: >- + Gate fixture. A minimal clean skill bundle that proves the delivery gate + accepts valid content, including a donor name inside a Provenance section. +--- + +# Sample skill + +> A reference project shows you a shape, not a contract. + +## 1. Norm + +State the field, name its consumer, and give a recipe that reproduces the check +in the reader's own tree. A norm that cannot be re-derived is decoration. + +## Provenance + +Measured on Epic's Lyra Starter Game for Unreal Engine 5.6. Source addresses +stay in the research archive; the recipes below reproduce each finding in your +own project. This section is the only place the donor may be named. + +## 2. Failure modes + +Detection recipes and guardrails: [failure modes](references/failure-modes.md). diff --git a/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/references/failure-modes.md b/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/references/failure-modes.md new file mode 100644 index 0000000..b77e761 --- /dev/null +++ b/plugins/ue-design-skills/_gate/fixtures/clean/skills/sample-skill/references/failure-modes.md @@ -0,0 +1,37 @@ +# Failure modes: sample skill + +Gate fixture. One complete entry, in the exact shape the gate accepts. + +An entry is valid only if a reader who has never seen the donor project can run +`Detect` in their own tree and get an answer. The three middle fields are the +ones that cannot be written by someone who did not do the work. + +--- + +### SS-01 - Editable property with no writer + +**Mechanism.** A tunable float is compared against a timestamp field, but the +timestamp is never assigned anywhere in the codebase. + +**Why it is silent.** Time since an unset stamp is time since world start, which +always exceeds the threshold. The comparison is permanently true, so the gate it +was meant to impose never rejects anything and no error is ever raised. + +**Why the obvious check misses it.** The field is read twice and participates in +a comparison, so any "is this used?" search answers yes. What is absent is the +**writer**, not the reader. + +**Symptom.** A designer tunes the property for days and reports that it changes +nothing. + +**Detect.** For each editable property, search for an assignment rather than a +mention, and subtract the declaration -- a member declared with an initializer +matches every assignment pattern you can write: + + rg -n "\bPropertyName\b\s*(=|\+=|-=)" Source/ \ + | rg -v "\b(bool|u?int\d+|float|double|F[A-Z]\w+)\s+PropertyName\b" + +If reads exist and this comes back empty, the property is a stub. + +**Guardrail.** Do not ship an editor-exposed property whose value has no writer. +Every tunable needs one owning assignment site and one test that moves it. diff --git a/plugins/ue-design-skills/_gate/gate.py b/plugins/ue-design-skills/_gate/gate.py new file mode 100644 index 0000000..e06e3b6 --- /dev/null +++ b/plugins/ue-design-skills/_gate/gate.py @@ -0,0 +1,592 @@ +"""Delivery gate for the ue-design-skills plugin bundle. + +Checks the shipping surface of a plugin tree against rules that the +depersonalization work must satisfy. Output is ASCII only. + + python _gate/gate.py [plugin_root] + python _gate/gate.py --list-rules + +Exit code 0 when clean, 1 when any rule fires. + +Design notes +------------ +The scan surface is a WHITELIST, not "everything minus exclusions". A blacklist +lets an unscanned directory appear by accident; rule P11 closes the remaining +hole by rejecting any top-level entry that is not on the known list. So +`_gate/` is unscanned because it is not shipped content, and nothing else can +quietly join it. + +This gate cannot detect a fabricated claim. Green means the form is right, not +that the entry is true. +""" +from __future__ import annotations + +import json +import re +import sys +from dataclasses import dataclass +from pathlib import Path + +# -------------------------------------------------------------------------- +# Scan surface +# -------------------------------------------------------------------------- + +SCAN_DIRS = ('.claude-plugin', 'skills') +SCAN_ROOT_FILES = ('README.md', 'LICENSE') +ALLOWED_TOP_LEVEL = {'.claude-plugin', 'skills', '_gate', 'README.md', + 'LICENSE', '.gitignore', 'catalog.json'} + +# Harness neutrality. The skills are plain Markdown with YAML frontmatter and +# must be usable by any agent or runner. A vendor manifest is an adapter that +# sits outside skills/ -- content that names one harness cannot be run by +# another, so naming one inside skills/ is a portability defect, not a style +# preference. +HARNESS_TOKENS = ( + r'\$\{CLAUDE_[A-Z_]*\}', + r'\bClaude\b', + r'\bCursor\b', + r'\bCopilot\b', + r'(?= 16, 'rule set shrank; a rule was lost' +assert len({r.id for r in RULES}) == len(RULES), 'duplicate rule id' + + +# -------------------------------------------------------------------------- +# Helpers +# -------------------------------------------------------------------------- + +def _read(path: Path) -> str: + return path.read_text(encoding='utf-8', errors='replace') + + +def _shipped_files(root: Path): + """Every file on the shipping surface, whitelist order.""" + for name in SCAN_ROOT_FILES: + p = root / name + if p.is_file(): + yield p + for d in SCAN_DIRS: + base = root / d + if not base.is_dir(): + continue + for p in sorted(base.rglob('*')): + if p.is_file(): + yield p + + +def _is_text(root: Path, path: Path) -> bool: + """Whether the line-level rules should read this file. + + Extension is the usual signal, but LICENSE has none. It sat in + SCAN_ROOT_FILES and was skipped by every line rule for as long as it did + not exist: a LICENSE carrying an absolute path, a citation, the donor name, + Cyrillic and a wikilink passed the gate green. A root file on the declared + surface is text whatever its suffix. + """ + if path.suffix.lower() in TEXT_SUFFIXES: + return True + return path.parent == root and path.name in SCAN_ROOT_FILES + + +def _banner_provenance_lines(text: str) -> set[int]: + """1-based line numbers inside a banner-style PROVENANCE section. + + Markdown marks a section with '## Provenance'; a plain-text LICENSE marks + it with a title between rules of '='. The donor carve-out is about the + section, not about the syntax that happens to delimit it, so a LICENSE + stating where the material came from must not be forced to omit the name. + """ + allowed: set[int] = set() + lines = text.splitlines() + inside = False + i = 0 + n = len(lines) + while i < n: + # A banner heading is three lines: rule, title, rule. Reading them one + # at a time makes the closing rule look like a new heading with an + # empty title, which closed the section on the line that opened it. + if (i + 2 < n + and RE_BANNER_RULE.match(lines[i]) + and RE_BANNER_RULE.match(lines[i + 2]) + and lines[i + 1].strip()): + inside = lines[i + 1].strip().upper() == 'PROVENANCE' + i += 3 + continue + if inside: + allowed.add(i + 1) + i += 1 + return allowed + + +def _provenance_lines(text: str) -> set[int]: + """1-based line numbers that sit inside a Provenance section.""" + allowed: set[int] = set() + lines = text.splitlines() + depth = None + for i, line in enumerate(lines, 1): + m = RE_HEADING.match(line) + if m: + level = len(m.group(1)) + if depth is not None and level <= depth: + depth = None + if RE_PROVENANCE.match(line): + depth = level + allowed.add(i) + continue + if depth is not None: + allowed.add(i) + return allowed + + +def _frontmatter(text: str) -> dict[str, str] | None: + if not text.startswith('---\n'): + return None + end = text.find('\n---', 4) + if end == -1: + return None + block = text[4:end] + out: dict[str, str] = {} + key = None + for line in block.splitlines(): + m = re.match(r'^([A-Za-z_][A-Za-z0-9_-]*):\s*(.*)$', line) + if m: + key = m.group(1) + out[key] = m.group(2).strip() + elif key and line.strip(): + out[key] = (out[key] + ' ' + line.strip()).strip() + return out + + +# -------------------------------------------------------------------------- +# Line-level rules +# -------------------------------------------------------------------------- + +def check_text_rules(root: Path) -> list[Violation]: + out: list[Violation] = [] + for path in _shipped_files(root): + if not _is_text(root, path): + continue + rel = path.relative_to(root).as_posix() + text = _read(path) + suffix = path.suffix.lower() + if suffix == '.md': + prov = _provenance_lines(text) + elif not suffix: + # LICENSE and friends: no extension, no Markdown headings, but + # they still carry a PROVENANCE section and it still needs the + # donor name to say anything honest. + prov = _banner_provenance_lines(text) + else: + prov = set() + for n, line in enumerate(text.splitlines(), 1): + if RE_ABS_PATH.search(line): + out.append(Violation('P01', rel, n, line.strip())) + m = RE_CITATION.search(line) + if m: + out.append(Violation('P02', rel, n, m.group(0))) + if n not in prov: + for token in DONOR_TOKENS: + if re.search(r'\b' + re.escape(token), line, re.I): + out.append(Violation('P03', rel, n, line.strip())) + break + if RE_CYRILLIC.search(line): + out.append(Violation('P04', rel, n, line.strip()[:60])) + if RE_WIKILINK.search(line): + out.append(Violation('P05', rel, n, line.strip())) + return out + + +# -------------------------------------------------------------------------- +# Structural rules +# -------------------------------------------------------------------------- + +def check_harness_neutral(root: Path) -> list[Violation]: + """P13: nothing under skills/ may name a specific agent harness. + + The vendor manifest in .claude-plugin/ is an adapter and is exempt. Skill + content is not: a recipe that says "run this Claude command" cannot be run + by another agent, and the reader has no way to translate it. + """ + out: list[Violation] = [] + base = root / 'skills' + if not base.is_dir(): + return out + patterns = [re.compile(p) for p in HARNESS_TOKENS] + for path in sorted(base.rglob('*')): + if not path.is_file() or path.suffix.lower() not in TEXT_SUFFIXES: + continue + rel = path.relative_to(root).as_posix() + for n, line in enumerate(path.read_text( + encoding='utf-8', errors='replace').splitlines(), 1): + for pat in patterns: + m = pat.search(line) + if m: + out.append(Violation('P13', rel, n, + f'harness-specific token: {m.group(0)}')) + break + return out + + +def check_catalog(root: Path) -> list[Violation]: + """P14: catalog.json describes every skill without vendor vocabulary. + + A harness that does not read .claude-plugin/ still needs to know what is + here and when to reach for it. Without this file the bundle is only usable + by the one runner whose manifest format we happened to write. + """ + out: list[Violation] = [] + cat = root / 'catalog.json' + skills_dir = root / 'skills' + present = sorted(p.name for p in skills_dir.iterdir() + if p.is_dir()) if skills_dir.is_dir() else [] + if not cat.is_file(): + if present: + out.append(Violation('P14', 'catalog.json', 0, + 'missing; skills are not discoverable ' + 'without a vendor manifest')) + return out + try: + data = json.loads(_read(cat)) + except json.JSONDecodeError as exc: + out.append(Violation('P14', 'catalog.json', 0, f'invalid JSON: {exc}')) + return out + + entries = data.get('skills') + if not isinstance(entries, list): + out.append(Violation('P14', 'catalog.json', 0, + 'no "skills" array')) + return out + + listed = [] + for i, e in enumerate(entries): + if not isinstance(e, dict): + out.append(Violation('P14', 'catalog.json', 0, + f'entry {i} is not an object')) + continue + name = e.get('id', '') + listed.append(name) + for field in ('id', 'path', 'description', 'use_when'): + if not str(e.get(field, '')).strip(): + out.append(Violation('P14', 'catalog.json', 0, + f'{name or i}: empty {field}')) + p = str(e.get('path', '')) + if p and not (root / p / 'SKILL.md').is_file(): + out.append(Violation('P14', 'catalog.json', 0, + f'{name}: path does not hold a SKILL.md: {p}')) + for missing in sorted(set(present) - set(listed)): + out.append(Violation('P14', 'catalog.json', 0, + f'skill not in catalog: {missing}')) + for ghost in sorted(set(listed) - set(present)): + out.append(Violation('P14', 'catalog.json', 0, + f'catalog names a skill that is absent: {ghost}')) + return out + + +def check_id_prefixes(root: Path) -> list[Violation]: + """P16: entry-id prefixes must be unique across skills. + + The archive-side resolver keys entries by bare id. Two skills sharing a + prefix therefore overwrite each other in a dict, and the loss is silent -- + the resolver reports a smaller total and still says "0 unresolved". Found + the hard way when a settings skill and an ability-system skill both claimed + GS, and nine entries vanished from the map without any check firing. + """ + out: list[Violation] = [] + skills_dir = root / 'skills' + if not skills_dir.is_dir(): + return out + owners: dict[str, list[str]] = {} + for skill in sorted(p for p in skills_dir.iterdir() if p.is_dir()): + fm = skill / 'references' / 'failure-modes.md' + if not fm.is_file(): + continue + for line in _read(fm).splitlines(): + m = re.match(r'^###\s+([A-Z]{2,4})-\d+\b', line) + if m: + owners.setdefault(m.group(1), []) + if skill.name not in owners[m.group(1)]: + owners[m.group(1)].append(skill.name) + break + for prefix, skills in sorted(owners.items()): + if len(skills) > 1: + for name in skills: + out.append(Violation( + 'P16', f'skills/{name}/references/failure-modes.md', 0, + f'prefix {prefix} also used by: ' + f'{", ".join(s for s in skills if s != name)}')) + return out + + +def check_top_level(root: Path) -> list[Violation]: + out: list[Violation] = [] + for entry in sorted(root.iterdir()): + if entry.name not in ALLOWED_TOP_LEVEL: + out.append(Violation('P11', entry.name, 0, + 'not part of the plugin shipping surface')) + return out + + +def check_junk(root: Path) -> list[Violation]: + out: list[Violation] = [] + for path in _shipped_files(root): + rel = path.relative_to(root).as_posix() + if path.suffix.lower() in JUNK_SUFFIXES or '__pycache__' in rel: + out.append(Violation('P12', rel, 0, 'disallowed file type')) + return out + + +def check_skills(root: Path) -> list[Violation]: + out: list[Violation] = [] + skills_dir = root / 'skills' + if not skills_dir.is_dir(): + return out + for skill in sorted(p for p in skills_dir.iterdir() if p.is_dir()): + rel = f'skills/{skill.name}/SKILL.md' + md = skill / 'SKILL.md' + if not md.is_file(): + out.append(Violation('P06', rel, 0, 'missing SKILL.md')) + continue + text = _read(md) + front = _frontmatter(text) + if front is None: + out.append(Violation('P06', rel, 1, 'missing YAML frontmatter')) + continue + + name = front.get('name', '') + if not name: + out.append(Violation('P06', rel, 1, 'no name field')) + elif not RE_KEBAB.match(name): + out.append(Violation('P06', rel, 1, f'name not kebab-case: {name}')) + elif name != skill.name: + out.append(Violation('P06', rel, 1, + f'name {name} != folder {skill.name}')) + + desc = front.get('description', '').strip() + if not desc or desc in ('>-', '>', '|'): + out.append(Violation('P07', rel, 1, 'no description field')) + + fm = skill / 'references' / 'failure-modes.md' + if not fm.is_file(): + out.append(Violation('P08', rel, 0, + 'missing references/failure-modes.md')) + elif 'references/failure-modes.md' not in text: + out.append(Violation('P08', rel, 0, + 'SKILL.md does not link failure-modes.md')) + else: + out.extend(check_entries(root, fm)) + + out.extend(check_links(root, skill, md, text)) + refs_dir = skill / 'references' + if refs_dir.is_dir(): + for ref in sorted(refs_dir.glob('*.md')): + out.extend(check_links(root, skill, ref, _read(ref))) + return out + + +def check_entries(root: Path, fm: Path) -> list[Violation]: + """Every '### ' entry in failure-modes.md carries all required fields.""" + out: list[Violation] = [] + rel = fm.relative_to(root).as_posix() + lines = _read(fm).splitlines() + entries: list[tuple[int, str, list[str]]] = [] + cur = None + for n, line in enumerate(lines, 1): + m = RE_HEADING.match(line) + if m and len(m.group(1)) == 3: + cur = (n, m.group(2), []) + entries.append(cur) + elif cur is not None: + fm_field = RE_ENTRY_FIELD.match(line) + if fm_field: + cur[2].append(fm_field.group(1).strip()) + if not entries: + out.append(Violation('P09', rel, 0, 'no "### " entries found')) + return out + + # P15: identifiers must run 01, 02, 03... in the order a reader meets them. + # Reordering sections without renumbering leaves gaps that read as deleted + # entries, and a stable id that moved is worse than one that never existed. + # Caught twice by hand during authoring, hence a rule. + seq = [] + for start, title, _ in entries: + m = re.match(r'^([A-Z]{2,4})-(\d+)\b', title) + if m: + seq.append((start, m.group(1), int(m.group(2)))) + if seq: + prefixes = {p for _, p, _ in seq} + if len(prefixes) > 1: + out.append(Violation('P15', rel, seq[0][0], + f'mixed id prefixes: {sorted(prefixes)}')) + for i, (start, prefix, num) in enumerate(seq, 1): + if num != i: + out.append(Violation( + 'P15', rel, start, + f'{prefix}-{num:02d} is entry {i} in document order')) + break + for start, title, fields in entries: + present = {f.lower() for f in fields} + missing = [f for f in REQUIRED_ENTRY_FIELDS + if f.lower() not in present] + if missing: + out.append(Violation('P09', rel, start, + f'{title}: missing {", ".join(missing)}')) + return out + + +def check_links(root: Path, skill: Path, src: Path, text: str) -> list[Violation]: + out: list[Violation] = [] + rel = src.relative_to(root).as_posix() + skill_res = skill.resolve() + for n, line in enumerate(text.splitlines(), 1): + for target in RE_MD_LINK.findall(line): + t = target.strip() + if '://' in t or t.startswith('#') or t.startswith('mailto:'): + continue + clean = t.split('#', 1)[0] + if not clean: + continue + dest = (src.parent / clean).resolve() + try: + dest.relative_to(skill_res) + except ValueError: + out.append(Violation('P10', rel, n, + f'link escapes skill dir: {t}')) + continue + if not dest.exists(): + out.append(Violation('P10', rel, n, f'broken link: {t}')) + return out + + +def check_manifest(root: Path) -> list[Violation]: + out: list[Violation] = [] + mf = root / '.claude-plugin' / 'plugin.json' + if not mf.is_file(): + out.append(Violation('P11', '.claude-plugin/plugin.json', 0, + 'manifest missing')) + return out + try: + data = json.loads(_read(mf)) + except json.JSONDecodeError as exc: + out.append(Violation('P11', '.claude-plugin/plugin.json', 0, + f'invalid JSON: {exc}')) + return out + name = data.get('name', '') + if not name or not RE_KEBAB.match(name): + out.append(Violation('P11', '.claude-plugin/plugin.json', 0, + f'manifest name not kebab-case: {name!r}')) + for entry in sorted((root / '.claude-plugin').iterdir()): + if entry.name != 'plugin.json': + out.append(Violation('P11', f'.claude-plugin/{entry.name}', 0, + '.claude-plugin holds only the manifest')) + return out + + +CHECKS = (check_text_rules, check_top_level, check_junk, check_skills, + check_manifest, check_harness_neutral, check_catalog, + check_id_prefixes) + + +def run(root: Path) -> list[Violation]: + out: list[Violation] = [] + for check in CHECKS: + out.extend(check(root)) + return sorted(out, key=lambda v: (v.rule, v.path, v.line)) + + +def main(argv: list[str]) -> int: + if '--list-rules' in argv: + for r in RULES: + print(f'{r.id} {r.what}') + return 0 + args = [a for a in argv if not a.startswith('--')] + root = Path(args[0]).resolve() if args else Path(__file__).resolve().parents[1] + if not root.is_dir(): + print(f'ERROR: not a directory: {root}') + return 2 + violations = run(root) + for v in violations: + loc = f'{v.path}:{v.line}' if v.line else v.path + print(f'{v.rule} {loc} {v.text}') + fired = sorted({v.rule for v in violations}) + print(f'\n{len(violations)} violations, ' + f'{len(fired)} of {len(RULES)} rules fired') + if violations: + print('rules fired: ' + ' '.join(fired)) + return 1 if violations else 0 + + +if __name__ == '__main__': + sys.exit(main(sys.argv[1:])) diff --git a/plugins/ue-design-skills/_gate/test_gate.py b/plugins/ue-design-skills/_gate/test_gate.py new file mode 100644 index 0000000..5866cff --- /dev/null +++ b/plugins/ue-design-skills/_gate/test_gate.py @@ -0,0 +1,191 @@ +"""Poisoned-fixture corpus for the delivery gate. + + python _gate/test_gate.py + +Every rule in gate.RULES must have exactly one poison here, and that poison must +make its rule fire. A rule that never reddens on a fixture does not exist -- the +previous validator in this project ran "successfully" for months while checking +a literal that could not occur. This file is the answer to that. + +The clean baseline lives in _gate/fixtures/clean and must be green. Each variant +is that tree copied to a temp directory plus exactly one mutation. + +Output is ASCII only; the Cyrillic poison is written as escapes and never +printed. +""" +from __future__ import annotations + +import json +import shutil +import sys +import tempfile +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import gate # noqa: E402 + +CLEAN = Path(__file__).resolve().parent / 'fixtures' / 'clean' +SKILL = 'skills/sample-skill/SKILL.md' +FM = 'skills/sample-skill/references/failure-modes.md' + +# rule id -> (operation, path, payload, replacement) +POISONS: dict[str, tuple] = { + 'P01': ('append', SKILL, + '\nDonor tree lives at D:\\Work\\NG\\overmind on the build box.\n'), + 'P02': ('append', SKILL, + '\nSee RangedWeaponInstance.cpp:194 for the original.\n'), + 'P03': ('append', SKILL, + '\nThe donor project Lyra shipped this pattern unchanged.\n'), + 'P04': ('append', SKILL, + '\n\u041f\u0440\u0438\u043c\u0435\u0447\u0430\u043d\u0438\u0435.\n'), + 'P05': ('append', SKILL, '\nRelated: [[ue-modular-gameplay]]\n'), + 'P06': ('replace', SKILL, 'name: sample-skill', 'name: renamed-skill'), + 'P07': ('replace', SKILL, 'description: >-', 'summary: >-'), + 'P08': ('delete', FM), + 'P09': ('replace', FM, '**Guardrail.**', 'Guardrail:'), + 'P10': ('append', SKILL, + '\nSee [the archive](../../../notes/00-index.md).\n'), + 'P11': ('mkfile', 'notes/leak.md', '# a note from the archive\n'), + 'P12': ('mkfile', 'skills/sample-skill/references/stale.pyc', 'junk\n'), + 'P13': ('append', SKILL, + '\nRun this through Claude Code with ${CLAUDE_PLUGIN_ROOT} set.\n'), + 'P14': ('delete', 'catalog.json'), + # The single clean entry is SS-01 and sits first. Renaming it to SS-02 + # leaves position 1 holding number 02, which is exactly the gap that + # reordering sections during authoring produces. + 'P15': ('replace', FM, '### SS-01', '### SS-02'), + # A second skill reusing the SS- prefix. This is the defect that let nine + # shipped entries vanish from the archive resolver without any count + # disagreeing loudly enough to notice. + 'P16': ('copyskill', 'skills/sample-skill', 'sample-twin'), +} + + +# Surface coverage is a separate question from rule coverage. POISONS proves +# each rule can fire; this proves each file the gate CLAIMS to scan actually +# reaches the line-level rules. Both were needed: every rule above had a +# working fixture while a LICENSE carrying all five text violations passed +# green, because the fixtures all poison .md and LICENSE has no extension. +# Declaring a file on the surface and filtering it out one line later is the +# same defect class as a rule with no fixture -- coverage asserted, not had. +SURFACE_POISON = ( + 'Donor tree at D:\\Work\\NG\\overmind\n' + 'See RangedWeaponInstance.cpp:194\n' + 'The donor project Lyra shipped this\n' + '\u041f\u0440\u0438\u043c\u0435\u0447\u0430\u043d\u0438\u0435\n' + 'Related: [[ue-modular-gameplay]]\n') +SURFACE_EXPECT = {'P01', 'P02', 'P03', 'P04', 'P05'} + + +def check_surface(failures: list[str]) -> None: + """Every name in gate.SCAN_ROOT_FILES must reach the line-level rules.""" + for name in gate.SCAN_ROOT_FILES: + with tempfile.TemporaryDirectory() as td: + root = materialize(Path(td) / 'surface') + (root / name).write_text(SURFACE_POISON, encoding='utf-8') + fired = {v.rule for v in gate.run(root) if v.path == name} + missing = SURFACE_EXPECT - fired + print(f'surface {name:<12} ' + f'{"scanned" if not missing else "NOT SCANNED"}') + if missing: + failures.append( + f'{name} is on the scan surface but line rules ' + f'{sorted(missing)} never saw it') + + +def materialize(tmp: Path) -> Path: + root = tmp / 'plugin' + shutil.copytree(CLEAN, root) + return root + + +def poison(root: Path, spec: tuple) -> None: + op = spec[0] + target = root / spec[1] + if op == 'append': + target.write_text( + target.read_text(encoding='utf-8') + spec[2], encoding='utf-8') + elif op == 'replace': + text = target.read_text(encoding='utf-8') + if spec[2] not in text: + raise AssertionError(f'poison anchor absent: {spec[2]!r}') + target.write_text(text.replace(spec[2], spec[3], 1), encoding='utf-8') + elif op == 'delete': + target.unlink() + elif op == 'mkfile': + target.parent.mkdir(parents=True, exist_ok=True) + target.write_text(spec[2], encoding='utf-8') + elif op == 'copyskill': + # Duplicate a skill under a new folder name, keeping its entry-id + # prefix. The frontmatter name and the catalog entry are fixed up so + # that P06 and P14 stay green and only the prefix collision fires. + src, dst = root / spec[1], root / spec[1].rsplit('/', 1)[0] / spec[2] + shutil.copytree(src, dst) + md = dst / 'SKILL.md' + md.write_text( + md.read_text(encoding='utf-8').replace( + f'name: {src.name}', f'name: {spec[2]}'), encoding='utf-8') + cat = root / 'catalog.json' + data = json.loads(cat.read_text(encoding='utf-8')) + twin = dict(data['skills'][0]) + twin['id'] = spec[2] + twin['path'] = f'skills/{spec[2]}' + twin['entry'] = f'skills/{spec[2]}/SKILL.md' + twin['references'] = [r.replace(src.name, spec[2]) + for r in twin.get('references', [])] + data['skills'].append(twin) + cat.write_text(json.dumps(data, indent=2) + '\n', encoding='utf-8') + else: + raise AssertionError(f'unknown poison op: {op}') + + +def main() -> int: + failures: list[str] = [] + rule_ids = {r.id for r in gate.RULES} + + missing = rule_ids - set(POISONS) + extra = set(POISONS) - rule_ids + if missing: + failures.append(f'rules with no fixture: {sorted(missing)}') + if extra: + failures.append(f'fixtures for unknown rules: {sorted(extra)}') + + with tempfile.TemporaryDirectory() as td: + root = materialize(Path(td) / 'clean') + found = gate.run(root) + if found: + failures.append('clean baseline is not green:') + for v in found: + failures.append(f' {v.rule} {v.path}:{v.line} {v.text}') + else: + print('clean baseline green') + + for rule in sorted(POISONS): + with tempfile.TemporaryDirectory() as td: + root = materialize(Path(td) / rule) + poison(root, POISONS[rule]) + fired = sorted({v.rule for v in gate.run(root)}) + ok = rule in fired + others = [f for f in fired if f != rule] + note = f' (also {" ".join(others)})' if others else '' + print(f'{rule} poison ' + f'{"fires" if ok else "SILENT"}{note}') + if not ok: + failures.append( + f'{rule}: poison did not fire the rule; fired={fired}') + + check_surface(failures) + + print() + if failures: + for f in failures: + print('FAIL: ' + f) + return 1 + print(f'{len(POISONS)} rules, each reddens on its own fixture, ' + f'{len(gate.SCAN_ROOT_FILES)} root files reach the line rules, ' + 'baseline clean') + return 0 + + +if __name__ == '__main__': + sys.exit(main()) diff --git a/plugins/ue-design-skills/catalog.json b/plugins/ue-design-skills/catalog.json new file mode 100644 index 0000000..8b97395 --- /dev/null +++ b/plugins/ue-design-skills/catalog.json @@ -0,0 +1,199 @@ +{ + "schema": "ue-skills-catalog/1", + "name": "ue-design-skills", + "version": "0.1.0", + "description": "Design-review skills for Unreal Engine 5.6 projects. Norms plus reproducible detection recipes for silent architectural failures.", + "author": "MagentaDolphin", + "publisher": "Kodlo.art", + "license": "CC-BY-ND-4.0", + "license_note": "Prose is CC BY-ND 4.0. This file, .claude-plugin/plugin.json, everything under _gate/ and the shell commands inside Detect fields are Apache-2.0. Format adapters, translations and private adaptation are permitted; see LICENSE for the exact grants.", + "about": "Harness-neutral index of this bundle. Every skill is a directory holding SKILL.md (YAML frontmatter plus Markdown body) and a references/ folder. No agent runtime is required to read them: point any harness at the entry path, or load the whole directory as context. The vendor manifest under .claude-plugin/ is one adapter over this catalog, not the source of truth.", + "skills": [ + { + "id": "ue-reference-project-adoption", + "path": "skills/ue-reference-project-adoption", + "entry": "skills/ue-reference-project-adoption/SKILL.md", + "description": "Classify what you borrow from a reference or sample project as defect, showcase stub, deliberate scope narrowing, architecture tax or foreign legacy, then decide copy/close/skip per element.", + "use_when": "Starting a project from a template, lifting a subsystem out of a sample, reviewing code copied from one, or explaining why a setting copied from the reference does nothing.", + "references": [ + "skills/ue-reference-project-adoption/references/patterns.md", + "skills/ue-reference-project-adoption/references/failure-modes.md" + ] + }, + { + "id": "ue-multiplayer-authority", + "path": "skills/ue-multiplayer-authority", + "entry": "skills/ue-multiplayer-authority/SKILL.md", + "description": "Design and review the authority model of a networked game: who owns each fact, what is predicted versus replicated versus validated, RPC and validation discipline, hit registration, replication modes, relevancy and the cheat surface.", + "use_when": "Starting a multiplayer feature, porting single-player code to network, reviewing trust boundaries, or investigating desync, rubber-banding, \"works in the editor but not on a dedicated server\", or suspected cheating.", + "references": [ + "skills/ue-multiplayer-authority/references/patterns.md", + "skills/ue-multiplayer-authority/references/failure-modes.md" + ] + }, + { + "id": "ue-gameplay-cues", + "path": "skills/ue-gameplay-cues", + "entry": "skills/ue-gameplay-cues/SKILL.md", + "description": "Design and review the GameplayCue presentation layer: cue forms and addressing routes, static versus actor notifies, the dedicated-server rule, cue discovery and loading cost, registering cue paths from feature plugins, teardown, and the tag-to-asset validation gap.", + "use_when": "Adding VFX or audio feedback to abilities or effects, decoupling presentation from gameplay assets, debugging a cue that never fires or never stops, or auditing cue loading cost.", + "references": [ + "skills/ue-gameplay-cues/references/patterns.md", + "skills/ue-gameplay-cues/references/failure-modes.md" + ] + }, + { + "id": "ue-gameplay-tag-governance", + "path": "skills/ue-gameplay-tag-governance", + "entry": "skills/ue-gameplay-tag-governance/SKILL.md", + "description": "Govern gameplay tags as a typed vocabulary rather than loose strings: declaration sources, namespace design, constraining tag inputs, exact versus hierarchical matching, redirects and renaming, replicating a tag registry, tag-driven counters, and auditing an existing inventory.", + "use_when": "Adding a tag namespace, wiring systems through tags, reviewing tag usage, or debugging a tag-driven feature that silently does nothing.", + "references": [ + "skills/ue-gameplay-tag-governance/references/patterns.md", + "skills/ue-gameplay-tag-governance/references/failure-modes.md" + ] + }, + { + "id": "ue-gameplay-messaging", + "path": "skills/ue-gameplay-messaging", + "entry": "skills/ue-gameplay-messaging/SKILL.md", + "description": "Design and review a tag-addressed publish/subscribe message bus: bus scope and lifetime, typed payloads and runtime type checks, exact versus partial channel matching, listener handles and ownership, parity between native and visual-script APIs, and the boundary between local messaging and replication.", + "use_when": "Decoupling gameplay from UI, feedback, telemetry or processors, or debugging duplicate, stale, mismatched or wrongly replicated messages.", + "references": [ + "skills/ue-gameplay-messaging/references/patterns.md", + "skills/ue-gameplay-messaging/references/failure-modes.md" + ] + }, + { + "id": "ue-modular-gameplay", + "path": "skills/ue-modular-gameplay", + "entry": "skills/ue-modular-gameplay/SKILL.md", + "description": "Design and review modular gameplay: experience definitions, feature plugins, reusable action sets, feature actions and their rollback, actor extension handlers, component init-state chains, asset bundles by role, and runtime activation and deactivation.", + "use_when": "Adding game modes, seasonal or downloadable features, modular components, abilities, input or UI, or fixing initialization-order and feature teardown bugs.", + "references": [ + "skills/ue-modular-gameplay/references/patterns.md", + "skills/ue-modular-gameplay/references/failure-modes.md" + ] + }, + { + "id": "ue-gas-architecture", + "path": "skills/ue-gas-architecture", + "entry": "skills/ue-gas-architecture/SKILL.md", + "description": "Design and review a production Gameplay Ability System layer: ability packages and revocation receipts, buffered ability input, activation policies and exclusivity groups, attribute set responsibilities, effect contexts, tag relationship policy, match phases as abilities, and the health/death seam.", + "use_when": "Adding combat, abilities, equipment-granted powers, phases or rounds, damage and healing, or diagnosing ability activation and lifecycle bugs.", + "references": [ + "skills/ue-gas-architecture/references/patterns.md", + "skills/ue-gas-architecture/references/failure-modes.md" + ] + }, + { + "id": "ue-input-architecture", + "path": "skills/ue-input-architecture", + "entry": "skills/ue-input-architecture/SKILL.md", + "description": "Design and review Enhanced Input architecture: physical mapping contexts, semantic input tags, native versus ability paths, per-frame buffering, player-mappable settings, device switching, and runtime injection and removal of input by feature plugins.", + "use_when": "Adding controls, rebinding, gamepad or touch support, ability input, mode-specific input, or debugging an action that binds correctly and never fires.", + "references": [ + "skills/ue-input-architecture/references/patterns.md", + "skills/ue-input-architecture/references/failure-modes.md" + ] + }, + { + "id": "ue-ui-architecture", + "path": "skills/ue-ui-architecture", + "entry": "skills/ue-ui-architecture/SKILL.md", + "description": "Design and review scalable game UI: one root layout per local player, tagged layer stacks for modality, extension points that let features contribute widgets without depending on the shell, input and focus policy per screen, and centrally arbitrated loading screens.", + "use_when": "Building HUD, menu or modal architecture, adding mode or plugin UI without base-class dependencies, supporting gamepad or local multiplayer, or debugging focus, widget ordering, duplicate widgets and teardown.", + "references": [ + "skills/ue-ui-architecture/references/patterns.md", + "skills/ue-ui-architecture/references/failure-modes.md" + ] + }, + { + "id": "ue-data-driven-architecture", + "path": "skills/ue-data-driven-architecture", + "entry": "skills/ue-data-driven-architecture/SKILL.md", + "description": "Design and review a data-driven definition layer: composition layers, data assets versus zero-node class defaults versus instanced fragments, reference semantics and asset bundles, the code/data boundary, and the validation a data graph needs because it compiles regardless.", + "use_when": "Adding game modes, pawn archetypes, ability packages, item or equipment definitions, playlists or feature plugins, or when a manager is accumulating a branch per variant.", + "references": [ + "skills/ue-data-driven-architecture/references/patterns.md", + "skills/ue-data-driven-architecture/references/failure-modes.md" + ] + }, + { + "id": "ue-cosmetics-and-teams", + "path": "skills/ue-cosmetics-and-teams", + "entry": "skills/ue-cosmetics-and-teams/SKILL.md", + "description": "Design and review modular character cosmetics and team systems: replicated cosmetic intent with local realization, controller-owned selection versus pawn-owned realization, dedicated-server exclusion, tag-driven mesh and animation selection, one authoritative team identifier with subscribed mirrors, and a single damage policy shared by authority and prediction.", + "use_when": "Building skins, modular characters, team assignment, friendly fire, hit-marker filtering, or team-coloured UI and effects.", + "references": [ + "skills/ue-cosmetics-and-teams/references/patterns.md", + "skills/ue-cosmetics-and-teams/references/failure-modes.md" + ] + }, + { + "id": "ue-game-settings-architecture", + "path": "skills/ue-game-settings-architecture", + "entry": "skills/ue-game-settings-architecture/SKILL.md", + "description": "Design and review a settings layer: setting models separated from widgets, reflection-backed property paths and their cost, declarative edit conditions, apply and cancel as a transaction, machine-local versus player-shared persistence, scalability and device profiles, benchmark wiring, and sampled performance statistics.", + "use_when": "Adding video, audio, input or accessibility settings, cloud saves, quality presets, frame-rate modes, device tuning, or performance overlays.", + "references": [ + "skills/ue-game-settings-architecture/references/patterns.md", + "skills/ue-game-settings-architecture/references/failure-modes.md" + ] + }, + { + "id": "ue-asset-loading-and-memory", + "path": "skills/ue-asset-loading-and-memory", + "entry": "skills/ue-asset-loading-and-memory/SKILL.md", + "description": "Audit asset loading and CPU memory residency: reference semantics and transitive closure, asset bundles and cook rules, load handle ownership, synchronous load hitches, startup sequencing, garbage-collection reachability, retention leaks and unload paths.", + "use_when": "Reducing load times or memory, choosing reference types, adding content-pulling definitions, diagnosing boot or transition hitches, or investigating objects that are never collected.", + "references": [ + "skills/ue-asset-loading-and-memory/references/patterns.md", + "skills/ue-asset-loading-and-memory/references/failure-modes.md" + ] + }, + { + "id": "ue-streaming-and-platform-budgets", + "path": "skills/ue-streaming-and-platform-budgets", + "entry": "skills/ue-streaming-and-platform-budgets/SKILL.md", + "description": "Audit GPU and streaming budgets: separating disk, CPU memory, video memory and the streaming pool; console-variable priority and device profiles; texture groups and scalability buckets; quality clamps; and the lifetime of streamed level content.", + "use_when": "Targeting a platform memory ceiling, tuning quality tiers, adding a background or preview level, or investigating why a quality setting has no effect on some hardware.", + "references": [ + "skills/ue-streaming-and-platform-budgets/references/patterns.md", + "skills/ue-streaming-and-platform-budgets/references/failure-modes.md" + ] + }, + { + "id": "ue-runtime-allocation-and-caching", + "path": "skills/ue-runtime-allocation-and-caching", + "entry": "skills/ue-runtime-allocation-and-caching/SKILL.md", + "description": "Audit runtime allocation, pooling and caching: pointers stored into container storage, unbounded caches, pool lifecycle, per-frame churn in paint and tick paths, protocol identifier discipline, and which cost axis a replication change actually moves.", + "use_when": "Profiling hitches, chasing memory that grows during a session, reviewing effect or UI churn, or deciding how to reduce replicated object count.", + "references": [ + "skills/ue-runtime-allocation-and-caching/references/patterns.md", + "skills/ue-runtime-allocation-and-caching/references/failure-modes.md" + ] + }, + { + "id": "ue-architecture-guardrails", + "path": "skills/ue-architecture-guardrails", + "entry": "skills/ue-architecture-guardrails/SKILL.md", + "description": "Cross-system review of a modular architecture: seven invariants, an ownership ledger, lifecycle as a transaction, context keys, data-graph validation, temporal dependencies, observability, the false-genericity test, and a risk-based integration matrix with a go/no-go rule.", + "use_when": "Before adopting a modular architecture, at design review, before enabling runtime feature switching, during an engine upgrade, or when indirect systems have become hard to debug.", + "references": [ + "skills/ue-architecture-guardrails/references/patterns.md", + "skills/ue-architecture-guardrails/references/failure-modes.md" + ] + }, + { + "id": "ue-evidence-discipline", + "path": "skills/ue-evidence-discipline", + "entry": "skills/ue-evidence-discipline/SKILL.md", + "description": "How to produce and check claims about a codebase so they can be trusted later: separating measured from derived from open, why an empty search proves the pattern rather than the absence, verifying a headline claim yourself, treating a delegated report as raw material, and only publishing recipes that have been run.", + "use_when": "Auditing a codebase, reviewing someone else's audit, adopting a reference project, or writing any document that will be quoted back at you. Read this before the other skills in this bundle.", + "references": [ + "skills/ue-evidence-discipline/references/failure-modes.md" + ] + } + ] +} diff --git a/plugins/ue-design-skills/skills/ue-architecture-guardrails/SKILL.md b/plugins/ue-design-skills/skills/ue-architecture-guardrails/SKILL.md new file mode 100644 index 0000000..a5afc83 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-architecture-guardrails/SKILL.md @@ -0,0 +1,319 @@ +--- +name: ue-architecture-guardrails +description: >- + Audit an Unreal Engine architecture for systemic failure modes across + data-driven definitions, feature plugins, ability systems, input, UI, + messaging, cosmetics and teams, settings and editor workflows. Use before + adopting a modular architecture, at design review, before enabling runtime + feature switching, during engine upgrades, or when indirect systems have + become hard to debug. Provides ownership, lifecycle, context, validation, + observability and integration-test gates. +--- + +# UE architecture guardrails + +This skill is for **cross-system review**. It does not cover the implementation +of any one subsystem — each of those has its own skill and its own detection +recipes, and this one tells you which of them to run and in what order. + +Systemic patterns and the go/no-go review: [patterns](references/patterns.md). +Fifteen cross-boundary failure modes: [failure modes](references/failure-modes.md). + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +The invariant behind all of it: + +> Every systemic failure in a modular architecture lives **between** two +> subsystems that are each individually correct. That is why no subsystem's +> tests find it, and why review by subsystem cannot either. + +--- + +## The seven invariants + +A modular, data-driven architecture is healthy only when all seven hold: + +1. **Justified abstraction** — every layer pays for real, named variability. +2. **Explicit ownership** — every dynamic resource has an owner, a receipt and + an inverse. +3. **Symmetric lifecycle** — activate/deactivate and create/destroy are round + trips, tested as round trips. +4. **Explicit context** — world, local player, authority and activation context + are never inferred from globals. +5. **Compiled data graph** — semantic joins, references and bundles are + validated by something that can fail a build. +6. **Observable indirection** — composition, blockers and listeners can be + dumped at runtime. +7. **Cross-boundary tests** — lifecycle is tested across subsystem seams, not + inside one manager. + +**If a design review cannot show evidence for an invariant, treat it as absent.** +Not "probably fine" — absent. Every finding in the failure modes was in a +codebase where the invariant was assumed rather than demonstrated. + +--- + +## 1. Abstraction pressure test + +For each proposed layer, fill this in before writing code: + +| Layer | Concrete variability #1 | #2 | Operations it removes | New lifecycle and tests it adds | +|---|---|---|---|---| + +Reject or defer a layer when it has fewer than two real variants, when direct +code would be smaller and equally safe, when the project will not use dynamic +activation or a packaging boundary, or when the team cannot maintain the +validation and observability the layer requires. + +**Do not adopt a large-studio architecture as an identity marker.** The cost is +paid per layer, per release, by whoever debugs it. Recipe: AG-01. + +--- + +## 2. Ownership audit + +Build a ledger for every dynamically added resource: + +```text +resource · owner · scope · receipt · apply · revoke · sharing policy · rollback +``` + +Red flags, each of which has been observed in shipping code: + +- a weak pointer presented as ownership; +- a global subsystem that "owns everything"; +- cleanup that clears **all** resources rather than owned ones; +- a handle discarded immediately after registration; +- an owner that can die before the resource; +- two owners that can both remove one shared resource. + +### The kill-owner test + +Destroy each owner at the worst moment: during an async load, after apply but +before deactivation, after the target actor is gone, during map travel. Every +owned resource must disappear or transfer explicitly. + +Subsystem recipes: `ue-modular-gameplay` MG-01 to MG-03, +`ue-input-architecture` IN-03, `ue-gas-architecture` GA-01. Recipe: AG-02. + +--- + +## 3. Lifecycle as a transaction + +```text +validate → load → activate 1..N → committed +failure at K → reverse K-1..1 → failed, with a reason +``` + +Three tests, all mandatory: + +```text +baseline → activate → deactivate → baseline +baseline → activate → fail each step → baseline +baseline → activate → deactivate → reactivate → exactly one copy +``` + +The third is the one that finds duplicates, and it is the one nobody runs during +development, because in development you restart the editor instead. Recipe: +AG-03. + +--- + +## 4. Context audit + +For every mutable map and every subscription, ask which key is actually +required: process, game instance, world handle, activation context, local +player, actor, or net role. **A global key is valid only when the behaviour is +genuinely global.** + +Required scenario matrix: + +| Scenario | What it catches | +|---|---| +| multi-world editor play | process-global leaks | +| dedicated server + client | local-broadcast and presentation assumptions | +| listen server | double execution of local paths | +| two local players | UI, input and settings context leaks | +| map travel | stale listeners and state on long-lived scopes | + +Recipes: AG-04, plus `ue-game-settings-architecture` GS-03 and +`ue-gameplay-messaging` MB-08. + +--- + +## 5. Compile the data graph + +Generate and validate the whole chain: production roots → identifiers → +references → plugin boundaries → bundles → semantic joins → leaf assets. + +Failures that must fail a build, not a log line: unresolved identifiers, null +required fields, plugin dependency cycles, role-mismatched bundles, orphaned +producers or consumers of a semantic tag, prototype content under a production +root, an input tag with no granted consumer, a UI contribution with no matching +point, an incompatible payload family on a parent channel. + +**Editor-only validation is feedback, not a gate.** Recipes: AG-05, plus +`ue-data-driven-architecture` DD-09 and DD-10. + +--- + +## 6. Temporal dependency audit + +Draw initialization prerequisites as a directed graph and reject: any cycle; a +prerequisite produced only by an optional or disabled path; an anonymous +next-tick delay; a generic "extension added" event used where semantic readiness +is required; a transition that can block with no diagnostic reason. + +Then permute: feature activation before and after actor spawn, data before and +after possession, replication arrival order, contribution before and after its +point, listener before and after event. **Correct systems converge to the same +state.** Recipes: AG-06, plus `ue-modular-gameplay` MG-14 and MG-15. + +--- + +## 7. Observability requirements + +At minimum, expose read-only dumps for: the selected composition and its load +state; pending assets, plugins and actions with reasons; active feature plugins +and who requires them; per-actor readiness state and blockers; active input +contexts and semantic joins; granted capabilities by owner receipt; UI layers, +points and contributions with owners; message channels and listener counts; +authoritative identity and its mirrors; requested versus effective settings. + +Logs need correlation fields: world, local player, actor, composition, plugin, +action, channel, owner. + +**If diagnosis requires stepping through many assets with no state dump, the +indirection is operationally incomplete** — the architecture works and cannot be +supported. Recipe: AG-07. + +--- + +## 8. The false-genericity test + +For every public field or hook that implies behaviour — an ordering value, a +viewer parameter, an automatic mode, a removal function, a policy suffix: + +1. find its consumer; +2. mutate the value in a fixture; +3. assert an observable delta; +4. test both the native and the scripting API where both are exposed. + +No consumer and no delta means: implement it, remove it, or mark it +unsupported. **Never let a configurable no-op survive as stable API.** + +This single test has the highest yield of anything in this skill. In the audited +material it independently found an ordering field that never sorts, a viewer +parameter that is ignored, an automatic quality pipeline with no caller, a +removal function with an empty body, and a composition part with no writer. +Recipe: AG-08, plus `ue-ui-architecture` UI-01, `ue-cosmetics-and-teams` CT-10, +`ue-game-settings-architecture` GS-06 to GS-08. + +--- + +## 9. Event, state and command + +Classify every interaction: + +- **state** — a queryable authoritative model; +- **command** — one accountable service with a result and a failure; +- **event** — zero or many optional observers, a past-tense fact. + +Red flags: UI reconstructing state from transient messages; a message asking an +unknown listener to perform required work; a local bus assumed to replicate; +presentation treated as authority. Recipes: AG-09, plus +`ue-gameplay-messaging` MB-12 to MB-14. + +--- + +## 10. Version and adoption audit + +Before copying any framework or sample, classify each element: architectural +invariant, sample scaffolding, project-specific behaviour, incomplete path, +confirmed defect, prototype content, version-sensitive API. Then pin the engine +version, the sample version, plugin versions and any serialized schema version, +and re-run behaviour probes after every upgrade. + +**A current documentation page does not prove the content you copied is +current.** See `ue-reference-project-adoption` for the full classification and +its four tests. Recipe: AG-10. + +--- + +## 11. Risk-based integration matrix + +Do not attempt a Cartesian product. Require this spine: + +1. standalone baseline; +2. dedicated server + remote client; +3. listen server; +4. two local players; +5. feature activates **before** receiver readiness; +6. feature activates **after** receiver readiness; +7. deactivate and reactivate; +8. actor destroyed before feature teardown; +9. late join; +10. map travel or composition change; +11. packaged clean client; +12. engine or plugin upgrade fixture. + +**Every critical invariant needs at least one test crossing two subsystem +boundaries.** Unit tests inside one manager cannot find anything in this skill. + +--- + +## 12. Severity triage + +**Critical — stop feature work.** Ownership or rollback unknown; authority +duplicated; world or local-player context leaked; production data graph +unvalidated; persistent state reconstructed from messages; client presentation +affecting authority. + +**High — fix before runtime switching or shipping.** Initialization cycle or +timing dependency; missing bundle with no fallback; native/scripting semantic +mismatch; undocumented ordering dependency; insufficient observability; version +mismatch. + +**Medium — track with a bound.** Abstraction tax; linear scans acceptable at +current data size; editor-only gaps; duplicated scaffolding. + +--- + +## 13. Go/no-go + +- [ ] Every layer has two named variability pressures. +- [ ] Ownership ledger complete, kill-owner test passes. +- [ ] Lifecycle rollback and reactivation tests pass. +- [ ] Context matrix passes. +- [ ] Data graph validation runs where it can fail a build. +- [ ] Prerequisite graph acyclic and observable. +- [ ] State dumps and log correlation exist. +- [ ] Every public generic hook has a consumer test. +- [ ] Event, state and command classifications are correct. +- [ ] Replicated intent has a packaged-client realization test. +- [ ] Version set pinned. +- [ ] Cross-boundary matrix automated or scheduled. + +**No-go rule: three unchecked Critical gates mean stop adding abstraction.** +Repair ownership, context and validation first — further indirection multiplies +the failure surface rather than the capability. + +--- + +## Provenance + +The systemic patterns and their severities come from a cross-system audit of +Epic's Lyra Starter Game on Unreal Engine 5.6 — the same audit that produced the +subsystem skills in this bundle. This skill exists because the same shapes +appeared in every subsystem: the individual findings are in those skills, and +what is here is the structure they share. + +Source addresses stay in the research archive. Each entry carries a stable +identifier (`AG-01`, `AG-02`, …) resolving back to the audited material there. + +## Evidence boundary + +Cross-system conclusions from one project at one engine version, read as source +rather than run. Where a finding is a general principle rather than a measured +instance, the failure modes say so explicitly. The subsystem recipes referenced +above are the measured half; this skill is the ordering. diff --git a/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/failure-modes.md new file mode 100644 index 0000000..ada8f6f --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/failure-modes.md @@ -0,0 +1,379 @@ +# Failure modes: cross-system architecture + +Ten failures that belong to **no single subsystem**, which is why each one +survived review of every subsystem it passes through. + +The shared property here is different from the other documents in this bundle. +Elsewhere the mechanism is "a check exists and cannot fire". Here it is: **every +individual component is correct, and the defect lives in the seam.** Nobody owns +a seam. A reviewer of the messaging layer sees a correct bus; a reviewer of the +UI sees a correct widget; the failure is that the UI reconstructs state from the +bus, and that fact is written down nowhere. + +That has a practical consequence for how these are detected. Most recipes below +are not a single search — they are a **join between two searches**, and the +finding is in the difference between the two result sets. + +Where a mechanism is fully described by another skill, the entry says so and does +not restate it. Identifiers (`AG-01` and up) are stable and resolve back to the +audited material in the research archive. + +--- + +## Abstraction + +### AG-01 - A layer adopted for completeness rather than pressure + +**Mechanism.** A modular architecture is copied whole from a reference project +because it is the reference's architecture, without asking which variability each +layer absorbs. Every layer is correctly implemented and paid for in full. + +**Why it is silent.** Nothing fails. The cost is not a defect but a permanent tax: +every new mechanic now needs a class, a definition asset, a tag, an action, a +plugin, a validation rule and a UI extension. The team experiences this as "the +engine is like this". + +**Why the obvious check misses it.** Review asks "is this implemented correctly?" +and the answer is yes at every layer. The question that finds it — "what are the +two concrete variants this layer separates?" — is not part of any code review, +because it is not about code. + +**Symptom.** Feature velocity that falls as the project matures, with no single +slow component. Estimates that are consistently wrong in the same direction. + +**Detect.** Count the axes against the variants they serve: + +```bash +rg -c "class \w+ : public UGameFeatureAction" --glob "*.h" . # action types +fd -e uasset -p "Experiences|GameFeature" | wc -l # variants shipped +``` + +The finding is arithmetic, not textual: **planned variants fewer than +architectural axes.** In the audited reference, five feature plugins out of +eighty-one total carried the modular machinery — which pays for that project and +would not pay for a single-mode title. + +**Guardrail.** For each layer, name two real variabilities it separates before +adopting it. If you cannot, defer the layer. See `ue-reference-project-adoption` +for the full classification and the adoption budget. + +--- + +### AG-02 - Distributed control flow with no trace + +**Mechanism.** A request travels definition → feature → action → extension event → +component → tag → message → widget. Each hop is decoupled by design. + +**Why it is silent.** Decoupling is working exactly as intended. The absence of a +call stack is the feature, not a defect. + +**Why the obvious check misses it.** A debugger shows one hop. Each hop's owner +can explain their hop. Nobody can explain the path, because reconstructing it +requires reading assets, config and code in three modules — and it has to be +reconstructed again next time. + +**Symptom.** "Who triggered this?" costs an afternoon. Bugs reproduce reliably +while no class looks responsible. New engineers take months rather than weeks. + +**Detect.** Test the property directly rather than searching for it: take a +recent bug and ask an engineer who did not write the feature to reconstruct the +path **from logs and dumps alone**, without opening assets. Then measure what +they needed and did not have. + +Structurally, the precondition is visible: + +```bash +rg -c "GetSubsystem<|BroadcastMessage|SendGameFrameworkComponentExtensionEvent" --glob "*.cpp" . +rg -c "UE_LOG.*Verbose.*(World|LocalPlayer|Experience|Feature)" --glob "*.cpp" . +``` + +A large first number with a small second is the finding: heavy indirection, no +correlated logging. + +**Guardrail.** Structured lifecycle logs carrying world, player, experience, +plugin, action and owner identifiers; a generated composition graph; one index of +semantic tags and their consumers. Indirection without observability is not +architecture, it is a maze. + +--- + +## Ownership and lifecycle + +### AG-03 - Activation reviewed, deactivation assumed + +**Mechanism.** Adding a component, a binding, a widget or a grant is easy and +visible. The inverse operation is written from memory, or not at all. + +**Why it is silent.** Development restarts the editor instead of deactivating. +The happy path is exercised hundreds of times a day; the reverse path is +exercised by nobody until a player switches modes. + +**Why the obvious check misses it.** Both functions usually exist and look +symmetric. The asymmetry is in what each enumerates — see `ue-modular-gameplay` +recipes MG-01 through MG-06 for the six distinct shapes this takes, each with its +own recipe. + +**Symptom.** The second activation produces duplicates. Reported long after the +first, and usually attributed to the feature that was activated second. + +**Detect.** The behavioural test is stronger than any search, and it is the one +gate this whole skill exists to insist on: + +```text +baseline → activate → verify additions → deactivate → verify baseline + → reactivate → verify exactly one copy +``` + +Run it with the actor existing before activation, spawned after activation, and +destroyed before deactivation. + +**Guardrail.** An ownership ledger — resource, receipt, inverse — filled in +**before** implementation. A blank middle column is a rejected design, not a +follow-up task. + +--- + +### AG-04 - A subsystem's lifetime mistaken for resource ownership + +**Mechanism.** A long-lived subsystem holds registrations for resources owned by +short-lived objects. Weak references prevent crashes, so nothing appears wrong. + +**Why it is silent.** Weak pointers do their job: the dead object is not called. +The **record** remains, and records are not visible in any profiler view that +answers "is this leaking?". + +**Why the obvious check misses it.** The code is defensively written and looks +careful. Weak references read as evidence that ownership was considered — when +they are precisely the mechanism that lets the bookkeeping rot silently. + +**Symptom.** Registration lists that grow across a session; cleanup that resorts +to clearing everything because per-owner removal was never possible; the same +resource removed twice by two owners. + +**Detect.** For every registry, compare adds against removes and check the key: + +```bash +rg -n "\.Add\(|\.Emplace\(|\.FindOrAdd\(" --glob "*.cpp" . | rg -i "listener|extension|handle|request" +rg -n "\.Remove\(|\.RemoveSwap\(|Unregister" --glob "*.cpp" . | rg -i "listener|extension|handle|request" +``` + +A registry with adds and no owner-keyed removal is the finding. A registry whose +only removal is a full clear is the same finding, one step later. + +**Guardrail.** Every dynamic resource has exactly one named owner and one receipt. +Shared resources use reference counts or leases. "The subsystem owns it" is not an +answer; subsystems outlive the things they track. + +--- + +## Context + +### AG-05 - Global state where the scope is world or player + +**Mechanism.** A registration, cache or setting is keyed globally, while the +things it describes belong to a world, a local player or an activation context. + +**Why it is silent.** With one world and one player — the configuration in which +almost all testing happens — global and scoped are indistinguishable. The code is +correct in the case you run. + +**Why the obvious check misses it.** The accessor reads naturally: a player asking +for its own settings, a subsystem holding its own registry. The scope error is one +line inside an accessor, or a missing key in a map declaration. + +**Symptom.** Multi-world editor sessions cross-contaminate; split-screen players +share what should be per-player; a listen server processes an event twice. Each +appears as an unrelated bug in a different subsystem. + +**Detect.** For every mutable registry, ask what the minimum sufficient key is, +then check what it actually is: + +```bash +rg -n "TMap<.*>\s+\w+;" --glob "*.h" . | rg -v "FObjectKey|FGameFeatureStateChangeContext|ULocalPlayer" +rg -n -A4 "::Get\w*Settings\(\)" --glob "*.cpp" . | rg "::Get\(\)|GEngine->" +``` + +The second search is the specific case documented in +`ue-game-settings-architecture` GS-03: a per-player accessor returning a global +singleton, whose tell is a proliferation of "primary player only" conditions +elsewhere. + +**Guardrail.** Key every mutable record by the minimum context that makes it +correct — activation context, world handle, local player where applicable. Then +run the required matrix: multi-world editor, dedicated server plus client, listen +server, two local players, map travel. + +--- + +### AG-06 - Editor behaviour that differs from the shipped configuration + +**Mechanism.** A branch keyed on running in the editor loads more, validates less, +or guesses identifiers that the packaged build resolves strictly. + +**Why it is silent.** Both branches are correct for their environment. The editor +branch is usually more permissive, so everything works better where you are +looking. + +**Why the obvious check misses it.** The branch is a single condition in a +subsystem nobody reads while working on a feature. Its consequences appear in +memory profiles, cook results and packaged-only failures — three places that are +each somebody else's job. + +**Symptom.** Memory numbers that describe no shipping configuration. Features that +work in the editor and silently do nothing when packaged. Asset identifiers that +resolve in one and not the other. + +**Detect.** Enumerate every editor divergence and judge each one deliberately: + +```bash +rg -n "GIsEditor|WITH_EDITOR|IsRunningCommandlet|GIsPlayInEditorWorld" --glob "*.cpp" . -A3 +rg -n "bShouldGuessTypeAndNameInEditor|PreloadInEditor|bOnlyCookProduction" Config/ +``` + +In the audited reference this finds an editor branch that loads **both** role +bundles — which alone invalidates in-editor residency measurement — and a +configuration that guesses asset identifiers in the editor and not in the build. + +**Guardrail.** Keep a written list of editor divergences and their justification. +Any measurement taken in the editor states which divergences apply to it. +See `ue-asset-loading-and-memory` AL-10. + +--- + +## Data and validation + +### AG-07 - A data graph with no compiler + +**Mechanism.** Null references, wrong identifiers, cross-plugin cycles, prototype +content and semantic mismatches are all valid data. They load, they cook, they run. + +**Why it is silent.** Data does not compile. There is no stage that can reject it +except one somebody chose to write — and that validation is typically +editor-only, so it does not run where it would matter. + +**Why the obvious check misses it.** The validation *exists*, which satisfies the +question "is the data validated?". What it does not do is run in the build. In the +audited reference, eleven of twelve validation implementations were compiled out +of non-editor builds. + +**Symptom.** A production playlist pointing at a test mode; a health pickup +granting a weapon definition; a missing bundle discovered only in a packaged +build. + +**Detect.** Count validation, then count how much of it survives the build: + +```bash +rg -c "IsDataValid" --glob "*.cpp" . +rg -n -B6 "IsDataValid" --glob "*.cpp" . | rg -c "WITH_EDITOR" +rg -n "class \w*ValidationCommandlet|UEditorValidatorBase" --glob "*.h" . +``` + +A ratio close to one, with no commandlet or automation path, means the data graph +is unvalidated where it ships. See `ue-data-driven-architecture` DD-09 and DD-10. + +**Guardrail.** Run the same validation in automation that you run in the editor. +Add a production-root allow list so prototype paths cannot reach a shipped +playlist. + +--- + +### AG-08 - A generic hook with no consumer + +**Mechanism.** A public field or extension point is stored, copied and threaded +through an API — and never read by anything that changes behaviour. + +**Why it is silent.** Every "is this used?" check answers yes, because the value +*is* used: passed, assigned, copied. What is missing is the comparison, the +branch, or the sort. + +**Why the obvious check misses it.** This is the single most repeated shape in +this whole bundle, and it earns its own cross-system entry because it recurs in +every subsystem independently: an ordering field that never sorts, a viewer +identity that is ignored, a benchmark decision with no caller, a profile suffix +that is never populated, a removal function with an empty body, a flag written in +a constructor and never read. + +**Symptom.** A designer configures a documented setting and observes no effect, +concludes their data is wrong, and works around it. + +**Detect.** The general form — mentions minus comparisons: + +```bash +F='Priority' +rg -c "\b$F\b" --glob "*.cpp" --glob "*.h" . # mentions +rg -n "\b$F\b\s*(<|>|<=|>=|==)|Sort.*\b$F\b|\b$F\b.*Sort" --glob "*.cpp" . +``` + +Mentions without comparisons is the finding. Per-subsystem instances have their +own recipes: `ue-ui-architecture` UI-01, `ue-cosmetics-and-teams` CT-10, +`ue-game-settings-architecture` GS-06 and GS-07, `ue-gas-architecture` GA-06, +`ue-modular-gameplay` MG-01. + +**Guardrail.** Every public setting gets a consumer test: mutate it in a fixture, +assert an observable delta. Anything without one is implemented, removed, or +marked unsupported — never left as configurable decoration. + +--- + +## Verification + +### AG-09 - Single-process testing of a distributed property + +**Mechanism.** A listen-server host shares memory with its client, so the +client/server split that produces a whole class of defects does not exist during +the test that would have caught it. + +**Why it is silent.** The tests pass. They are real tests exercising real code; +they simply cannot express the failure. + +**Why the obvious check misses it.** Coverage looks good and the feature demonstrably +works. The missing dimension is a **configuration**, not a code path, so no +coverage tool reports it. + +**Symptom.** Features that ship having never run in the configuration they will +run in. Bugs that appear at first playtest and are attributed to the network layer. + +**Detect.** This one is answered by an inventory rather than a search: list the +failure modes that require a separate process, and confirm each has a test in a +configuration that has one. From this bundle, the ones that cannot occur in a +single process include replicated-versus-validated confusion, call-site authority +guards, multicast used where state belongs, incomplete replicated-array callbacks, +readiness gates that assume a controller, and prediction with no correction path +(`ue-multiplayer-authority` NA-01, NA-04, NA-11, NA-13, NA-17, NA-19). + +**Guardrail.** A dedicated server with two remote clients, one under latency and +packet loss, as the **minimum** configuration for accepting a networked feature. +Plus a mid-match joiner. + +--- + +### AG-10 - A version set that drifts silently + +**Mechanism.** Engine, reference sample and plugins evolve independently. Code +copied from a sample at one version keeps running against another. + +**Why it is silent.** Compilation succeeds. Serialized fields still load. A stub +that changed behaviour between versions still returns something plausible. + +**Why the obvious check misses it.** Documentation for the current version is +easy to find and describes the current version — not the vendored copy in the +project. Reading the docs actively produces false confidence. + +**Symptom.** A method that "works differently now"; a structure size assertion +that fails after an upgrade; content that disagrees with the plugin that reads it. + +**Detect.** Pin and verify rather than search. Where code manually enumerates the +members of an engine structure, a size assertion is the correct tripwire: + +```bash +rg -n "static_assert\(sizeof\(" --glob "*.cpp" --glob "*.h" . +``` + +In the audited reference this appears five times against one engine structure. A +failing assertion after an upgrade is a **feature**: it means a new field would +otherwise have been silently ignored by five hand-written functions. + +**Guardrail.** Pin the engine, sample and plugin versions. Keep compile-time +guards where you enumerate engine structures by hand. Re-run structural and +behavioural probes after every upgrade, and never treat a documentation page as +evidence about the code in your tree. diff --git a/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/patterns.md b/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/patterns.md new file mode 100644 index 0000000..5473eb8 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-architecture-guardrails/references/patterns.md @@ -0,0 +1,330 @@ +# Patterns: reviewing an architecture across its seams + +This is the cross-system half of the bundle. The other fifteen skills each audit +one subsystem; this one exists because the most expensive failures are not inside +a subsystem at all. + +[Failure modes](failure-modes.md) carries the detection recipes. This file +carries the review procedure — what to ask, in what order, and what evidence to +demand before believing an answer. + +--- + +## 1. Why seams are where the cost is + +A subsystem has an owner. Someone can be asked whether the ability system is +correct, and they can answer. + +A seam has no owner. "The UI reconstructs state from the message bus" is not a +fact about the UI or about the bus; it is a fact about their relationship, and +relationships do not appear in any file. Every finding in this skill was reviewed +— twice, once from each side — and survived. + +That produces the review technique this whole document is built on: + +> **Ask each question of the pair, not of the component.** Not "is this registry +> correct?" but "which context is this registry keyed by, and which context do its +> entries belong to?" The defect lives in the difference. + +--- + +## 2. The seven invariants + +An architecture of this kind is healthy only when all seven hold: + +1. **Justified abstraction** — every layer pays for real variability. +2. **Explicit ownership** — every dynamic resource has an owner, a receipt and an + inverse. +3. **Symmetric lifecycle** — activate/deactivate and create/destroy are round + trips, tested as round trips. +4. **Explicit context** — world, local player, authority and activation context + are never inferred from globals. +5. **Compiled data graph** — semantic joins, references and bundles are validated + somewhere that runs in the build. +6. **Observable indirection** — composition, blockers and listeners can be dumped. +7. **Cross-boundary tests** — lifecycle is tested across subsystem seams. + +**If a review cannot produce evidence for an invariant, record it as absent.** Not +"probably fine" — absent. Every finding in the audit that produced this skill was +in a system whose author would have said it was fine. + +--- + +## 3. Abstraction pressure test + +Fill this before adopting a layer, not after: + +| Layer | Variability #1 | Variability #2 | Operations removed | Lifecycle and tests added | +|---|---|---|---|---| +| feature plugin | | | | | +| experience / mode definition | | | | | +| semantic input | | | | | +| UI extension points | | | | | +| message bus | | | | | + +Defer a layer when it has fewer than two real variants or owners; when direct code +would be smaller and equally safe; when the project will not use dynamic +activation or a packaging boundary; or when the team cannot maintain the +validation and observability the layer requires. + +**The arithmetic that decides it:** planned variants versus architectural axes. If +axes exceed variants, the architecture is aspirational. Recipe: AG-01. + +--- + +## 4. The ownership ledger + +For every dynamically added resource: + +```text +resource +owner +scope process / game instance / world / local player / actor +receipt the handle that proves this owner added it +apply the call that adds +revoke the call that removes exactly this +sharing what happens when two owners want the same resource +rollback what happens if apply fails halfway +``` + +Red flags, each seen in the audited reference: + +- a weak reference presented as ownership; +- a global subsystem that "owns everything"; +- cleanup that clears all resources rather than owned ones; +- a handle discarded at the moment it is returned; +- an owner that can die before its resource; +- two owners able to remove the same shared resource. + +### The kill-owner test + +For each owner, destroy it at the worst moment: during an async load, after apply +but before normal deactivation, after the target actor is destroyed, and during +map travel. Every owned resource must disappear or transfer explicitly. + +This is a cheap test that finds expensive defects, and almost nobody runs it. + +--- + +## 5. Lifecycle as a transaction + +```text +validate → load → activate 1..N → committed + +failure at step K + → reverse K-1 .. 1 + → failed, with a reason that names the step +``` + +Three mandatory sequences: + +```text +baseline → activate → deactivate → baseline +baseline → activate → fail each step in turn → baseline +baseline → activate → deactivate → reactivate → exactly one copy +``` + +The third is the one that finds duplicates, and the one development never runs +because development restarts the editor instead. Recipe: AG-03. + +--- + +## 6. Context audit + +For every mutable map or subscription, ask which key is **required**: process, +game instance, world handle, activation context, local player, actor, or net role. + +A global key is correct only when the behaviour is genuinely global. The +temptation is that a global key is simpler and, with one world and one player, +indistinguishable. + +### The required matrix + +| Scenario | What it catches | +|---|---| +| multi-world editor session | process-global leaks | +| dedicated server plus remote client | local-broadcast and presentation assumptions | +| listen server | doubled local paths | +| two local players | UI, input and settings context leaks | +| map travel | stale listeners surviving a world | + +Recipes: AG-05, AG-06. + +--- + +## 7. The data graph needs a compiler + +Generate and validate the path from production roots through identifiers, +references, plugin boundaries, bundles and semantic joins to leaf assets. + +Fail the build on: unresolved identifiers or soft references; null required +fields; plugin dependency cycles; role-bundle mismatches; orphan producers or +consumers of a semantic tag; prototype or test paths reachable from a production +root; an input tag with no granted consumer; a UI extension with no point or +contract; an incompatible payload family under a parent channel. + +**The crucial property is where it runs.** Editor-only validation is useful +feedback and is not a safety boundary. Recipe: AG-07. + +--- + +## 8. Temporal dependencies + +Draw initialization prerequisites as a directed graph — each feature state and +what it requires. Reject: cycles; a prerequisite produced only by an optional or +disabled path; an anonymous next-tick delay; a generic "extension added" event +used where semantic readiness is meant; a blocked transition with no diagnostic. + +### Permutation test + +Vary activation before and after actor spawn, archetype data before and after +possession, replication arrival order, UI point before and after extension, and +listener before and after event. **Correct systems converge to the same state +from every order.** Systems that pass by luck have one order that works. + +--- + +## 9. Observability is a requirement, not tooling polish + +Expose read-only dumps for: the selected mode and its load state; pending assets, +plugins and actions **with reasons**; active feature plugins and who requires +them; per-feature actor readiness and blockers; active input contexts and their +semantic joins; granted capabilities by owner receipt; UI layers, points, +extensions and owners; message channels with listener counts and owners; the +authoritative team identifier and its mirrors; requested versus effective +settings with clamps. + +Logs carry correlation fields: world, local player, actor, experience, plugin, +action, channel, owner. + +**The acceptance test:** an engineer who did not build the feature diagnoses an +injected failure from logs and dumps alone, without opening assets, within a +bounded time. If that is impossible, the indirection is operationally incomplete +regardless of how clean the code is. Recipe: AG-02. + +One thing worth stealing outright: the audited reference ships console variables +that inject artificial delay into mode loading. A reproducible way to exercise +slow-load, timeout and cancellation paths is worth more than it costs, and almost +nobody builds one. + +--- + +## 10. The false-genericity sweep + +For every public field or hook implying behaviour — an ordering value, a viewer +identity, an automatic-benchmark decision, a removal function, a profile suffix, +a policy flag: + +1. find its consumer; +2. mutate the value in a fixture; +3. assert an observable delta; +4. test both native and visual-script surfaces where exposed. + +No consumer or no delta means: implement it, remove it, or mark it explicitly +unsupported. Never let a configurable no-op survive as stable API. + +This sweep is worth running as a scheduled activity rather than a review step, +because the shape recurs independently in every subsystem. Recipe: AG-08. + +--- + +## 11. Event, state and command + +Classify every interaction: + +- **state** — a queryable authoritative model; +- **command** — one accountable handler, with a result and a failure; +- **event** — a past-tense fact with zero or many optional observers. + +Red flags: UI reconstructing state from transient messages; a message asking an +unknown listener to perform required work; a local bus assumed to replicate; +presentation treated as authority. + +The test that settles it: **if a listener subscribes one second late, must it know +the current value?** Yes means state. No means an event is acceptable. + +--- + +## 12. Version and adoption + +Before copying from a reference, classify each element: architectural invariant +(adopt), sample scaffolding (replace), project-specific behaviour (evaluate), +unfinished path (do not claim supported), confirmed defect (fix with a regression +test), prototype content (exclude from production), version-sensitive API (verify +live). + +Pin the engine version and changelist, the sample version, plugin versions, and +any serialized schema version. Re-run probes after every upgrade. + +**A current documentation page is not evidence about the code in your tree.** +Recipe: AG-10. The full classification is `ue-reference-project-adoption`. + +--- + +## 13. The integration matrix + +Do not attempt the full product of configurations. Require this spine: + +1. standalone baseline; +2. dedicated server plus remote client; +3. listen server; +4. two local players; +5. feature activates **before** receiver readiness; +6. feature activates **after** receiver readiness; +7. deactivate and reactivate; +8. actor destroyed before feature teardown; +9. late join; +10. map travel or mode change; +11. packaged clean client; +12. engine or plugin upgrade fixture. + +Every critical invariant needs at least one test crossing **two** subsystem +boundaries. Unit tests inside one manager cannot express these failures. +Recipe: AG-09. + +--- + +## 14. Triage + +**Critical — stop feature work.** Resource ownership or rollback unknown; +authority duplicated; world or player context leaking; production data graph +unvalidated; persistent state reconstructed from messages; client presentation +affecting authority. + +**High — fix before runtime switching or shipping.** Initialization cycles or +timing dependencies; missing packaged-asset fallback; native and visual-script +semantics diverging; undocumented ordering dependencies; insufficient +observability; version mismatch. + +**Medium — track with a bound.** Abstraction tax; linear scans acceptable at +current data size; editor-only polish gaps; duplicated sample scaffolding. + +### The no-go rule + +Three unchecked Critical gates mean **stop adding abstraction and features**. +Repair ownership, context and validation first. More indirection on an +unvalidated base multiplies the failure surface rather than organising it. + +--- + +## Provenance + +The invariants and gates here are generalized from a full-project audit of Epic's +Lyra Starter Game on Unreal Engine 5.6 — its data-driven layer, feature plugins, +ability system, input, UI, messaging, cosmetics and teams, settings, loading and +streaming — read as source and configuration rather than run. + +Every cross-system failure mode in [failure modes](failure-modes.md) was observed +in that project, in a subsystem whose local implementation was correct. That is +the reason this skill exists as a separate document rather than as a section in +each of the others. + +Source addresses stay in the research archive that produced this skill. Entry +identifiers resolve back to the audited locations, so any specific claim can be +produced on request. + +## Evidence boundary + +One project, one engine version, one workspace. The invariants are transferable; +the specific instances are evidence, not guarantees about other versions. Re-run +the detection recipes against your own tree before acting on any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/SKILL.md b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/SKILL.md new file mode 100644 index 0000000..5f4f232 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/SKILL.md @@ -0,0 +1,298 @@ +--- +name: ue-asset-loading-and-memory +description: >- + Design or audit asset loading and CPU memory residency in Unreal Engine: the + asset manager and primary asset identifiers, asset bundles and cook rules, + hard versus soft versus weak reference semantics, streamable handle ownership, + synchronous load hitches, startup job sequencing, garbage-collection + reachability, retention leaks and unload paths. Use when reducing load times + or memory, choosing reference types, adding content-pulling definitions, + diagnosing hitches, or investigating objects that are never collected. +--- + +# UE asset loading and CPU memory + +The invariant: + +> An asset stays in memory because something reachable **owns** it. Loading is a +> policy decision; residency is an ownership fact. They are set by different code +> and must be reviewed separately. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes: [failure modes](references/failure-modes.md). + +Related skills: `ue-data-driven-architecture`, `ue-modular-gameplay`, +`ue-streaming-and-platform-budgets`, `ue-runtime-allocation-and-caching`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Separate the five budgets before optimizing anything + +| Budget | What it is | Measured by | +|---|---|---| +| A. Disk / cook | package bytes shipped | build size, chunk sizes | +| B. CPU memory | object graph plus bulk data | memory report, object list | +| C. GPU memory | live render resources | platform GPU profiler | +| D. Streaming pool | a **limit**, not consumption | streaming stats | +| E. Net objects | actors, channels, bandwidth | network profiler | + +This skill owns A and B. C and D belong to +`ue-streaming-and-platform-budgets`; E to `ue-runtime-allocation-and-caching`. + +**The most common analytical error is treating D as B.** A pool setting grants +permission to use memory. It never reports use. + +--- + +## 2. Reference semantics decide load cost + +Three mechanisms, routinely conflated: + +| Mechanism | Effect | +|---|---| +| reflected property, manual reference collection, a retained streamable handle | **owns** — keeps resident | +| weak pointer, object key | **observes** — keeps nothing | +| soft object or class pointer | **a path, not a reference** | + +The third line is where most reasoning fails. **A soft pointer saves nothing by +itself.** Residency is decided by where the loaded result is stored — a soft +reference resolved into a reflected property is a hard reference with extra +steps. + +### Measure the transitive closure, not the direct count + +Direct reference counts are misleading. What matters is how many packages load +when this one loads. In the audited reference the same architectural layer +spanned three orders of magnitude depending on reference type alone: an archetype +holding hard references pulled over a thousand packages, while a catalog entry +addressing by identifier pulled two. + +That single contrast is the argument for treating reference type as +architecture rather than style. Recipe: AL-01. + +### Choosing + +| Need | Reference | +|---|---| +| address before load; cook and bundle identity | primary asset identifier | +| optional, feature-scoped or async content | soft object or class pointer | +| required whenever the owner is loaded, and small | hard pointer | +| data selects an implementation class | class reference | +| owner-exclusive polymorphic fragment | instanced object | + +Rule of thumb: a definition that is a **catalog entry** references by identifier +or softly; a definition that is an **archetype** may reference hard — but its +closure must be measured and budgeted, not assumed. + +--- + +## 3. Bundles express role, not grouping + +Declare the load role beside the reference, then request bundles per runtime role +at load time. + +Gates: + +- every soft reference carries an intentional bundle tag; +- server-irrelevant presentation is client-only; +- prediction-relevant gameplay classes are client and server; +- **a missing bundle tag means "always loaded with the owner"** — decide it, do + not default into it; +- cook rules control budget **A**, never budget **B**. + +A label asset with thousands of soft references and no hard references determines +what ships and retains nothing at runtime. Do not read cook rules as residency. + +Watch for a declared bundle name that no property ever uses: it is requested at +load time, matches nothing, and looks like a working tier. Recipe: AL-02. + +--- + +## 4. Startup discipline + +### Never synchronously load a graph you have not measured + +The dominant boot hitch in a data-driven project is one synchronous load whose +hard closure is large. An elegant async bundle pipeline that runs *after* it is +decoration. + +- [ ] every synchronous load on the boot path is justified; +- [ ] its transitive closure is measured, not assumed; +- [ ] an async alternative was considered, with an explicit wait state; +- [ ] the loading screen carries a reason string; +- [ ] cancellation is distinguishable from completion. + +### Startup jobs must own handles and report truthfully + +A progress bar that never moves is a bug, not a cosmetic issue — it hides which +job is slow. Three independent layers can each break it: a macro that never +fills the handle, a throttle comparison with swapped operands, and an empty +progress receiver. Any one of them is enough. Recipe: AL-03. + +### Global data assets need bundles too + +Soft class pointers on a globally-loaded data asset with **no** bundle tags load +synchronously at first use — typically in the frame of first damage. Tag them and +preload them with the mode. Recipe: AL-04. + +--- + +## 5. Retention: what actually keeps objects alive + +Audit six owners explicitly: reflected containers on long-lived subsystems; +manual reference collection in native code; retained streamable handles; delegate +bindings capturing strong references; component and actor ownership chains; class +default objects reached through class references. + +### A convenience loader must not default to permanent + +The worst shape in this area is a helper that loads an asset and adds it to a +reflected set, with the keep-in-memory parameter **defaulting to true** and no +release API anywhere. Every casual call then synchronously loads *and permanently +roots* the asset. The soft pointer was chosen for deferred loading; the API turns +it into a process-lifetime hard reference. + +Required design: the flag is explicit at every call site or defaults to false; a +paired release exists and is used; the retained set is inspectable from a console +command; retention has an owner and a scope. Recipe: AL-05. + +### Strong keys pin object graphs — and raw pointers lose them + +```text +map one missed unregister pins the actor, + its meshes, materials, sounds, effects + +map + garbage collection cannot see the + values; the key can dangle +``` + +Both errors can coexist in one codebase. **Audit direction as well as presence.** +Prefer object keys with explicitly-owned values plus a periodic prune. Recipe: +AL-06. + +--- + +## 6. Unloading must exist before you claim modularity + +If a project cannot answer "what is released when a mode ends", its memory +profile only grows. Either implement release, or state plainly that mode changes +require map travel — that is a legitimate scope decision, and leaving it +undefined is not. + +The measurement is blunt and takes one command: count unload calls, bundle +releases and explicit collection requests across the whole project. In the +audited reference every one of those counts was **zero**. Recipe: AL-07. + +### Avoid collection ping-pong + +A forced collection at a transition boundary must not collect what the next +transition immediately needs. Options: keep transition-critical classes rooted, +defer collection past the next load, or do not force it and let the budget drive +it. + +--- + +## 7. Failure handling changes memory behaviour + +- a cancelled load must not run the success path; +- asynchronous results must be read, not discarded; +- partially loaded state must be releasable; +- waiting without a timeout converts a stall into a hang; +- a permanent loading screen with no reason string is not error handling. + +The sharpest instance: binding the cancel delegate to the **same callback** as +completion. A cancelled load then reports success, downstream code dereferences +unloaded assets, and each of those dereferences is individually silent. Recipes: +AL-08, AL-09. + +--- + +## 8. Review checklist + +### References + +- [ ] reference type chosen per lifecycle and documented in metadata; +- [ ] transitive hard closure measured for every archetype-level definition; +- [ ] no soft reference resolved into an unbounded reflected cache; +- [ ] bundle tags on all soft references, and every requested bundle is used; +- [ ] cross-plugin edges intentional and acyclic. + +### Loading + +- [ ] no unmeasured synchronous load on boot or in a hot path; +- [ ] async loads own their handles and can be cancelled; +- [ ] progress reporting is verified to move, not merely implemented; +- [ ] cancellation and failure paths are distinct from success; +- [ ] global data assets are bundled, not lazily loaded at first use. + +### Retention + +- [ ] every long-lived container has an owner, a bound and a prune policy; +- [ ] strong keys justified; object keys used where identity suffices; +- [ ] non-reflected object pointers eliminated or documented as weak by design; +- [ ] register and unregister pairs verified across teardown paths; +- [ ] a release API exists for anything deliberately kept resident. + +--- + +## 9. Measurement plan + +Static first — it is cheap and finds structural problems: + +1. build the dependency graph from the asset registry: direct hard and soft + counts, transitive hard closure per definition, cross-plugin edges; +2. rank definitions by closure and compare against their architectural layer; +3. search for synchronous loads on boot and hot paths; +4. inventory retention owners by type. + +Runtime, once a packaged build exists: memory reports before and after a mode +cycle; object listings on suspected roots — root-graph evidence, not process +memory; fifty to a hundred spawn and travel cycles looking for monotonic growth; +an allocation profiler for attribution. + +### Do not measure residency in the editor + +Editor builds commonly load **both** client and server bundles, so an in-editor +memory profile represents neither shipping role. Residency numbers require a +packaged build; use the editor for relative structural comparison only. Recipe: +AL-10. + +--- + +## 10. When to stop + +Optimize loading and residency when boot or transition time is a product problem, +a platform ceiling is real and near, growth is monotonic across a session, or a +specific closure is measurably oversized. + +Do not restructure references on aesthetics. A fourteen-package closure is not a +problem; a thousand-package synchronous closure on the boot path is. Measure +first, and **keep the measurement** so the next change can be compared against +it. + +--- + +## Provenance + +The findings come from a source and dependency-graph audit of Epic's Lyra Starter +Game on Unreal Engine 5.6. The engine's own asset manager and streamable manager +are not part of that project, so every claim about their internals is inferred +from call sites and marked as such rather than presented as measured. + +No profiler was run and no timing was captured: this material is structural by +construction, and it deliberately contains no performance numbers. What it +contains is counts — of call sites, of unload paths, of packages in a closure — +which is a different kind of evidence and a more portable one. + +Source addresses stay in the research archive. Each entry carries a stable +identifier (`AL-01` and up) resolving back to the audited location there. + +## Evidence boundary + +One project, one engine version. Binary assets were not read, so the contents of +any specific closure are structural rather than byte-measured. Re-run the recipes +against your own tree; the counts quoted show the shape of a contrast, not a +target. diff --git a/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/failure-modes.md new file mode 100644 index 0000000..656437d --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/failure-modes.md @@ -0,0 +1,389 @@ +# Failure modes: asset loading and CPU memory + +Ten ways a project loads more than it meant to, keeps it forever, and reports +nothing. + +The shared property here is different from the other skills in this bundle, and +it is worth stating because it changes what the recipes look like: **memory +defects are not events.** There is no moment at which something goes wrong. An +asset is loaded because a reference exists; it stays because a reference still +exists. Both are correct behaviour at every instant, and the failure is a +property of an aggregate that no single frame contains. + +That has two consequences for detection. First, **counting is the primary +instrument** — call sites, unload paths, packages in a closure. Second, a count of +**zero** is often the finding: no release API, no unload call, no bundle tag. +Several recipes below are therefore searches you expect to come back empty, and +the empty result is the evidence rather than the absence of it. + +Recipes use `rg` from a project source root, and were executed against the +audited project while this file was written. + +--- + +## References + +### AL-01 - A reference type chosen by convenience, costing three orders of magnitude + +**Mechanism.** A definition holds hard references to the things it selects — a +class, a data asset, a package of capabilities — because hard references are +simpler and always resolve. Loading the definition therefore loads its entire +transitive closure. + +**Why it is silent.** Everything works, immediately and correctly. Hard +references never fail, never arrive late and never need a fallback. The cost is +paid once, at load, in a place nobody attributes to this asset. + +**Why the obvious check misses it.** Reviewing the definition shows a handful of +properties — five or six references, all necessary. The cost is not the direct +count but the closure, and the closure lives in assets this file never mentions. +Nothing at the declaration hints at its size. + +**Symptom.** A boot or transition hitch attributed to "the level" or "shaders". +In the audited reference the same architectural family ranged from **two** +packages when addressing by identifier to **over a thousand** when holding a hard +chain — a difference produced entirely by reference type. + +**Detect.** Direct counts are not the measurement; build the closure. From +source you can only find the candidates: + +```bash +rg -n "TObjectPtr<|TSubclassOf<" --glob "*.h" . | rg -i "definition|data|archetype" +rg -n "TSoftObjectPtr<|TSoftClassPtr<" --glob "*.h" . | rg -i "definition|data" +``` + +Then, from the editor or a registry dump, compute the transitive hard closure per +definition and rank by size. **Compare each against its architectural layer**: a +catalog entry with a large closure is the finding; an archetype with one may be +correct and budgeted. + +**Guardrail.** Write the intended reference type into the property metadata and +validate it. Re-measure the closure of every archetype-level definition when it +changes; a closure without a recorded baseline cannot be reviewed. + +--- + +### AL-02 - A bundle name that no property uses + +**Mechanism.** A bundle tier is declared as a constant and requested at load time +alongside the role bundles. No property in the codebase is annotated with it. + +**Why it is silent.** Requesting a bundle that matches nothing is not an error. +The load succeeds, the other bundles arrive, and the tier appears to be part of +the loading strategy. + +**Why the obvious check misses it.** The constant is declared, referenced at the +request site, and named meaningfully. "Is this bundle used?" answers yes at two +places. The absent half is the **annotation**, and nothing connects a request to +the properties it is meant to gather. + +**Symptom.** No runtime symptom at all. The cost is a false model: the team +believes there is an equipped-content tier, plans around it, and the loading +strategy has one fewer dimension than it appears to. + +**Detect.** Compare declared bundle names against annotated ones: + +```bash +rg -n "AssetBundles=\"[^\"]*\"" --glob "*.h" . -o | sed 's/.*AssetBundles=//' | sort -u +rg -n "BundlesToLoad.Add|LoadStateClient|LoadStateServer|FName\(TEXT\(\"" --glob "*.cpp" . \ + | rg -i "bundle" +``` + +Any name in the second list and not the first is requested and unpopulated. In +the audited project one such tier is requested on every load regardless of net +mode, and appears in no annotation in the entire source tree. + +**Guardrail.** Assert at startup that every requested bundle name matches at +least one annotated property, or remove the tier. A bundle nobody fills is a +plan, not a mechanism. + +--- + +## Startup + +### AL-03 - Progress scaffolding where every layer is independently broken + +**Mechanism.** A startup framework reports per-job progress. Three things must +work: the job must hand back a load handle, a throttle must permit an update, and +a receiver must render it. + +**Why it is silent.** A progress bar that does not move looks like a fast load or +a coarse-grained one. There is no error state for "progress was never reported", +and the framework's logs still print per-job timings, so it looks instrumented. + +**Why the obvious check misses it.** Each layer is individually plausible. The +macro compiles and runs the job. The throttle has a sensible-looking comparison. +The receiver is a real function with a real name. Only reading all three together +shows that the handle is never assigned, the comparison can never be true, and +the receiver's body is a comment. + +**Symptom.** No visible progress during startup, so the slowest job is unknown. +The team optimizes what it guesses rather than what it measured. + +**Detect.** Check the three layers separately, and expect each to look fine: + +```bash +# 1. does anything assign the out-parameter handle? +rg -n -B4 -A8 "TSharedPtr& \w+" --glob "*.h" --glob "*.cpp" . +# 2. is the throttle comparison the right way round? +rg -n -B2 -A6 "LastUpdate|LastReport" --glob "*.h" --glob "*.cpp" . +# 3. does the receiver have a body? +rg -n -A4 "UpdateInitial\w*Percent|::UpdateProgress" --glob "*.cpp" . +``` + +In the audited project the throttle reads `LastUpdate - Now > interval`, where +the elapsed value is monotonically increasing and the stored value starts at +zero — so the difference is never positive and the branch is unreachable. + +**Guardrail.** Test that progress fires at all before trusting it. A callback +that has never been observed to run is indistinguishable from one that cannot. + +--- + +### AL-04 - Global data with soft references and no bundle tags + +**Mechanism.** A globally-loaded data asset holds soft class pointers to effects +or classes used across the game. The asset is loaded at startup; its soft targets +are not, because nothing tags them into a bundle. + +**Why it is silent.** The pointers resolve on demand and everything works. The +first resolution is a synchronous load, and it lands in whatever frame first +needs it — commonly the first damage event, which is already a busy frame. + +**Why the obvious check misses it.** The asset is loaded at startup, which +answers "is the global data preloaded?" with yes. Whether its *contents* are +preloaded is a different question, decided by annotations that are absent — and +absent annotations look exactly like properties that do not need them. + +**Symptom.** A hitch at first damage, first dynamic tag, or first use of any +global class. Reproduces once per session, which makes it easy to dismiss. + +**Detect.** Find globally-loaded data assets and check their soft properties for +tags: + +```bash +rg -n -A20 "class \w*GameData" --glob "*.h" . | rg "TSoftClassPtr|TSoftObjectPtr|AssetBundles" +``` + +Soft properties with no bundle annotation on a startup-loaded asset is the +finding. Then find where they first resolve: + +```bash +rg -n "GetSubclass\(|GetAsset\(" --glob "*.cpp" . | rg -i "gamedata" +``` + +**Guardrail.** Tag them and load them with the mode. A soft reference on a global +asset without a bundle is a deferred synchronous load with an unpredictable +trigger. + +--- + +## Retention + +### AL-05 - A keep-in-memory pool with no way out + +**Mechanism.** A convenience loader resolves a soft pointer, and adds the result +to a reflected set so it stays alive. The keep-in-memory parameter **defaults to +true**. No removal API exists. + +**Why it is silent.** It is a cache doing its job. Every call is faster than the +last, nothing is ever missing, and the set has no size at which it complains. + +**Why the obvious check misses it.** The parameter is visible in the signature +and reads as a considered choice — which it is, at the declaration. At every call +site it is invisible, because it is defaulted. Reviewing a call shows a load; the +retention is in the default argument of a function in another file. + +**Symptom.** Residency that only grows across a session, unaffected by mode +changes, map travel or anything else. Because the growth is monotonic and +attributable to nothing, it is usually first noticed as "we are over budget" +rather than as a leak. + +**Detect.** Count additions against removals, and read the default: + +```bash +rg -n "LoadedAssets|KeptAssets|RetainedAssets" --glob "*.h" --glob "*.cpp" . +rg -n "bKeepInMemory" --glob "*.h" . | head +``` + +In the audited project the set is added to at one site, read at two — both inside +a dump command — and **never removed from anywhere**, with the flag defaulting to +true and no call site in the project passing false. + +**Guardrail.** Default to false, or make the flag explicit at every call site. +Ship a paired release API and a console command that prints the set's size. A +cache without an eviction policy is a leak with a nicer name. + +--- + +### AL-06 - Strong keys pin graphs; raw keys lose them + +**Mechanism.** Two opposite errors in the same area. A map keyed by a **strong +object pointer** keeps its key alive, so one missed unregister pins an actor and +everything it owns. A map holding object pointers inside a **non-reflected** +struct is invisible to garbage collection, so the values can be collected while +still referenced. + +**Why it is silent.** The first has no symptom until memory is measured — the +actor is dead in gameplay terms and alive in memory. The second has no symptom +until collection happens to run at the wrong moment, which is rare and +non-deterministic. + +**Why the obvious check misses it.** Both look like ordinary containers. The +difference between a reflected and a non-reflected struct is one macro, and the +difference between a strong and a weak key is one wrapper type. Neither reads as +a memory decision at the declaration site. + +**Symptom.** Actors that never disappear from a memory report, or — from the +other error — a crash on a pointer that was valid a moment ago. Both are usually +investigated far from the container. + +**Detect.** Audit direction as well as presence: + +```bash +# strong keys +rg -n "TMap|TMap|U\w+\* \w+;" --glob "*.h" . | rg -v "UPROPERTY" +``` + +Both patterns appearing in one codebase is common and is worth stating plainly in +a review: they are not variants of one mistake, they are opposite mistakes, and a +fix for one does not address the other. + +**Guardrail.** Prefer object keys with explicitly-owned values plus a periodic +prune. Every object pointer that must survive collection is reflected; every one +that must not is weak. There is no third category. + +--- + +### AL-07 - No unload path anywhere + +**Mechanism.** Content is loaded per mode through bundles. Nothing releases it — +no unload call, no bundle removal, no explicit collection request. + +**Why it is silent.** For a session-based game with a process restart between +matches, it is correct and cheap. The absence only becomes a defect when modes +change within one process, which is a product decision made later. + +**Why the obvious check misses it.** There is nothing to see. Reviewing the +loading code finds a complete, well-built loading path; the absence of a +symmetric release path is not a line anyone reads. The authors of the audited +project recorded it themselves, in comments, and shipped it. + +**Symptom.** Memory that never returns to baseline across mode changes. The first +mode's content is resident for the whole session alongside the second's. + +**Detect.** One command, and the expected answer is a row of zeros: + +```bash +for p in UnloadPrimaryAsset bRemoveAllBundles CollectGarbage FlushAsyncLoading TrimMemory; do + printf "%-24s %s\n" "$p" "$(rg -c "$p" --glob '*.cpp' --glob '*.h' . | wc -l)" +done +``` + +In the audited project every one of those is **zero**. That table is the finding, +and it takes ten seconds to produce in any project. + +**Guardrail.** Either implement release, or write down that mode changes require +map travel. An undefined middle is what grows. + +--- + +## Failure handling + +### AL-08 - Cancellation bound to the completion callback + +**Mechanism.** An asynchronous load binds a completion delegate, and binds the +**same** delegate to cancellation. + +**Why it is silent.** Cancellation is rare — it needs a mode change or a world +teardown mid-load. When it does happen, the system reports success and proceeds, +and every downstream consumer of a missing asset fails silently in its own way. + +**Why the obvious check misses it.** Both delegates are bound, which is more +than most code does. The lambda is short and reads as deliberate. Nothing +distinguishes "handle cancellation" from "handle cancellation *correctly*" at the +binding site. + +**Symptom.** A mode reports loaded with partially loaded content. Downstream, +unloaded classes resolve to null and are skipped without logging, so the result +is a mode missing several features and no diagnostic anywhere. + +**Detect.** Compare the two bindings: + +```bash +rg -n -B6 -A6 "BindCancelDelegate" --glob "*.cpp" . +``` + +If the cancel lambda executes the completion delegate, that is this defect. In +the audited project it does, in three lines, immediately below the completion +binding. + +**Guardrail.** Cancellation enters a distinct state — failed, or cancelled — that +the loading screen and the consumers can observe. Success is a claim; make it +provable. + +--- + +### AL-09 - An asynchronous result that is never read + +**Mechanism.** A plugin or asset load completes with a result parameter carrying +success or failure. The callback decrements a counter and ignores the parameter. + +**Why it is silent.** The counter reaches zero either way, so the pipeline +completes. A failed load is indistinguishable from a successful one at every +point downstream. + +**Why the obvious check misses it.** The callback has the right signature and is +bound correctly. An unused parameter in a delegate signature produces no warning, +because the signature is imposed by the API rather than chosen. + +**Symptom.** A feature that silently does not activate, in a build where one of +its assets failed to load. The failure is attributed to the feature. + +**Detect.** Find completion callbacks whose result parameter is unused: + +```bash +rg -n -A8 "::On\w*LoadComplete\(const \w+::\w+& Result\)" --glob "*.cpp" . \ + | rg -c "Result" +``` + +A count of one — the signature only — is the finding. + +**Guardrail.** Read the result, log the failure with the identifier that failed, +and enter a failed state. An ignored result is a decision to be surprised later. + +--- + +## Measurement + +### AL-10 - Measuring residency in the editor + +**Mechanism.** The editor loads both client and server bundles, because it must +be able to run either role. + +**Why it is silent.** The measurement completes and produces a number. The number +is real; it just describes a configuration that ships to nobody. + +**Why the obvious check misses it.** It is intentional and correct — nothing to +find in the memory system. The inflating condition is a boolean expression in the +bundle-selection code, and nobody profiling memory reads the loading policy. + +**Symptom.** Memory budgets built on editor numbers, wrong in a direction that +feels safe until a platform limit is real. + +**Detect.** Read the bundle-selection condition before trusting any in-editor +number: + +```bash +rg -n -B2 -A6 "bLoadClient|bLoadServer|LoadStateClient" --glob "*.cpp" . | rg "GIsEditor" +``` + +An editor branch that unions both sets means editor residency is directional +only. The audited project does this in two places, and one of them carries a +comment from its authors noting the resulting hitching. + +**Guardrail.** Residency numbers require a packaged build of the target role. Use +the editor for relative structural comparison, and say so whenever an in-editor +number is quoted. diff --git a/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/patterns.md b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/patterns.md new file mode 100644 index 0000000..3ff2484 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-asset-loading-and-memory/references/patterns.md @@ -0,0 +1,360 @@ +# Patterns: asset loading and residency in a measured reference + +A worked reading of one shipped project's loading path, from asset-manager +startup to feature-plugin activation, audited as source. + +Individual defects are in [failure modes](failure-modes.md) with detection +recipes. This file is about the shape: which decisions determine load cost, which +determine residency, and why those are two different reviews. + +Markers: **[measured]** — read in source or config; **[derived]** — conclusion +from measured facts; **[open]** — not answerable without a packaged build or +engine source, and left open. + +--- + +## 1. The five budgets, and the one that is not a budget + +| Budget | What it is | Measured by | +|---|---|---| +| A. Disk / cook | package bytes shipped | build and chunk sizes | +| B. CPU RAM | object graph plus bulk data | memory report, object list | +| C. GPU VRAM | live render resources | platform GPU profiler | +| D. Streaming pool | **a limit, not consumption** | streaming stats | +| E. Net objects | actors, channels, bandwidth | network profiler | + +This skill owns A and B. + +**The most common analytical error in this area is reading D as B.** A pool +setting grants permission to use memory. It never reports use, it never causes +use, and lowering it does not free anything that is already resident. + +--- + +## 2. Reference type is the largest single lever + +Three mechanisms, routinely conflated: + +| Mechanism | Effect | +|---|---| +| a reflected property, an explicit GC reference, a retained load handle | **owns** — keeps resident | +| a weak pointer, an object key | **observes** — keeps nothing | +| a soft pointer | **a path, not a reference** | + +The third line is where reasoning usually fails. A soft pointer saves nothing by +itself. Residency is decided by **where the loaded result is stored** — and a +soft reference resolved into a strong property is a hard reference with extra +steps. + +### The measured spread + +From the dependency graph of the audited project — 3 874 packages, 16 639 hard +edges, 9 333 soft edges **[measured]**: + +| Asset | Packages pulled by loading it | +|---|---:| +| pawn archetype with a hard class-and-object chain | **1 015** | +| weapon pickup definition | 800 | +| ability set | 121 | +| input config | 14 | +| action set referencing by identifier and soft pointer | **2** | +| playlist referencing by primary asset identifier | **2** | + +**[derived]** Three orders of magnitude, from one decision made per property. +That is the whole argument for treating reference type as architecture rather +than as style — and the reason a "make it soft" refactor has to start from the +closure measurement rather than from the property list. + +### Choosing + +| Need | Reference | +|---|---| +| address before load; cook and bundle identity | primary asset identifier | +| optional, feature-scoped or async content | soft object or class pointer | +| required whenever the owner is loaded, and small | hard pointer | +| data selects an implementation class | class reference | +| owner-exclusive polymorphic fragment | instanced object | + +Rule of thumb: a definition that is a **catalog entry** references by identifier +or soft pointer; a definition that is an **archetype** may reference hard, but its +closure must be measured and budgeted. + +--- + +## 3. Bundles declare role, and their absence declares nothing + +Put the load role beside the reference: + +```cpp +UPROPERTY(EditAnywhere, meta=(AssetBundles="Client")) +TSoftClassPtr WidgetClass; + +UPROPERTY(EditAnywhere, meta=(AssetBundles="Client,Server")) +TSoftClassPtr AbilityType; +``` + +Then request bundles per runtime role at load time. + +**[measured]** In the audited project the bundle vocabulary is small — client +and client-plus-server annotations only — and one declared bundle name is +requested on every load while appearing on **no** property in source at all. It +is either satisfied entirely by content, or it is dead. **[derived]** A bundle +name that no property claims is indistinguishable from a typo, and neither the +compiler nor the cooker will say which it is. + +Gates worth adopting: + +- every soft reference has an intentional bundle annotation; +- server-irrelevant presentation is client-only; +- prediction-relevant gameplay classes are client-and-server; +- a missing annotation means "loaded whenever the owner loads" — decide it, do + not default into it; +- cook rules control budget A, not budget B. A label asset with thousands of + soft references and no hard ones determines what ships and retains nothing. + +--- + +## 4. Startup: the shape that works, and the shape that only looks like it + +The audited project has a well-formed startup-job framework **[measured]**: named +jobs, weights, per-job timing logs, a progress delegate, and boot-timing scopes +that integrate with the engine's trace. + +**[measured]** It also has: a macro that never fills the load handle it is +designed to return, a progress throttle whose comparison can never be true, and a +progress receiver whose body is a comment. + +**[derived]** Three independent layers, each individually harmless, combining +into a startup with no progress reporting at all — inside a framework built to +report it. This is the single most instructive thing in the audit, and the +general form is worth memorizing: **scaffolding is not evidence of function.** +Recipes: AL-02, AL-03. + +What to take from it anyway: + +- **name every startup job and log its duration.** Per-job timing is nearly free + and it is the only thing that answers "what is slow at boot". +- **wrap boot phases in the engine's timing scopes**, so the data lands in the + same trace as everything else. +- **fail fatally on genuinely required global data.** The audited project does + this for its global data asset, with a comment explaining that a soft failure + here would be harder to diagnose than a hard one **[measured]**. That is the + correct call, and it is rare. + +--- + +## 5. The synchronous load that dominates everything else + +**[measured]** The root definition of a game mode is resolved with a synchronous +load in the game thread, and its own authors marked it for async conversion. +That definition holds a hard pointer to a pawn archetype, which holds hard +references to the pawn class, its ability sets, its input config, its tag policy +and its camera mode — all hard. + +**[derived]** One line therefore pulls the closure measured at over a thousand +packages, **before** the elegant asynchronous bundle pipeline underneath it runs +at all. An async pipeline that executes after the bulk of the work is decoration. + +The general lesson is a review question rather than a rule: *for every +synchronous load on a boot or transition path, what is its transitive closure?* +If nobody has measured it, the loading architecture is unverified regardless of +how much of it is asynchronous. + +**[measured]** Note the asymmetry that makes this easy to miss: on a client the +same definition arrives by replication and is resolved by the net driver, so the +synchronous cost exists on one side only, and profiling the wrong side finds +nothing. + +--- + +## 6. Handle ownership: the difference between loading and controlling + +**[measured]** In the audited experience loader, every load handle is a local +variable. The component stores none of them. + +**[derived]** Three consequences, none of which produce an error: + +1. **Cancellation is impossible.** If the mode changes or the world dies mid-load, + there is no reference to cancel. +2. **Release is impossible.** Residency is held inside the asset manager rather + than by the component, and unloading requires the API that nothing in the + project calls. +3. **The load survives only incidentally** — because the completion delegate is + bound to the handle and the streamable manager keeps it alive until it fires. + +**[measured]** The same codebase contains the correct discipline elsewhere: UI +async actions cancel their handles on teardown, and an async helper mixin cancels +in its destructor. The pattern was known; it was not applied at the place with the +largest closure. + +**Take the shape from the UI code, not from the loader.** + +--- + +## 7. Retention: six owners, audited by name + +Anything that stays resident is held by one of: + +1. reflected containers on long-lived subsystems; +2. explicit GC references added by native code; +3. retained load handles; +4. delegate bindings capturing strong references; +5. component and actor ownership chains; +6. class default objects reached through class references. + +### The convenience loader that defaults to permanent + +**[measured]** A helper resolves a soft pointer, and a boolean parameter +defaulting to **true** adds the result to a reflected set. There is no removal +API, no call passing false, and no clearing anywhere in the project. + +**[derived]** Every casual call therefore loads synchronously *and* roots the +asset for the lifetime of the process. The soft pointer was chosen to defer +loading; the helper converts it into a permanent hard reference at the call site, +invisibly, by default. + +The design that avoids it: make retention explicit at the call site or default it +to false, provide a paired release, make the retained set inspectable from a +console command, and give retention an owner and a scope rather than assigning it +to "the asset manager". Recipe: AL-05. + +### Both directions of the pointer error + +Strong keys pin more than intended — a map keyed by actor pointers keeps every +key actor and its whole presentation graph alive on one missed unregister. + +Raw pointers in a non-reflected struct do the opposite: the collector cannot see +them, so the values can be collected while referenced and the keys can dangle. + +**[derived]** Both errors appear in real codebases, sometimes in the same one. +Audit *direction*, not just presence. + +--- + +## 8. Unloading, or the six "never"s + +**[measured]** In the audited project, searched across all source: + +| Mechanism | Occurrences | +|---|---:| +| primary asset unload | 0 | +| bundle removal | 0 | +| explicit collection | 0 | +| async flush | 0 | +| memory trim | 0 | + +The one forced collection in the project is in the loading-screen hide path +**[measured]**, and the experience loader's own teardown carries a comment +admitting it deactivated without unloading **[measured]**. + +**[derived]** The residency table for this project has "never" in the release +column for six of its rows. For a session-based game restarted between matches +that is an acceptable engineering position. For a service title that swaps modes +in place it is not — and the difference is a product decision that must be made +explicitly rather than discovered from a memory graph. + +**The honest options are two:** implement release, or state plainly that mode +changes require travel. What does not work is claiming runtime modularity while +the release column is empty. + +### Watch for collection ping-pong + +**[measured]** Hiding the loading screen forces a full purge; showing it again +synchronously loads the widget class that purge just collected. A forced +collection at a transition boundary must not collect what the next transition +immediately needs. Context: AL-07. + +--- + +## 9. Failure handling is a memory concern + +- a cancelled load must not run the success path; +- the result of an asynchronous activation must be read; +- partially loaded state must be releasable; +- an unbounded wait converts a stall into a hang; +- a loading screen with no reason string is not error handling. + +**[measured]** In the audited project the cancel delegate invokes the *same* +callback as completion, and the plugin-activation result parameter is never read. +**[derived]** A cancelled or failed load therefore transitions to "loaded", and +downstream code resolves soft references that are not there — where, by design, +a failed resolve is silently skipped. The result is graceful degradation with no +diagnostic, which is the most expensive kind. + +### One tool worth copying outright + +**[measured]** The project ships console variables that inject artificial delay +into mode loading. That is a reproducible way to exercise slow-load and +cancellation paths without a network or a cold cache, and almost nobody builds +one. + +--- + +## 10. Measurement plan + +Static first, because it is cheap and finds structural problems: + +1. build the dependency graph from the asset registry: direct hard and soft + counts, transitive hard closure per definition, cross-plugin edges; +2. rank definitions by closure and compare against their architectural layer — + a catalog entry with a large closure is the finding; +3. search for synchronous loads on boot and hot paths; +4. inventory retention owners by type (§7). + +Runtime, once a packaged build exists: + +- memory reports before and after a full mode cycle; +- object lists and reference queries on suspected roots — root-graph evidence, + not process memory; +- fifty to a hundred spawn, despawn and travel cycles, looking for monotonic + growth; +- allocation attribution from the engine's profiler; +- the allocator's dangling-pointer mode when corruption is suspected. + +### Do not measure residency in the editor + +**[measured]** Editor builds can load **both** role bundles, by an explicit +editor branch in bundle selection — a fact the project's own comments +acknowledge, noting the resulting hitches. + +**[derived]** An in-editor memory profile therefore represents neither shipping +role. Use it for relative structural comparison only; absolute residency requires +a packaged build. + +--- + +## 11. When to stop + +Optimize loading and residency when boot or transition time is a product problem, +when a platform memory ceiling is real and near, when growth is monotonic across +a session, or when a specific closure is measurably oversized. + +Do not restructure references on aesthetics. A fourteen-package closure is not a +problem. A thousand-package synchronous closure on the boot path is. Measure +first, and keep the measurement, so the next change has something to be compared +against. + +--- + +## Provenance + +The analysis and every number above come from a source and configuration audit of +Epic's Lyra Starter Game on Unreal Engine 5.6, plus a dependency graph extracted +from its asset registry through the editor. + +Two boundaries applied to that reading and bound these claims: the engine's own +asset manager, streamable manager and feature subsystem are not part of the +project, so statements about their internals are inference from call sites rather +than from source; and no profiler was run, so there are no timing or +memory-footprint numbers here by construction — the package counts are static +graph measurements. + +Source addresses stay in the research archive that produced this skill. Entry +identifiers in [failure modes](failure-modes.md) resolve back to the audited +locations, so any specific claim can be produced on request. + +## Evidence boundary + +One project, one engine version, one workspace. The structure is transferable; +the specific defects are evidence, not guarantees about other versions. Re-run +the detection recipes against your own tree before acting on any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/SKILL.md b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/SKILL.md new file mode 100644 index 0000000..01975d7 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/SKILL.md @@ -0,0 +1,315 @@ +--- +name: ue-cosmetics-and-teams +description: >- + Design or review modular character cosmetics and team systems in Unreal + Engine: replicated cosmetic intent with local realization, controller-owned + selection versus pawn-owned realization, fast-array entries, dedicated-server + exclusion, tag-driven mesh and animation selection, one authoritative team + identifier with subscribed mirrors, team relationship and damage policy, and + data-driven material presentation. Use when building skins, modular + characters, team assignment, friendly fire, hit-marker filtering or + team-coloured UI and effects. +--- + +# UE cosmetics and teams + +Two invariants, one per half: + +> Replicate cosmetic **intent**, never the spawned presentation. + +> Store team membership **once**, authoritatively; every other copy is a mirror +> with a subscription. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes: [failure modes](references/failure-modes.md). + +These two systems are in one skill because they meet at exactly one point — +presentation — and the most common design error is letting them meet anywhere +else. + +Related skills: `ue-data-driven-architecture`, `ue-gas-architecture`, +`ue-gameplay-messaging`, `ue-modular-gameplay`, `ue-multiplayer-authority`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +# Part I — Modular cosmetics + +## 1. Split persistent intent from pawn realization + +A controller-level component owns the selection, so it survives pawn death and +respawn. A pawn-level component realizes it on the current body. + +```text +controller component pawn component +├─ desired part list ├─ replicated part list +├─ listens for pawn changes ├─ spawns local presentation components +└─ applies and removes on the ├─ computes cosmetic tags + current pawn, keeping └─ notifies body and animation selection + handles +``` + +If the only record of a player's appearance lives on the pawn, it dies with the +pawn. That is the whole reason for the split. + +### Possession transfer + +On a possession change, in this order: remove owned parts from the old pawn; +find the component on the new pawn; add owned parts; retain the returned handles. +**Each owner removes only its own parts** — cheats, developer settings and real +content coexist in one list, distinguished by a source field rather than by +separate lists. + +--- + +## 2. Define a part as data, and make it incapable of gameplay + +A minimal part: a class to spawn, a socket to attach to, and a collision mode +that defaults to none. + +Default presentation safety: collision off unless explicitly requested; no +ability system, attributes or authority inside the cosmetic actor; no replicated +gameplay state; validated attach target; locally tracked handle. + +### The guarantee is the net-mode guard, not the convention + +Nothing in the type system prevents putting a gameplay-capable actor into a part +class. The real isolation comes from one line: **presentation is not spawned at +all on a dedicated server.** The server therefore cannot trace against, collide +with, or be influenced by cosmetics, because they do not exist there. + +Be precise about the limit of that guarantee: on a listen server or in +standalone, the parts *do* exist on the authority. Recipe: CT-01. + +--- + +## 3. Replicate intent with a fast array + +The replicated entry holds the intent — class and socket. It must not hold the +local handle or the spawned component. Put both of those in the same struct as +**non-replicated** fields and let the annotation express the split, rather than +maintaining two parallel types. + +Callbacks realize changes locally: post-add spawns, pre-remove destroys, +post-change rebuilds. + +### Decide what "change" means before you need it + +A common and defensible implementation makes post-change destroy and respawn the +actor, because propagating a diff into an already-spawned actor is hard. That is +fine and it is a decision with a cost: every edit is a full rebuild, so a system +that edits parts frequently will hitch. Recipe: CT-02. + +--- + +## 4. Tag-driven selection: rules plus a default + +Parts contribute tags — body style, animation style, armour set. A selection set +maps required tags to a mesh or an animation layer, with a fallback: + +1. gather tags from all active parts; +2. walk the rules in order; +3. the first rule whose required tags are **all** present wins; +4. otherwise use the default. + +This microscopic pattern — an array of (asset, required tags) plus a default, +first match wins — is about ten lines and is worth copying wholesale. It is the +same shape whether it selects a mesh, a physics asset or an animation layer. + +### Its one sharp edge + +**Priority is array order.** There are no weights and no specificity ranking, so +a broad rule placed above a narrow one silently shadows it, and the symptom is a +wrong mesh rather than an error. Document the ordering requirement next to the +array, and validate that no rule's tag set is a subset of an earlier rule's. +Recipe: CT-03. + +### Do not let cosmetic tags become gameplay tags + +They may share a type. They must not share authority. A tag contributed by a +cosmetic part is client-realizable data; nothing that decides gameplay may read +it without an explicit, reviewed promotion. + +--- + +## 5. Cosmetic lifecycle tests + +Add and remove one part; several owners contributing; controller changes pawn; +pawn destroyed before controller cleanup; replicated add, remove and change; late +join reconstructs everything; **dedicated server spawns zero presentation +actors**; listen server does not realize twice; invalid part class logs rather +than crashes; incomplete tag set falls back. + +--- + +# Part II — Teams + +## 6. One authoritative source, and loud mirrors + +Put the team identifier on the player state: it survives respawn, already +replicates, and belongs to player identity rather than to a body. + +Controllers, pawns and local players may mirror it. The rules that make mirroring +safe: + +- mirrors **subscribe** to the source and re-subscribe when the source object + changes; +- a write to a mirror is **rejected loudly**, not silently ignored; +- a pawn may hold its own value only while unpossessed, and must say so in the + error when it refuses. + +The second rule is the one teams get wrong. A silent no-op setter on a mirror is +indistinguishable from a working setter until something depends on it. Log an +error naming the real owner. Recipe: CT-06. + +--- + +## 7. Resolve team from an arbitrary object, in one place + +Centralize the cascade so damage, UI, AI and targeting cannot each grow their own +version: + +1. the object implements the team interface; +2. a pawn resolves through its controller or player state; +3. an actor resolves through its **instigator** — this is what lets projectiles + and effects carry the shooter's team without owning one; +4. otherwise unassigned. + +### Return a relationship, not a boolean + +```text +same team / different teams / indeterminate +``` + +A boolean conflates "enemy" with "unknown", and the caller then has to guess. +Every caller must handle the third case explicitly — and "indeterminate means +allowed" is the failure that ships. See `ue-multiplayer-authority` for why an +error path that permits is worse than one that denies. + +--- + +## 8. One damage policy, used by both authority and prediction + +One function answers "may this instigator damage this target". The damage +calculation multiplies by its result as a scalar rather than branching, so the +team rule participates in the formula like any other modifier. + +**The same function decides whether to show a hit marker.** That single choice +removes the entire class of "the client showed a hit and the server dealt no +damage" bugs, because there is no second implementation to disagree with. + +### Make the policy configurable before you need it + +Friendly fire, self-damage, damage to unassigned actors and per-mode overrides +are policy inputs. If they are hard-coded inside the damage calculation, a mode +that wants different rules requires editing the calculation. + +Watch for the shape where a header comment promises settings that the body does +not implement — the comment is a design that was not built, and it reads as a +feature. Recipe: CT-07. + +--- + +## 9. Presentation as a named parameter bag + +A team display asset should hold **named material parameters** — scalars, +colours, textures, a short name — rather than a set of assets. Consumers ask for +a parameter; materials are authored to expect the names. Adding a team then costs +one asset instead of a duplicated material tree. + +Apply it to meshes, effects and UI. Applying to an actor should include child +actors by default — that is precisely what carries team colour onto cosmetic +parts, and it is the only place the two halves of this skill touch. + +### Two traps in the applying code + +**Colour conversion drops alpha.** Converting a linear colour to a three-element +vector to feed a material parameter silently discards the fourth channel. If any +team parameter encodes opacity, it is gone. Recipe: CT-09. + +**A viewer-relative parameter that is ignored.** If the API accepts a viewer +identity — "always show my own team as blue" — then the implementation must use +it. An accepted-and-ignored parameter is worse than an absent one: it makes the +feature look supported and unimplementable at the same time. Recipe: CT-10. + +--- + +## 10. Replicate intent, realize locally — the shared principle + +| Domain | Replicated intent | Local realization | +|---|---|---| +| cosmetics | part definitions | spawned components, meshes, materials | +| teams | team identifier | colours, UI, effects, nameplates | +| composition | selected definition | activated plugins and actions | + +Realization must be **idempotent**: applying the same intent twice produces one +result. This is what makes late join deterministic from compact state. + +--- + +## 11. Validation checklist + +### Cosmetics + +- [ ] part class is presentation-only and valid for the body; +- [ ] socket exists, or a fallback is defined; +- [ ] collision disabled by default; +- [ ] handle and spawned component are non-replicated; +- [ ] dedicated server realizes nothing; +- [ ] every owner retains removal handles; +- [ ] possession transfer removes then applies, exactly once; +- [ ] rule ordering documented, fallback present, no shadowed rules; +- [ ] no gameplay decision reads a cosmetic tag. + +### Teams + +- [ ] one authoritative source, authority-only writes; +- [ ] mirrors subscribe, re-subscribe, and refuse writes loudly; +- [ ] resolution is centralized and includes the instigator step; +- [ ] the indeterminate relationship is handled explicitly at every caller; +- [ ] damage and hit feedback call the same policy function; +- [ ] friendly fire and self-damage are configurable and tested; +- [ ] "private" team data is actually filtered, or is renamed; +- [ ] display parameters exist in the target materials; +- [ ] a viewer-relative API uses the viewer. + +--- + +## 12. Keep them separate + +They meet at one seam: + +```text +team identifier → display asset → material parameters on cosmetic parts +``` + +Do not let cosmetics own the team identifier, and do not let the team system own +part selection. Separation is what allows skins independent of teams, players +with cosmetics and no team, team recolour over any skin, and viewer-relative +presentation without touching authoritative state. + +Combine them only inside a small presentation adapter — never in authority or +replication ownership. + +--- + +## Provenance + +The findings behind the failure modes come from a source audit of Epic's Lyra +Starter Game on Unreal Engine 5.6 — roughly 1300 lines of cosmetics and 1800 of +teams, read rather than run. Several of the wiring call sites live in Blueprint +graphs, which are binary and were not read; those are marked open rather than +concluded, and the skill is written so that none of its claims depend on them. + +Source addresses stay in the research archive. Each entry carries a stable +identifier (`CT-01`, `CT-02`, …) resolving back to the audited location there. + +## Evidence boundary + +One project, one engine version, one workspace. The limits are specific and worth +repeating: the components are added to actors from data assets that were not +read, the team colour application has no C++ caller, and the animation-layer +selection is invoked from a Blueprint. Absence of a C++ caller is **not** absence +of a caller — see the failure modes for how to check properly before deleting +anything. diff --git a/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/failure-modes.md new file mode 100644 index 0000000..8c8b736 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/failure-modes.md @@ -0,0 +1,406 @@ +# Failure modes: cosmetics and teams + +Eleven ways an appearance system or a team system produces the wrong result +without producing an error. + +The shared property differs slightly between the two halves, and both are worth +stating. + +**Cosmetics fail by looking almost right.** A wrong mesh, a missing part or a +duplicated accessory is a visual outcome, and every visual outcome is plausible +to code. Nothing can assert that a character looks correct, so these defects are +found by people, late, and reported as "the skin is broken". + +**Teams fail by permitting.** Every interesting team defect resolves to +"indeterminate", and the branch that handles indeterminate almost always allows +the action, because denying it would have blocked something during development. + +Recipes use `rg` from a project source root and were executed against the audited +project while this file was written. + +--- + +## Cosmetics + +### CT-01 - The isolation guarantee is one line, and it is narrower than it looks + +**Mechanism.** Cosmetic presentation is excluded from the dedicated server by a +single net-mode check at the spawn site. That check is the entire reason cosmetic +content cannot influence authoritative simulation. + +**Why it is silent.** It works. On a dedicated server no presentation actor +exists, so none can collide, block a trace or tick. The guarantee holds exactly +where it is claimed. + +**Why the obvious check misses it.** The guarantee is usually restated as +"cosmetics cannot affect gameplay", which is a stronger claim than the code +makes. On a listen server or in standalone the parts **do** exist on the +authority, with whatever collision and components their class carries. Nothing in +the type system prevents a part class from containing an ability system or +enabled collision — the constraint is content discipline plus one `if`. + +**Symptom.** A cosmetic accessory that blocks a shot, or a part actor that ticks +expensively, in exactly the configuration used for local playtesting — and never +on the dedicated server where the "real" testing happens. + +**Detect.** Find the guard, confirm there is only one, and check what part +classes are permitted to contain: + +```bash +rg -n "NM_DedicatedServer|IsNetMode" --glob "*.cpp" / +rg -n "CollisionMode|SetActorEnableCollision" --glob "*.cpp" --glob "*.h" \ + / +``` + +In the audited project the first command returns **exactly one line** for the +whole module. That is elegant and it is also the entire safety boundary, which is +worth knowing before relying on it. + +**Guardrail.** State the guarantee accurately: "not present on a dedicated +server", not "cannot affect gameplay". Validate part classes at author time — +reject any that contain an ability system component, replicated properties or +enabled collision by default. + +--- + +### CT-02 - Every change rebuilds the actor + +**Mechanism.** The replication callback for a changed entry destroys the spawned +presentation and spawns a new one, because propagating a diff into a live actor +is hard and rarely needed. + +**Why it is silent.** It is correct. The result after a change is exactly the +result after an add, which is what the system promises. The cost is a rebuild, +and a rebuild looks identical to a first build. + +**Why the obvious check misses it.** The implementation is deliberate and +commented as such. Reviewing it shows a decision, not a defect. The consequence +only appears at a usage frequency nobody has yet — a system that changes parts +rarely and one that changes them per second are the same code. + +**Symptom.** A hitch on every cosmetic edit. Discovered when a feature starts +mutating parts at runtime — a preview screen, a colour cycler, a progression +system — long after the mechanism was reviewed. + +**Detect.** Read what the change callback does, and compare it with the add path: + +```bash +rg -n -A12 "PostReplicatedChange" --glob "*.cpp" . | rg -i "destroy|spawn|update" +``` + +If change is implemented as destroy plus spawn, treat the mechanism as +add/remove only and design callers accordingly. + +**Guardrail.** Document the cost where authors will see it. If frequent edits are +planned, add a real update path for the fields that can change in place, and keep +the rebuild for the ones that cannot. + +--- + +### CT-03 - Rule order is priority, and nothing says so + +**Mechanism.** A selection set is an array of rules, each with a set of required +tags, plus a default. The first rule whose tags are all present wins. Rules are +not sorted by specificity. + +**Why it is silent.** Every possible input produces a result — either a rule or +the default. There is no "no match" state to report, and the result is always a +valid asset. + +**Why the obvious check misses it.** The array looks like a set, and sets have no +order. A reviewer checking "are the rules correct?" checks each rule in +isolation, where each is correct. The defect exists only in the relationship +between a broad rule and a narrower one below it, which requires comparing every +pair. + +**Symptom.** A character with a specific tag combination gets the generic mesh. +Nobody notices for weeks, because it is the *right kind* of mesh. + +**Detect.** Find selection sets and check for shadowing — a rule whose tag set is +a subset of an earlier rule's: + +```bash +rg -n -B2 -A6 "SelectBest|RequiredTags" --glob "*.h" --glob "*.cpp" . +``` + +Then, for each set, verify the shadowing property in a test rather than by +reading: for every pair of rules `i < j`, assert that rule `i`'s tags are not a +subset of rule `j`'s. + +**Guardrail.** Write the ordering rule as a comment on the array and enforce the +subset check in validation. A rule that can never win is a content bug the +compiler cannot see. + +--- + +### CT-04 - The realization path re-applies unconditionally + +**Mechanism.** Recomputing appearance sets the mesh on every change broadcast, +relying on the setter to be cheap when the value has not changed — while passing +a flag that forces a pose reinitialization. + +**Why it is silent.** The result is correct. The mesh is right, the pose is +right, and the cost is invisible at the frequency the system is normally used. + +**Why the obvious check misses it.** The code carries a comment explaining that +the setter is a no-op when the mesh is unchanged, which answers the question a +reviewer would ask. The forced flag is a separate argument on the same call, and +it is not what the comment is about. + +**Symptom.** A visible pose pop or an animation reset whenever anything about +cosmetics changes, including changes that do not affect the mesh at all. + +**Detect.** Check the arguments of the re-application call, not just the call: + +```bash +rg -n -B6 "SetSkeletalMesh\(" --glob "*.cpp" . | rg "bReinit|true|Broadcast" +``` + +A hardcoded reinitialization flag on a path that runs on every change is the +finding. + +**Guardrail.** Compare before applying, and pass the reinitialization flag only +when the mesh actually changed. Idempotent realization means "applying twice +costs nothing", not merely "applying twice is correct". + +--- + +### CT-05 - Everything is a hard reference + +**Mechanism.** Part classes and selection-set assets are hard references, so the +entire cosmetic catalogue is loaded with whatever holds the rules. + +**Why it is silent.** Hard references always resolve. Nothing is ever missing, +nothing ever fails to load, and every asset is available the moment it is needed. + +**Why the obvious check misses it.** Reference correctness is what reviews check, +and hard references are maximally correct. The cost is memory residency, which +belongs to a different discipline and a different person, and which does not +appear in any test of the cosmetics system. + +**Symptom.** Memory proportional to the size of the cosmetic catalogue rather +than to what is worn, discovered during a platform memory pass and attributed to +content rather than to a reference-type decision. + +**Detect.** Count hard versus soft references in the cosmetic data types: + +```bash +rg -n "TSubclassOf<|TObjectPtr<" --glob "*.h" / | wc -l +rg -n "TSoftClassPtr<|TSoftObjectPtr<" --glob "*.h" / | wc -l +``` + +A large first number with a zero second is the finding. See +`ue-asset-loading-and-memory` for what this costs and how to measure it. + +**Guardrail.** Cosmetic catalogues are the textbook case for soft references and +bundles: large, optional, and mostly unworn. Decide per field and record the +decision in metadata. + +--- + +## Teams + +### CT-06 - The mirror setter that silently does nothing + +**Mechanism.** Team identity is authoritative on one object and mirrored on +several others. The mirrors implement the interface's setter because they must, +and the implementation does nothing. + +**Why it is silent.** Doing nothing is correct — the mirror is not the owner. The +call succeeds, returns, and the value is unchanged. + +**Why the obvious check misses it.** The setter exists and is implemented. Any +review asking "does this type support setting a team?" answers yes. The caller +has no return value to check and no reason to suspect the write did not land. + +**Symptom.** Code that sets a team on a controller or a pawn and observes the old +value. The investigation goes to replication, then to ordering, then eventually to +the setter. + +**Detect.** Read the bodies of every team setter and classify them: + +```bash +rg -n -A4 "::SetGenericTeamId|::SetTeamId" --glob "*.cpp" . \ + | rg -B1 "UE_LOG|ensure|^\s*\}" +``` + +The audited project does this **correctly** and is worth copying: the mirror +setters log an error naming the real owner rather than returning quietly. That +one line converts an invisible failure into a searchable one. + +**Guardrail.** A mirror's setter logs an error identifying the authoritative +owner. Never implement a no-op setter to satisfy an interface. + +--- + +### CT-07 - The header promises a policy the body does not implement + +**Mechanism.** The damage-permission function is documented as taking friendly +fire settings into account. The body contains no settings and no lookup. + +**Why it is silent.** The behaviour is coherent — allies are always safe — and +that is the right default for most modes. Nothing malfunctions. + +**Why the obvious check misses it.** The comment *is* the documentation. A team +adopting the system reads the header, concludes friendly fire is supported, and +plans a mode around it. Discovering otherwise requires reading a body that looks +finished. + +**Symptom.** A mode that needs friendly fire is scoped as a configuration change +and turns out to be a design change, in a function every damage path depends on. + +**Detect.** Compare the promise with the implementation: + +```bash +rg -n -i "friendly fire" --glob "*.h" --glob "*.cpp" . +rg -n -i "bFriendlyFire|bAllowFriendly|FriendlyFireSetting" --glob "*.h" --glob "*.cpp" . +``` + +In the audited project the first command finds the promise in a header comment +and the second returns **zero results project-wide**. A promise with no +implementing symbol is the finding, and this recipe generalises to any header +comment describing configurability. + +**Guardrail.** Treat a comment describing behaviour as a claim requiring a +symbol. If the setting does not exist, the comment describes a plan and must say +so. + +--- + +### CT-08 - Indeterminate resolves to permitted + +**Mechanism.** Team comparison correctly returns three states. The damage rule +treats the indeterminate case as allowed when the target satisfies an unrelated +condition — in the audited project, having an ability system component. + +**Why it is silent.** It exists to make something work: an unassigned training +target must be damageable. It succeeds at that, and the permissive branch is +never reached by any assigned actor during normal play. + +**Why the obvious check misses it.** The function has three branches and handles +all three, so it passes any review asking whether the indeterminate case is +handled. It *is* handled — permissively — and the marker in the code says +"temporary". + +**Symptom.** Any actor with an ability system and no team assignment is damageable +by anyone. Harmless while the only such actor is a target dummy; a rule violation +the moment a neutral faction, a destructible objective or an NPC exists. + +**Detect.** Find every consumer of the relationship enum and read its +indeterminate branch: + +```bash +rg -n -B4 -A10 "InvalidArgument|Indeterminate|NoTeam" --glob "*.cpp" . \ + | rg "return true|= true|Allow" +rg -n -i "//\s*@?TODO.*(temporary|until)" --glob "*.cpp" . +``` + +The second command is the high-yield one: a permissive branch that its own author +marked temporary is the strongest possible confirmation that the finding is real +and known. + +**Guardrail.** Failure to determine must deny. Give the target dummy a team +instead of giving every unassigned actor an exemption — solve the specific +problem in data, not the general rule in code. + +--- + +### CT-09 - A colour conversion that drops alpha + +**Mechanism.** A four-channel colour is converted to a three-element vector to be +passed as a material parameter. The fourth channel is discarded. + +**Why it is silent.** Three of four channels arrive correctly, so the colour is +right. Alpha is frequently unused in team colours, making the loss invisible for +as long as nobody encodes anything in it. + +**Why the obvious check misses it.** The conversion is one token inside an +otherwise correct call. Both types are colours, both are correct, and the review +question — "is the parameter set?" — is answered yes. + +**Symptom.** A team parameter that encodes opacity or a blend factor silently +reads as its default. In the same file, the effects-system path preserves all four +channels, so the same asset behaves differently in two places. + +**Detect.** Find lossy conversions at parameter-setting sites: + +```bash +rg -n "SetVectorParameterValue\w*\(.*FVector\(" --glob "*.cpp" . +rg -n "SetVariableLinearColor|SetVectorParameterValue" --glob "*.cpp" . +``` + +Two applicators for the same data with different channel counts is the finding. +In the audited project the material paths convert and the effects path does not. + +**Guardrail.** Pass four channels where the API accepts them. If a conversion is +unavoidable, assert or document that alpha is unused. + +--- + +### CT-10 - An accepted parameter that is ignored + +**Mechanism.** The display-asset accessor takes a viewer identity so that +presentation can be relative — "my team always looks blue". The implementation +does not read it. + +**Why it is silent.** The function returns the correct absolute asset. Every +caller gets a valid result, and the feature the parameter implies has simply +never been exercised. + +**Why the obvious check misses it.** The signature and the header comment +describe the feature completely. Confirming absence requires reading the body, +where the parameter is unused — and an unused parameter produces no warning +because it is part of a virtual-looking public API. + +**Symptom.** A team implementing viewer-relative colours writes correct calling +code and observes absolute colours. The bug appears to be in their code. + +**Detect.** For any parameter that names a feature, check that the body reads it: + +```bash +P='ViewerTeamId' +rg -n "\b$P\b" --glob "*.h" --glob "*.cpp" . +``` + +Occurrences only in the declaration and the definition's signature — with none in +the body — is the finding. In the audited project the body carries a comment +stating the parameter is currently ignored, which is honest and still ships an +API that cannot do what it claims. + +**Guardrail.** Do not accept a parameter you do not use. Remove it, or implement +it, or make the function name state the limitation. This is the same defect class +as an ordering field that never sorts — see `ue-ui-architecture`, UI-01. + +--- + +### CT-11 - "Private" that is not filtered + +**Mechanism.** Team data is split into a public and a private actor to express a +replication boundary. The private one has no filtering; it is an empty subclass +with a note that privacy is not implemented. + +**Why it is silent.** Everything works. Both actors replicate, both carry their +data, and the split is structurally correct — it is only the *filtering* that is +absent. + +**Why the obvious check misses it.** The architecture is visible and right: two +types, two names, a clear intent. Reading the class names answers the question. +Confirming means noticing that the private subclass has no body. + +**Symptom.** Data placed in the private actor because it is private is replicated +to every client. Nothing indicates this until someone inspects network traffic — +or until a competitor does. + +**Detect.** Check that the privacy-named type actually filters: + +```bash +rg -n -A15 "class \w*PrivateInfo|class \w*Private\w*" --glob "*.h" . +rg -n "IsNetRelevantFor|GetLifetimeReplicatedProps|COND_" --glob "*.cpp" . \ + | rg -i "private" +``` + +An empty subclass, or one with no relevancy or condition logic, is the finding. + +**Guardrail.** A name is not a mechanism. Either implement the filtering — via +relevancy, replication conditions or a replication graph — or rename the type to +what it is. In the meantime, do not put anything in it that matters. diff --git a/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/patterns.md b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/patterns.md new file mode 100644 index 0000000..c8cba46 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-cosmetics-and-teams/references/patterns.md @@ -0,0 +1,288 @@ +# Patterns: cosmetics and teams in a measured reference product + +A worked example of both halves of `SKILL.md` against one reference project read +as source. The two systems are presented together because the audit found what +the skill claims: they touch at exactly one point, and both are well built on +either side of it. + +This file is unusually positive. That is the finding, not a tone: the +architecture here is the strongest in the audited project, and the defects are +narrow and specific. Separating "this design is worth copying" from "this +implementation is finished" is the entire job. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from +measured facts; **[open]** — not answerable from source, because the call site is +in a binary asset. + +--- + +## 1. The two-level split, and what it buys + +```text +controller component owns the desired part list, survives respawn + │ on possession change: remove from old, apply to new, keep handles + ▼ +pawn component replicated list, spawns presentation, computes tags +``` + +The controller component subscribes to possession changes **only on the +authority**, and if a pawn already exists at subscription time it invokes its own +handler with a null previous pawn **[measured]**. That second detail is the +pattern's whole robustness: the same code path handles "pawn already here" and +"pawn arrives later", so ordering stops mattering. + +Transfer removes by handle from the old pawn and re-applies to the new, guarding +against double-add by checking whether the handle is already valid **[measured]**. + +**[derived]** This generalises well beyond appearance. Any state that belongs to +the *player* and is realized on the *body* — loadouts, perks, temporary buffs — +has the same shape and the same failure mode if you skip it: the state dies with +the pawn. + +The link from controller to pawn is found dynamically each time, with no cache +**[measured]**. If the pawn has no cosmetics component the lookup returns null and +nothing happens, silently — a tolerance that is correct for optional cosmetics and +would be wrong for anything required. + +--- + +## 2. Replicating intent, with the split declared in the type + +One structure carries both server-only and client-only fields, distinguished by +annotation rather than by having two types **[measured]**: + +| Field | Replicated | Lives on | +|---|---|---| +| the part itself: class plus socket | yes | server → clients | +| the handle returned to the caller | **no** | server only | +| the spawned presentation component | **no** | client only | + +**[derived]** This is the cleanest available demonstration of "replicate intent, +realize locally". The wire carries a class reference and a socket name; the +receiving machine reconstructs the appearance. Late join is deterministic because +the state is compact and complete. + +Change handling is deliberately coarse: the change callback destroys and respawns +rather than propagating a diff, and the authors say so in a comment +**[measured]**. Recipe: CT-02. + +Part identity is defined as the pair (class, socket) and deliberately **excludes** +the collision mode from the comparison **[measured]** — so two requests differing +only in collision are the same part. **[derived]** That is a real decision about +what identity means, made once, in the comparison function where it belongs. + +--- + +## 3. The isolation guarantee, stated precisely + +One line excludes presentation from the dedicated server, and it is the only such +line in the module **[measured]**. + +**[derived]** The consequence is strong and worth stating exactly: on a dedicated +server the cosmetic actors do not exist, so they cannot collide, block a trace, +tick, or influence the authoritative simulation in any way. That is a real +architectural guarantee obtained for one `if`. + +It is also narrower than the sentence people repeat. On a listen server or in +standalone the parts exist on the authority, with whatever their class contains, +and nothing in the type system prevents a part class from carrying an ability +system or enabled collision. The other protections — collision defaulting to off, +authority-only mutators, no ability-system includes anywhere in the module +**[measured]** — are conventions and content discipline. Recipe: CT-01. + +**The honest form:** *not present on a dedicated server*, not *cannot affect +gameplay*. + +--- + +## 4. The ten-line pattern worth stealing + +An array of rules, each an asset plus a set of required tags, plus a default. +Gather tags, walk the rules in order, first rule whose tags are **all** present +wins, otherwise the default **[measured]**. + +It appears at least three times in the audited project — body mesh, forced +physics asset, and weapon animation layer **[measured]** — and costs roughly ten +lines each time. + +**[derived]** Its sharp edge is that priority is array order, with no weights and +no specificity ranking. A broad rule above a narrow one silently shadows it, and +the failure produces a plausible wrong asset rather than an error. Recipe: CT-03. + +The tag namespace is declared in config and is deliberately small — a body style +and two animation styles **[measured]**. **[derived]** Small is the right size: +every tag in this namespace multiplies the rule matrix that has to be reasoned +about. + +--- + +## 5. Where the two systems touch + +Exactly one place, and the code says so. The cosmetics component broadcasts a +change with a comment noting that observers may need to apply team colouring +**[measured]**; the team display asset applies material parameters to an actor +**and its child actors by default** **[measured]** — and cosmetic parts are child +actors. + +**[derived]** That default is what carries team colour onto modular parts. It is +one boolean default, and it is the seam. + +**[open]** The actual call is made from a Blueprint: the applying function has no +C++ caller. Which leads to the most important methodological point in this file. + +--- + +## 6. What could not be answered, and why that matters + +Four things in this subsystem have **no C++ call site** **[measured]**: + +- who applies team colours to an actor; +- who selects the weapon animation layer from cosmetic tags; +- where the cosmetics component is added to a pawn; +- where the controller component is added to a controller. + +All four are wired in Blueprint or data assets, which are binary and were not +read. + +**[derived]** The correct conclusion is *"the wiring is in a place this audit +cannot see"*, and the incorrect one — available, tempting, and wrong — is *"this +function is unused"*. Both components are marked as spawnable from Blueprint +**[measured]**, which is positive evidence for the first reading. + +This is the general rule from `ue-reference-project-adoption`, met in its natural +habitat: **absence of a C++ caller is not absence of a caller.** Anyone deleting +"dead" code in a project with a large Blueprint surface needs the editor's +reference viewer, not a text search. + +--- + +## 7. Teams: one source, mirrors that refuse loudly + +Membership lives on the player state, replicated push-based, authority-only, with +a non-authority write logged as an error **[measured]**. + +Every mirror — controller, bot controller, local player, pawn — reads through to +the source, and their setters **log an error naming the real owner** rather than +silently doing nothing **[measured]**. + +**[derived]** This is the single most copyable line in the team system. A silent +no-op setter on a mirror is indistinguishable from a working one until something +depends on it; an error line converts an invisible failure into a searchable one, +at a cost of one statement. Recipe: CT-06. + +The pawn mirror is more subtle and also correct: it accepts a direct write **only +while unpossessed**, and refuses with an explanatory error once a controller owns +it **[measured]**. On possession it takes the controller's value and subscribes to +its change delegate; on unpossession it unsubscribes and calls an overridable hook +whose default is "no team" **[measured]**. + +**[derived]** That hook is an extension point placed exactly where a neutral +faction would need it, which is the difference between a design and an +implementation that happens to work. + +--- + +## 8. Resolution, comparison, and the one permissive branch + +Resolution is a four-step cascade — interface, instigator, team-info actor, +associated player state **[measured]**. **[derived]** The instigator step is what +lets projectiles and area effects carry the shooter's team without owning one; a +design that omits it grows a team field on every projectile. + +Comparison returns three states rather than a boolean **[measured]**, which is +correct and is what makes the next finding visible at all. + +The damage rule permits four ways **[measured]**, and the third is the finding: +when the relationship is indeterminate, damage is allowed if the target has an +ability system component — with an author's comment marking it temporary until a +training dummy gets a team assignment. + +**[derived]** It exists to make one specific actor damageable and it generalises +to every unassigned actor with an ability system. Harmless while the only such +actor is a dummy; a rule violation the moment a neutral faction or a destructible +objective exists. **Solve it in data — give the dummy a team — rather than in the +rule.** Recipe: CT-08. + +Friendly fire is promised by a header comment and implemented nowhere: a +project-wide search for any friendly-fire symbol returns **zero** **[measured]**. +Allies are unconditionally safe. Recipe: CT-07. + +--- + +## 9. One policy function, two consumers — the pattern to copy + +The same permission function is called by the authoritative damage calculation +and by the client-side decision to show a hit marker **[measured]**. + +**[derived]** This removes an entire bug class by construction. "The client showed +a hit and the server dealt no damage" requires two implementations of one rule to +disagree, and here there is only one implementation. Any project with predicted +feedback should copy this arrangement before copying anything else in the team +system. + +The damage calculation consumes the result as a **0-or-1 multiplier** in the +formula rather than as an early return **[measured]**. **[derived]** The team rule +becomes an ordinary participant in the arithmetic, alongside distance and surface +attenuation, which keeps the execution pipeline uniform and makes the rule easy to +soften later — a scalar between zero and one is already expressible. + +--- + +## 10. Presentation as named parameters + +The display asset is a bag of named material parameters — scalars, colours, +textures, a short name — not a set of assets **[measured]**. Materials are +authored to expect the names, so adding a team costs one asset rather than a +duplicated material tree. + +Two defects in the applying code, both narrow: + +- **Alpha is dropped** by the material paths, which convert a four-channel colour + to a three-element vector, while the effects path preserves all four + **[measured]**. The same asset therefore behaves differently in two consumers. + Recipe: CT-09. +- **A viewer identity parameter is accepted and ignored**, with a comment saying + so **[measured]**. The signature and header describe viewer-relative colouring; + the implementation returns the absolute asset. Recipe: CT-10. + +Reactive binding is done well: an observation action broadcasts **immediately on +activation** so a late subscriber gets current state rather than waiting for the +next change, and a second action re-subscribes correctly when the observed team +changes **[measured]**. **[derived]** Both are small patterns that eliminate a +race, and both are reusable anywhere state is observed through a key that can +itself change. + +--- + +## 11. What to copy, in order + +1. **One policy function for authority and prediction** (§9). Highest value, + lowest cost, removes a whole bug class. +2. **Loud mirrors** (§7). One log line per setter. +3. **Controller-owned intent, pawn-owned realization** (§1), including the + already-possessed case. +4. **Non-replicated fields inside a replicated entry** (§2) — one type instead of + two. +5. **Rules-plus-default selection** (§4), with a shadowing check added. +6. **Immediate first broadcast on subscription** (§10). + +What to fix while copying: the indeterminate branch (CT-08), the friendly-fire +gap (CT-07), the alpha conversion (CT-09), the ignored viewer parameter (CT-10), +and the privacy that is a name rather than a mechanism (CT-11). + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6: roughly 1300 +lines of cosmetics across 11 files and 1800 lines of teams across 22, read as +source in a single workspace. Source addresses stay in the research archive that +produced this skill; each `CT-` identifier resolves back to the audited location +there. + +## Evidence boundary + +Four wiring call sites are in Blueprint or data assets and were not read; they are +marked open above and nothing here depends on them. Everything else describes one +project on one engine version. Re-run the recipes against your own tree before +trusting any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-data-driven-architecture/SKILL.md b/plugins/ue-design-skills/skills/ue-data-driven-architecture/SKILL.md new file mode 100644 index 0000000..840e023 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-data-driven-architecture/SKILL.md @@ -0,0 +1,413 @@ +--- +name: ue-data-driven-architecture +description: >- + Design or review data-driven architecture in Unreal Engine: definition layers, + data asset versus zero-node Blueprint class defaults versus instanced + fragments and actions, primary asset identifiers, soft and hard and class + references, asset bundles, modular composition, validation, and the boundary + between code and data. Use when adding game modes, pawn archetypes, ability + packages, item and equipment definitions, playlists or feature plugins, or + when a manager is growing per-variant branching. +--- + +# UE data-driven architecture + +Data-driven does **not** mean "put fields in a data asset". It means: + +> Code owns lifecycle, algorithms, replication and invariants. Data selects and +> composes implementations at explicit architectural layers. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes for silent data-graph failures: +[failure modes](references/failure-modes.md). + +**The characteristic failure of a data graph is that it compiles.** A wrong +reference, an unregistered identifier, a missing bundle and a copy-pasted +prototype are all valid data. Validation is the only compiler this layer will +ever have, and building it is part of the architecture rather than a follow-up +task. + +Related skills: `ue-modular-gameplay`, `ue-gas-architecture`, +`ue-input-architecture`, `ue-ui-architecture`, `ue-cosmetics-and-teams`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Start with layers, not asset classes + +Before creating an asset, name the level of composition it owns. A robust +product usually needs several narrow definitions, not one omniscient +configuration: + +```text +catalog / playlist +├─ where and how to launch (map, session, display metadata) +└─ gameplay definition + ├─ feature plugins to enable + ├─ reusable action sets + ├─ mode-specific delta + └─ pawn archetype + ├─ pawn class + ├─ capability packages + ├─ input semantics + ├─ policy matrix + └─ camera +``` + +Items form a separate pipeline: + +```text +pickup presentation +└─ item definition + ├─ inline fragments (display, stats, equippable, reticle) + └─ equipment definition + ├─ runtime instance class + ├─ capability packages + └─ actors and components to spawn +``` + +### The layer test + +For every field ask: + +1. Does it describe **launch and catalog UI**, or **runtime gameplay**? +2. Is it common to several modes, or a mode-specific delta? +3. Does it define an archetype, a capability, or presentation? +4. Who consumes it, and at which lifecycle phase? + +If one asset answers all four, split it. + +--- + +## 2. Pick the container by lifecycle + +The storage type is an architectural decision, not a preference. + +| Need | Use | Why | +|---|---|---| +| Standalone immutable record referenced as an object | data asset | object identity, editable in Details | +| Asset-manager discovery, bundles, async loading | primary data asset | stable identifier, bundle rules | +| A class default, a class reference, inheritance, or a runtime instance of that class | zero-node Blueprint subclass | class identity plus editable defaults | +| Polymorphic optional facet embedded in a definition | instanced object, edit-inline | composition without asset explosion | +| Polymorphic load/activate/deactivate operation | instanced action object | behaviour in code, targets in data | +| Dense homogeneous tabular records | data table | rows, import and export, bulk editing | + +### A Blueprint with zero nodes is a data container + +This is the most consequential and least visible fact in this skill. An entire +definition layer can live in Blueprint classes that contain no graph at all, +with values on their class defaults. + +**An inventory that searches only for data-asset instances will not see any of +it.** Count both, or your census of the data layer is wrong by however much that +layer holds. Recipe: DD-01. + +### When not to use a Blueprint class default + +Do not use one merely because the editor offers it. Prefer a data asset when +consumers need an object rather than a class, when no runtime instance of the +subclass is created, when inherited defaults add nothing, and when asset-manager +identity matters more than a class reference. + +--- + +## 3. Pick reference semantics explicitly + +Every reference chooses loading and ownership behaviour. + +| Reference | Use when | Failure mode | +|---|---|---| +| primary asset identifier | the target must be addressed before it is loaded | type or name not registered; resolves to nothing | +| soft object or class pointer | async load, optional content, role-specific content | not loaded when dereferenced; omitted from the build | +| hard pointer | the target is required whenever the owner is loaded | dependency graph grows transitively | +| class reference | data selects an implementation class or its defaults | construction and lifecycle become class-coupled | +| instanced object | the owner exclusively owns a small polymorphic facet | duplication; cannot be shared independently | + +Write the decision into the property metadata and the validation, not only into +a design document. + +### Asset bundles belong beside the soft reference + +```cpp +UPROPERTY(EditAnywhere, meta=(AssetBundles="Client")) +TSoftClassPtr WidgetClass; + +UPROPERTY(EditAnywhere, meta=(AssetBundles="Client,Server")) +TSoftClassPtr AbilityType; +``` + +The definition then declares both **composition** and the **load and cook +boundary** in one place. Do not maintain a separate, undocumented table of which +assets belong to which role. + +Two failure modes follow directly, and both are silent: + +- a soft reference with **no** bundle annotation is not pulled into the build by + its owner, so it resolves to null in a packaged game and works in the editor; +- a bare accessor on a soft reference assumes the bundle worked, and returns null + without complaint when it did not. + +Recipes: DD-03, DD-04. + +### Cross-plugin dependency gate + +For every reference crossing a feature-plugin boundary: is the target plugin +active before this asset resolves? Does the owning bundle pull it into the +build? Can the source deactivate while the target is still referenced? Is the +dependency one-way, or did you just create a cycle? + +--- + +## 4. Code consumer, data selection + +A healthy definition has a narrow code consumer. The consumer owns when loading +happens, authority and client rules, replication, initialization order, apply and +rollback, validation and error handling. + +Data owns which class or asset or tag is selected, numeric and presentation +parameters, variant composition, and role flags where the consumer supports them. + +### The switch test + +If adding a new mode, weapon or pawn requires another branch in a manager, a +definition, action or fragment boundary is probably missing. + +### But do not move an invariant into data + +These stay in code, always: + +- "only the authority grants capabilities"; +- "the archetype must exist before this readiness state"; +- replication rules; +- teardown order; +- "one grant revokes only its own handles". + +**The variant is data; the safety rule is code.** A rule that can be edited by +opening an asset is a rule that will be edited by someone who does not know it is +a rule. + +--- + +## 5. Compose reuse; do not hide it in inheritance + +Prefer flat composition: + +```text +experience = shared input + standard components + standard HUD + mode delta +``` + +over an inheritance chain: + +```text +base → shooter → team → deathmatch +``` + +Composition makes dependencies visible, reusable and diffable. Inheritance +accumulates defaults silently and creates ordering questions nobody asked. + +If your definition type supports Blueprint subclassing, consider **rejecting it +in validation** and pointing the author at composition instead. A reference +product does exactly this, with an error message naming the alternative +— a validator that teaches rather than only refuses. Recipe: DD-05. + +### The symmetry gate for actions + +Before approving any action that adds something at runtime, fill this table: + +| Resource added | Handle retained | Exact rollback | +|---|---|---| +| extension handler | request handle | release handle | +| component | component request or instance | remove only the owned component | +| capability grant | granted handles | revoke only those handles | +| input binding | bind handles | remove each binding | +| mapping context | player plus context | remove from the same player | +| widget or layout | extension or layout handle | unregister or remove | + +**A blank middle column means the action cannot be rolled back.** Do not +advertise runtime-deactivatable modularity until the table is complete and +tested. See `ue-modular-gameplay` for the failure modes this prevents. + +--- + +## 6. Use fragments for optional facets + +Avoid one definition carrying every possible field for every possible item. +Use polymorphic inline fragments and let consumers query by class: + +```cpp +UCLASS(DefaultToInstanced, EditInlineNew, Abstract) +class UItemFragment : public UObject {}; + +UPROPERTY(EditDefaultsOnly, Instanced) +TArray> Fragments; +``` + +The definition stays open to new facets without changing its base class, +irrelevant fields do not appear on unrelated items, and designers compose facets +in the editor. + +Fragment rules: one responsibility each; state whether duplicates are legal; +validate required combinations; never depend on array order unless the order is +named and documented; fragments select and configure behaviour, they do not own +lifecycle. + +--- + +## 7. Put cross-cutting policy in a policy definition + +When many capabilities repeat the same relationships, do not copy the same +containers into all of them. Use a matrix selected by the archetype, so the same +capabilities can run under different policies. + +Good candidates: capability relationships, presentation parameter bags, +sensitivity curves, class-to-widget mappings, quality thresholds. + +Bad candidates: authority decisions, anti-cheat rules, irreversible lifecycle +transitions, and the validation logic itself. + +--- + +## 8. Validation is part of the architecture + +A data graph creates failure classes that code-only systems do not have. Every +definition family needs validation, and it needs to run somewhere other than the +editor. + +### Universal checks + +- required references non-null; +- arrays with no effective entries rejected; +- class derives from the required base or implements the required interface; +- tags valid and constrained to the expected category; +- identifiers resolve; +- soft targets belong to available, cooked plugins; +- duplicate semantic keys detected; +- client-only assets not bundled for the server, and the reverse; +- every action has a complete rollback path; +- production definitions reference no prototype or test content. + +### Cross-family checks + +These are the ones that catch real defects, because they cross an ownership +boundary that no single reviewer owns: + +```text +playlist identifier → resolves to a registered gameplay definition +gameplay definition → archetype has a class, or is explicitly empty +archetype capabilities → input tags match a reachable input config +item equippable fragment → resolves to an equipment definition +equipment definition → capability packages valid for its instance type +pickup presentation → matches the item definition it claims to present +``` + +The last row is not theoretical. In the audited reference, two health pickups +reference the pistol item definition. It compiles, cooks, and ships. Recipe: +DD-07. + +### Editor-only validation is feedback, not a boundary + +If invalid data can reach a packaged build through an automated cook, the +validation must run there too. In the audited reference **eleven of twelve +validation entry points are compiled out of non-editor builds** — which is +normal, correct, and exactly why a commandlet or automation test has to exist +beside them. Recipe: DD-09. + +--- + +## 9. Auditing an existing project + +### A. Inventory both containers + +1. Query the asset registry for instances deriving from the data-asset base. +2. Query the class registry for zero-node classes whose parent is a definition + type. +3. Include every mount root — the primary content root alone is incomplete in a + plugin-based project. +4. Exclude world-partition external actor packages from asset counts. + +**Do not read binary asset files from disk.** Use the asset registry and live +editor APIs. This is a hard constraint, not a preference: a text search over +binary assets produces confident nonsense. + +### B. Build the reference graph + +For each definition record the owner layer, the consumer, the identifier, the +reference kinds, the bundle metadata, and the incoming and outgoing plugin +boundaries. + +### C. Measure behaviour density + +A definition class should have zero or very few graph nodes. A definition +carrying a large graph is secretly both data and manager. + +### D. Trace one complete variant end to end + +Do not stop at class declarations. Pick one real variant and follow it: + +```text +playlist → definition → action sets → archetype +→ capability and input policy → item → equipment → grant +``` + +Record the actual configured values. Generic reflection frequently hides +inherited class-default fields and reports zero properties on a populated +object; read known fields by name from the schema instead of concluding the +object is empty. Recipe: DD-02. + +--- + +## 10. Review checklist + +- [ ] The definition's architectural layer has a name. +- [ ] It has exactly one consumer owner. +- [ ] Container type matches lifecycle. +- [ ] Reference kinds are intentional and written in metadata. +- [ ] Bundle metadata matches client and server use. +- [ ] Reuse is composition, not a chain of inherited defaults. +- [ ] Actions and fragments have focused responsibilities. +- [ ] Apply and rollback symmetry is proven, not assumed. +- [ ] Required links and semantic joins are validated. +- [ ] Cross-plugin references are acyclic and cooked. +- [ ] A real instance was traced end to end. +- [ ] Production data contains no prototype or copy-pasted references. + +If more than two boxes are unknown, the architecture is not data-driven yet — it +is data-shaped. + +--- + +## 11. When to stop + +Introduce a definition layer when at least one is true: several modes reuse the +same composition; content ships after launch; variants are authored by people who +do not write code; a manager is accumulating per-variant branches. + +Do not introduce one because the reference project has it. Every layer costs an +asset type, a validation surface, a loading path and a debugging hop. The +measured benefit of this architecture is that a mode becomes a list rather than a +subclass — if you have one mode, you are paying for an answer to a question you +do not have. + +--- + +## Provenance + +The measured material comes from an inspection of Epic's Lyra Starter Game on +Unreal Engine 5.6, performed through a live editor bridge rather than by reading +files: 84 native data assets across 21 runtime classes, 23 zero-node definition +classes, and 37 inline actions across ten definitions and five action sets. That +method matters — see DD-01, which exists because a purely text-based census of +this layer is structurally incapable of seeing most of it. + +Source addresses and asset paths stay in the research archive that produced this +skill. Each entry carries a stable identifier (`DD-01`, `DD-03`, …) resolving +back to the audited location there. + +## Evidence boundary + +Values obtained through the editor describe one project at one moment. Graph +assets whose internals no safe reader exposed — environment queries, state trees, +blackboards — were listed but not read, and are marked as such rather than +summarised. Re-run the recipes against your own project; the counts here show the +shape of a contrast, not a target. diff --git a/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/failure-modes.md new file mode 100644 index 0000000..c776087 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/failure-modes.md @@ -0,0 +1,401 @@ +# Failure modes: data-driven architecture + +Ten ways a data graph is wrong while remaining perfectly valid. + +The shared property: **data does not compile, so nothing rejects it.** A null +reference, an unregistered identifier, a missing bundle annotation, a +copy-pasted prototype and a field nobody reads are all well-formed assets. They +load. They cook. The only thing that can reject them is validation somebody +chose to write, and that validation is usually compiled out of the build where +it would matter. + +A second property specific to this area, and the reason the first recipe is +first: **most of this layer is invisible to text search.** Definitions live in +binary assets; the values live on class defaults. A recipe that greps source +answers a question about the *schema*, not about the *data*. Several recipes +below therefore require the editor, and say so rather than pretending a grep +will do. + +--- + +## Census + +### DD-01 - Half the definition layer is invisible to the census + +**Mechanism.** A definition layer can be built from two containers at once: data +asset instances, and Blueprint classes with no graph nodes whose values live on +the class default object. An inventory that queries for one container reports a +complete-looking number. + +**Why it is silent.** The number is real, plausible, and internally consistent. +Nothing indicates it is partial — an inventory has no way to report the +containers it did not think to look in. + +**Why the obvious check misses it.** "Find the data assets" is the correct +phrasing of the wrong question. The second container is not a data asset by class, +is not named like one, and appears in the content browser as a Blueprint. In the +audited reference this was not a corner case: **every** gameplay definition, item +definition and equipment definition lived in the second container. + +**Symptom.** Audits, rename sweeps, memory budgets and dependency graphs that are +each individually correct and collectively wrong. Decisions get made on a census +that missed the most important layer. + +**Detect.** Count both containers, from the editor — this cannot be done from +source: + +```python +# editor-side, via the asset registry +ar = unreal.AssetRegistryHelpers.get_asset_registry() +# 1. instances deriving from the data asset base +# 2. blueprint classes whose native parent is a definition type, +# then check node count on the generated class +``` + +From source you can only establish the *schema* — which types exist and which +are meant to be subclassed: + +```bash +rg -n "class \w+ : public UPrimaryDataAsset|class \w+ : public UDataAsset" \ + --glob "*.h" . +rg -n "UCLASS\([^)]*Blueprintable" --glob "*.h" . | rg -i "definition|data" +``` + +**Guardrail.** Define "the data layer" as a set of *types*, not a set of asset +classes, and enumerate instances per type. Record the count per container so a +future reader can see which containers were searched. + +--- + +### DD-02 - Reflection reports an empty object that is full + +**Mechanism.** A generic property dump over a class default object returns +nothing, or almost nothing, for an object whose inherited fields are populated. + +**Why it is silent.** An empty result is a legitimate answer — plenty of objects +genuinely have no properties set. The tool did not fail; it succeeded and +returned nothing. + +**Why the obvious check misses it.** The check *is* the dump. Its output is +authoritative-looking structured data, and the natural conclusion — "this +definition is empty, so the values must live elsewhere" — sends the investigation +in a direction with no bottom. + +**Symptom.** Hours spent looking for where the configuration "really" lives, +followed by a confident and wrong claim that a definition carries no data. + +**Detect.** Never conclude from a generic dump. Read known fields by name, +driven by the schema you extracted from source: + +```bash +# 1. get the field names from the C++ declaration +rg -n -A20 "class \w*ExperienceDefinition" --glob "*.h" . | rg "UPROPERTY" -A1 +``` + +```python +# 2. editor-side, read each named field explicitly rather than enumerating +for field in ("GameFeaturesToEnable", "DefaultPawnData", "Actions", "ActionSets"): + value = cdo.get_editor_property(field) +``` + +A field that a generic dump omitted and a named read returns is the finding, and +it invalidates every conclusion drawn from the dump. + +**Guardrail.** Treat generic reflection as a discovery aid and named reads as +evidence. When reporting "this object has no data", state which method produced +that answer. + +--- + +## Loading and boundaries + +### DD-03 - Soft reference with no bundle annotation + +**Mechanism.** A definition holds a soft reference to content needed only by one +role. Without a bundle annotation, nothing tells the asset manager to pull that +content in when the owner loads. + +**Why it is silent.** In the editor everything is loaded anyway — commonly both +role bundles at once — so the reference resolves and the feature works. The +failure requires a build where only what was asked for is present. + +**Why the obvious check misses it.** The property is correct C++ and the +reference is correct data. The annotation is optional metadata whose absence +looks like every other property that legitimately does not need one. Reviewing +the definition shows nothing missing. + +**Symptom.** A feature that works in the editor and silently does nothing in a +packaged build, because a class pointer resolved to null and the code path that +uses it treats null as "nothing to do". + +**Detect.** List soft references, list annotations, and compare: + +```bash +rg -n "TSoftObjectPtr<|TSoftClassPtr<" --glob "*.h" . | wc -l +rg -no 'AssetBundles="[^"]*"' --glob "*.h" . | sed 's/.*AssetBundles=//' \ + | sort | uniq -c +``` + +The gap between the two counts is your exposure. In the audited reference the +whole project carried **ten** bundle annotations — two client-only and eight for +both roles — against a much larger population of soft references. Most of those +are loaded by other means; the point of the recipe is that it produces the short +list worth checking rather than a verdict. + +**Guardrail.** Annotate at the declaration, beside the reference, and validate +that every soft reference in a definition family either carries an annotation or +is documented as loaded by its consumer. + +--- + +### DD-04 - Bare accessor on a soft reference + +**Mechanism.** The consumer resolves a soft reference with an accessor that +returns whatever is already loaded, and uses the result without checking. + +**Why it is silent.** Inside the pipeline the annotation was designed for, the +asset is always loaded. The null path is unreachable in every configuration +anyone tests. + +**Why the obvious check misses it.** The bundle annotation is visible on the +property, so the loading question appears answered. The accessor call is one +word and reads as a dereference, not as a decision. + +**Symptom.** One feature's content silently missing, in one configuration, with +no log line naming the asset. + +**Detect.** Find bare accessors on soft references in consumer code: + +```bash +rg -n "\.Get\(\)" --glob "*.cpp" . | rg -i "class|widget|ability|component" +rg -n -A2 "\.Get\(\)" --glob "*.cpp" . | rg -c "if \(|ensure|check" +``` + +The finding is a resolve with no null branch. Compare against the same file's +other resolves — in the audited reference one loop tested for null and the +adjacent loop did not, which is the clearest possible evidence that the check was +intended. + +**Guardrail.** Either load explicitly with a failure path, or check and log at +warning level with the asset name. "The bundle guarantees it" is a claim about +one pipeline, not about the code. + +--- + +## Composition + +### DD-05 - Inheritance used where composition was intended + +**Mechanism.** A definition type is Blueprint-subclassable, so authors subclass +one definition from another to reuse its values, producing a chain of inherited +defaults. + +**Why it is silent.** It works. Inherited defaults are a supported feature, and +the child behaves as the sum of the chain. + +**Why the obvious check misses it.** Each subclass is individually reasonable and +locally minimal — it changes two fields. The cost is structural: values now come +from several assets that must be opened in sequence, and a change to a parent +silently alters every descendant. + +**Symptom.** A definition whose effective configuration cannot be read from the +definition. Diffs that show nothing while behaviour changes. + +**Detect.** Find whether the type permits subclassing, and whether validation +objects: + +```bash +rg -n "UCLASS\([^)]*Blueprintable" --glob "*.h" . | rg -i "definition" +rg -n -B2 -A6 "IsDataValid" --glob "*.cpp" . | rg -i "parent|subclass|inherit" +``` + +The audited reference does this well and is worth copying exactly: its +definition validation **rejects Blueprint subclasses of Blueprint definitions** +and the error text names the alternative — use composition via action sets. A +validator that says what to do instead is worth several that only refuse. + +**Guardrail.** Decide per definition type whether subclassing is legal. If it is +not, reject it in validation with a message naming the composition mechanism. + +--- + +### DD-06 - An action set that is really a parent class + +**Mechanism.** Reusable slices are introduced, and then one slice starts +depending on being applied after another, or on values from a third. + +**Why it is silent.** Order is stable in practice, because it comes from a list +that rarely changes. The coupling produces correct behaviour for as long as +nobody reorders anything. + +**Why the obvious check misses it.** Composition was adopted specifically to +avoid ordering questions, so nobody looks for one. The dependency is not +expressed anywhere — it exists only in the fact that the current order works. + +**Symptom.** Reordering a list, or reusing one slice without its neighbour, +breaks a feature that has no visible relationship to the change. + +**Detect.** Look for actions that read state another action writes: + +```bash +rg -n "class UGameFeatureAction_\w+" --glob "*.h" . -A30 \ + | rg "FindComponentByClass|GetSubsystem|Find\w+ForActor" +``` + +Every hit is an action that consumes something it did not create — a candidate +ordering dependency. For each, ask what happens if it runs first. + +**Guardrail.** An action set is a set, not a sequence. If order matters, make it +explicit — a phase, a prerequisite declaration, or a single combined action — and +never leave it implied by list position. + +--- + +## Content correctness + +### DD-07 - Production data references the wrong definition + +**Mechanism.** A presentation asset is duplicated to make a new one, and the +reference inside it is not updated. The copy points at the original's definition. + +**Why it is silent.** Both assets are valid, both load, and the reference is a +real object of the right type. Nothing in the graph is broken; it simply says +something untrue. + +**Why the obvious check misses it.** Validation checks that references are +non-null and correctly typed, which this passes. Confirming it means comparing +the *meaning* of two assets — that a health pickup should reference a health +item — and no type system encodes that. + +**Symptom.** Picking up one thing and receiving another. In the audited +reference, two health pickups reference the pistol item definition. It ships. + +**Detect.** This one is a cross-family check, and it needs the editor: + +```python +# editor-side: for each pickup definition, compare its presentation identity +# against the item definition it references +for pickup in pickup_definitions: + item = pickup.get_editor_property("InventoryItemDefinition") + # heuristic that actually works: does the pickup asset's name share a token + # with the item definition's name? +``` + +The heuristic is crude and that is the point — a crude cross-family check finds +copy-paste errors that no type check can. Add it to validation and let it produce +warnings a human triages. + +**Guardrail.** Every cross-family reference gets a semantic check, even a weak +one. Also: keep prototype content out of production roots so that a naming-based +check has a chance. + +--- + +### DD-08 - Fields that exist, are authored, and are never read + +**Mechanism.** A definition exposes extension fields for future use. They are +never wired to a consumer. + +**Why it is silent.** An unread field costs nothing and looks like configuration. +Its default is usually the correct behaviour, so leaving it empty is right. + +**Why the obvious check misses it.** The field is declared, is editable, and +appears next to fields that do work. "Is this used?" answers yes on the +declaration. Only searching for a *reader* answers correctly — and unlike the +same failure in C++, the reader may be in a Blueprint graph that no text search +can inspect, so a negative result is weaker here. + +**Symptom.** A designer configures a field and observes no effect. Because the +consumer may be in a binary asset, the investigation cannot conclude from source +alone, and often stops with "it must do something". + +**Detect.** From source, find declaration-without-reader candidates; then confirm +in the editor: + +```bash +F='ExtraArgs' +rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" . +``` + +A single hit — the declaration — makes the field a candidate. **Do not conclude +it is dead**: check the editor's reference viewer for Blueprint consumers before +deleting. In the audited reference several such fields were empty in every +instance, which is evidence about usage but not about wiring. + +**Guardrail.** Do not ship an editable field with no consumer. If it exists for +forward compatibility, mark it as such where the designer sees it, not only in a +commit message. + +--- + +## Validation + +### DD-09 - Validation compiled out of the build that needs it + +**Mechanism.** Data validation is written as editor-only, so it runs when an +author opens the asset and never during an automated cook. + +**Why it is silent.** It is the correct default and it works — authors do get +feedback. Nothing degrades; the check simply is not present in the pipeline where +bad data would otherwise be caught. + +**Why the obvious check misses it.** Searching for validation finds it, in +quantity, well written. The question that matters is narrower: *does any of it +run outside the editor?* — and answering that means reading preprocessor guards +rather than validation logic. + +**Symptom.** Invalid data reaches a packaged build. The team's belief that "we +validate our data" is true and irrelevant. + +**Detect.** Count validation entry points, then count the ones that survive a +non-editor build, then look for a runner: + +```bash +rg -n "IsDataValid" --glob "*.h" . | wc -l +rg -n -B3 "IsDataValid" --glob "*.h" . | rg -c "WITH_EDITOR" +rg -n "Commandlet|AutomationTest|ValidateAssets" --glob "*.h" --glob "*.cs" . +``` + +In the audited reference this gave twelve entry points, **eleven** of them +editor-only, plus a content-validation commandlet and a validator base class — +so the mechanism to run them in automation exists and the question becomes +whether the pipeline invokes it. + +**Guardrail.** Editor validation is feedback; a commandlet or automation test in +the build pipeline is the boundary. Ensure the runner **fails the build** on +error — a validator whose result is only printed is a log line, not a gate. + +--- + +### DD-10 - The validator reports and the pipeline does not care + +**Mechanism.** Validation runs in automation and prints errors. The process exit +status does not reflect them, so the pipeline continues. + +**Why it is silent.** Everything works: the check runs, the errors are found, the +log contains them. The only missing link is the one nobody looks at until a bad +asset ships. + +**Why the obvious check misses it.** Both halves are present and correct — there +is validation, and there is a pipeline step running it. The defect is in the +contract between them, which is a single integer nobody reads. + +**Symptom.** A green build over a log full of validation errors. Confidence +inversely proportional to actual coverage. + +**Detect.** Check what the runner returns and what the pipeline does with it: + +```bash +rg -n -A20 "class UContentValidationCommandlet|int32 Main\(" --glob "*.h" --glob "*.cpp" . \ + | rg "return |ExitCode|GIsCriticalError" +rg -n "Commandlet" .github/ *.yml *.yaml 2>/dev/null +``` + +A commandlet whose `Main` returns zero unconditionally is the finding. Confirm by +deliberately breaking one asset and running the pipeline step: if it stays green, +the gate does not exist. + +**Guardrail.** Prove the gate by breaking something on purpose. **A check that +has never failed is indistinguishable from a check that cannot fail** — this is +the same principle as the poisoned-fixture rule for any linter, applied to your +content pipeline. diff --git a/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/patterns.md b/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/patterns.md new file mode 100644 index 0000000..74e4029 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-data-driven-architecture/references/patterns.md @@ -0,0 +1,268 @@ +# Patterns: a measured data-driven layer + +A worked example of the layering rules from `SKILL.md`, measured in one reference +product. Unlike most files in this bundle it was produced **through a live editor +session**, not by reading source — because the thing being measured is data, and +data lives in binary assets. + +That method is the first lesson and it is not a technicality. Everything below +would have been invisible, or wrong, from a text search. + +Markers: **[measured]** — read from live objects through the editor; +**[derived]** — conclusion from measured facts; **[open]** — not readable with +the tools available. + +--- + +## 1. The census, and why the obvious census is wrong + +| Container | Count **[measured]** | Visible to a data-asset query | +|---|---:|---| +| Native data-asset instances, 21 runtime classes | 84 | yes | +| Zero-node Blueprint classes used as definitions | 23 | **no** | +| **Total inspected definition objects** | **107** | | + +The 23 break down as ten gameplay definitions, seven item definitions and six +equipment definitions **[measured]**. In other words: **every gameplay, +item and equipment definition in the project is in the container that a +data-asset query does not return.** + +**[derived]** A census that asks "which data assets exist?" returns 84, is +internally consistent, and misses the layer that actually composes the game. This +is the single most portable warning in this file. Recipe: DD-01. + +The same structure produced a second trap. A generic property dump over one of +those class defaults reported almost nothing, because inherited fields were not +enumerated **[measured]**. Reading the *named* fields from the schema returned +full configuration. **[derived]** An empty reflection result is evidence about the +reflection call, not about the object. Recipe: DD-02. + +--- + +## 2. The layer split, as actually built + +```text +playlist (8 instances) map + definition + session and menu metadata + └─ gameplay definition (10) plugins + archetype + action sets + delta + ├─ action set (5) reusable slice + └─ archetype (6) pawn class, capabilities, input, policy, camera + ├─ capability package (12) + ├─ input config (5) + └─ policy matrix (1, nine rules) +``` + +Two properties of this split are worth copying. + +**The playlist knows about maps and menus; the definition does not.** One mode +can run on several maps, one map can host several modes, and menu metadata never +enters the runtime composition **[measured]** — all eight playlists were verified +to carry map, definition, visibility, default flag and player count, and nothing +gameplay-facing. + +**Reuse is a list, not a parent.** Two shooter modes differ only by a +mode-specific delta — a different capability package, different scoring and music +components, different score widgets — while sharing input, components and HUD as +action sets **[measured]**. Adding a mode is editing a list. + +That is enforced rather than merely encouraged: validation **rejects a Blueprint +subclass of a Blueprint definition** and its error text names composition as the +alternative **[measured]**. A validator that says what to do instead is worth +several that only refuse. Recipe: DD-05. + +--- + +## 3. Composition is dependency injection, spelled as data + +37 inline actions across ten definitions and five action sets **[measured]**. The +shape of the most common one: + +```text +(target actor class, component class, apply on client?, apply on server?) +``` + +One mode adds scoring on both roles, music on the client only, and bot spawning, +team setup and spawn rules on the server only — **without modifying the game state +class** **[measured]**. A second mode swaps three of those and keeps the rest. + +**[derived]** This is the payoff of the whole architecture, and it is worth +stating plainly: the base classes contain no branch per mode, because the modes +are not known to them. The switch test from the skill is passed by construction. + +The same mechanism carries UI: one action pushes a layout into a layer, another +registers eleven widgets into named slots **[measured]**. Mode-specific UI is two +more rows in a list. + +--- + +## 4. Where the role boundary is declared + +Soft references in the action data carry bundle annotations, so the definition +declares composition **and** the cook boundary in one place **[measured]**. + +The measured population is small and the number is instructive: **ten annotations +in the entire project — two for the client, eight for both roles** **[measured]**. + +**[derived]** That is not a criticism; most content is loaded by other means. It +is a statement about how narrow the annotated surface is, which matters because +the consumers of those references use a bare accessor that returns null when the +bundle did not deliver, and one of the two loops checks the result while its +neighbour does not **[measured]**. The annotation and the null check are the two +halves of one contract, and only one half is consistently present. Recipes: +DD-03, DD-04. + +--- + +## 5. The capability layer, and why input is not one-to-one with it + +12 capability packages **[measured]**, ranging from an eleven-capability hero +package down to packages that grant **only effects** and no capabilities at all. + +The instructive detail: **not every granted capability has an input tag** +**[measured]**. Death, spawn effects, auto-reload and auto-respawn are activated +by events and policies, not by a button. + +**[derived]** This is the concrete reason to keep the semantic input layer +separate from the capability layer rather than folding one into the other. A +design that assumes "capability implies binding" has no place to put the four +above, and will grow a special case for each. + +Six archetypes reuse the same capability packages across different pawn classes +and input configurations **[measured]** — one archetype reuses the shooter +capability package while changing pawn class and input config, which demonstrates +that implementation and capability are genuinely independent axes. + +One archetype has a **null** pawn class **[measured]**. **[derived]** It is a +sentinel used as a configuration fallback, not an archetype — worth recognising +because a validator that requires a non-null class must special-case it, and a +validator that does not will be silently satisfied by every broken archetype. + +--- + +## 6. Fragments: composition inside a definition + +Item definitions carry a display name and a list of instanced polymorphic +fragments; consumers query by fragment class **[measured]**. The measured pipeline: + +```text +pickup presentation (data asset) +└─ item definition (zero-node class) + ├─ display, quickbar, pickup, reticle fragments + ├─ ammo statistics + └─ equippable fragment → equipment definition (zero-node class) + ├─ runtime instance class + ├─ capability package + └─ actor to spawn, at a named socket +``` + +Three weapons differ only in these values **[measured]** — different instance +class, capability package, spawned visual, reticle fragment and ammo numbers — +with no change to the inventory or equipment managers. + +**[derived]** The fragment list is what keeps the item definition from growing a +boolean and five fields for every item type that will ever exist. It is the same +move as the action list one layer up, applied inside a single asset. + +--- + +## 7. The two content defects, and what they teach + +**A health pickup that references the pistol item definition.** Two of them +**[measured]**. Both assets are valid, correctly typed and non-null; the graph is +intact and says something untrue. + +**[derived]** No type check can catch this, and no validation in the project does. +The only mechanism that would is a weak semantic check — does a pickup's identity +share anything with the item it presents? — which feels too crude to write until +you see what it catches. Recipe: DD-07. + +**Prototype content in the production tree.** One pickup carries a display name +and visual assets belonging to a different weapon entirely **[measured]**. + +**[derived]** Prototype content in a production root is what makes the crude check +above hard to write, because naming stops being a signal. The two defects +compound: keep prototypes out of production roots and a naming heuristic becomes +viable. + +Both are recorded here for the reason the whole bundle exists: **these shipped in +a reference project that many teams copy from.** Copying the architecture is +correct; copying the content is not. + +--- + +## 8. Fields that exist and are never set + +All eight playlists carry extra launch arguments and a replay flag. In every one +of the eight, the arguments are empty and the flag is false **[measured]**. + +**[open]** Whether anything reads them cannot be settled from source, because a +consumer could be in a Blueprint graph. Recorded as open rather than concluded — +which is the honest form of this answer and the same discipline as any empty +search result. + +**[derived]** Regardless of wiring, a field that is empty in every instance is an +extension surface rather than a feature, and a reader who assumes otherwise will +look for behaviour that is not there. Recipe: DD-08. + +--- + +## 9. Validation: present, well built, and mostly not in the pipeline + +| Fact **[measured]** | Reading | +|---|---| +| 12 validation entry points across the definition families | The discipline exists and is applied broadly | +| **11 of them compiled out of non-editor builds** | Feedback for authors, not a gate on the build | +| A content-validation commandlet and a validator base class exist | The mechanism to run them in automation is present | + +**[derived]** This is the correct default arrangement, and it means the whole +question reduces to one that source cannot answer: does the build pipeline invoke +the commandlet, and does it fail on error? A validator whose result is printed +rather than returned is a log line. + +Worth being precise about, because an earlier form of this material overstated +it: **structural validation exists in this project.** What does not exist is +semantic cross-family validation — nothing checks that a pickup presents the item +it claims to. The correct criticism is narrow, and the broad version ("the data +graph has no compiler") is wrong. Recipes: DD-09, DD-10. + +--- + +## 10. What to copy, and in what order + +1. **Separate the launch catalog from the runtime definition.** One asset type, + large payoff, no dependencies on anything else here. +2. **Reusable slices as flat lists, and reject definition inheritance in + validation** with a message naming the alternative. +3. **Fragments for optional facets**, so item definitions do not grow a field per + item type. +4. **Bundle annotations beside soft references**, plus a null check at every + consumer — the two halves of one contract. +5. **A cross-family semantic check**, however crude, run in automation with a + non-zero exit status. + +Items 1–3 are architecture and transfer directly. Items 4–5 are the parts the +audited reference left incomplete, and are therefore the parts most likely to be +inherited by anyone copying it. + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6 through a live +editor bridge on a single day: 84 native data assets across 21 runtime classes, +23 zero-node definition classes, 37 inline actions, all eight playlists, six +archetypes, twelve capability packages and five input configurations, read as +live objects and class defaults. No binary asset was parsed from disk; no +gameplay session was run. + +Asset paths and configured values stay in the research archive that produced this +skill. Each `DD-` identifier resolves back to the audited location there. + +## Evidence boundary + +Graph assets — environment queries, state trees, blackboards — were enumerated +but their internals were **not** read, because no safe reader was available. +They are listed as present and nothing is claimed about their contents. + +Everything here describes one project at one moment through one tool. Counts are +quoted to show the shape of a contrast, not as targets. Re-run the census against +your own project, and count both containers. diff --git a/plugins/ue-design-skills/skills/ue-evidence-discipline/SKILL.md b/plugins/ue-design-skills/skills/ue-evidence-discipline/SKILL.md new file mode 100644 index 0000000..1119c79 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-evidence-discipline/SKILL.md @@ -0,0 +1,242 @@ +--- +name: ue-evidence-discipline +description: >- + How to produce and check claims about an Unreal Engine codebase so they can be + trusted later: separating measured from derived from open, why an empty search + result proves nothing, verifying a headline claim yourself before repeating + it, treating a delegated report as raw material, and writing detection recipes + that have been run. Use when auditing a codebase, reviewing someone else's + audit, adopting a reference project, or writing any document that will be + quoted back at you. +--- + +# Evidence discipline for codebase work + +The invariant this whole bundle rests on: + +> A claim without a source you can point at is not a result. A skill that is +> confident and wrong is worse than a missing one. + +Every other skill here is normative — it tells you how to design something. This +one is methodological: it tells you how to know whether what you just wrote is +true. Read it once before using the others, and again before writing an audit +someone else will act on. + +Ten ways this discipline fails in practice, each with a detection recipe: +[failure modes](references/failure-modes.md). They are all mistakes made during +the work that produced this bundle, which is the only reason they are specific +enough to be useful. + +--- + +## 1. Three kinds of statement, never mixed in one sentence + +| Marker | Meaning | Test | +|---|---|---| +| **measured** | read in the source, config or a live object | someone else can open the same place and see it | +| **derived** | a conclusion from measured facts plus engine semantics | the facts are cited and the inference is stated | +| **open** | not answerable with the tools used | says what would answer it | + +**Do not mix two of these in one sentence.** The overwhelming majority of false +conclusions in codebase work are born exactly there: a measured fact and an +inference joined by a comma, where the reader inherits the confidence of the +first half for the second. + +Concretely, this is wrong: + +> The pool is only configured for one tier, so the other tiers are broken. + +and this is right: + +> The pool is configured for one tier **[measured]**. The other tiers therefore +> inherit engine values **[derived]** — whether that is wrong for this product is +> **[open]** until someone measures the target hardware. + +The second version is longer and it is the only one you can act on. + +### Corollary: an absence has a marker too + +"There is no lag compensation in this project" is a measured claim **only** if +you say which searches you ran. Otherwise it is derived from your search +strategy, which brings us to the next rule. + +--- + +## 2. An empty search result proves the pattern, not the codebase + +This is the single most expensive mistake available in this work, because it +produces confident negative claims that survive review. + +A search returns nothing when the code is absent — and also when the pattern was +wrong, when the path was wrong, when the file type was excluded, when ignore +rules filtered the directory, or when the identifier lives in a binary asset. +The tool reports the same empty result for all six. + +**Before writing "this does not exist":** + +1. widen the pattern and confirm it can match something you know is there; +2. check the path and include plugins, not only the primary source directory; +3. check file-type filters in both directions — too narrow hides code, too wide + counts compiled debug symbols as source; +4. for anything that could live in a binary asset, state that the search cannot + see it. + +Two real corrections from the audit behind this bundle, both caught by re-running: + +- a search for one declaration form found 99 entries; the correct pattern found + **212 across five files**, because plugin configuration uses a syntax that + differs by one character; +- a search for a suspected dead field returned four hits; three were compiled + debug symbol files and the source truth was **one**. + +Both would have shipped as facts. + +--- + +## 3. Verify a headline claim yourself before repeating it + +Any statement that will lead a section, appear in a summary, or be quoted gets +re-read at its source **by the person publishing it** — not accepted from a +teammate, a tool, or an earlier version of the same document. + +This is not distrust. It is that a claim changes shape as it travels: a specific +finding about one function becomes a general statement about a subsystem in one +retelling, and nobody notices because each step was small. + +Three corrections that came from exactly this re-reading, during the work behind +this bundle: + +- "a field that is never read" was actually **written once in a constructor and + never read** — the precise version names a real defect class; the loose version + hides it; +- "the data graph has no compiler" was wrong: structural validation existed, in + quantity. The true statement was narrower — **semantic cross-family validation + is absent** — and the narrow version is the one that tells you what to build; +- a count of definitions "spread over 40+ files" was **31 files**, and the + original number had never been measured. + +Note that in all three the correction made the claim *more* useful, not less. A +weaker true statement beats a stronger false one, and usually points at the fix. + +--- + +## 4. A delegated report is raw material + +Output from a subagent, a script, or a search tool is an input to your judgement, +not a result. Re-verify every headline claim by reading the source directly +before it reaches a document or a conclusion. + +The same rule with a sharper edge: **a report that agrees with what you expected +deserves more scrutiny, not less.** During the work behind this bundle a delegated +search reported three instances of a defect; direct measurement found nineteen. +The report was not wrong in kind, only in scale, which is the hardest error to +notice because nothing about it looks incorrect. + +--- + +## 5. A recipe you have not run is a hypothesis + +Every detection recipe in this bundle was executed against a real codebase before +being written down. That is not thoroughness for its own sake — it is the only +way to find the recipes that are confidently wrong. + +Two examples from this bundle, both caught by running them: + +- a search for "is this field ever assigned" matched the **declaration** + `double LastFireTime = 0.0;`, so it reported a writer where none exists — a + false negative, in the exact case the recipe was written to catch; +- a search for "is this ordering value ever compared" matched the arrow operator + in `Entry->Priority`, so two assignments were counted as comparisons. + +Both recipes looked correct. Both would have told the reader "no problem here" at +precisely the moment there was one. + +**A detector that fails silently in the case it exists for is worse than no +detector**, because it converts an open question into a wrong answer. Run the +recipe against a codebase where you already know the answer, both when the answer +is yes and when it is no. + +--- + +## 6. A check that has never failed may not be a check + +Write the failing case for every gate you add, and confirm it fails. + +The worked example is from this bundle's own tooling. A validator ran green for +months. It checked exactly one condition: that skill documents did not contain an +absolute path from **a different workspace than the one it ran in** — a literal +that could never appear. The check existed, the consumer existed, the report said +zero errors, and nineteen real leaks were present. + +That is category C2 from `ue-reference-project-adoption` — a mechanism whose +consumer exists and whose condition can never become true — found in our own +code rather than in someone else's. It is the reason the delivery gate for this +bundle ships with one poisoned fixture per rule and refuses to pass if any rule +has never reddened. + +Apply the same rule to content-validation commandlets, asserts, editor +validation and CI steps: **break something on purpose and confirm the pipeline +notices.** A gate whose failure path has never executed is a log line with +ambitions. + +--- + +## 7. Write findings down where they survive + +A finding that exists only in a conversation is lost. Context ends; files do not. + +- write the observation at the moment it is made, not at the end of the session; +- put the citation next to the claim, not in a separate index; +- record open questions **as open**, rather than closing them with a plausible + guess; +- when a later measurement contradicts an earlier note, add the correction and + keep both — the fact that the number changed is itself information about how + it was obtained. + +That last point matters more than it looks. Two of the corrections in §3 were +only findable because the original number had been written down precisely enough +to be checked. + +--- + +## 8. Publication checklist + +Before an audit, a review or a skill leaves your hands: + +- [ ] every headline claim re-read at its source by you; +- [ ] measured, derived and open marked, and never mixed in one sentence; +- [ ] every negative claim states the searches that produced it; +- [ ] every recipe executed against a real codebase, in both directions; +- [ ] every gate has a failing case that has been observed to fail; +- [ ] open questions listed as open; +- [ ] counts re-measured rather than carried over from an earlier draft; +- [ ] corrections recorded rather than silently applied. + +--- + +## 9. What this does not give you + +None of the above makes a claim true. It makes a false claim **findable** — by a +reader, by a later measurement, by the person who inherits the document. That is +the whole ambition, and it is worth stating plainly because the alternative +belief is dangerous: a document full of markers and citations can still be +wrong, and a green gate means the form is right, not the content. + +The one thing that reliably catches the rest is the habit in §3: read it yourself +before you say it. + +--- + +## Provenance + +Every rule here was written after violating it. The examples are real, they come +from the audit that produced this bundle, and they are included specifically +because rules stated without the failure that motivated them do not survive +contact with a deadline. + +## Evidence boundary + +This skill contains no claims about any codebase. Its examples describe mistakes +made during one research project and the corrections that followed; they are +offered as illustrations of a class, not as findings about the project they came +from. diff --git a/plugins/ue-design-skills/skills/ue-evidence-discipline/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-evidence-discipline/references/failure-modes.md new file mode 100644 index 0000000..febd8f2 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-evidence-discipline/references/failure-modes.md @@ -0,0 +1,327 @@ +# Failure modes: evidence discipline + +Ten ways an audit produces a confident wrong answer. + +Every entry below happened during the work that produced this skill bundle. Not +"could happen" — happened, was caught, and cost real time. That is the only +provenance this document has and the only one it needs: a method document written +from imagination describes a method nobody has used. + +The shared property is uncomfortable: **each of these produces output that looks +like evidence.** A count, a file list, a quoted line, a green check. The failure +is never a missing answer; it is a well-formed answer to a question you did not +ask. + +Identifiers (`ED-01` and up) are stable. + +--- + +## Searching + +### ED-01 - An empty result read as absence + +**Mechanism.** A search returns nothing, and "nothing" is recorded as "this does +not exist in the codebase". + +**Why it is silent.** Zero hits is a definite-looking answer. It has no error +state, no partial result, and nothing that suggests the pattern rather than the +tree was at fault. + +**Why the obvious check misses it.** The obvious check *is* the search. Verifying +it requires a second, differently-shaped search — which feels redundant precisely +when it is most needed, because the first one was so clear. + +**Symptom.** A claim of the form "there is no X here" that turns out to be "my +pattern did not match X". In this project it happened three times: a structure +assertion missed because the namespace prefix was omitted; a subsystem file +missed because the path was guessed rather than found; and — the one that would +have shipped — a leaked-path audit whose pattern used doubled backslashes and +returned exactly one hit, which was nearly reported as "only one leak" when the +real count was nineteen across two donors. + +**Detect.** Before recording an absence, widen deliberately: + +```bash +rg -n "Scalability::FQualityLevels" . # first attempt: 0 hits +rg -n "FQualityLevels" . # widened: 5 hits +rg -c "PatternPart" . ; rg -c "OtherPart" . # both halves separately +``` + +Then check what your tool excludes by default — ignore files, binary files, +hidden directories — and whether the path you searched is the path that exists. + +**Guardrail.** **An empty result proves the pattern, not the absence.** Record +absences only after a widened search, a path check, and a statement of what the +search could not see. + +--- + +### ED-02 - Build output counted as source + +**Mechanism.** An unrestricted search matches compiled artefacts — debug symbols, +binaries, intermediate files — and their hits are counted alongside source. + +**Why it is silent.** The matches are real. The count is arithmetically correct. +Nothing marks a hit as coming from a file that is generated rather than written. + +**Why the obvious check misses it.** The result list is usually long enough that +nobody reads every line, and the summary count is what gets quoted. The tool even +labels binary matches — in a line most people skip. + +**Symptom.** A field described as "appearing four times" when it appears once. In +this project exactly that: three of four hits were debug symbol files, and the +source truth was a single declaration — which is a much stronger finding, since a +field declared once and never read is a cleaner defect than one used four times. + +**Detect.** Restrict to source globs, always, and check the difference: + +```bash +rg -c "Symbol" . # everything +rg -c "Symbol" --glob "*.cpp" --glob "*.h" . # source only +``` + +If the two numbers differ, the first one was never a fact about your code. + +**Guardrail.** Default to source globs. When a count matters, state which file +types it covers. + +--- + +### ED-03 - A count taken over the wrong root + +**Mechanism.** A census runs over the primary content directory in a project whose +plugins mount their own roots, and reports a total. + +**Why it is silent.** The number is real and internally consistent. Nothing in +the result indicates which roots were not visited. + +**Why the obvious check misses it.** The primary root is the obvious root, and it +is where nearly all hand-authored content lives in a small project. The error only +appears at a scale where checking is expensive. + +**Symptom.** Budgets, dependency graphs and audits that are individually correct +and collectively wrong. In this project the primary root held 2 838 assets while +the full set of 105 mount roots held 17 256 — a factor of six. Separately, a +native-symbol count restricted to the main source directory undercounted by +eighteen for the same reason. + +**Detect.** Enumerate roots before counting within them: + +```python +roots = ar.get_sub_paths("/", False) # returns all mount roots +``` + +```bash +rg -c "PATTERN" Source/ # one root +rg -c "PATTERN" Source/ Plugins/ # all of them +``` + +**Guardrail.** A count states its scope in the same sentence as its value. "3 876 +project-owned assets across 105 mount roots" is a fact; "3 876 assets" is a number +waiting to be misused. + +--- + +## Reasoning + +### ED-04 - Measured and inferred mixed in one sentence + +**Mechanism.** A sentence contains something read in source and something +concluded from it, with no marker separating them. + +**Why it is silent.** The sentence is true. Both halves are defensible. The reader +inherits the conclusion with the same confidence as the observation, which is one +level of confidence too many. + +**Why the obvious check misses it.** Review checks whether claims are correct, not +whether their epistemic status is labelled. A correct inference passes. + +**Symptom.** An inference propagating into other documents as a measurement, +where it can no longer be traced back and questioned. Most false conclusions in +this project's audit originated at exactly this seam. + +**Detect.** This one is answered by format, not by search. Require a marker per +claim — measured, derived, open — and treat an unmarked claim as unreviewed. +Then check the derived ones for the strongest available failure: what would have +to be true for this inference to be wrong, and did anyone check? + +**Guardrail.** Three categories, never mixed within a sentence. An open question +stays written as open rather than closed with a plausible guess — a plausible +guess is indistinguishable from a finding six months later. + +--- + +### ED-05 - A subagent's headline accepted without re-reading + +**Mechanism.** Delegated investigation returns a confident summary. The summary is +used directly. + +**Why it is silent.** Reports are fluent, structured and usually mostly right. +Errors arrive in the same register as correct findings. + +**Why the obvious check misses it.** Reading the report *is* the check, and the +report is internally coherent. Detecting the error requires re-reading the source +the report was derived from — that is, redoing the delegated work at the points +that matter. + +**Symptom.** In this project: an exploration agent reported three sites carrying +absolute workstation paths; the real count was nineteen. Separately, a claim that +the game-phase system lived in a feature plugin, when it lives in the main module. + +**Detect.** For each headline claim in a report, open the cited location and read +it. If the report cites no location, the claim is unverifiable and is dropped +rather than softened. + +**Guardrail.** Delegated output is **raw material**, not a result. Every headline +claim is re-read at the source before it reaches a document or a conclusion. + +--- + +### ED-06 - A number carried forward without re-measurement + +**Mechanism.** A count measured once is quoted in later documents. The tree +changes, or the original measurement was scoped differently, and the number +persists. + +**Why it is silent.** Numbers do not expire visibly. A stale count looks exactly +like a fresh one, and having a number at all suppresses the impulse to take one. + +**Why the obvious check misses it.** Review checks whether the document is +coherent, and it is. The original measurement was correct when taken. + +**Symptom.** In this project, an audit stated "59 native definitions across 40+ +files". Re-measured during depersonalization: **94 definitions across 31 files**, +with the discrepancy caused by ED-03 in the original. The correction was recorded +alongside the original rather than replacing it, because which of the two readings +was direct is itself information. + +**Detect.** Re-run the measurement when quoting it in a new context, and diff: + +```bash +rg -c "PATTERN" Source/ Plugins/ | awk -F: '{s+=$2} END {print s}' +``` + +**Guardrail.** A quoted number carries the command that produced it, so the next +reader can re-run it in one paste. Corrections are recorded as corrections — an +archive that silently self-heals loses the trail of which reading was direct. + +--- + +## Recording + +### ED-07 - A finding that never left the conversation + +**Mechanism.** Something real is discovered, discussed, and not written down. + +**Why it is silent.** At the moment of discovery it feels known. The cost arrives +later, when the context holding it is gone. + +**Why the obvious check misses it.** There is nothing to check. The absence of a +note is not visible from anywhere except the future. + +**Symptom.** The same investigation performed twice. A conclusion remembered +without its evidence, which then cannot be defended or corrected. + +**Detect.** At the end of a session, diff what was concluded against what was +written. Anything in the first list and not the second is already lost — writing +it down later is a reconstruction, not a record. + +**Guardrail.** **Found and not written down is lost.** Write the observation +immediately, with its address, before continuing. A conversation is not a store. + +--- + +### ED-08 - A check whose condition cannot become true + +**Mechanism.** A validator tests for a literal that could never appear — a path +from a different workspace, a name from a previous project, a case the current +tree cannot produce. + +**Why it is silent.** The check runs, passes, and reports success. Green is the +outcome it was built to produce, and it produces it forever. + +**Why the obvious check misses it.** The validator exists, is invoked, and is +listed in the process documentation. Everything about it says "this is checked". +Nobody re-reads a passing check. + +**Symptom.** In this project: a skill validator searched for one hardcoded +absolute path from a *different* project, and only inside one file type. It ran +green for months while nineteen leaked paths from two donors sat in the tree. +That validator became this bundle's worked example — our own code, not the +reference's. + +**Detect.** For every check, produce the input that makes it fail. If you cannot +construct one, the check does not exist: + +```bash +python validate.py # green +# now poison a fixture with the exact thing the rule forbids +python validate.py # must be red +``` + +**Guardrail.** **A rule with no fixture that reddens it is not a rule.** Keep one +poisoned fixture per rule, plus a clean baseline, and assert the rule count so a +rule cannot be dropped silently. + +--- + +### ED-09 - A namespace collision that collapses two sets + +**Mechanism.** Two independent collections use the same identifier prefix. A tool +that merges them by key silently overwrites, and reports success on the survivors. + +**Why it is silent.** Every remaining entry resolves. The tool's output is a +consistent, complete-looking mapping — of a smaller set than it was given. + +**Why the obvious check misses it.** The check verifies that everything present +resolves, which is true. Nothing verifies that everything given is still present. +The count is the only tell, and only if someone compares it against the source. + +**Symptom.** In this project, two skills adopted the same two-letter prefix. Nine +entries vanished from the resolver, which reported "173 shipped, 173 mapped, 0 +unresolved" — a perfectly green run against a set that had lost nine members. It +was caught only because the total failed to grow after adding nine entries. + +**Detect.** Count both sides independently and compare: + +```bash +rg -c "^### [A-Z]{2,4}-" skills/*/references/*.md | awk -F: '{s+=$2} END {print s}' +# compare against what the merging tool reports +``` + +Then check prefix uniqueness directly, and make it a rule rather than a habit. + +**Guardrail.** Any tool that merges by key asserts that its output cardinality +equals its input cardinality. A resolver that cannot lose entries is worth more +than one that reports zero failures. + +--- + +### ED-10 - A recipe published without being run + +**Mechanism.** A detection recipe is written from understanding of the defect +rather than from executing it against a real tree. + +**Why it is silent.** The recipe is plausible, well-formed, and would work if the +world were slightly simpler. It is published in a document whose whole purpose is +to be trusted. + +**Why the obvious check misses it.** Reading the recipe confirms it expresses the +right idea. Only running it reveals that the pattern matches something it should +not, or misses something it should catch. + +**Symptom.** In this project, three recipes were wrong on first draft and each was +wrong in a way that produced a **false negative** — the worst direction. A search +for a field's writer matched the declaration's own initializer and reported "there +is a writer" in exactly the case the recipe exists to catch. A character class +intended to find comparisons matched the arrow operator and reported assignments +as comparisons. A search for a namespace matched nothing because real names carry +a prefix — and zero hits would have read as "this problem is absent here". + +**Detect.** Run every recipe against a tree where you already know the answer, and +check both directions: does it find the known instance, and does it stay quiet +where there is none? + +**Guardrail.** **A recipe that has not been run is a hypothesis.** Publish the +corrected form, and where the first draft failed in an instructive way, publish +that too — the trap is often more useful than the recipe. diff --git a/plugins/ue-design-skills/skills/ue-game-settings-architecture/SKILL.md b/plugins/ue-design-skills/skills/ue-game-settings-architecture/SKILL.md new file mode 100644 index 0000000..43c2472 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-game-settings-architecture/SKILL.md @@ -0,0 +1,321 @@ +--- +name: ue-game-settings-architecture +description: >- + Design or review game settings and performance architecture in Unreal Engine: + setting model trees separated from widgets, reflection-backed property paths, + declarative edit conditions, apply and cancel as a transaction, machine-local + versus player-shared persistence, scalability and console variables and device + profiles, hardware benchmark wiring, and sampled performance statistics. Use + when adding video, audio, input or accessibility settings, cloud saves, + quality presets, frame-rate modes, device tuning or performance overlays. +--- + +# UE game settings architecture + +The invariant: + +> A setting is a **model** with persistence, validation and application +> semantics. A widget is one view of that model. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes: [failure modes](references/failure-modes.md). + +Settings look like the least architectural part of a game and are one of the +most: they span a UI framework, a reflection layer, two persistence stores, the +console-variable system, device profiles and the engine's scalability tables. +Every one of those boundaries fails silently. + +Related skills: `ue-data-driven-architecture`, `ue-ui-architecture`, +`ue-input-architecture`, `ue-modular-gameplay`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Separate model, registry and presentation + +```text +registry (per local player) +└─ collections and pages + └─ setting models + ├─ data source: getter and setter + ├─ display text and description + ├─ edit conditions + ├─ options and validation + └─ apply / store-initial / restore lifecycle + +visual data asset +└─ setting model class → entry widget class +``` + +The screen renders a tree of models through a class-to-widget map; it does not +hand-build a widget per option. The payoff: one model renders in several screens, +search and grouping operate on models, edit conditions are testable without any +UI, and settings code depends on no concrete widget class. + +Resolve the widget by walking the model's **superclass chain**, so mapping one +base class covers every derived setting and a specific override remains possible. + +--- + +## 2. Reflection-backed data sources, and their one cost + +A setting stores a **path** to a getter and setter rooted at the local player, +and resolves it through reflection: + +```text +local player → GetLocalSettings() → GetShadowQuality / SetShadowQuality +``` + +This removes a class per setting. It also removes the compiler from the +relationship. + +### The cost, stated precisely + +A macro that stringifies a function name checks that the **name** exists. It does +not check the return type, the parameter count, or constness. Those failures +surface as a runtime assertion during registry initialization — **and assertions +are compiled out of shipping**, so in a shipping build a broken path is a setting +that silently does nothing. + +Three consequences worth designing around: + +1. validate every path during registry initialization and **fail loudly with the + setting's developer name**, not with a reflection error; +2. run that validation in an automated test, not only when a developer opens the + screen; +3. prefer explicit binding where conversion is complex, where several values + change atomically, or where the setter has side effects worth reading. + +Recipe: GS-01. + +--- + +## 3. Edit conditions are declarative dependencies + +Availability depends on platform traits, primary-player status, other settings, +capabilities, privileges, input device and build configuration. Express it in the +model, never in widget visibility. + +Return more than a boolean: + +```text +editable +disabled (reason shown to the player) +hidden (reason recorded for the developer) +option disabled (one entry removed from a list) +killed (hidden, not resettable, not reported) +``` + +A condition also declares which settings it depends on, so that changing one +setting re-evaluates the others. + +**The two-reason split is the part worth copying.** A disabled setting shows the +player why; a hidden one records why for whoever later asks where it went. Most +hand-built settings screens have neither. + +--- + +## 4. Persistence: machine-local versus player-shared + +| | Machine-local | Player-shared | +|---|---|---| +| Holds | resolution, quality, frame rate, device profile, benchmark result | accessibility, sensitivities, subtitles, language, bindings | +| Backed by | engine user settings, local ini | per-player save game, cloud-syncable | +| Scope | one per machine | one per player | + +### The ownership test + +1. Should this follow the user to another machine? +2. Can two local players need different values at once? +3. Does it change global renderer or audio device state? +4. Can a secondary local player safely apply it? + +### The leak to watch for + +A per-player accessor that returns a **process-global singleton** is the standard +way this abstraction breaks. It compiles, it reads naturally at every call site, +and it silently makes split-screen impossible for every value behind it. The +symptom in a codebase that has it: a proliferation of "only the primary player +may edit this" conditions, which is the architecture apologising for itself. + +Recipe: GS-03. If you find it, the fix is not to remove the conditions. + +--- + +## 5. Apply and cancel is a transaction + +1. store initial values on open; +2. edits update the model, and optionally preview; +3. apply commits and invokes application; +4. cancel restores initial values **and reapplies the prior runtime state**; +5. reset-to-defaults is another staged mutation, not an immediate save. + +### Classify application timing deliberately + +| Setting | Timing | +|---|---| +| volume, UI preview | immediate, restored on cancel | +| scalability channel | preview or apply, per policy | +| resolution and window mode | apply with a confirmation countdown | +| key binding | staged profile plus collision check | +| restart-required | save now, mark pending | + +Applying a heavyweight global change from inside a change notification is a +decision, not a default. If it is done because staging was not written yet, mark +it as such — and treat a comment saying "for now" on a global apply path as a +finding rather than a note. Recipe: GS-05. + +--- + +## 6. From setting to console variable + +The preferred path keeps a typed object in the middle: + +```text +model → typed settings setter → staged quality levels +→ apply → engine scalability API → expanded console variables +→ clamped by device profile +``` + +A typed object gives range validation, persistence, transactional restore, +platform clamping, change notification and a test seam. A widget writing a +console variable directly gives none of them. + +When a setting genuinely maps to a custom variable: resolve and cache safely, +set a defined priority source, clamp, restore on cancel, and document shipping +availability. Never write to a read-only variable and assume it took. + +--- + +## 7. Device profiles are constraints, not presets + +```text +platform default → base profile → mode suffix → user request +→ platform clamp → effective value +``` + +Keep **requested** and **effective** separate, and be able to explain to the +player why a request was clamped. + +Profile selection commonly composes a name from parts — base, mode, user choice — +and falls back through progressively shorter combinations until one exists. That +is a good pattern with one specific hazard: a composition part that is declared, +read, and never assigned. The fallback then silently always takes the shorter +branch, and the feature the part represents does not exist while appearing to. +Recipe: GS-07. + +--- + +## 8. Wire the benchmark, do not merely implement it + +```text +can run? → run benchmark → map indices through thresholds +→ set quality and resolution scale → save and apply +``` + +A function named "should this run at startup" with **no caller** means startup +autodetection does not exist, however complete the rest of the pipeline is. + +Decide explicitly: automatic on first launch or manual only; how hardware changes +are detected; whether user overrides survive; what happens on failure. Then add a +**call-site test**, not only a unit test of the threshold mapping — the mapping +was never the part that was missing. Recipe: GS-08. + +--- + +## 9. Performance statistics: collect once, present many + +```text +engine frame data → one consumer → latest values + sampled ring buffers → widgets +``` + +Rules that keep this honest: + +- one collector, never one per widget; +- a fixed-capacity ring buffer with a **documented sample cadence** — a buffer + size without a cadence describes no window at all; +- units written down: milliseconds, frames per second, bytes per second, percent; +- an unavailable statistic returns *unavailable*, never zero; +- availability in shipping is a deliberate setting. + +A frame-rate counter is not evidence about rendering cost. Use it to notice, not +to conclude. + +--- + +## 10. Reapply after composition changes + +Global scalability and frame pacing may need reapplying after a game mode or +feature set finishes loading, because features can change device profile, +console variables and frontend policy. + +Do it **once, at one named lifecycle point, after the composition is fully +known** — not from each feature. The same entry point should serve the other +trigger that changes the same state, such as a hotfixed device profile. Two +causes, one reaction, one idempotent function. + +--- + +## 11. Guard engine-struct assumptions + +Code that enumerates every member of an engine struct — copying, clamping or +maxing quality levels field by field — should carry a compile-time size guard +with a message telling the upgrader what to review. + +A failing size assertion after an engine upgrade is a **feature**. It is the only +thing standing between a newly added quality channel and being silently ignored +by five functions. Update the guard only after reviewing every one of them, and +know how many there are before you start. Recipe: GS-09. + +--- + +## 12. Review checklist + +- [ ] Setting is a model; the widget is a view. +- [ ] Stable developer name is separate from localized display text. +- [ ] Every property path is validated at initialization and in a test. +- [ ] Machine-local versus player-shared ownership is explicit per value. +- [ ] No per-player accessor returns a global singleton. +- [ ] Apply and cancel restore runtime state, not only stored values. +- [ ] Expensive apply operations have deliberate, documented timing. +- [ ] Requested and effective values are distinguishable. +- [ ] Direct console-variable writes are justified and restorable. +- [ ] The benchmark has a real caller and a stated policy. +- [ ] Statistics are collected once, sampled at a documented cadence. +- [ ] Unavailable is distinct from zero. +- [ ] Reapplication after composition change is centralized and idempotent. +- [ ] Engine-struct assumptions carry compile-time guards. + +--- + +## 13. When to simplify + +A prototype can bind widgets straight to engine user settings. Adopt a model +registry when at least one holds: platforms hide or clamp settings differently; +apply and cancel are required; several local users need separate preferences; +search, categories or dynamic conditions matter; gamepad navigation needs +reusable rows; settings span ini, save games, console variables and subsystems. + +Do not adopt a five-thousand-line framework for three toggles. Adopt it when the +settings screen has become an application subsystem rather than a panel. + +--- + +## Provenance + +The findings come from a source audit of Epic's Lyra Starter Game on Unreal +Engine 5.6 and the settings plugin it ships with — roughly 5700 lines of +framework plus 2100 of project-specific settings backends, read rather than run. +The framework itself contains no reference to the game and is the most portable +component examined in the whole research project; the findings below are about +its **wiring**, not its design. + +Source addresses stay in the research archive. Each entry carries a stable +identifier (`GS-01`, `GS-03`, …) resolving back to the audited location there. + +## Evidence boundary + +One project, one engine version. Concrete widget assets are binary and were not +read, so claims here concern the model layer and its bindings. Values quoted from +configuration describe that project on that day. diff --git a/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/failure-modes.md new file mode 100644 index 0000000..5c5e6bd --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/failure-modes.md @@ -0,0 +1,348 @@ +# Failure modes: game settings and performance + +Nine ways a settings layer stores a value the player chose and does nothing with +it. + +The shared property is unusual and worth stating carefully: **in this area the +framework is usually right and the wiring is usually missing.** Every finding +below sits at a seam — between a model and a property path, between a stored +value and the code that applies it, between a complete pipeline and the caller it +never got. None of them is a design flaw in the mechanism they belong to. + +That matters for adoption, which is what this skill is for: the correct response +to this list is *copy the framework, audit the wiring*, not *write your own*. + +Recipes use `rg` from a project source root and were executed against the audited +project while this file was written. + +--- + +## Binding + +### GS-01 - A property path the compiler does not check + +**Mechanism.** Settings bind to storage through a stringified path to a getter +and setter, resolved by reflection at runtime. A macro checks that the function +*name* exists; nothing checks the return type, the parameter count or constness. + +**Why it is silent.** A mismatched path fails at initialization with an +assertion — and assertions are compiled out of shipping builds. In shipping, a +broken binding is a setting that reads a default and writes nowhere. + +**Why the obvious check misses it.** The macro's name contains "checked", and it +genuinely does check something. Review sees a compile-time-looking construct and +stops. The remaining gap is invisible until the types diverge, which happens +during a refactor rather than at authoring time. + +**Symptom.** In development, an assertion when the settings screen opens — often +dismissed as unrelated. In shipping, a setting that silently does nothing, which +is indistinguishable from a setting that is not implemented. + +**Detect.** Find the path-building macros and confirm the failure mode of the +resolver, then check whether anything validates paths outside the editor: + +```bash +rg -n "GET_FUNCTION_NAME_STRING_CHECKED|FCachedPropertyPath|PropertyPathHelpers" \ + --glob "*.h" --glob "*.cpp" . +rg -n -B4 "ensure|checkf" --glob "*.cpp" . | rg -i "getter|setter|propertypath" +rg -n "UE_BUILD_SHIPPING" --glob "*.cpp" . | rg -i "ensure" +``` + +**Guardrail.** Validate every path during registry initialization, report the +**setting's developer name** rather than a reflection error, and run that +validation in an automated test so it fails a build rather than a play session. + +--- + +### GS-02 - Uniqueness enforced on the identifier, nothing enforced on the pair + +**Mechanism.** Settings carry a stable developer name, and the registry checks it +is unique. Nothing checks that the name still matches the property path it binds. + +**Why it is silent.** Both halves are individually valid: a unique name and a +resolvable path. A setting named for one property and bound to another works +perfectly and is wrong. + +**Why the obvious check misses it.** The uniqueness check exists and passes, +which answers "are the identifiers sane?". Copy-paste between two settings +changes the display text and leaves the path — and the resulting setting looks +correct in every view except behaviour. + +**Symptom.** Two settings that move together, or one that changes something the +player did not select. Reported as "the wrong option is wired up", investigated +in the widget layer. + +**Detect.** Extract name/path pairs and eyeball the mismatches: + +```bash +rg -n -B2 -A6 "SetDevName\(" --glob "*.cpp" . \ + | rg "SetDevName|SetDynamicGetter|SetDynamicSetter" +``` + +Read the triples. A developer name and a getter that share no token is the +finding — crude, and it catches the copy-paste error that produces this. + +**Guardrail.** Assert the correspondence in a test that walks the registry, or +generate the name from the path. A convention nothing enforces is a convention +that has already been broken somewhere. + +--- + +## Ownership + +### GS-03 - A per-player accessor that returns a global singleton + +**Mechanism.** The local player exposes an accessor for its settings; the +implementation returns a process-wide singleton. + +**Why it is silent.** With one local player it is correct in every observable +way. The abstraction is only wrong when a second player exists, which is a +configuration most testing never enters. + +**Why the obvious check misses it.** Every call site reads correctly — a player +asking for its settings. The defect is one line inside an accessor that nobody +reads twice, because accessors are the least interesting code in a file. + +**Symptom.** In split-screen, the second player edits the first player's +settings. Usually discovered not as a bug but as a growing population of "only +the primary player may change this" conditions, added one at a time by people +working around a cause nobody named. + +**Detect.** Read the body of every per-player settings accessor: + +```bash +rg -n -A4 "::Get\w*Settings\(\) const|::Get\w*Settings\(\)" --glob "*.cpp" . \ + | rg -i "::Get\(\)|GEngine->|StaticClass" +rg -n "PlayingAsPrimaryPlayer|IsPrimaryPlayer" --glob "*.cpp" . | wc -l +``` + +The second count is the corroborating signal: a large number of primary-player +conditions in a project that claims per-player settings means the claim is not +true. In the audited project the accessor returns a global object, and the +conditions are numerous. + +**Guardrail.** Decide per value whether it is machine-scoped or player-scoped, +and let the accessor's return type say which. If a value must be global, name the +accessor globally — the honest name prevents the workarounds. + +--- + +### GS-04 - Player preferences stored in machine-scoped storage + +**Mechanism.** A value that is clearly a personal preference — volume, for +instance — lives in the machine-local store because that is where the subsystem +applying it happens to read from. + +**Why it is silent.** It works for one player on one machine, which is the +overwhelmingly common case. The value persists, applies, and survives restart. + +**Why the obvious check misses it.** The placement follows the *implementation* +rather than the *semantics*, and the implementation reason is real: the applying +code is machine-global. A reviewer asking "does this work?" gets yes. + +**Symptom.** Preferences that do not follow the player to another machine, and +two local players who cannot have different values. Discovered when cloud saves +or split-screen are added, long after the placement was set. + +**Detect.** Classify the fields of both stores against the ownership test: + +```bash +rg -n "UPROPERTY\(Config\)" -A2 --glob "*.h" .h | rg "\w+ \w+;" +rg -n "UPROPERTY\(\)" -A2 --glob "*.h" .h | rg "\w+ \w+;" +``` + +For each field ask the four ownership questions from the skill. In the audited +project the audio volumes sit in machine-local storage, and the registry that +builds the audio screen reaches into both stores from adjacent lines. + +**Guardrail.** Place by semantics, then solve the application problem. If the +applying subsystem is global, the setting can still be player-scoped with the +primary player's value applied — that is a deliberate policy rather than an +accident of storage. + +--- + +## Application + +### GS-05 - A global apply invoked from a change notification + +**Mechanism.** Changing a setting immediately calls the heavyweight global apply +path from inside the change handler, rather than staging the value until the +player presses apply. + +**Why it is silent.** The result is correct and even feels responsive — the +player sees the change instantly. Cancel still restores the value, so the +transaction is not broken, only leaky. + +**Why the obvious check misses it.** It is three lines in an edit condition, +placed where a reviewer expects a condition rather than an action. The cost only +appears with a setting a player can scrub — a slider makes it one global reapply +per tick. + +**Symptom.** Stutter while adjusting a quality setting; a visible flash on values +the player then cancels; and, for expensive channels, a hitch per change. + +**Detect.** Find application calls inside notification paths, and check for the +tell: + +```bash +rg -n -B6 "ApplyScalabilitySettings|ApplySettings|SetQualityLevels" --glob "*.cpp" . \ + | rg "SettingChanged|OnSettingChanged|NotifySettingChanged" +rg -n -i "//\s*TODO.*(for now|immediately)" --glob "*.cpp" . +``` + +The second command is the high-yield one: in the audited project the immediate +apply carries a comment saying "for now", which converts a judgement call into a +confirmed finding. + +**Guardrail.** Stage by default; preview deliberately and cheaply; apply on +apply. Treat "for now" on a global apply path as a scheduled item, not a note. + +--- + +### GS-06 - A stored preference with no application code + +**Mechanism.** A value is exposed, edited, persisted and loaded. The function +that would apply it has an empty body. + +**Why it is silent.** Every visible part of the round trip works: the slider +moves, the value saves, it comes back after restart. Only the effect is missing, +and the effect is subjective — sensitivity, in the audited case. + +**Why the obvious check misses it.** The apply function exists and is called. A +review of the settings layer finds a complete implementation; the emptiness is in +a body that the caller has no reason to open. And because the value *is* readable, +game code may consume it directly elsewhere, making the empty function neither +wrong nor sufficient. + +**Symptom.** A player changes a preference, the UI confirms it, and nothing feels +different. Support cannot reproduce it because the value really is stored. + +**Detect.** Find apply functions with empty bodies: + +```bash +rg -n -A3 "void \w+::Apply\w+\(\)" --glob "*.cpp" . | rg -B2 "^\s*\}" +``` + +Then, for each stored preference, find its consumer: + +```bash +V='MouseSensitivityX' +rg -n "\b$V\b" --glob "*.cpp" --glob "*.h" . | rg -v "UPROPERTY|Set$V|Get$V" +``` + +No consumer outside the accessors means the value goes nowhere. + +**Guardrail.** A setting ships with its consumer, in the same change. If +application is intentionally the game's responsibility rather than the +framework's, delete the empty function — an empty function is a claim. + +--- + +## Wiring + +### GS-07 - A composition part that is read six times and never assigned + +**Mechanism.** A device-profile name is composed from parts — base platform, a +mode suffix, a user choice — with fallback through progressively shorter +combinations. One part is declared as a local, read repeatedly, and never +written. + +**Why it is silent.** The fallback is designed to handle a missing part, and it +does. Every path produces a valid profile name; the composition simply never +takes the branches that use the absent part. + +**Why the obvious check misses it.** The variable appears in six places, +including a log line that prints it. Any "is this used?" search answers yes +loudly. The absent half is the **writer**, and no default tool asks that +question — this is the canonical shape from `ue-reference-project-adoption`, +found here in a subsystem where the consequence is a whole feature that does not +exist. + +**Symptom.** Documentation and code structure both indicate that a game mode can +select its own device profile. It cannot. A team planning per-mode performance +profiles discovers this after designing around it. + +**Detect.** For any composed name, count reads against writes: + +```bash +V='ExperienceSuffix' +rg -n "\b$V\b" --glob "*.cpp" . +rg -n "\b$V\b\s*=[^=]" --glob "*.cpp" . +``` + +Reads with no assignment is the finding. In the audited project the first command +returns six lines — a declaration, four reads and a log — and the second returns +nothing but the declaration itself, which the authors annotated with a comment +saying nothing sets it. + +**Guardrail.** Do not ship a composition part with no producer. If it is +scaffolding for a planned feature, say so where a reader of the *composition* +will see it, not only at the declaration. + +--- + +### GS-08 - A complete pipeline with no caller + +**Mechanism.** Automatic quality detection is fully implemented: a capability +check, a benchmark, threshold mapping, application and save. The function that +decides whether to run it at startup has no caller anywhere. + +**Why it is silent.** The manual path works. A player who presses the button gets +correct auto-detection, so the feature demonstrably functions — just never on its +own. + +**Why the obvious check misses it.** Every component is present and testable, and +unit tests of the threshold mapping pass. "Do we support automatic quality?" +answers yes from any angle except the one that matters: who starts it. + +**Symptom.** First-run experience uses default quality on every machine. Nobody +notices, because developers' machines default acceptably and the button exists. + +**Detect.** For the decision function, distinguish declaration from call: + +```bash +F='ShouldRunAutoBenchmarkAtStartup' +rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" . +``` + +Exactly two hits — the declaration and the definition — is the finding. Extend to +content if the project can call functions from assets, and record the result as +open if you cannot inspect them. + +**Guardrail.** Every entry-point function gets a **call-site test**, not only a +unit test of what it computes. The mapping was never the part at risk. + +--- + +### GS-09 - A size guard whose count nobody knows + +**Mechanism.** Code enumerates every member of an engine struct — copying, +clamping, maximising quality levels field by field — and protects itself with a +compile-time assertion on the struct's size. + +**Why it is silent.** It is correct, and it is good practice: after an engine +upgrade adds a channel, the build breaks instead of silently ignoring it. + +**Why the obvious check misses it.** Nothing is wrong here. The hazard is +procedural: the guard fires during an upgrade, under time pressure, and the +person who bumps the number needs to know **how many functions** must be reviewed. +That number is not written anywhere near the assertion. + +**Symptom.** An engine upgrade where the size constant is updated and one of the +five enumerating functions is not, silently dropping a quality channel — exactly +the failure the guard was written to prevent. + +**Detect.** Count the guards and the functions they protect before upgrading: + +```bash +rg -n "static_assert\(sizeof\(" --glob "*.cpp" --glob "*.h" . +rg -c "static_assert\(sizeof\(Scalability::FQualityLevels\)" --glob "*.cpp" . +``` + +In the audited project this returns **five** sites for one struct. Five is the +number an upgrader needs and would otherwise have to discover by grepping mid-fix. + +**Guardrail.** Write the count and the list of functions into the assertion +message. The guard is only as good as the instruction it gives the person it +stops. diff --git a/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/patterns.md b/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/patterns.md new file mode 100644 index 0000000..0900593 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-game-settings-architecture/references/patterns.md @@ -0,0 +1,229 @@ +# Patterns: a measured settings and performance layer + +A worked example of `SKILL.md` against one reference product. The conclusion is +unusual enough to state at the top, because it changes what you should do with +the rest of the file: + +> **The framework is the best thing in the audited project. The wiring around it +> is where every defect is.** + +Roughly 5700 lines of settings framework contain no reference to the game that +ships it, and no finding in this file is about its design. The nine failure modes +are all in the 2100 lines of project-specific backend and registry that bind it — +paths that the compiler cannot check, a pipeline with no caller, a composition +part with no writer. That distinction is the adoption advice: **take the +framework, audit the wiring**. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from +measured facts; **[open]** — not answerable from source. + +--- + +## 1. What the framework gives you, concretely + +| Capability **[measured]** | Why it is hard to rebuild | +|---|---| +| Setting model fully separated from widgets | The separation is what makes search, grouping and testing possible at all | +| Widget resolved by walking the model's **superclass chain** | Map one base class, cover every derived setting, keep specific overrides | +| Apply / cancel / reset as a real transaction over a dirty set | Correct restore of *runtime* state, not only stored values | +| Edit conditions with five outcomes, not a boolean | disabled-with-player-reason, hidden-with-developer-reason, single option disabled, killed | +| Cascading dependencies between settings | Changing one setting re-evaluates the others without hand-written wiring | +| Full-text search over rich text reduced to plain text | Falls out of the model separation | +| Asynchronous readiness before the screen renders | Audio device lists and cloud values arrive late; the screen waits | +| A debug mode that shows developer names and classes in the UI | Makes every other finding in this file findable | + +**[derived]** The five-outcome edit condition is the piece most worth +internalising even if you write your own framework. Hand-built settings screens +almost universally have a boolean, and the two missing things — *tell the player +why it is disabled* and *tell the developer why it is hidden* — are exactly the +two questions asked later at cost. + +--- + +## 2. The trade the framework makes, and where it lands + +Bindings are **string paths resolved by reflection** — a getter and setter path +rooted at the local player **[measured]**. That removes a class per setting and +removes the compiler from the relationship. + +The failure mode is precise: the stringification macro checks that a function +*name* exists, not its type, arity or constness; mismatches surface as an +assertion at registry initialization; **assertions are compiled out of shipping** +**[measured]**. + +**[derived]** So the worst case is a shipping build where a setting reads a +default and writes nowhere, with no diagnostic anywhere. That is not an argument +against the design — it is an argument for one automated test that walks the +registry and resolves every path. Recipes: GS-01, GS-02. + +--- + +## 3. The ownership split, and the leak inside it + +The design is a clean two-store split **[measured]**: + +| | Machine-local | Player-shared | +|---|---|---| +| Base | engine user settings, ini | per-player save game, cloud-syncable | +| Holds | resolution, quality, frame limits, device profile, benchmark result | colour-blind mode, sensitivities, subtitles, language, haptics | + +The authors state the rationale in a comment: shared settings are safe to put in +the cloud and are stored per player, so controller preferences belong there +rather than in local settings where every user would receive them **[measured]**. + +Then the accessor breaks it. The per-player getter for local settings returns a +**process-global singleton** **[measured]**. + +**[derived]** "This player's local settings" is a fiction; there is one object per +process. The corroborating evidence is structural rather than a second defect: +the registry is dense with "only while playing as the primary player" conditions, +which is the architecture apologising for the accessor. Recipe: GS-03. + +A second, milder leak in the same area: audio volumes live in machine-local +storage **[measured]** — a pure player preference placed by where the applying +subsystem reads, not by what the value means. The audio registry reaches into +both stores from adjacent lines. Recipe: GS-04. + +**[derived]** Both are worth seeing together, because they are the same mistake at +two scales: **storage placed by implementation convenience rather than by +semantics**, then papered over at the call sites. + +--- + +## 4. Three findings that are the same shape + +This is the most transferable observation in the file. Three unrelated features, +one structure: + +| Feature | Present **[measured]** | Missing **[measured]** | +|---|---|---| +| Automatic quality detection at first launch | capability check, benchmark, threshold mapping, apply, save | **any caller** of the decision function | +| Per-game-mode device profiles | name composition, fallback chain, four candidate forms, logging | **any writer** of the mode suffix | +| Input sensitivity application | storage, dirty tracking, save, load, an apply function | **a body** in the apply function | + +**[derived]** Each is complete except for one link, and in each case the missing +link is invisible to every check that the rest of the feature passes. The +decision function has two occurrences — its declaration and its definition. The +suffix has six occurrences, of which one is a log line printing it and none is an +assignment. The apply function is called correctly from the right place and +contains nothing. + +The suffix case deserves emphasis because it is the canonical shape from +`ue-reference-project-adoption` appearing in a subsystem where the consequence is +architectural: **a variable read four times and written zero times**, with the +authors' own comment noting that nothing sets it. Any search asking "is this +used?" answers yes, loudly. Recipes: GS-06, GS-07, GS-08. + +--- + +## 5. Where the transaction leaks + +Video settings apply **immediately from inside a change notification**, not on +apply, and the code says so with a comment reading "for now" **[measured]**. + +**[derived]** Cancel still restores the value, so the transaction is not broken — +it is leaky in a way the player sees: a preview they did not ask for, on a global +apply path, potentially once per slider tick. + +The comment is what turns this from a judgement call into a finding. **A "for +now" on a global apply path is a scheduled item.** Recipe: GS-05. + +--- + +## 6. The upgrade guard, and the number it does not tell you + +Manual enumeration of an engine struct's members — copying, clamping and +maximising quality levels field by field — is protected by a compile-time +assertion on the struct's size, at **five separate sites** for one struct +**[measured]**. + +**[derived]** This is good practice and it works: after an engine upgrade adds a +quality channel, the build breaks rather than silently ignoring it. The gap is +procedural. The person who hits the assertion during an upgrade needs to know +*how many functions* must be reviewed, and that number is nowhere near the +message. Recipe: GS-09. + +**Put the count in the assertion text.** A guard is only as useful as the +instruction it gives the person it stops. + +--- + +## 7. The reapplication point, and why it is where it is + +Global scalability is reapplied at the **last line** of the mode-load completion, +after the load state is set and after all three waves of load notification +**[measured]**. The same function serves the other trigger that changes the same +state — a hotfixed device profile **[measured]**. + +Four reasons, all measured, and together they are a small lesson in ordering: + +1. feature plugins activated by the mode can change console variables and + frontend policy, so effective values are unknown until they have all run; +2. the device profile itself may change, and the reapplication reselects it; +3. scalability is process-global, not per world — the authors note this in a + comment explaining why they do not track it per world in multi-instance + editor play; +4. it is guarded off for servers, where render settings are meaningless. + +**[derived]** Two different causes — mode loaded, profile hotfixed — reduced to +one idempotent reaction at one named point. That arrangement is worth copying +directly, and it is the opposite of letting each feature reapply what it thinks +it changed. + +--- + +## 8. Performance statistics: one collector, and the honest limits + +Statistics come from the engine's own frame-data consumer interface rather than +from hand-rolled timers **[measured]**, feeding a ring buffer for graphing. +Eighteen statistics, with a compile-time count assertion in three places +**[measured]**. + +**[derived]** The design point worth copying is that presentation and collection +are separate: widgets read a cache, and no widget collects. The design point worth +questioning before copying is that a buffer size is meaningless without a +documented sample cadence — a hundred-odd samples describes a window only once you +know the interval. + +Two hard couplings to be aware of before lifting the subsystem: one statistic +reads a game-specific state class, and latency statistics are gated on a specific +GPU vendor because the project integrates one vendor's plugin **[measured]**. +Both are small and both are visible; neither is a reason to rewrite the rest. + +--- + +## 9. What to copy, in order + +1. **The framework itself**, if the UI dependency is acceptable. Five-outcome edit + conditions, superclass-chain widget resolution, the change tracker, async + readiness. +2. **The reapplication point** (§7) — one idempotent function, two causes. +3. **The two-store split** (§3) — but decide placement by semantics and check + your accessor's body before trusting the word "local". +4. **The size-guard discipline** (§6), improved by writing the count into the + message. +5. **Collector/presenter separation** for statistics (§8). + +Audit before shipping: every property path resolves (GS-01, GS-02); the accessor +is not a singleton (GS-03); every apply function has a body (GS-06); every +composition part has a writer (GS-07); every pipeline has a caller (GS-08). + +**[derived]** That list is short, mechanical, and would have caught every defect +in this file. None of it requires understanding the framework — which is the +point of shipping recipes rather than conclusions. + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6 and the settings +plugin shipped with it: roughly 5700 lines of framework, 2100 of project settings +backends, 2000 of registry, and 800 of performance code, read as source in a +single workspace. Source addresses stay in the research archive that produced +this skill; each `GS-` identifier resolves back to the audited location there. + +## Evidence boundary + +Widget assets are binary and were not read, so nothing here describes the +presentation layer beyond its base classes. Configuration values quoted describe +that project on that day. Re-run the recipes against your own tree. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-cues/SKILL.md b/plugins/ue-design-skills/skills/ue-gameplay-cues/SKILL.md new file mode 100644 index 0000000..00d6866 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-cues/SKILL.md @@ -0,0 +1,376 @@ +--- +name: ue-gameplay-cues +description: >- + Design, implement or review the GameplayCue presentation layer in Unreal + Engine GAS: what a cue is and when it is the right mechanism, Executed versus + OnActive/WhileActive/OnRemove forms, the three addressing paths, Static versus + Actor notifies, the dedicated-server rule, cue asset discovery and loading + cost, registering cue paths from feature plugins, teardown and leaked looping + cues, and the tag-to-asset validation gap. Use when adding VFX or audio + feedback to abilities or effects, decoupling presentation from gameplay + assets, debugging a cue that never fires or never stops, or auditing cue + loading cost. +--- + +# UE gameplay cues + +The invariant: + +> A cue is a one-way, cosmetic, tag-addressed broadcast. Gameplay must produce +> exactly the same result if every cue in the project fails to fire. + +Worked analysis of a real cue subsystem, including the parts that were built and +never reached: [patterns](references/patterns.md). +Twenty detection recipes for silent cue failures: [failure modes](references/failure-modes.md). + +Read the failure modes before adopting or reviewing a cue-based design. Each entry +names the mechanism, why it stays silent, why the obvious check misses it, and a +command you can run against your own tree. + +Related skills: `ue-gas-architecture`, `ue-gameplay-tag-governance`, +`ue-modular-gameplay`, `ue-asset-loading-and-memory`, `ue-gameplay-messaging`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. What a GameplayCue is, and the problem it solves + +A GameplayCue is a **presentation event addressed by a GameplayTag** rather than +by an asset reference. + +Without it, the chain looks like this: + +```text +GE_Fire_Damage --hard reference--> NS_Fire_Impact (Niagara) + SW_Fire_Impact (audio) +``` + +The gameplay asset now owns the VFX and the audio. Loading the effect loads the +particles. A dedicated server, which needs the damage numbers and nothing else, +loads them too. An artist changing the impact effect touches a gameplay asset. + +With a cue, the chain is cut: + +```text +GE_Fire_Damage --tag--> "GameplayCue.Fire.Impact" <--tag-- GCN_Fire_Impact +``` + +The gameplay asset references a **string-like identifier**. A separate registry +maps that tag to a notify asset, discovered by scanning content directories. The +gameplay asset has no idea the presentation exists. + +What this buys: + +- gameplay assets carry no presentation payload, so they are small and + server-safe; +- presentation can be added, replaced or shipped in a separate plugin without + touching gameplay; +- the whole presentation layer can be absent on a dedicated server. + +What it costs — and this is the part that gets skipped: + +- **the join is no longer checked by anything.** A tag with no notify is silence. + A notify with no matching tag is a never-loaded asset. Neither is an error. +- discovery is directory-scan based, so cue assets must be found, indexed and + loaded by a manager that has its own lifecycle and its own failure modes; +- the mechanism is one-way and unreliable by design, which makes it wrong for + anything that must be correct. + +Treat cues as an application of `ue-gameplay-tag-governance` with a dedicated +loading subsystem attached. Every rule about tag joins applies here first. + +--- + +## 2. The three forms + +| Form | Fires when | Use for | +|---|---|---| +| **Executed** | a one-shot event | impacts, hit reactions, one-off flashes | +| **OnActive** | a duration effect starts, on clients present at that moment | start of a looping effect | +| **WhileActive** | an actor becomes relevant while the effect is already running | joining/relevancy catch-up for the same loop | +| **OnRemove** | the duration effect ends | stopping the loop, cleanup | + +### Rules + +1. **`OnActive` and `WhileActive` must produce the same visual state.** A client + who was present and a client who arrived late must see the same thing. Putting + the "start" logic only in `OnActive` means late joiners see nothing. +2. **Everything `OnActive`/`WhileActive` spawns, `OnRemove` destroys.** Looping + cues are the main source of leaked VFX in GAS projects. +3. **Never use `Executed` for anything with a duration**, and never use the + duration forms for a one-shot — the recovery paths differ. +4. **Assume `OnRemove` may run without `OnActive` having run** on that client, and + the reverse. Write both defensively. + +--- + +## 3. Static versus Actor notifies + +| Notify type | Instanced | State | Use for | +|---|---|---|---| +| **Static** | no — a CDO handles the call | none | one-shot effects, sounds, decals | +| **Actor** | yes — an actor is spawned per instance | yes | looping effects, anything requiring cleanup | + +### Rules + +1. **Default to Static.** It allocates nothing and cannot leak. +2. **A Static notify must not store per-instance state.** It runs on the class + default object; a member written during one invocation is shared by every + subsequent one, including other players' invocations. +3. **Use an Actor notify only when there is something to tear down** — a looping + particle system, an attached component, a timeline. +4. **Actor notifies are actors**: they have relevancy, they cost replication + consideration if not marked otherwise, and they can outlive their owner if + removal is missed. + +--- + +## 4. The three ways a cue gets addressed + +```text +1. GameplayEffect GE lists cue tags; ASC fires them with the effect +2. Direct ASC call ExecuteGameplayCue / AddGameplayCue / RemoveGameplayCue +3. IGameplayCueInterface the actor handles cue tags itself +``` + +### Rules + +1. **Prefer route 1 for anything tied to an effect's lifetime.** The effect and + its presentation then start and stop together by construction. +2. **Route 2 is for presentation with no effect behind it** — a UI-driven flash, a + client-local confirmation. It is manual, so its removal is manual too. +3. **Route 3 is for actor-specific handling** that no generic notify can express. + It couples the actor to the cue vocabulary; use sparingly. +4. **Do not mix routes for one cue tag.** Two owners of the same presentation + produces double effects that appear only when both paths trigger. + +--- + +## 5. The network rule + +**Cues never run on a dedicated server.** The presentation layer does not exist +there. + +Consequences that must be designed for, not discovered: + +1. **A cue may not carry gameplay side effects.** Damage, score, state changes, + spawning — none of it may live in a notify, because on a dedicated server that + code never executes. This is the single most important rule in this document. +2. **Cue parameters are a lossy channel.** They are packed for replication; + whatever your conversion helper drops is dropped silently on every client. + Verify the round trip for every field you rely on, especially tag containers. +3. **Cue delivery is unreliable.** Packet loss, relevancy changes and joining + mid-effect all produce missed cues. Design for a missed cue being invisible, + not desynchronising. +4. **Cue assets belong in the client asset bundle only.** Putting them in the + server bundle loads presentation content on a machine that will never use it. + +### Test that the rule holds + +Run the feature on a dedicated server with the presentation content deliberately +absent. Gameplay results must be byte-identical. + +--- + +## 6. Registering cue paths from feature plugins + +A feature plugin that ships cue assets must register its content directory with +the cue manager, or its cues are invisible. + +The important structural point: **cue path registration must happen earlier than +feature activation.** The manager builds its tag-to-asset index during its object +library scan; a path added after that scan is not indexed until the next rescan. +This is earlier than any of the standard activation hooks. + +The pattern that solves it: + +> **The feature action is a pure data declaration; a lifecycle observer is the +> executor.** + +The action object declares only the directory list and validates it in the +editor. A separate observer, registered with the feature policy, listens for the +earliest lifecycle phase and performs the registration for every action of that +type it finds. + +### Rules + +1. **Register on the earliest available phase**, not on activation. +2. **Batch the rescan.** Add all paths, then rebuild the index once — one rescan + per feature, not one per directory. +3. **Refresh the primary-asset registration when the path set changes**, on the + way in *and* on the way out. Skipping it on removal leaves the asset manager + describing directories that no longer exist. +4. **Implement removal and assert it.** Count the paths removed and compare + against the paths added; a mismatch means the registry is drifting across + activation cycles. +5. **Use the same manager instance on both sides of the lifecycle.** Registering + through your subclassed manager and unregistering through the engine base + class is a silent asymmetry — it compiles, runs, and leaves entries behind. + +--- + +## 7. Loading and cost + +Cue assets are discovered by scanning directories and are, by default, loaded on +demand — which means a hitch the first time each cue fires. + +The mitigation is preloading, and it is where most complexity in a cue manager +lives: async library loading, a preload set, always-loaded tags, reference +tracking, and hooks for GC and map transitions. + +### Rules + +1. **Decide the load policy explicitly and verify it is reachable.** A load-mode + switch that short-circuits every branch renders the entire preload subsystem + dead code — the machinery exists, is maintained, is never executed, and any + diagnostic command reports zero because there is genuinely nothing there. +2. **A constant is not a knob.** A local `const bool` inside a namespace whose + name ends in `Cvars` is not a console variable and cannot be changed at + runtime. If a policy must be tunable, register it as an actual console + variable; otherwise do not imply that it is one. +3. **Preload the small set that must never hitch** — the cues on the critical + path, hit feedback above all — and let the rest load on demand. +4. **Put cue paths in the client bundle only** (see §5). +5. **Measure the cue set size before optimising it.** Discovery cost scales with + the number of scanned directories; load cost scales with the assets actually + referenced. + +--- + +## 8. Teardown + +Cues attach to an avatar. When the avatar changes — death, possession change, +respawn — cues that are still active belong to nothing. + +The correct teardown sequence when unbinding an ability system from an avatar: + +```text +1. verify the ASC's avatar is actually this actor +2. cancel abilities +3. clear ability input bindings +4. remove all gameplay cues +5. clear the avatar +``` + +### Rules + +1. **Check avatar identity before tearing down.** A component that unbinds + unconditionally can rip the ASC away from an avatar it no longer owns. +2. **Remove all cues explicitly.** Do not rely on the removal of the effects that + added them; direct-call cues (route 2) have no effect behind them. +3. **Test the cycle twice.** Leaks and duplicates appear on the second respawn, + not the first. + +--- + +## 9. The validation gap, and what to do about it + +Nothing in the engine checks that: + +- a cue tag referenced by an effect has a notify asset; +- a notify asset's tag is referenced by anything; +- a cue tag is spelled the way the notify expects. + +All three failures present identically: no visual, no log, no error. + +### Minimum guardrails + +1. **Constrain cue tag properties** with `meta=(Categories="GameplayCue")` so + authoring uses a dropdown. +2. **Write a content validator** that resolves every cue tag in every effect + against the manager's index, and reports both directions of orphan. +3. **Log once, in development builds, when a cue fires with no handler.** This + converts the silent case into a searchable one. +4. **Keep cue tags in one namespace with one owner.** A cue tag placed in the + wrong namespace because "the notify happens to live there" is a rename waiting + to happen — and renames need redirects (see `ue-gameplay-tag-governance`). + +--- + +## 10. Review checklist + +- [ ] No gameplay logic inside any notify. +- [ ] Feature works identically with all cue assets absent. +- [ ] `OnActive` and `WhileActive` converge on the same visual state. +- [ ] Everything spawned in the duration forms is destroyed in `OnRemove`. +- [ ] Static notifies hold no per-instance state. +- [ ] Actor notifies exist only where there is something to tear down. +- [ ] One addressing route per cue tag. +- [ ] Cue assets are in the client bundle only. +- [ ] Cue parameter round-trip verified for every field used. +- [ ] Feature-plugin cue paths registered at the earliest lifecycle phase. +- [ ] Path removal implemented, asserted, and using the same manager instance. +- [ ] Load policy is reachable, and any "knob" is a real console variable. +- [ ] Teardown checks avatar identity and removes all cues. +- [ ] A validator reports orphan tags and orphan notifies. + +--- + +## 11. Tests + +### Behaviour + +- one-shot cue fires once per execution, not per client-side prediction retry; +- looping cue starts on `OnActive`, is matched by `WhileActive` for a late + joiner, and stops on `OnRemove`; +- `OnRemove` without a preceding `OnActive` does not crash or leak; +- cue with no registered notify produces no error and no side effect. + +### Network + +- dedicated server + remote client: gameplay results identical with and without + cue content; +- client joining mid-effect sees the same state as a client present at the start; +- relevancy lost and regained during a looping cue converges to the correct state. + +### Lifecycle + +- death and respawn twice: no duplicated or orphaned cues; +- feature plugin activate, deactivate, activate: cue path count returns to its + original value; +- possession change mid-cue. + +### Cost + +- first-fire hitch measured for a non-preloaded cue; +- preloaded set size and residency reported by a diagnostic command that has been + verified to read a live code path. + +--- + +## 12. When not to use a cue + +Use something else when: + +- the effect must be guaranteed — it is gameplay, not presentation; +- the receiver needs a return value or acknowledgement — direct call; +- the event drives UI state that must survive a missed message — replicated + state plus a local message (`ue-gameplay-messaging`); +- exactly one known actor reacts, in the same module — direct call or delegate. + +A cue earns its cost when presentation is genuinely optional, genuinely +many-to-many, and genuinely absent on the server. Used for anything else, it is a +tag join with no validation and unreliable delivery — which is a description of a +bug that has not happened yet. + +--- + +## Provenance + +The failure modes and the worked analysis in `references/` come from a +line-by-line audit of Epic's Lyra Starter Game on Unreal Engine 5.6, read as +source rather than run. The cue subsystem there is unusually instructive because +much of its preload machinery is present, maintained, and unreachable in the +shipped configuration — a shape that reading the class list will never reveal. + +Source addresses stay in the research archive that produced this skill; what ships +is the detection recipe. Each entry carries a stable identifier (`GC-01` and up) +that resolves back to the audited location, so any specific claim can be produced +on request. + +## Evidence boundary + +The measured material refers to one project on one engine version. It is evidence +of failure modes, not a guarantee about other versions. Re-run the detection +recipes against your own tree before acting on any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-cues/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-gameplay-cues/references/failure-modes.md new file mode 100644 index 0000000..5d2f20b --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-cues/references/failure-modes.md @@ -0,0 +1,665 @@ +# Failure modes: GameplayCues + +Twenty ways a cue-based presentation layer fails without saying so. + +The shared property of this list, stated once so it is not repeated twenty times: +**a cue defect is silent by construction.** The mechanism is a tag join with no +validation, delivering an event that is *allowed* to be missed, to a layer that +does not exist on the server. Every diagnostic has to be built deliberately, +because the engine will not volunteer one. + +Each entry gives the mechanism, why it stays silent, why the obvious check misses +it, the symptom a human reports, a `Detect` recipe you can run against your own +tree, and the guardrail. Identifiers (`GC-01` and up) are stable and resolve back +to the audited source in the research archive. + +Detection recipes use `rg` (ripgrep). They were run against the audited project +while this file was written; where a first-draft recipe gave a wrong answer, the +corrected form is the one printed here and the trap is called out. + +--- + +### GC-01 - Gameplay logic inside a notify + +**Mechanism.** Damage, scoring, state changes or spawning placed in a cue notify, +because the notify is where the designer was already working. + +**Why it is silent.** In the editor and on a listen server the notify runs on the +same process that owns gameplay, so the write lands and the feature appears +correct. Nothing anywhere declares that a notify is presentation-only. + +**Why the obvious check misses it.** Code review of the notify reads as reasonable +gameplay code, and it *is* reasonable — in the wrong file. The defect is not in the +statement, it is in the file the statement is in. No compiler, linter or test that +runs in-editor can see it. + +**Symptom.** The feature works for everyone during development and simply does not +happen on a dedicated server. Reported as "the server build is broken". + +**Detect.** Look for writes inside notify classes: + +```bash +rg -n --type cpp -g '*Cue*' \ + "(ApplyGameplayEffect|AddScore|->Server|SpawnActor|SetHealth|=\s*true;)" \ + Source/ +``` + +Stronger, and the one that actually settles it: run the feature on a dedicated +server with every cue asset removed. Gameplay results must be identical. + +**Guardrail.** A notify may read state and produce audiovisual output, nothing +else. Review every notify for writes, and keep the dedicated-server-without-cues +run in the acceptance list for any feature that ships a cue. + +--- + +### GC-02 - Cue tag with no notify asset + +**Mechanism.** An effect references a cue tag; no notify is registered for it — +a typo, an unscanned directory, or a plugin that failed to register its path. + +**Why it is silent.** An unhandled cue is a legal runtime state: the engine +delivers the event to a registry that has no entry, and returns. Missing +presentation is exactly what "presentation is optional" means, so there is nothing +for the engine to complain about. + +**Why the obvious check misses it.** The tag exists, the effect compiles, the asset +cooks. Both halves of the join are individually valid; only their *relation* is +broken, and nothing in the toolchain owns that relation. + +**Symptom.** Silence. No visual, no log, no error. The artist is told the effect +"doesn't work" and starts editing a notify that was never being called. + +**Detect.** Compare the two sides of the join. Tags referenced by content versus +tags claimed by notifies: + +```bash +rg -no "GameplayCue\.[A-Za-z0-9_.]+" Content/ Source/ | sort -u +``` + +Feed both sets into a diff. Then prove your tooling works by deliberately +misspelling one cue tag and confirming the check reports it — a validator nobody +has seen fail is not known to work. + +**Guardrail.** Constrain cue tag properties with `meta=(Categories="GameplayCue")` +so authoring uses a dropdown; add a content validator that resolves every +referenced tag against the runtime index; log once per unhandled cue in +development builds. + +--- + +### GC-03 - Notify asset that no tag references + +**Mechanism.** The reverse orphan: a notify exists for a tag nobody fires, +usually left behind by a rename or a cancelled feature. + +**Why it is silent.** An unreferenced notify is still scanned, indexed and +sometimes preloaded. It behaves exactly like a notify waiting for its first +trigger, which is a normal state. + +**Why the obvious check misses it.** Reference-viewer style checks ask "what does +this asset reference?", not "who addresses this asset by string?" Nothing +references a notify by hard pointer — that is the entire point of the design. + +**Symptom.** No symptom at all, which is why it accumulates. It surfaces as +unexplained package count and memory in the cue set. + +**Detect.** Same recipe as GC-02, read in the other direction: tags claimed by +notify assets minus tags referenced by effects. Run the validator both ways in one +pass, and report both orphan directions separately. + +**Guardrail.** One validator, two directions, in CI. Orphan notifies are deleted +in the same commit that finds them, or the list stops being read. + +--- + +### GC-04 - Per-instance state on a Static notify + +**Mechanism.** A Static notify runs on the class default object. A member written +during one invocation persists into every later invocation, including other +players'. + +**Why it is silent.** Single-player and single-source testing never overlaps two +invocations, so the shared member always holds the value the current caller just +wrote. The bug requires concurrency to appear. + +**Why the obvious check misses it.** The class looks like an ordinary object with +ordinary members. Nothing at the declaration site says "this is a singleton" — the +CDO behaviour comes from how the cue manager dispatches, which is in a different +file entirely. + +**Symptom.** One player's cue reads another player's data. Effects that are correct +alone and wrong in multiplayer, intermittently. + +**Detect.** Find mutable members on Static notify classes: + +```bash +rg -n -A40 "class .*GameplayCueNotify_Static" Source/ \ + | rg -n "^\s*(float|int32|bool|F[A-Z]\w+|T\w+<)[^()]*;" +``` + +Any non-const, non-config member on a Static notify is the defect. + +**Guardrail.** Static notifies hold no mutable state. If state is needed, it is an +Actor notify — that is what the distinction is for. + +--- + +### GC-05 - Looping cue that is never removed + +**Mechanism.** `OnActive`/`WhileActive` spawns a persistent effect; `OnRemove` +does not destroy it, or never runs because the owning effect was removed by a path +that skipped cue removal. + +**Why it is silent.** A spawned particle system is a legitimate long-lived object. +Nothing distinguishes "still playing because the effect is active" from "still +playing because nobody stopped it". + +**Why the obvious check misses it.** The first death-respawn cycle usually looks +clean, because the leak needs a second cycle to become visible as duplication. +Most manual testing stops at one cycle. + +**Symptom.** Accumulating VFX and audio across respawns; frame time degrading over +a match; a bug report that says "it gets worse the longer you play". + +**Detect.** For each duration-form notify, check that the removal path destroys +what the activation path spawned: + +```bash +rg -n "OnActive|WhileActive|OnRemove" Source/ -g '*Cue*' -A15 \ + | rg -n "(SpawnSystem|SpawnEmitter|SpawnSound|Destroy|Deactivate)" +``` + +Then count active cue actors before and after two full death-respawn cycles. The +count must return to baseline, not merely stop growing. + +**Guardrail.** Everything the duration forms spawn, `OnRemove` destroys. Write +`OnRemove` to be safe when `OnActive` never ran on that client. + +--- + +### GC-06 - `OnActive` without `WhileActive` + +**Mechanism.** Start-of-loop logic placed only in `OnActive`, which fires only for +clients present at the moment the effect began. + +**Why it is silent.** For every client who was present, the behaviour is perfect. +The failure exists only in the experience of a client who was not, and that client +sees an absence — the hardest thing to notice. + +**Why the obvious check misses it.** Both callbacks exist on the class, so a +structural review ticks the box. The difference between them is a delivery-timing +property of the engine, not a property of the code you are reading. + +**Symptom.** Players who join, or who lose and regain relevancy, see nothing while +everyone else sees the effect. Reproduces only with a real client join, never in a +single-process test. + +**Detect.** Find notifies that implement one callback and not the other: + +```bash +rg -l "OnActive" Source/ -g '*Cue*' > /tmp/active.txt +rg -l "WhileActive" Source/ -g '*Cue*' > /tmp/while.txt +comm -23 <(sort /tmp/active.txt) <(sort /tmp/while.txt) +``` + +Anything listed implements the start and not the catch-up. + +**Guardrail.** `OnActive` and `WhileActive` must converge on the same visual state. +Test by joining a match mid-effect, and by losing and regaining relevancy during a +loop. + +--- + +### GC-07 - Two addressing routes for one tag + +**Mechanism.** The same cue tag is fired both by a gameplay effect and by a direct +ability-system call, or additionally handled through the cue interface on the +actor. + +**Why it is silent.** Each route is individually correct and each was added by +someone solving a real problem. The duplication is a property of the pair, which +no single author sees. + +**Why the obvious check misses it.** Searching for the tag finds both sites, but +they look like they belong to different features — one in an effect asset, one in +C++. Nothing marks a cue tag as having an owner. + +**Symptom.** Doubled effects, appearing only when both paths trigger, often only in +one game mode. Reported as "the explosion is twice as loud in Team Deathmatch". + +**Detect.** For each cue tag, enumerate every firing route: + +```bash +TAG="GameplayCue.Fire.Impact" +rg -n "$TAG" Source/ Content/ Config/ +rg -n "(ExecuteGameplayCue|AddGameplayCue|RemoveGameplayCue)" Source/ -A2 | rg -n "$TAG" +``` + +More than one owner is the defect, even when the current behaviour looks right. + +**Guardrail.** One owner per cue tag, documented at the tag declaration. Search +every route before adding a firing site. + +--- + +### GC-08 - Manual cue added without a manual removal + +**Mechanism.** A direct add call has no effect behind it, so no effect removal will +ever clean it up. + +**Why it is silent.** The add succeeds and the presentation appears. Removal is a +separate call that nobody is forced to write, and the object model does not record +that a removal is owed. + +**Why the obvious check misses it.** The code that adds is present and correct. The +defect is the absence of a second call on every exit path, and absence is what +review is worst at seeing. + +**Symptom.** A looping cue that survives the state that created it — including +across death, possession change and disconnect. + +**Detect.** Count adds against removes: + +```bash +rg -c "AddGameplayCue" Source/ +rg -c "RemoveGameplayCue" Source/ +``` + +An imbalance is a lead, not a verdict — then check each add site for a paired +removal on *every* exit path, not just the happy one. + +**Guardrail.** Direct-call cues need an explicit paired removal on every exit path. +Prefer attaching the cue to an effect's lifetime instead, so removal is structural +rather than remembered. + +--- + +### GC-09 - Teardown without an avatar identity check + +**Mechanism.** A component unbinds an ability system component unconditionally, +without checking that the ASC's current avatar is still this actor. + +**Why it is silent.** The unbind succeeds. The damage is done to a *different* +actor's state — the one the ASC has already moved to — so the error surfaces far +from its cause. + +**Why the obvious check misses it.** The teardown code is correct in isolation and +correct in the common case. It is wrong only in the ordering where the avatar +changed first, which single-process testing rarely produces. + +**Symptom.** Abilities and cues on the *current* avatar are cancelled when an old +one is cleaned up. Looks like a spontaneous ability cancellation with no cause. + +**Detect.** Find unbind paths and check for an identity guard above them: + +```bash +rg -n -B8 "(ClearAbilityInput|CancelAbilities|RemoveAllGameplayCues|SetAvatarActor\(nullptr)" Source/ \ + | rg -n "GetAvatarActor\(\)\s*==|IsValid\(.*Avatar" +``` + +An unbind path with no `GetAvatarActor()` comparison above it is the defect. + +**Guardrail.** Verify the ASC's avatar is this actor before cancelling abilities, +clearing input and removing cues. The audited project does this correctly; the +guard is one comparison and it is worth copying verbatim. + +--- + +### GC-10 - Cue assets in the server bundle + +**Mechanism.** Cue content is assigned to an asset bundle that the dedicated +server loads. + +**Why it is silent.** Loading extra assets is never an error. The server runs +correctly, just fatter, and nothing reports "you loaded a particle system you can +never play". + +**Why the obvious check misses it.** Bundle assignment lives in asset manager +configuration, not next to the cue content. Reviewing the cue does not show you +which bundle it is in. + +**Symptom.** Server memory spent on particles and audio. Discovered, if ever, +during a memory investigation months later. + +**Detect.** Find where cue paths are registered and which bundle constant is used: + +```bash +rg -n "AddGameplayCueNotifyPath|GameplayCueRefs|LoadState" Source/ -B3 -A3 +``` + +The bundle name at that site should be the client bundle. Cross-check against +the asset manager's bundle definitions. + +**Guardrail.** Cue paths belong to the client bundle only. Make the bundle constant +explicit at the registration site so review can see it without opening +configuration. + +--- + +### GC-11 - Load policy that short-circuits the whole subsystem + +**Mechanism.** A load-mode constant causes an early return at every switch site, +including the one that binds the subsystem's delegates. + +**Why it is silent.** Every branch is a legal configuration. "Load everything +upfront" is a sensible mode; the fact that choosing it disables an entire preload +subsystem is an emergent property of three separate switch statements, none of +which says so. + +**Why the obvious check misses it.** The preload machinery exists, is maintained, +and is full of correct code. Reading it tells you nothing about whether it runs. +Worse, the diagnostic command built for it reports **zero**, which reads as "there +is nothing to preload" rather than "this code never executes". + +**Symptom.** A named feature that has no effect. Engineers investigate the preload +implementation for a bug that is not in it. + +**Detect.** Find the mode constant, then find every switch on it, then check which +branch each falls into: + +```bash +rg -n "static .*LoadMode\s*=" Source/ +rg -n "switch\s*\(.*LoadMode" Source/ -A12 +``` + +If the configured value hits an early `return` at a site that also performs +delegate binding, everything below that binding is unreachable. In the audited +project three switch sites short-circuit, and the third returns before its +delegate bindings. + +**Guardrail.** Verify the load policy reaches its intended branch. Treat "the dump +command shows nothing" as a question about reachability, not about content. + +--- + +### GC-12 - A constant presented as a runtime knob + +**Mechanism.** A local `const bool` or a `static` variable placed in a namespace +named after console variables, with no console-variable registration. + +**Why it is silent.** It compiles and behaves as the constant it is. The naming is +the only thing that lies, and naming is not checked by anything. + +**Why the obvious check misses it.** The namespace name contains the word an +engineer greps for. Searching for the knob finds the declaration, which looks +exactly like the seven real registrations elsewhere in the project. + +**Symptom.** Engineers try to change behaviour from the console and cannot. Time is +lost concluding that the console, the build, or the module is broken. + +**Detect.** Compare pseudo-knobs against real registrations: + +```bash +rg -n "^namespace \w*Cvar\w*" Source/ # namespaces that promise knobs +rg -c "FAutoConsoleVariableRef" Source/ # registrations that deliver them +``` + +Note the trap: searching for the literal `namespace Cvars` finds nothing, because +real names are prefixed (`FooManagerCvars`). The first draft of this recipe +returned zero hits and would have been read as "no such problem here". In the +audited project the corrected recipe finds one such namespace against eight real +registrations, and inside that namespace only a console *command* is registered — +the two variables beside it are plain constants. + +**Guardrail.** If it must be tunable, register it as a real console variable. +Otherwise do not name its scope after one. + +--- + +### GC-13 - A collection that is iterated but never filled + +**Mechanism.** A local container is declared, immediately looped over, and never +populated — usually with a comment where the population belongs. + +**Why it is silent.** A loop over an empty container is a no-op, which is +indistinguishable from a loop whose input happens to be empty right now. + +**Why the obvious check misses it.** The loop body is real code, often including +error handling and logging, so the feature reads as implemented. There is even a +warning path for invalid entries — which can never fire, because there are no +entries. + +**Symptom.** A named feature ("always loaded cues") that has no inputs and +therefore no effect, while appearing complete in review. + +**Detect.** Find containers declared and iterated within a few lines, with no +insertion between: + +```bash +rg -n -A6 "T(Array|Set)<\w+>\s+\w+;" Source/ | rg -n "for\s*\(" +``` + +For each hit, search the same scope for `Add|Emplace|Append|Insert` on that name. +None means the loop is decorative. + +**Guardrail.** A loop over an empty local is a review blocker. If the population is +future work, the loop is future work too. + +--- + +### GC-14 - Feature-plugin cue path registered too late + +**Mechanism.** Registration is performed on feature activation, after the manager +has already built its tag-to-asset index. + +**Why it is silent.** Registration succeeds. The path is in the list. It is simply +not in the *index*, and no error distinguishes those two states. + +**Why the obvious check misses it.** In the editor, incidental rescans — asset +saves, hot reloads, content browser operations — rebuild the index often enough +that the cues appear to work. The failure is specific to a cold packaged run. + +**Symptom.** A plugin's cues are invisible in a packaged build and fine in the +editor. The classic "works on my machine" that is actually "works in my editor". + +**Detect.** Find where cue paths are added and which lifecycle phase the call sits +in: + +```bash +rg -n "AddGameplayCueNotifyPath" Source/ -B12 \ + | rg -n "(OnGameFeatureRegistering|OnGameFeatureLoading|OnGameFeatureActivating|Activate)" +``` + +`Registering` is early enough; `Activating` is not. Also check the rescan flag: one +rescan per feature, not one per directory. + +**Guardrail.** Register on the earliest lifecycle phase available, via a policy +observer rather than an activation hook. Batch the rescan: add all paths, rebuild +once. The audited project does exactly this, and its feature action deliberately +has no activation body at all. + +--- + +### GC-15 - Asymmetric manager access across the lifecycle + +**Mechanism.** Registration goes through a subclassed manager; removal goes through +the engine base class. + +**Why it is silent.** Both calls compile and both succeed, because the subclass +inherits the base method. As long as the subclass adds no bookkeeping, the two +paths are equivalent — today. + +**Why the obvious check misses it.** The two lines are in different functions, +often hundreds of lines apart, and both read as "get the manager, call the +method". The asymmetry is only visible when you put the two resolution expressions +side by side. + +**Symptom.** Nothing, until the subclass gains state. Then removal silently leaves +entries behind and the registry drifts across activation cycles. + +**Detect.** Put the add and remove sites next to each other: + +```bash +rg -n "(AddGameplayCueNotifyPath|RemoveGameplayCueNotifyPath)" Source/ -B4 \ + | rg -n "(GetGameplayCueManager|UAbilitySystemGlobals|Cast<)" +``` + +Different resolution expressions on the two sides is the defect. + +**Guardrail.** Resolve the manager the same way on both sides, and assert the +removal count. The audited project asserts the count — that assertion is the +mitigating factor and the part worth copying. + +--- + +### GC-16 - Registry refresh on the way in but not on the way out + +**Mechanism.** The primary-asset registration is refreshed when paths are added and +not when they are removed. + +**Why it is silent.** A stale registration describes directories that no longer +exist. Asset lookups against them fail the same way a genuinely empty directory +fails: nothing found, no error. + +**Why the obvious check misses it.** The refresh call on the add side is visible +and correct. Its absence on the remove side is one missing line in a function that +otherwise does everything right, including asserting its own removal count. + +**Symptom.** Divergence that accumulates across feature activation cycles. Path +counts do not return to baseline after activate, deactivate, activate. + +**Detect.** Count refresh calls against path-set mutations: + +```bash +rg -n "RefreshGameplayCuePrimaryAsset" Source/ -B15 \ + | rg -n "(Registering|Unregistering|Add|Remove)" +``` + +A refresh on the registering path with no counterpart on unregistering is the +defect. Confirm behaviourally: activate, deactivate, activate a feature and +compare path counts. + +**Guardrail.** Refresh on every path-set change, in both directions. + +--- + +### GC-17 - Lossy payload conversion helper + +**Mechanism.** A helper that converts between a message struct and cue parameters +drops fields. + +**Why it is silent.** The conversion produces a valid, well-formed result. The +missing field arrives as its default value, which for a tag container is empty — +and empty is a legal value that means "no tags". + +**Why the obvious check misses it.** The function has a name that promises a +conversion, a signature that mentions both types, and a body that assigns most +fields. Review confirms "yes, this converts". Only a field-by-field round trip +shows the hole. + +**Symptom.** A field populated at the sender and empty at the receiver, on every +client, with no warning. Investigated as a replication bug for as long as it takes +to read the helper. + +**Detect.** Read every conversion function field by field, and treat every +unfinished-work comment inside one as a declaration that the field is not carried: + +```bash +rg -n -A20 "(ToCueParameters|FromCueParameters|VerbMessageTo|CueParametersTo)" Source/ \ + | rg -n "(TODO|FIXME|\?\?\?)" +``` + +In the audited project this finds the same field dropped in both directions, each +marked with its own comment. + +**Guardrail.** Verify the round trip field by field before relying on any field. +Treat any unfinished-work comment in a conversion function as "this field is not +carried". + +--- + +### GC-18 - Concluding "unused" from absent C++ callers + +**Mechanism.** A Blueprint-callable function's call sites can live in binary asset +graphs, which are not text-searchable. + +**Why it is silent.** The search is clean. Zero hits is a definite-looking answer, +and definite-looking answers are the ones people act on. + +**Why the obvious check misses it.** The obvious check *is* the problem: `rg` +searches text, and the callers are not text. The tool answers the question it was +asked — "where does this string appear in source?" — which is not the question the +engineer meant. + +**Symptom.** Code deleted as dead breaks a Blueprint at runtime, discovered after +the fact, usually by a designer. + +**Detect.** For any Blueprint-exposed symbol, treat source search as inconclusive +by construction: + +```bash +rg -n "UFUNCTION\(.*Blueprint(Callable|ImplementableEvent|NativeEvent)" Source/ -A2 +``` + +Anything on that list needs the editor's reference viewer before removal. An empty +`rg` result proves the pattern, not the absence. + +**Guardrail.** "No C++ callers" is not evidence of disuse for a Blueprint-exposed +symbol. Record it as an open question, not as dead code — which is exactly what the +audit did with the helpers from GC-17. + +--- + +### GC-19 - Cue tag parked in the wrong namespace + +**Mechanism.** A tag is declared next to the notify asset that happens to implement +it, rather than in the namespace that owns the concept. + +**Why it is silent.** The tag works. Its location is a naming decision, and naming +decisions have no runtime behaviour — until someone tries to fix one. + +**Why the obvious check misses it.** Nothing validates that a tag's namespace +matches its meaning. The only record that the placement is wrong is, at best, a +comment written by the person who did it. + +**Symptom.** Moving it later requires a redirect. Without redirects, every asset +referencing it silently loses the reference — the rename is not blocked, it is +merely destructive. + +**Detect.** Read the tag declarations for confessions, then check whether the +project has any migration mechanism at all: + +```bash +rg -n "DevComment=.*(needs to move|wrong|temporary|should be)" Config/ +rg -c "GameplayTagRedirects" Config/ +``` + +In the audited project the first search finds a tag whose own comment says it is +misplaced, and the second returns zero — the tag cannot be moved without breaking +references. + +**Guardrail.** Cue tags follow the concept, not the asset location. Add the +redirect in the same commit as any move, and make sure a redirect mechanism exists +before you need it. + +--- + +### GC-20 - Editor measurements treated as representative of cue loading cost + +**Mechanism.** The editor loads both the client and the server asset bundles. + +**Why it is silent.** The measurement completes and produces a number. Numbers +from a profiler carry an authority that their provenance does not. + +**Why the obvious check misses it.** The condition that causes it is one clause in +one branch of the bundle-selection logic, in a different subsystem from the one +being measured. Nobody profiling cue residency reads the experience loader. + +**Symptom.** Memory and load-time figures that describe no shipping configuration. +Cue residency in particular is overstated, and budgets built on it are wrong in the +safe direction until the day they are not. + +**Detect.** Read the bundle-selection branch before trusting any in-editor asset +memory number: + +```bash +rg -n "GIsEditor\s*\|\|" Source/ -A4 +``` + +If the editor branch unions both bundle sets, editor measurements are usable for +direction of change only. + +**Guardrail.** Use editor measurements for direction of change. Absolute cue +loading cost requires a packaged build, and any budget quoted without one carries +that caveat in writing. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-cues/references/patterns.md b/plugins/ue-design-skills/skills/ue-gameplay-cues/references/patterns.md new file mode 100644 index 0000000..f781f92 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-cues/references/patterns.md @@ -0,0 +1,200 @@ +# Patterns: a cue subsystem read end to end + +A worked reading of one real GameplayCue subsystem, audited as source. It is here +because the cue layer has a property that makes reading a class list useless: +**large parts of it can be present, correct, maintained, and unreachable**, and +nothing in the structure says which parts. + +Individual failures are in [failure-modes.md](failure-modes.md), one entry each, +with a recipe. This file is about the shapes — what the subsystem is for, where +its cost lives, and which of its parts were actually running. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from measured +facts; **[open]** — not settled by source reading, and left open. + +--- + +## 1. The trade the mechanism makes + +A cue replaces an asset reference with a string-like tag, and a registry resolves +the tag by scanning content directories. + +What you buy is real: gameplay assets carry no presentation payload, presentation +ships in a separate plugin, and the whole layer is absent on a dedicated server. + +What you pay is one specific thing, and it is worth naming precisely: + +> **You have exchanged a reference the toolchain checks for a join the toolchain +> does not check.** + +A hard pointer that dangles is a cook error. A tag with no notify is silence. +Both halves of the join stay individually valid — the tag exists, the asset +exists — and no tool owns the relation between them. Everything in this file +follows from that one exchange. + +This is why cues are an application of tag governance with a loading subsystem +attached, rather than a feature in their own right. Every rule about string joins +applies here first, and the loading subsystem is where the surprises live. + +--- + +## 2. Where the complexity actually is + +Reading the audited manager, the code divides into three very unequal parts: + +| Part | Size | Status in the shipped configuration | +|---|---|---| +| Dispatch: tag arrives, notify is invoked | small | running | +| Discovery: scan directories, build the tag-to-asset index | moderate | running | +| Preload: async library load, preload set, always-loaded tags, reference tracking, GC and map-transition hooks | **most of the file** | **unreachable** | + +The third row is the finding, and it is a shape rather than a bug. + +A load-mode constant is set to "load everything upfront" **[measured]**. That +value short-circuits three separate switch sites **[measured]**, and the third of +them returns before the statements that bind the subsystem's delegates +**[measured]**. + +The consequence **[derived]**: tag-loaded, post-garbage-collect and post-load-map +handlers are never bound; the preload set, the always-loaded set and the reference +tracker are never populated; and the project's own diagnostic command reports +zero, forever. + +**That last detail is the trap.** A diagnostic reporting zero reads as "there is +nothing to preload". It actually means "this counter is dead". Anyone +investigating cue memory starts from a number that is not measuring anything, and +the natural next step — reading the preload implementation for a bug — is a day +spent in correct code. Recipe: GC-11. + +The transferable rule is not about cues at all: + +> Before debugging a subsystem, establish that its code runs. A diagnostic that +> reports zero is a claim about reachability until proven otherwise. + +--- + +## 3. Knobs that are not knobs + +In the same file, a namespace named after console variables contains three +declarations **[measured]**. Exactly one — a console *command* — is registered +with the engine. The other two, including the load mode from §2, are plain +statics **[measured]**. + +Project-wide the ratio is one such namespace against eight real console-variable +registrations **[measured]**. + +Two things make this worth an entry rather than a footnote: + +1. **The naming is the entire defect.** The code is correct; it is a constant and + behaves as one. What lies is the scope name, and nothing checks scope names. +2. **The obvious search fails.** Grepping for the literal `namespace Cvars` finds + nothing, because real names are prefixed with the owning type. The first draft + of the detection recipe for this returned zero hits and would have been read as + "we do not have this problem". The corrected recipe searches for any namespace + whose name *contains* the token. This is recorded in GC-12 with the trap + spelled out, because the trap generalises further than the finding does. + +--- + +## 4. Feature-plugin registration: the one thing done right + +Cue path registration from a feature plugin has a timing constraint that is easy +to get wrong and invisible when you do: the manager builds its index during its +object library scan, so a path added after that scan is not indexed until +something triggers a rescan. In the editor, incidental rescans hide it. In a cold +packaged run, nothing does. Recipe: GC-14. + +The audited project solves it with a pattern worth copying verbatim +**[measured]**: + +> **The feature action is a pure data declaration; a lifecycle observer is the +> executor.** + +The action object declares the directory list and validates it in the editor — and +has *no activation body at all* **[measured]**. A separate observer registered +with the feature policy listens for the earliest lifecycle phase and performs the +registration for every action of that type it finds. Paths are added with the +rescan flag off, and a single index rebuild follows **[measured]**. + +The empty activation body is the part people delete when adopting this, because an +action with no `Activate` looks unfinished. It is not: it is the point. + +### And the asymmetry that survived inside it + +The same function pair is also the audit's best example of a near-miss +**[measured]**: + +| Direction | Manager resolution | Refresh call | +|---|---|---| +| Registering | project's own manager subclass | present | +| Unregistering | engine base class | **absent** | + +Both calls compile and both work, because the subclass inherits the method. They +are equivalent exactly as long as the subclass adds no bookkeeping **[derived]**. +And the refresh performed on the way in has no counterpart on the way out, so the +asset manager keeps describing directories that have been removed **[derived]**. + +Recipes: GC-15, GC-16. + +**What redeems it, and what to actually copy:** the removal path counts what it +removed and asserts the count against what was added **[measured]**. In an audit +whose dominant finding across the whole project was "add works, remove is +incomplete", this was the one place where teardown was both implemented and +asserted. Copy the assertion; fix the asymmetry it sits next to. + +--- + +## 5. The network rule, and why it is first + +Cues never run on a dedicated server, because the presentation layer does not +exist there. + +Every consequence of that is a design constraint rather than a caution: + +- a notify may not carry gameplay side effects — on a dedicated server that code + never executes (GC-01); +- cue parameters are a lossy channel, and what a conversion helper drops is + dropped silently on every client. The audited helpers drop a context tag + container in **both** directions, each with its own unfinished-work comment + **[measured]** (GC-17); +- delivery is unreliable by design: packet loss, relevancy changes and mid-effect + joins all produce missed cues; +- cue assets belong in the client bundle only. + +The acceptance test that covers all four at once is cheap and nobody runs it: **run +the feature on a dedicated server with every cue asset removed, and require +byte-identical gameplay results.** + +--- + +## 6. What source reading could not settle + +Two limits, stated because they bound the claims above: + +1. **Blueprint call sites are invisible to text search.** The conversion helpers + in §5 have zero C++ callers and are Blueprint-callable **[measured]**. That is + recorded as **[open]** — an open question, not dead code. Deleting them on the + strength of an empty `rg` result is GC-18, committed by the person who wrote + the audit. (It was not.) +2. **Editor memory figures do not describe shipping.** The experience loader's + bundle selection unions client and server bundles when running in the editor + **[measured]**, so in-editor cue residency is overstated by whatever the server + set contains. Editor measurements are good for direction of change and nothing + else (GC-20). + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source +rather than run, in a single workspace. Source addresses stay in the research +archive that produced this skill; each `GC-` identifier resolves back to the +audited location there, so any specific claim above can be produced on request. + +## Evidence boundary + +One project, one engine version, one workspace. These are examples and failure +evidence, not guarantees about other engine versions or other samples. Several +claims are explicitly marked open and stay open. Re-run the recipes in +[failure-modes.md](failure-modes.md) against your own tree before acting on +anything here. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-messaging/SKILL.md b/plugins/ue-design-skills/skills/ue-gameplay-messaging/SKILL.md new file mode 100644 index 0000000..a5c2db0 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-messaging/SKILL.md @@ -0,0 +1,435 @@ +--- +name: ue-gameplay-messaging +description: >- + Design, implement or review a typed tag-addressed pub/sub message bus in + Unreal Engine: subsystem scope and lifetime, struct type checks, exact versus + hierarchical channel matching, listener handles and ownership, parity between + the C++ and Blueprint APIs, the local-versus-network boundary, and migration + from direct dependencies. Use when decoupling gameplay from UI, feedback, + telemetry or processors, or when debugging duplicate, stale, mismatched or + wrongly replicated messages. +--- + +# UE gameplay messaging + +The invariant: + +> Messages announce facts inside one process. They do not own state, ordering, +> persistence, authority or replication. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Seventeen detection recipes for silent bus failures: +[failure modes](references/failure-modes.md). + +Read the failure modes before adopting a bus implementation. Each entry names the +mechanism, why it stays silent, why the obvious check misses it, and a command you +can run against your own tree. + +Related skills: `ue-gameplay-tag-governance`, `ue-ui-architecture`, +`ue-gas-architecture`, `ue-multiplayer-authority`, `ue-data-driven-architecture`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Choose the bus scope explicitly + +A bus implemented as a GameInstance subsystem: + +- is shared by every world belonging to that GameInstance; +- survives map travel; +- is destroyed at GameInstance teardown; +- is local to one process; +- does **not** replicate. + +That scope suits: + +- gameplay event → HUD and feedback; +- local processors and telemetry; +- decoupling feature plugins from the base module; +- cross-system notifications where no receiver owns the sender. + +It is unsuitable as the source of truth for: + +- inventory, team or score state; +- guaranteed ordered workflows; +- cross-process communication; +- persistence; +- anything an RPC should carry. + +### The state-versus-event test + +Ask: **"if a listener subscribes one second late, must it know the current +value?"** + +- **yes** → replicated property or state component, plus a change notification; +- **no** → a transient message may be appropriate. + +"An elimination occurred" is an event. "The current team score" is state. A bus +that is asked to answer the second question will appear to work for every +listener that happened to be subscribed early, which is every listener during +development. + +--- + +## 2. Storage model and handle identity + +```text +map: channel tag -> channel listener list + +channel listener list +├─ monotonically increasing handle ID (local to THIS channel) +└─ listeners[] + ├─ callback + ├─ weak expected struct type + ├─ had-valid-type flag + ├─ handle ID + └─ exact / partial match +``` + +Registration returns a handle carrying the subsystem, the channel, and the +channel-local ID. **The handle is the ownership receipt.** The caller, or an RAII +wrapper, retains it and unregisters when its lifecycle ends. + +### The handle identity rule + +If IDs are allocated per channel, then **every internal removal must use both the +registration channel and the ID.** Never search another channel by ID alone. + +This is not a stylistic point. A hierarchical broadcast walks several channels in +one call; code written inside that walk has two plausible variables in scope — the +channel the message was sent on, and the channel currently being visited — and +they are the same value only on the first iteration. Recipe: MB-01. + +--- + +## 3. Typed broadcast + +Template the C++ API and resolve the struct type at compile time on both sides. +Store the expected type at registration. + +On broadcast, the payload is deliverable when: + +```text +broadcast struct IsChildOf listener expected struct +``` + +which permits a listener registered for a base struct to receive derived +payloads, and makes exact equality a special case. + +### Type mismatch policy + +- log the channel, the sent struct and the expected struct; +- skip that listener; +- never reinterpret incompatible bytes; +- never abort the rest of the broadcast; +- add development-only assertions and tests. + +A path that disables the type check entirely — registering with a null expected +type — must be private and documented. If it is reachable from Blueprint through +an empty type pin, it is not private. Recipe: MB-03. + +### Payload lifetime + +The bus should pass a pointer to the caller's stack object rather than copying. +That is the right performance decision and it creates one hard rule: +**a listener may not retain the payload pointer after its callback returns.** +Copy what you need. Recipe: MB-04. + +--- + +## 4. Exact versus hierarchical matching + +Channels are tags, so a broadcast can walk from the sent channel up through its +parents: + +```text +Event.Combat.Elimination +→ Event.Combat +→ Event +``` + +On the first tag every listener is called regardless of match type; on ancestors, +only listeners that registered for partial matching. + +**Default to exact.** It minimises accidental traffic and type ambiguity. Use +partial matching for analytics, debug logging, and aggregators that genuinely +want a subtree. + +### The hierarchy type gate + +Every descendant channel observed through one parent must share a compatible base +payload struct. Otherwise partial matching converts a semantic hierarchy into a +stream of type errors. Document the payload contract at each parent tag — this is +the same rule as parent-tag contracts in `ue-gameplay-tag-governance`, applied to +payload types instead of semantics. + +### Cost + +Walking parents costs one map lookup per tag depth on **every** broadcast, even +when no partial listeners exist anywhere. If the listener array is copied per +non-empty bucket for mutation safety, that copy includes a callable object per +listener, which may allocate. On a high-frequency channel this is measurable. +Recipe: MB-16. + +--- + +## 5. Broadcast safely under mutation + +Callbacks may unregister themselves or others. Copy the listener array before +iterating, then invoke the copy — and then **define the resulting semantics +explicitly**, because they are observable: + +- a listener removed during this broadcast may still receive this one message; +- a listener added during this broadcast receives from the next one; +- ordering is not guaranteed and must not be assumed. + +Do not let registration order become an accidental priority contract. If removal +uses a swap-with-last strategy, order changes as a side effect of unrelated +unsubscribes. Recipe: MB-11. + +--- + +## 6. Registration forms and ownership + +Provide three forms: a raw callable, a UObject member bound through a weak +reference, and a parameter struct for match type and future options. + +The trap is that these three are not equally expressive. If the most convenient +form — the UObject member overload — cannot express the match type, then the +match type is effectively unavailable to the code that would most benefit from +it, and a project can ship a fully implemented hierarchical matching feature that +nothing uses. Recipe: MB-05. + +### Ownership patterns + +| Owner | Retain | Release | +|---|---|---| +| C++ component | handle as a member | `EndPlay` / deinit | +| subsystem | handles in a container | `Deinitialize` | +| feature action | per activation context | on deactivation | +| async operation | inside the action object | completion, cancel or destruction | + +**A weak callback is not an unsubscription.** It prevents a call into a dead +object; the listener record and its callable stay in the map, and accumulate. +Recipe: MB-08. + +### Teardown must not assert + +An accessor that asserts when there is no world is the wrong tool inside +`EndPlay`, `NativeDestruct` or any other teardown path, where the world may +already be gone. Prefer unregistering through the handle, which resolves the +subsystem weakly and no-ops when it is dead. Recipe: MB-09. + +--- + +## 7. Blueprint parity + +A wildcard struct pin lets Blueprint send arbitrary payloads while C++ receives a +type plus bytes. An async listen node stores the channel, payload type and match +type, registers on activation, copies the payload into typed output storage, and +unregisters on cancel or destruction. + +### Keep the two APIs semantically identical + +This is the rule most often broken, because the two paths are written at +different times by different people: + +- if C++ accepts derived payloads, Blueprint must too; +- if C++ logs a mismatch, Blueprint must log it as well. + +A silent Blueprint skip produces "the node never fires", which every developer +first investigates as a wrong tag. Pick one contract, document it, and test both +APIs against the same matrix. Recipe: MB-02. + +--- + +## 8. Messaging is not replication + +No broadcast crosses the network. For networked facts the shape is always: + +```text +server-authoritative state change +→ replicated property / RPC / replicated fast array +→ receiving side broadcasts a LOCAL message +→ local UI, feedback and processors react +``` + +This keeps transport and local fan-out separate, and it is the only arrangement +in which a late joiner is correct: the state replicates, and the message is +merely a notification that it changed. + +**Never treat a client-local message as proof of a server-side fact.** A message +may request UI behaviour. It is not evidence of damage, score, ownership or team +membership. Recipe: MB-15. + +--- + +## 9. Payload design + +A message struct is a compact immutable fact: + +```text +verb / channel context +instigator +target +magnitude +context tags +optional IDs or weak references +``` + +Rules: + +- prefer IDs and weak references over large object graphs; +- carry enough context that a listener needs no callback into the sender; +- do not embed mutable ownership; +- initialise every field — a payload never passes through serialisation, so + there are no free defaults; +- version deliberately if the same struct is persisted or replicated elsewhere; +- one parent channel hierarchy shares one base payload. + +### Do not send commands disguised as events + +```text +bad: message "please give the player a weapon" +good: service call performs the grant; message "weapon equipped" +``` + +A command has exactly one accountable handler and a failure result. An event has +zero or many listeners and no failure path at all. Sending a command as an event +means nobody is responsible for it happening, and nobody notices when it does +not. Recipe: MB-14. + +--- + +## 10. Channel naming and the string contract + +Name channels so bus traffic is distinguishable from other tags: + +```text +..Message[.] +``` + +The word in the middle separates channels from the rest of the vocabulary and +gives a natural place for a partial-match subscription. + +**The channel string is the public contract, and the compiler does not check it.** +If publisher and subscribers each declare their own file-local tag constant for +the same string, the engine deduplicates by string and everything works — until +one of them is edited. Put shared channels in one header owned by the publisher. +Recipe: MB-10. + +--- + +## 11. Migration recipe + +When decoupling a direct dependency: + +1. Name the fact in the past tense. +2. Define its payload struct and channel tag. +3. Keep authoritative mutation in the existing owner. +4. Broadcast only **after** the mutation succeeds. +5. Move independent side effects to listeners, one at a time. +6. Retain and release every handle. +7. Test zero, one and multiple listeners. +8. Verify no late subscriber needs the state it missed. + +Good first migrations: combat event to HUD feed and audio; inventory delta to a +notification widget; phase transition to local presentation. + +Do not migrate tight request/response code merely to avoid a function reference. + +--- + +## 12. Review checklist + +- [ ] Message is an event, not durable state and not a command. +- [ ] Subsystem scope matches the required lifetime. +- [ ] Payload is a struct with a documented contract and initialised fields. +- [ ] Parent channels have a compatible payload hierarchy. +- [ ] Exact matching is the default; partial matching is justified per site. +- [ ] Registration returns a handle and the caller retains it. +- [ ] UObject callbacks bind weakly. +- [ ] Unregistration happens on every lifecycle exit, including teardown paths. +- [ ] Teardown does not call an asserting accessor. +- [ ] Broadcast tolerates mutation from inside a callback. +- [ ] C++ and Blueprint type semantics are identical and both log mismatches. +- [ ] No listener ordering is assumed without an explicit contract. +- [ ] Replication and RPC transport are separate from local fan-out. +- [ ] Shared channel tags live in one owned header. + +--- + +## 13. Tests + +### Type matrix + +- exact same struct; +- derived broadcast to base listener; +- base broadcast to derived listener — rejected; +- unrelated struct — logged and skipped; +- expired expected struct type; +- C++ and Blueprint produce the same result for every row above. + +### Channel matrix + +- exact listener, exact broadcast; +- parent with partial matching receives a child broadcast; +- parent with exact matching does not; +- sibling does not; +- a nested parent chain invokes each intended listener exactly once. + +### Lifecycle + +- unregister before broadcast; +- unregister from inside a callback; +- owner destroyed without an explicit unregister; +- async node cancelled; +- map travel with a GameInstance-scoped bus — listener count returns to baseline; +- feature activation and deactivation cycle, twice; +- duplicate registration is either documented or detected. + +### Network boundary + +- a server-side broadcast is invisible to clients; +- replicated transport triggers exactly one client-local rebroadcast; +- a listen-server process does not process both the server and client path. + +--- + +## 14. When not to add a bus + +Use a direct interface, service or delegate when: + +- exactly one receiver must handle the call; +- the caller needs a return value or a failure reason; +- ordering is part of correctness; +- sender and receiver share ownership and lifetime; +- the state must be queryable after the event. + +A bus reduces coupling only when the fact genuinely has independent, optional +observers. Used anywhere else it hides control flow and weakens correctness — +it converts a compile-time dependency into a runtime one that nothing verifies. + +--- + +## Provenance + +The findings behind the failure modes come from a line-by-line source audit of +the gameplay message router plugin shipped with Epic's Lyra Starter Game on +Unreal Engine 5.6, together with its use in that project. The plugin's runtime +module is roughly a thousand lines and was read in full, which is why several +claims here are negative results ("there is no networking in this module") stated +with confidence rather than hedged. + +Source addresses stay in the research archive that produced this skill; what +ships is the detection recipe. Each entry carries a stable identifier (`MB-01`, +`MB-02`, …) that resolves back to the audited location in that archive. + +## Evidence boundary + +The measured material refers to one plugin at one engine version in one +workspace. Blueprint graphs were not read — they are binary — so every statement +about Blueprint *usage* in that project is bounded by what C++ could show, and is +marked as open rather than concluded. Re-run the recipes against your own tree +before trusting any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/failure-modes.md new file mode 100644 index 0000000..904a154 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/failure-modes.md @@ -0,0 +1,646 @@ +# Failure modes: gameplay messaging + +Seventeen ways a tag-addressed message bus stops delivering, over-delivers, or +delivers into a void, without reporting anything useful. + +The shared property of this list: **a bus converts a compile-time dependency into +a runtime one, and then does not check it.** The publisher compiles whether or not +anyone listens; the listener compiles whether or not anyone publishes. Every +entry below is a consequence of that trade, and the observable is almost always +the same — nothing happens, or something happens twice. + +Recipes use `rg`. Two conventions worth stating once, because both cost real time +during the audit that produced this file: + +- **Restrict searches to source globs.** An unrestricted search of a built tree + matches compiled debug symbols, which are not code. In the measured reference a + search for a suspected dead field returned four hits; three were `.pdb` files + and the source truth was one. A count that includes build output is not a count + of your code. +- **Most recipes are a pair of searches.** The defect class here is "the thing is + present and the thing that would give it meaning is absent", and one search can + only show the first half. + +--- + +## Delivery correctness + +### MB-01 - Listener removed using the wrong channel key + +**Mechanism.** A hierarchical broadcast walks from the sent channel up through its +parents. Inside that walk, code that removes a listener addresses the removal by +the *original broadcast channel* rather than the ancestor tag currently being +visited. Handle IDs are allocated per channel, so the pair (channel, ID) is the +only unique identity. + +**Why it is silent.** On the first iteration of the walk the two variables hold +the same value, so exact-match listeners — which is nearly all of them in most +projects — are removed correctly. The bug requires a listener registered on a +parent with partial matching *and* an expired payload type, simultaneously. Both +conditions are rare, and neither is an error on its own. + +**Why the obvious check misses it.** The line reads correctly in isolation: +removing by "channel" is obviously right, and the variable is named `Channel`. +The defect is that the enclosing loop rebinds the meaning of "the channel we are +talking about" on every iteration, and the correct variable is three lines up in +the `for` header. Worse, the error-log line a few lines below uses the loop +variable correctly — so the file contains both the right and the wrong idiom, and +the wrong one is the one that mutates state. + +**Symptom.** Two outcomes, both bad and neither loud. Either the stale listener is +never removed and a warning is emitted on every subsequent broadcast forever, or — +worse — an unrelated listener that happens to hold the same channel-local ID in +the broadcast channel's bucket is removed instead, and its owner silently stops +receiving messages. + +**Detect.** Compare the loop variable against the key used for removal: + +```bash +rg -n -A12 "for\s*\(.*RequestDirectParent" --glob "*.cpp" . \ + | rg "UnregisterListenerInternal|Remove\(" +``` + +Read the output next to the `for` header. If the removal uses the loop's starting +value rather than the loop variable, that is this defect. In the measured +reference the walk iterated over an ancestor tag while removal passed the +broadcast channel. + +**Guardrail.** When handle IDs are channel-local, make the removal API take the +pair and nothing else, and assert inside it that the pair was found. A removal +that silently finds nothing is how this defect stays invisible. + +--- + +### MB-02 - Covariance in C++, equality in Blueprint + +**Mechanism.** The C++ delivery path accepts a payload whose struct derives from +the listener's expected struct. The Blueprint async listener compares types with +strict equality. + +**Why it is silent.** Both behaviours are individually reasonable and both are +implemented correctly. A Blueprint listener that does not match simply does not +fire, and not firing is the normal state of a listener most of the time. + +**Why the obvious check misses it.** The two implementations live in different +files, written for different consumers, and each is correct against its own local +reasoning. There is no shared test matrix, because the two APIs are usually tested +by different people — the C++ one by whoever wrote the bus, the Blueprint one by +whoever first used the node. Nothing in either file mentions the other. + +**Symptom.** A Blueprint listener subscribed to a base struct never fires for +derived payloads, while a C++ listener on the same base receives them. The +designer reports "the node never fires", which everyone investigates as a wrong +tag — the one hypothesis that is cheap to test and wrong. + +**Detect.** Find both type gates and compare the operators: + +```bash +rg -n "IsChildOf" --glob "*.cpp" . # covariant gate +rg -n "StructType\s*==|== \w*StructType" --glob "*.cpp" . # equality gate +``` + +Two different operators guarding the same conceptual check, in the same plugin, +is the finding. In the measured reference the C++ path used covariance and the +async Blueprint path used equality, on the same registration. + +**Guardrail.** Pick one contract, write it down, and run one test matrix through +both APIs. If they must differ, the difference belongs in the node's tooltip, not +only in the source. + +--- + +### MB-03 - Type check disabled by an empty type pin + +**Mechanism.** Registering with a null expected struct type sets a flag that +disables the type check entirely, so the listener receives any payload. The path +is documented as internal — and is reachable from a Blueprint node whose payload +type pin was left empty. + +**Why it is silent.** Nothing fails. The listener receives messages, possibly +several unrelated kinds, and the Blueprint graph downstream either uses the +payload or does not. There is no log, because from the bus's point of view the +listener asked for exactly this. + +**Why the obvious check misses it.** The comment on the branch says "for internal +use", and reviewers believe comments about intent. Confirming reachability means +tracing a Blueprint pin's default value into a factory function into a private +registration call — three hops across two modules, none of which look +interesting. + +**Symptom.** In Blueprint, usually nothing worse than a payload extraction that +returns false. In C++, the same path combined with a templated callback would +reinterpret unrelated bytes as the expected struct, which is undefined behaviour +with no diagnostic. + +**Detect.** Find the escape hatch and then find who can reach it: + +```bash +rg -n "bHadValidType|StructType\s*==\s*nullptr" --glob "*.cpp" --glob "*.h" . +rg -n -B4 "RegisterListenerInternal\(" --glob "*.cpp" . | rg -i "Get\(\)|nullptr" +``` + +Any caller that can pass a null type is a caller that can disable the check. + +**Guardrail.** If a code path must exist for internal reasons, make it +inaccessible rather than merely undocumented — a private overload, a passkey +type, or a compile-time gate. Reject an unresolved wildcard payload at Blueprint +compile time. + +--- + +### MB-04 - Payload pointer retained after the callback returns + +**Mechanism.** The bus passes a pointer to the sender's stack object rather than +copying. A listener stores that pointer. + +**Why it is silent.** The memory is still mapped and, for a short while, still +holds the right bytes. Reading it immediately after the broadcast usually works. +The failure requires the stack to be reused, which depends on what runs next. + +**Why the obvious check misses it.** The callback signature hands you a reference, +and references are normally safe to keep. Nothing in the type expresses "valid +until this function returns". The one place the constraint is written down is the +bus's own documentation, which the person writing the listener has no reason to +open. + +**Symptom.** Corrupted payload fields read at a later tick — values that are +plausible, occasionally correct, and change with unrelated code. This is one of +the few entries in this file that produces a crash, and the crash is far from the +cause. + +**Detect.** Find listeners that store rather than consume: + +```bash +rg -n -A8 "void .*\(FGameplayTag\s+\w+,\s*const\s+F\w+&\s*\w+\)" --glob "*.cpp" . \ + | rg "=\s*&|Ptr\s*=|AddRaw|CopyRef" +``` + +Any assignment of the payload's address into a member is the finding. A copy of +the payload by value is fine. + +**Guardrail.** State the lifetime in the callback's own documentation, and copy at +the boundary. Where the bus crosses into a scripting layer, copy by value — the +measured reference does exactly this for its Blueprint path, and nulls the stored +pointer immediately after the delegate returns. + +--- + +## Ownership and lifetime + +### MB-05 - The convenient overload cannot express the match type + +**Mechanism.** Three registration overloads exist. The one that binds a UObject +member weakly — the safest and by far the most used — delegates to the raw form +without forwarding the match type, so it is always exact. + +**Why it is silent.** Exact matching is the right default, so every listener +registered this way behaves correctly. The feature that is unavailable is one +nobody is currently trying to use, and its absence produces no error because you +cannot pass the argument at all. + +**Why the obvious check misses it.** The feature is fully implemented, documented +in the enum, exposed in the Blueprint node, and covered by the broadcast walk. A +review that asks "do we support hierarchical matching?" finds all of that and +answers yes. The gap is one unforwarded default argument in a header. + +**Symptom.** A project ships a complete hierarchical matching implementation that +nothing in its own C++ ever uses, and nobody notices, because using it requires +switching to a less convenient overload for reasons that are never stated. + +**Detect.** Count the feature's declaration sites against its use sites, and +exclude the implementation itself: + +```bash +rg -n "PartialMatch" --glob "*.h" --glob "*.cpp" . # declared +rg -n "PartialMatch" --glob "*.cpp" . | rg -v "MessageRouter|MessageSubsystem" +``` + +An empty second result with a rich first result is the finding. In the measured +reference the game code used partial matching zero times; the only similarly +named symbols elsewhere belonged to two unrelated subsystems that had copied the +pattern. + +**Guardrail.** Every overload of a registration API forwards every option, or the +option does not exist. If an overload deliberately restricts, name it so — +`RegisterExactListener` — rather than silently dropping an argument. + +--- + +### MB-06 - Parameter struct with no bound callback returns an invalid handle silently + +**Mechanism.** The options-struct registration form checks that a callback is +bound and, if not, returns a default-constructed invalid handle without +registering and without logging. + +**Why it is silent.** The caller receives a handle-shaped value. Storing it, +passing it around and eventually unregistering it all work. The only difference +is that nothing was ever registered. + +**Why the obvious check misses it.** The return type is not optional and not an +error code; it is a handle, and handles look like success. Checking validity +requires knowing that this overload can fail, which is stated nowhere at the call +site. + +**Symptom.** A listener that never receives anything, in code that reads as fully +wired. Because unregistration of an invalid handle is also silent, no part of the +lifecycle complains. + +**Detect.** Find registrations whose returned handle is never validity-checked: + +```bash +rg -n "RegisterListener\(" --glob "*.cpp" . -A3 | rg -v "IsValid|ensure|check" +``` + +**Guardrail.** Return an explicit failure, or log at warning level, or assert in +development builds. A silent failure path on a registration API is a listener +that cannot be debugged. + +--- + +### MB-07 - Async listener outlives its owner until the next message + +**Mechanism.** An async listen action detects its owner's death by noticing that +its dynamic delegate is no longer bound — which it can only do *while handling a +message*. Until the next message arrives, the action and its registration stay +alive. + +**Why it is silent.** The cleanup does eventually happen, and it happens before +anything observable goes wrong. On a busy channel the window is a frame. + +**Why the obvious check misses it.** There *is* cleanup, and reading it shows +correct logic. The gap is a timing property — cleanup is driven by traffic rather +than by the owner's destruction — and timing properties are invisible in a code +review that asks "is cleanup implemented?". + +**Symptom.** On a quiet channel, listener records accumulate for destroyed +owners. On a shutdown path, the last message of a session may be delivered into +an object graph that is already tearing down. + +**Detect.** Find cleanup that is triggered by message handling rather than by +destruction: + +```bash +rg -n -B6 "SetReadyToDestroy" --glob "*.cpp" . | rg -i "IsBound|HandleMessage" +rg -n -i "//\s*@?TODO.*(proactive|cleanup)" --glob "*.cpp" . +``` + +The second command matters: in the measured reference the authors had recorded +exactly this limitation as a TODO with a tracker reference, which is the +strongest possible confirmation that a finding is real and known. + +**Guardrail.** Tie the async action's lifetime to its owner's destruction +directly, not to the arrival of the next message. + +--- + +### MB-08 - Listener map is not cleared on map travel + +**Mechanism.** A GameInstance-scoped bus is cleared only at GameInstance +teardown. Listeners registered by objects belonging to a world that has been +travelled away from remain in the map. + +**Why it is silent.** Weak binding means the dead objects are never called. The +records are inert: they cost memory and a little iteration time, and produce no +wrong behaviour. + +**Why the obvious check misses it.** Deinitialisation exists and is correct for +the scope it was written for. "Does the bus clean up?" is answered yes. The +question that finds this is narrower — *at which lifecycle event?* — and the +answer is one that only matters in a project that travels between maps, which a +single-map test never does. + +**Symptom.** Listener counts that grow monotonically across a session of map +changes. Diagnosis is difficult because the leak is bounded by session length and +the objects themselves are correctly garbage collected. + +**Detect.** Find the scope and the reset point, and compare them: + +```bash +rg -n "public UGameInstanceSubsystem|public UWorldSubsystem" --glob "*.h" . +rg -n -A4 "::Deinitialize\(\)" --glob "*.cpp" . | rg "Reset\(|Empty\(|Clear\(" +``` + +A GameInstance scope whose only reset is in `Deinitialize` leaks across travel by +construction. Confirm at runtime: log the listener count, travel, log again. + +**Guardrail.** Either scope the bus to the world, or hook map change and drop +records whose weak owner has expired. A weak callback is not an unsubscription. + +--- + +### MB-09 - Teardown calls an accessor that asserts + +**Mechanism.** The static accessor resolves the world with an assert-on-failure +mode and then asserts on both the world and the subsystem. Teardown code calls it +to unregister. + +**Why it is silent.** In every ordinary shutdown the world is still present when +components tear down, so the asserts pass. The failure needs an unusual teardown +order — a subsystem destroyed first, PIE cancelled mid-initialisation, a level +unloaded during travel. + +**Why the obvious check misses it.** Unregistering in `EndPlay` is the correct +pattern and is what a reviewer is looking for. That the *accessor* is the +strictest of the several available ways to reach the subsystem is a property of a +different file. + +**Symptom.** An assert during shutdown, in a build configuration and teardown +order that nobody can reproduce on demand, at a call site that is doing the right +thing. + +**Detect.** Find teardown paths that use the asserting accessor: + +```bash +rg -n -B2 -A2 "::EndPlay|NativeDestruct|BeginDestroy" --glob "*.cpp" . \ + | rg "Subsystem::Get\(|::Get\(this\)" +rg -n -A6 "static .*& Get\(" --glob "*.h" . | rg "Assert|check\(" +``` + +The pairing is the finding: an asserting accessor called from a destruction path. +In the measured reference this pattern appeared in three separate consumer +classes, all of which could have unregistered through the handle instead. + +**Guardrail.** Unregister through the handle, which resolves the subsystem weakly +and no-ops when it is gone. Provide a non-asserting accessor and use it in every +teardown path. + +--- + +## Contract and vocabulary + +### MB-10 - The channel string is the contract, declared in several places + +**Mechanism.** Publisher and each subscriber declare their own file-local tag +constant for the same channel string. The engine deduplicates by string, so it +works. + +**Why it is silent.** It is correct. Every site resolves to the same tag, the join +succeeds, and the arrangement has a real benefit: modules do not depend on each +other's headers. + +**Why the obvious check misses it.** Each declaration is locally idiomatic and +each file reads well. Nothing links them. A rename in one file compiles, links, +and passes review — and the reviewer of that file has no way to see the other +three. + +**Symptom.** Editing one declaration silently detaches one participant from the +channel. This is the symmetric-typo failure from tag governance, arriving through +a different door: correctness means all sites agree, not that any one is right. + +**Detect.** Count independent declarations per channel string: + +```bash +rg -o -N 'UE_DEFINE_GAMEPLAY_TAG[_A-Z]*\([^,]+,\s*"([^"]+)"' -r '$1' \ + --glob "*.cpp" . | sort | uniq -c | sort -rn | head +``` + +Any string with a count above one is a contract with several owners. In the +measured reference one elimination channel was declared independently at four +sites — the publisher and three separate processors in a feature plugin. + +**Guardrail.** Shared channels live in one header owned by the publisher. Keep +file-local declarations only for channels that never leave their file. + +--- + +### MB-11 - Listener order assumed without a contract + +**Mechanism.** Nothing guarantees callback order, and removal by swap-with-last +actively reorders the array as a side effect of unrelated unsubscriptions. + +**Why it is silent.** An order exists on every run and is stable as long as +nothing unsubscribes. Code that depends on it works, and keeps working, until an +unrelated feature adds a listener or removes one. + +**Why the obvious check misses it.** The dependency is never written down — +that is what makes it an assumption. No search finds "this listener must run +first"; the knowledge lives in the fact that it currently does. + +**Symptom.** Ordering-dependent behaviour that changes when an unrelated system +subscribes, in a different module, in a different release. + +**Detect.** Find the removal strategy and the documented guarantee, and check for +listeners that mutate shared state: + +```bash +rg -n "RemoveAtSwap|RemoveSwap" --glob "*.cpp" . +rg -n -i "order.*not guaranteed|call order" --glob "*.h" . +``` + +If the second command finds an explicit warning from the bus authors, the +guarantee does not exist and any dependence on order is yours to remove. + +**Guardrail.** If ordering matters, it is not a bus concern — sequence the work in +one owner and broadcast the outcome. Never introduce priority through +registration order. + +--- + +### MB-12 - Reentrancy with no depth limit + +**Mechanism.** Broadcasting from inside a listener callback is legal and +supported. Nothing bounds the recursion. + +**Why it is silent.** It is a legitimate and useful pattern — an aggregator that +consumes one channel and publishes a derived fact on another does exactly this. +The dangerous case differs only in that the graph of channels contains a cycle. + +**Why the obvious check misses it.** Each listener is individually correct and +each broadcast is individually justified. The cycle exists in the graph formed by +several modules, which no file shows, and which a feature plugin can complete by +adding one processor. + +**Symptom.** Stack overflow, or a hang, with a call stack that is one repeating +pattern thousands of frames deep and no obvious origin. + +**Detect.** Map the graph: which listeners publish, and on which channels? + +```bash +rg -n -B10 "BroadcastMessage\(" --glob "*.cpp" . \ + | rg -B2 "OnMessageReceived|::Handle\w*Message" +``` + +Every hit is an edge from an input channel to an output channel. Draw them and +look for a cycle. In the measured reference this pattern was present and +deliberate — a processor consuming eliminations and publishing assists — with no +cycle, which is the correct state and worth confirming rather than assuming. + +**Guardrail.** Keep a broadcast-depth counter in development builds and assert +past a small bound. Document, per processor, which channels it reads and writes. + +--- + +## Scope and misuse + +### MB-13 - A message used where state is required + +**Mechanism.** A fact is published as an event only. A listener that subscribes +after the event has no way to learn it. + +**Why it is silent.** Every listener that existed at publication time is correct. +Subscription usually happens during initialisation, before anything interesting +has been published, so the whole system is correct for the entire development +period. + +**Why the obvious check misses it.** Testing subscribes early by construction — +you start the game, the widget is created, then things happen. Reproducing +requires a listener created *after* the fact, which means mid-match widget +creation, a late-joining client, or a feature activated at runtime. + +**Symptom.** A widget created mid-match shows a default value forever. A late +joiner's UI is empty while everyone else's is correct. + +**Detect.** For each channel, ask whether any subscriber can be created late: + +```bash +rg -n "RegisterListener\(" --glob "*.cpp" . -B6 | rg -i "NativeConstruct|BeginPlay|OnActivated" +``` + +Widget construction and feature activation are the late-creation paths. For each, +apply the test from the skill: if a listener subscribing one second late must know +the current value, the fact is state and the bus is the wrong carrier. + +**Guardrail.** Replicate or store the state; use the message purely as a change +notification. On subscription, read the current value directly from its owner — +which requires that the value have an owner, and that is the real design output. + +--- + +### MB-14 - A command disguised as an event + +**Mechanism.** A message is published to make something happen, rather than to +report that it has happened. + +**Why it is silent.** With exactly one listener it behaves identically to a +function call. It works, and it looks decoupled. + +**Why the obvious check misses it.** The mechanism is indistinguishable from +correct use — same API, same payload shape. Only the *name* reveals it, and names +in the imperative mood ("give", "apply", "spawn") pass review because they +describe what the sender wants. + +**Symptom.** Zero listeners means the command is silently dropped with no failure +result. Two listeners means it happens twice. Neither is reported, because a bus +has no concept of a required handler. + +**Detect.** Grep the channel vocabulary for the imperative mood: + +```bash +rg -o -N '"([A-Za-z]+\.[A-Za-z.]*Message[A-Za-z.]*)"' -r '$1' --glob "*.cpp" . \ + | sort -u | rg -i "^\w+\.(Add|Give|Set|Apply|Spawn|Request|Please|Do)" +``` + +A channel named for an action rather than for a completed fact is the finding. +Past-tense names — "equipped", "eliminated", "changed" — are the correct shape. + +**Guardrail.** Commands go to one accountable handler that can fail and say so. +Events report facts in the past tense and tolerate zero listeners. + +--- + +### MB-15 - Broadcast treated as replication + +**Mechanism.** A message is published on the server and expected to reach +clients. The bus is local to one process and has no networking at all. + +**Why it is silent.** In a single-process test — the editor's default multiplayer +mode — the host and its client share one process and, depending on the bus's +scope, may share one instance. The message appears to cross the network. + +**Why the obvious check misses it.** Everything about the code reads as +transport: a channel, a payload, publish and subscribe. The absence of networking +is a property of a module's dependency list, which is not where anyone looks when +reasoning about a message flow. + +**Symptom.** A feature that works in the editor and does nothing on a dedicated +server. Alternatively, a listen server processes both the server-side publication +and the client-side rebroadcast and everything happens twice. + +**Detect.** Prove the absence, then find who assumed otherwise: + +```bash +rg -n "UFUNCTION\([^)]*(Server|Client|NetMulticast)|NetSerialize|Replicated" \ + --glob "*.cpp" --glob "*.h" / # expect: nothing +rg -n -B4 "BroadcastMessage\(" --glob "*.cpp" . | rg "HasAuthority|WITH_SERVER_CODE|NM_" +``` + +The second command finds the correct pattern as well as the incorrect one: a +net-mode check before a rebroadcast is the guard that prevents double processing +on a listen server. Its absence next to a rebroadcast is the finding. + +**Guardrail.** Replicate state or send an RPC; rebroadcast locally on arrival; +guard the rebroadcast with a net-mode check so a listen server does not process +both paths. + +--- + +## Cost + +### MB-16 - Per-broadcast copy of the listener array + +**Mechanism.** For mutation safety the listener array of each non-empty bucket is +copied before iteration. The copied element type holds a callable object, whose +copy may allocate. + +**Why it is silent.** It is correct, and correctness was the reason for it. The +cost is proportional to broadcast frequency times listener count times tag depth, +and all three are small early in a project. + +**Why the obvious check misses it.** The line is a defensive copy with a comment +explaining why it is needed, which is exactly what a reviewer wants to see. The +cost is invisible unless you know that the element type is expensive to copy — +which requires reading a different struct in a different header. + +**Symptom.** Allocator churn proportional to message traffic, appearing in +profiles as generic allocation cost rather than as bus cost. Discovered when a +high-frequency channel — damage in a shooter — is added late. + +**Detect.** Find the copy and the depth walk, then measure: + +```bash +rg -n -B2 "TArray<\w*ListenerData>\s*\w+\(" --glob "*.cpp" . +rg -n "RequestDirectParent" --glob "*.cpp" . +``` + +Both present means cost per broadcast scales with tag depth even when no +hierarchical listener exists. Confirm with an allocation profile while +broadcasting on the highest-frequency channel. + +**Guardrail.** Copy handles rather than callables, or iterate an index with a +generation counter, or skip the ancestor walk when the bucket count for parent +tags is zero. Measure before optimising — this is a real cost, not a large one. + +--- + +### MB-17 - A declared field nobody uses, and the search that hides it + +**Mechanism.** A member is declared in a public struct, never written and never +read. It survives because removing it feels risky. + +**Why it is silent.** An unused field costs a few bytes and nothing else. It +serialises, copies and compiles like any other member. + +**Why the obvious check misses it.** This entry is here less for the dead field +than for the search that fails to prove it dead. A naive recursive search of a +built project matches compiled debug symbols, so a field with exactly one source +occurrence returns several hits and reads as "in use". The instrument reports the +build output as if it were code. + +**Symptom.** Dead API surface that persists across refactors, plus — more +expensively — a general loss of trust in the searches used to make deletion +decisions. + +**Detect.** Search source only, and say so explicitly: + +```bash +F='StateClearedHandle' +rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" . # source truth +rg -n "\b$F\b" . # unfiltered, for contrast +``` + +In the measured reference the filtered search returned exactly one line — the +declaration — while the unfiltered one returned four, three of them debug-symbol +files. Always compare the two before concluding either "dead" or "in use". + +**Guardrail.** Restrict every deletion-decision search by file type. Delete dead +fields in the same change that proves them dead, and record the search you ran. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/patterns.md b/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/patterns.md new file mode 100644 index 0000000..bb1cf2f --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-messaging/references/patterns.md @@ -0,0 +1,300 @@ +# Patterns: a message bus read line by line + +A worked reading of one real tag-addressed message bus and its use in the product +that ships it. The runtime module is about a thousand lines and was read in full, +which is why several claims below are **negative results stated with confidence** +rather than hedged — "there is no networking in this module" is a conclusion from +reading every line of it, not from a search that found nothing. + +Individual defects are in [failure-modes.md](failure-modes.md), one entry each, +with a recipe. This file is about the shapes: what the bus is, where its costs +sit, and how the surrounding product worked around what it does not do. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from measured +facts; **[open]** — not answerable from source, and left open. + +--- + +## 1. What the mechanism actually is + +A publish/subscribe bus where the **channel is a tag** and the **message is an +arbitrary struct**. Publisher and subscriber never reference each other; the +entire contract is the pair (channel tag, struct type). + +The implementation is one flat map from tag to listener list **[measured]**. +There is no tag tree in the data structure — hierarchy is walked at broadcast +time by asking each tag for its direct parent until the chain ends. + +Two properties fall out of that design and explain most of this file: + +1. **The join is a string and a reflection pointer.** Neither is checked by the + compiler. Everything in the failure-mode list is a consequence. +2. **Delivery is synchronous.** Callbacks run inside the broadcast call before it + returns **[measured]**. Good for debugging — the stack is continuous — and the + reason reentrancy is possible at all. + +The module's dependencies are the engine core, the engine, and the tag system +**[measured]**. Nothing from the game, nothing from the ability system. **[derived]** +That self-containment is why this kind of plugin is worth lifting wholesale, and +why its defects travel with it. + +--- + +## 2. Type safety is real, and has exactly one hole + +The payload crosses the bus as an untyped pointer plus a reflected struct type. +The gate is one line **[measured]**: deliver when the listener never had a valid +type, **or** when the sent struct derives from the listener's expected struct. + +That gives three branches, and they are not equally safe: + +| Branch | Behaviour | Assessment | +|---|---|---| +| Sent type derives from expected type | delivered | correct, and makes exact equality a special case | +| Types unrelated | logged as an error, listener skipped, broadcast continues | correct and loud | +| Listener registered with no expected type | **check bypassed entirely** | documented "for internal use" — and reachable from a Blueprint node with an empty type pin **[measured]** | + +The third row is the hole. What makes it interesting is not that it exists — an +internal escape hatch is a reasonable thing to have — but that the comment +asserting it is internal is the only thing protecting it **[derived]**. In the +Blueprint path a second, stricter check downstream limits the damage to a failed +payload extraction. In C++ the same registration combined with a templated +callback would reinterpret unrelated bytes with no diagnostic at all. Recipe: +MB-03. + +There is also a defensive feature worth copying: the expected struct type is held +weakly, with a flag recording that it was once valid, so a listener whose struct +was garbage collected is detected and removed at broadcast time with a warning +**[measured]**. Epic's own comment on those two fields says they were added in +response to real problems. **[derived]** That is a project that was bitten and +responded structurally — and the removal code it added is where MB-01 lives. + +--- + +## 3. The defect worth reading twice + +During the parent walk, a stale listener is removed using the **original +broadcast channel** rather than the ancestor tag currently being visited +**[measured]**. Handle IDs are unique only within a channel. + +This entry earns its place because of *where* the mistake is, not what it does: + +- the loop header three lines up rebinds what "the channel" means on every + iteration; +- the error-log statement a few lines **below** the bug uses the loop variable + correctly **[measured]**; +- so the same function contains both the right and the wrong idiom, and the wrong + one is the one that mutates state. + +**[derived]** Two outcomes, both quiet. Either nothing is found and the warning +repeats on every subsequent broadcast forever, or an unrelated listener holding +the same channel-local ID in the broadcast channel's bucket is removed instead, +and its owner silently stops receiving. + +It requires hierarchical matching plus an expired struct type simultaneously, so +the audited project never hits it — its own usage has neither. **That is exactly +why it matters to someone lifting the plugin**: the conditions that make it +harmless are properties of the original product, not of the code. Recipe: MB-01. + +--- + +## 4. A fully built feature that nothing uses + +Hierarchical matching is implemented in the broadcast walk, exposed in the enum, +and selectable in the Blueprint node **[measured]**. Game-code usage: **zero** +**[measured]**. + +The cause is one unforwarded default argument. The overload that binds a UObject +member weakly — the safest form, and the one used almost everywhere — delegates to +the raw form without passing the match type, so it is always exact **[measured]**. +Reaching hierarchical matching from C++ requires switching to a lambda or an +options struct, for reasons the API never states. + +**[derived]** The general shape is worth more than the instance: **a feature can +be complete, correct, tested and documented, and still be unreachable from the +path everyone actually uses.** Auditing for "is it implemented?" finds it. +Auditing for "who calls it?" finds the truth. Recipe: MB-05. + +Two neighbouring subsystems in the same product had independently copied the +channel-tag-plus-match-type pattern **[measured]** — evidence that the idea was +considered good enough to reproduce, in a project where the original was never +switched on. + +--- + +## 5. Blueprint parity is where the asymmetries live + +The scripting layer is a separate module, editor-only, with its own node +**[measured]**. It does several things well: a wildcard output pin retyped from +the selected payload struct, a compile-time error when a wildcard payload pin is +connected but untyped, and a copy of the payload by value at the boundary so +scripting never holds the sender's stack pointer **[measured]**. + +And then the two paths disagree about types **[measured]**: + +| | C++ | Blueprint | +|---|---|---| +| Type gate | derived accepted | strict equality | +| Mismatch | logged as an error | **silent skip** | + +**[derived]** A Blueprint listener on a base struct silently misses derived +payloads that its C++ equivalent receives, and reports nothing. The observable is +"the node never fires", which everyone investigates as a wrong tag first, because +that hypothesis is cheap to test. + +Neither behaviour is wrong on its own. They were written for different consumers +and each is correct against its own local reasoning; nothing in either file +mentions the other. **The rule that would have caught it is procedural, not +technical: one test matrix, run through both APIs.** Recipe: MB-02. + +--- + +## 6. Lifetime: three separate ways to leak + +| Shape | Measured | Consequence **[derived]** | +|---|---|---| +| Listener map cleared only at GameInstance teardown | reset appears in `Deinitialize` and nowhere else | Records from a travelled-away world persist for the session. Weak binding stops the calls, not the accumulation. MB-08 | +| Async node detects owner death only while handling a message | cleanup keyed on the delegate being unbound, inside the message handler; authors' own TODO records the limitation with a tracker ID | On a quiet channel, cleanup waits for traffic. MB-07 | +| Options-struct registration with no bound callback | returns a default invalid handle, no log | A listener that never fires, in code that reads as wired. MB-06 | + +The first row generalises past this plugin: **a weak callback is not an +unsubscription.** It prevents a call into a dead object and leaves the record. +Any bus that binds weakly and never sweeps will accumulate, and the accumulation +is invisible precisely because weak binding made it harmless. + +The second row is a good example of a finding you should *want* to discover: +the authors documented the limitation themselves, in a comment, with a reference. +A TODO that names the problem is the strongest confirmation available from +reading source. + +### And a teardown hazard on the consumer side + +The static accessor resolves the world in assert-on-failure mode and then asserts +on both the world and the subsystem **[measured]**. Three consumer classes in the +product call it from destruction paths **[measured]** — where the world may already +be gone — when unregistering through the handle would have resolved the subsystem +weakly and no-opped. **[derived]** Correct in every ordinary shutdown; an assert +in the unusual orders that nobody can reproduce on demand. Recipe: MB-09. + +--- + +## 7. How the product worked around no networking + +The bus has no networking at all: no server, client or multicast functions, no +custom serialisation, no replicated properties **[measured]**, in a module small +enough to have been read in full. The surrounding product needed networked facts +anyway, and solved it three ways — worth comparing, because they are not equally +good. + +**1. RPC wrappers on game state and player state.** A multicast or client-targeted +RPC receives the message struct and, on arrival, rebroadcasts it locally +**[measured]**. The rebroadcast is guarded by a net-mode check so a listen server +does not process both the server-side publication and the client-side +rebroadcast **[measured]**. That guard is the transferable part; without it every +listen-server host sees everything twice. The headers say plainly that these are +for notifications that can be lost. + +**2. A replicated fast array that rebroadcasts on arrival.** Fully implemented and +**used nowhere in the project** — no member of that type exists anywhere +**[measured]**. It is a worked example rather than working code, and it carries an +empty removal callback (see the network-authority skill for why that matters at +adoption time). + +**3. Replicated state plus a local broadcast on the replication callback.** The +cleanest and the most used **[measured]**. State replicates; the message is purely +a local notification that it changed. + +**[derived]** The third pattern is the one to copy, and the reason is the +state-versus-event test: only in that arrangement is a late joiner correct, +because the truth is in the replicated state and the message carries none of it. + +--- + +## 8. What the bus does not do, stated plainly + +Read as a list of things you will otherwise discover one at a time: + +- **No sticky or replay messages.** A subscriber cannot ask for the last value. + The product compensates by having new widgets read the owning component + directly on construction **[measured]** — which only works because the state has + an owner. MB-13. +- **No ordering guarantee.** The authors say so in a header comment, and removal + by swap-with-last actively reorders the array **[measured]**. MB-11. +- **No per-player filtering.** Consumers filter inside the callback, comparing a + target field against their own owner **[measured]**. +- **No cancellation.** Every matching listener receives every message. +- **No recursion guard.** Broadcasting from inside a callback is supported and + used deliberately by an aggregator in the product **[measured]**; nothing bounds + the depth if the channel graph ever contains a cycle. MB-12. +- **No thread safety.** Game thread only **[measured]**. +- **No networking.** §7. + +--- + +## 9. The architecture it made possible + +Worth stating, because the failure-mode list above is not an argument against the +pattern: + +```text +ability system --damage / elimination--> + processors in a feature plugin --assist / chain / streak--> + scripted relay --notification--> + UI widget +``` + +No layer references the next **[measured]**. The processors live in a feature +plugin that the base module has never heard of, and they are added and removed +with the feature. **[derived]** That is the whole case for a bus: it let +game-mode-specific logic move into a shippable plugin without creating a single +reference from the base module. + +Note also the vocabulary discipline that came with it: channels named so bus +traffic is distinguishable from other tags, tags declared next to the publisher, +payload fields initialised at declaration because a payload never passes through +serialisation and gets no free defaults **[measured]**. + +And the cost of that discipline, measured: one elimination channel string was +declared independently at **four** sites — the publisher and three processors +**[measured]**. It works because the engine deduplicates by string. **[derived]** +It means the string is the real contract and a single-site edit detaches one +participant silently. Recipe: MB-10. + +--- + +## 10. What source reading could not settle + +Stated because it bounds several claims above: + +- **Blueprint call sites are invisible to text search.** Three RPC wrappers and + one notification channel have zero C++ callers or publishers, and all are + script-callable **[measured]**. Whether they are used, and by what, is + **[open]**. They are recorded as open questions, not as dead code. +- **The extent of scripted bus usage is unknown.** A search for the listen node + across all source returns nothing **[measured]**, which means the entire + scripted traffic of the bus lives in binary assets. +- **The channel-key defect was found by reading, not by running.** It has not been + reproduced at runtime **[open]**, and its trigger conditions do not occur in the + audited project. + +--- + +## Provenance + +Measured against the gameplay message router plugin and its consumers in Epic's +Lyra Starter Game on Unreal Engine 5.6, read as source in a single workspace. The +runtime module was read in full, line by line; the negative results in §7 and §8 +rest on that rather than on a search. + +Source addresses stay in the research archive that produced this skill; each +`MB-` identifier resolves back to the audited location there, so any specific +claim above can be produced on request. + +## Evidence boundary + +One plugin, one engine version, one workspace. These are examples and failure +evidence, not guarantees about other versions. Binary assets were not read, so +every statement about scripted *usage* is bounded by what source could show and is +marked open where it is open. Re-run the recipes in +[failure-modes.md](failure-modes.md) against your own tree before acting on +anything here. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/SKILL.md b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/SKILL.md new file mode 100644 index 0000000..d835761 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/SKILL.md @@ -0,0 +1,386 @@ +--- +name: ue-gameplay-tag-governance +description: >- + Govern GameplayTags in Unreal Engine as a typed vocabulary rather than loose + strings: choosing declaration source (native macro, project ini, plugin ini, + DataTable, platform ini, runtime), namespace design, constraining tag inputs + with meta=(Categories=...), exact vs hierarchical matching, redirects and + renaming, replication of tag sets, tag-driven counters, and auditing an + existing tag inventory. Use when adding a tag namespace, wiring systems + through tags, reviewing tag usage, or debugging a tag-driven feature that + silently does nothing. +--- + +# UE gameplay tag governance + +The invariant: + +> A GameplayTag is a foreign key with no database behind it. Every tag join must +> be validated on both sides by something other than the compiler. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Eighteen detection recipes for silent tag failures: +[failure modes](references/failure-modes.md). + +Read the failure modes before adopting or reviewing a tag-driven design. Each +entry names the mechanism, why it stays silent, why the obvious check misses it, +and a command you can run against your own tree. + +Related skills: `ue-gameplay-messaging`, `ue-gas-architecture`, +`ue-modular-gameplay`, `ue-input-architecture`, `ue-architecture-guardrails`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Why tags need governance at all + +A tag replaces a compile-time reference with a runtime string comparison. That is +the whole point — and the whole cost. Once a system is joined by a tag: + +- the compiler no longer sees the dependency; +- "find all references" no longer finds it; +- a typo is not an error, it is an empty result; +- deleting the producer does not break the consumer's build. + +The characteristic failure is therefore never a crash. It is **nothing happens**, +reported weeks later by a designer, in a build nobody can bisect. + +Governance means deliberately re-adding the checks the compiler stopped doing. + +--- + +## 2. Decide the declaration source before the first tag + +Unreal accepts tags from at least six independent sources. They do not behave +the same and they are not equally visible to search. + +| Source | Form | Compile-time symbol | Found by grep pattern | +|---|---|---|---| +| Native, registered | `UE_DEFINE_GAMEPLAY_TAG_COMMENT` + `_EXTERN` in a header | yes | symbol name | +| Native, file-local | `UE_DEFINE_GAMEPLAY_TAG_STATIC` | file-local only | symbol name | +| Project config | `+GameplayTagList=(Tag="...")` in the project tag ini | no | `+GameplayTagList=` | +| Plugin config | `GameplayTagList=(Tag="...")` — **no leading `+`**, different section | no | `GameplayTagList=(Tag=` | +| DataTable | binary asset referenced by `+GameplayTagTableList` | no | not text-searchable | +| Runtime | `AddNativeGameplayTag` during module startup | no | call site only | + +### Rules + +1. **Default to native, registered tags** for anything C++ reads or writes. They + give Ctrl+Click, rename support, and a link error instead of a silent miss. +2. **Use ini only for namespaces owned by data**, where designers add entries + without engineers — cosmetic variants, surface types, platform traits. +3. **Never split one namespace across sources.** A namespace half in C++ and half + in ini has no single place to read its vocabulary. +4. **Record the source of every namespace** in one document. This is the only + defence against source #6, which no static inventory can see. +5. **Never trust a single grep pattern** when auditing. Plugin ini syntax differs + from project ini syntax by exactly one character, and a pattern anchored on + `+` returns a confident, wrong zero. Recipe: TG-05. + +--- + +## 3. Design the namespace as a tree that means something + +Hierarchy is not decoration. `MatchesTag` and partial matching turn parentage into +executable logic: a system asking "is this an action?" will answer yes for every +descendant, forever, including ones added next year by someone who never read that +call site. + +### Rules + +1. **A parent tag is a promise.** Adding a child enrols it in every hierarchical + check that already matches the parent. Only create a child if that enrolment + is intended. +2. **Do not encode two dimensions in one path.** `Weapon.Rifle.Silenced` conflates + class and modifier; the modifier cannot then be queried independently. Use two + tags. +3. **Depth is coupling.** Every additional level is another place a hierarchical + query can accidentally match. +4. **Reserve a leaf-only convention** for tags used as identities (channels, cue + names, state nodes). If a tag is an identity, nothing should ever be its child. +5. **Document each parent's contract** where the parent is declared — what a + descendant is required to satisfy. + +### Naming + +Pick one scheme and enforce it in review: + +```text +.. + +Ability.Type.Action.Reload +InitState.DataAvailable +Event.Combat.Elimination +Platform.Trait.Input.PrimaryController +``` + +Do not abbreviate inconsistently. Do not mix singular and plural at the same +level. These are unenforceable by tooling and therefore purely a review duty. + +--- + +## 4. Constrain every tag input + +An unconstrained `FGameplayTag` UPROPERTY offers the designer the entire project +vocabulary — thousands of entries, of which a handful are legal. This is the +single highest-leverage guardrail available, and it is one line: + +```cpp +UPROPERTY(EditAnywhere, meta = (Categories = "Ability.Type.Action")) +FGameplayTag ActionTag; + +UPROPERTY(EditAnywhere, meta = (Categories = "Event.Combat")) +FGameplayTagContainer ObservedChannels; +``` + +### Rules + +1. **Every editable tag property gets `Categories`.** No exceptions for "it's + obvious" — obvious is exactly what a typo defeats. +2. **`UPARAM(meta=(Categories=...))`** for `BlueprintCallable` tag arguments. +3. **Verify the constraint actually resolves.** `Categories` pointing at a + non-existent namespace produces an *empty* dropdown, not an error. It looks + like a working constraint and is worse than none, because it silently blocks + the designer while passing review. Recipe: TG-03. +4. **Re-verify `Categories` after any namespace rename.** It is a string; the + rename tool does not touch it. +5. **Measure coverage**: count tag properties versus properties carrying + `Categories`. Anything below full coverage is a list of known holes, and + should be tracked as such. Recipe: TG-04. + +### Verification procedure for a constraint + +Open the property in the editor. If the dropdown is empty, the namespace does not +exist. This takes five seconds and is the only check that exists. + +--- + +## 5. Choose exact versus hierarchical matching per call site + +Three different operations look alike and mean different things: + +| Call | Semantics | +|---|---| +| `A == B` | exact identity | +| `A.MatchesTagExact(B)` | exact identity, explicit | +| `A.MatchesTag(B)` | A is B **or a descendant of** B | +| `Container.HasTag(B)` | any element matches B or descends from it | +| `Container.HasTagExact(B)` | any element is exactly B | + +### Rules + +1. **State the intent at every call site.** Prefer the explicit `...Exact` form + even when `==` would do, so the next reader does not have to infer. +2. **Never mix the two forms across one feature's read and write paths.** If a + rule is stored hierarchically and queried exactly (or the reverse), the feature + works for the tags that were tested and fails for the rest — the worst possible + failure distribution. Recipe: TG-07. +3. **Hierarchical matching is a public API.** Once a system matches on a parent, + the parent's whole subtree is part of that system's contract. +4. **Default to exact** for identities; use hierarchical only where "any kind of + X" is genuinely the rule. + +--- + +## 6. Renaming: redirects are not optional + +A tag lives in binary assets, ini files, DataTables, and blueprint graphs. A +rename without a redirect silently zeroes every one of those references. + +```ini +[/Script/GameplayTags.GameplayTagsSettings] ++GameplayTagRedirects=(OldTagName="Ability.Type.Action.Fire",NewTagName="Ability.Type.Action.WeaponFire") +``` + +### Rules + +1. **Add the redirect in the same commit as the rename.** Not after; a build + between the two loses data permanently in any asset that gets resaved. +2. **Never delete a redirect** until every asset in every branch has been resaved + and verified. Redirects are cheap; lost designer work is not. +3. **A project with zero redirects and a non-trivial tag history has already lost + references** — or has never renamed a tag, which is itself worth confirming. + Recipe: TG-18. +4. **Typos are renames.** Fixing a misspelled tag requires a redirect exactly like + any other rename. + +### The symmetric-typo trap + +A misspelled tag that is misspelled *identically* in every location works +perfectly. It will keep working until one person fixes one of the spellings — +at which point the join breaks with no error at either end. Treat a discovered +typo as a coordinated multi-file change with a redirect, never as a local fix. +Recipe: TG-02. + +--- + +## 7. Replication of tag vocabulary + +Tags replicate as indices, not strings, when fast replication is enabled. That +requires **client and server to have byte-identical tag registries**. + +### Rules + +1. **Any source of tags that can differ between builds breaks fast replication.** + That includes plugin-supplied tags when the plugin loads on one side only, and + any runtime tag registration. +2. **Register tags before networking starts**, unconditionally on both sides. +3. **Watch the replicated container size cap.** It limits how many tags fit in one + replicated container; exceeding it is a runtime failure, not a compile error. + Recipe: TG-11. +4. **Do not gate tag registration on `IsRunningDedicatedServer()` or platform.** + Gate the *content* those tags refer to, never the vocabulary. + +--- + +## 8. Tag-driven counters and state + +Tag-to-count containers (stacks) are a common pattern for ammo, charges, +resources. They are arithmetic wearing a tag as a label, and inherit arithmetic's +problems. + +### Rules + +1. **Validate the amount, not only the tag.** A zero or negative delta that + silently returns is a lost write that reports as "the ability did nothing". + Recipe: TG-12. +2. **Guard the accumulator.** Unchecked `int32` addition on a stack count is an + overflow waiting for an idle-game or a modifier loop. Recipe: TG-13. +3. **Decide the contract for removing more than exists**, then implement it — + clamp to zero, fail loudly, or allow negative. Silence is not a contract. +4. **Keep the map and the array in sync in every replication callback** if the + container maintains both a replicated array and a lookup map. Every + `PreReplicatedRemove` / `PostReplicatedAdd` / `PostReplicatedChange` must + update the map, or clients read stale counts. + +--- + +## 9. String-to-tag lookup is a debug tool + +Resolving a tag from a runtime string — especially with partial/substring +matching — returns the first registry match in registry order. That order is not +stable against adding unrelated tags. + +### Rules + +1. **Confine partial matching to cheat/console/debug paths.** +2. **Do not export a partial-match helper with a permissive default** across a + module boundary. The first production caller who passes `true` gets a + non-deterministic tag and no warning. Recipe: TG-09. +3. **Exact lookup with explicit failure handling** is the only form allowed in + gameplay code. + +--- + +## 10. Auditing an existing tag inventory + +Run this before trusting any tag-driven system you did not write. Each step maps +to a recipe in the failure modes. + +```text +1. Count declarations per source (all six). Reconcile against the documented + list. Count with two patterns, not one. TG-05, TG-06 +2. Count editable tag properties; count those with Categories. + Report the ratio. TG-04 +3. Resolve every Categories string against the actual namespace list. + Empty dropdown = broken constraint. TG-03 +4. List every parent tag used in a hierarchical match; enumerate its + current descendants and confirm each one belongs. TG-08 +5. Search for tags declared but never read, and read but never declared. TG-16 +6. Count redirects. Zero redirects plus renamed tags in history = + data loss. TG-18 +7. Diff the tag registry between a client build and a server build. TG-10 +``` + +Steps 1 and 3 find most real defects. Step 4 finds the ones that will appear +later, when someone adds a child. + +--- + +## 11. Review checklist + +- [ ] Namespace has one documented declaration source. +- [ ] Tags read by C++ are native, not ini-only strings. +- [ ] Every editable tag property/parameter carries `Categories`. +- [ ] Every `Categories` string resolves to a namespace that exists. +- [ ] Exact vs hierarchical matching is explicit and consistent per feature. +- [ ] Parent tags used in hierarchical matches have a documented contract. +- [ ] Renames ship with redirects in the same commit. +- [ ] Tag registration is unconditional and identical on client and server. +- [ ] Tag counters validate amount and guard overflow. +- [ ] Partial string lookup exists only in debug paths. +- [ ] A tag-driven feature has a test that fails when the tag is wrong. + +--- + +## 12. Tests + +The purpose of these tests is to convert "nothing happened" into a red build. + +### Vocabulary + +- every tag referenced in C++ resolves at startup; +- every tag in ini is referenced somewhere, or is explicitly marked data-owned; +- registry contents identical between client and server targets. + +### Matching + +- exact query does not match a descendant; +- hierarchical query matches every current descendant; +- adding a new child does not change any exact query's result. + +### Constraints + +- each `Categories` string yields a non-empty candidate set; +- a property rejects a tag outside its category. + +### Counters + +- add zero, add negative, remove more than present, remove from empty; +- replication callbacks leave map and array consistent; +- large accumulations do not overflow. + +### Migration + +- renaming a tag with a redirect preserves asset references; +- loading an asset saved with the old name resolves to the new tag. + +--- + +## 13. When not to use a tag + +Use a stronger construct when: + +- the set of values is closed and owned by code → **enum**; +- the value selects an implementation → **class reference or DataAsset**; +- the join must be checked at compile time → **interface or typed pointer**; +- exactly one consumer exists and it is in the same module → **direct call**. + +A tag earns its cost when the vocabulary is open, the consumers are unknown to the +producer, and the join genuinely crosses a module or data boundary. Used anywhere +else, it converts a compile error into a silent runtime no-op — which is the only +thing it reliably does. + +--- + +## Provenance + +The findings behind the failure modes come from a line-by-line source audit of +Epic's Lyra Starter Game on Unreal Engine 5.6, read rather than run. Source +addresses stay in the research archive that produced this skill; what ships is the +detection recipe, because an address in someone else's tree is not something you +can act on and a recipe is. + +Each entry carries a stable identifier (`TG-01`, `TG-02`, …) that resolves back to +the audited location in that archive. If you need the original address to settle a +dispute, it exists and can be produced. + +## Evidence boundary + +The measured material refers to one specific reference project on one engine +version in one workspace. It is evidence of failure modes, not a guarantee about +other engine versions or other samples. Re-run the recipes against your own tree +before trusting any specific claim — including the counts, which are properties of +that project on that day and are quoted only to show the shape of the contrast a +recipe produces. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/failure-modes.md new file mode 100644 index 0000000..c3f39f5 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/failure-modes.md @@ -0,0 +1,670 @@ +# Failure modes: gameplay tag governance + +Eighteen ways a tag-driven system stops working without reporting anything. + +The common shape of every entry below: **the system does not fail, it declines to +act.** Tag defects do not produce crashes, log errors or failed builds. A missing +join and a correctly-evaluated "no" are the same observable. Design the gates +accordingly — the only reliable detection is a test that asserts the *positive* +behaviour, or a recipe that compares two counts. + +Recipes use `rg` and are written to be run from a project root. Most of them are +deliberately a **pair** of searches: the class of defect here is "the thing is +present, and the thing that would give it meaning is absent", and a single search +can only ever show you the first half. + +--- + +## Vocabulary and declaration + +### TG-01 - Typo in a string-declared tag + +**Mechanism.** A tag declared in ini, or authored into an asset by hand, is +misspelled. No compiler, linker or loader rejects it, because a tag is a string +until the moment something tries to match it. + +**Why it is silent.** The lookup succeeds in the sense that matters to the engine: +it returns "no match", which is a legal answer that the calling system is required +to handle gracefully. A misspelled tag is indistinguishable from a condition that +is legitimately not met right now. + +**Why the obvious check misses it.** Review reads the tag and it looks right — +misspellings survive review precisely because they are near-misses. Compile, +cook and startup all pass. The only artefact that would reveal it is a runtime +comparison that nobody logs, because logging every failed tag match would drown +the log. + +**Symptom.** The feature does nothing. No log line, no error. Usually reported by +a designer, weeks later, as "this never worked" — by which time the change that +introduced it is far outside any bisect window. + +**Detect.** Resolve the declared vocabulary against the referenced one: + +```bash +# every tag string declared in config +rg -o -N 'Tag="([^"]+)"' -r '$1' Config/ Plugins/*/Config/ | sort -u > /tmp/declared +# every tag string mentioned in source +rg -o -N '"([A-Z][A-Za-z]+(\.[A-Za-z][A-Za-z0-9]*)+)"' -r '$1' Source/ | sort -u > /tmp/used +comm -13 /tmp/declared /tmp/used # referenced but never declared +``` + +The third command is the finding: a tag string that code compares against and +that no source declares. + +**Guardrail.** Native tag macros for anything C++ reads, so a typo becomes a link +error. `Categories` constraints on every editable tag property so authoring uses a +dropdown rather than a keyboard. + +--- + +### TG-02 - Symmetric typo + +**Mechanism.** The same misspelling exists at both ends of a join. The feature +works. A later cleanup corrects one side. + +**Why it is silent.** It is not merely silent — it is *correct*. Both sides agree, +the join resolves, the feature ships and works for years. Nothing is wrong until +someone makes it right. + +**Why the obvious check misses it.** Every check that exists passes, because the +system is working. The defect has no runtime signature at all; it is latent in the +spelling, and it converts into a failure only through a well-intentioned edit +whose diff looks obviously correct — which is why it also survives review of the +fix. + +**Symptom.** A previously working feature breaks, and the commit that broke it is +a one-word spelling correction that every reviewer approves. + +**Detect.** Before correcting any tag spelling, count the sites: + +```bash +T='Primarly' # the suspected misspelling +rg -n --no-heading "$T" Config/ Source/ Plugins/ | tee /tmp/sites | wc -l +``` + +On the measured reference this returned six sites across four files: two project +tag declarations, two platform-specific config files, and a native tag macro plus +its use in one widget class. Correcting any one of them would have broken the +join with no error at either end. + +**Guardrail.** Treat a discovered typo as a coordinated multi-site rename plus a +redirect, never as a local edit. The rule generalises: **when a string literal is +the join, correctness means "all sides agree", not "the string is right".** + +--- + +### TG-03 - `Categories` filter pointing at a namespace that does not exist + +**Mechanism.** The `Categories` metadata string is not validated against the tag +tree. A stale or mistyped namespace yields an empty candidate set. + +**Why it is silent.** An empty dropdown is a legitimate state — it is what a +correctly-configured filter shows before anyone has created tags under that +namespace. The editor has no way to distinguish "this filter is wrong" from "this +namespace is empty today". + +**Why the obvious check misses it.** The property has a constraint, which is the +thing reviewers check for, and it is spelled plausibly. Confirming it requires +opening the property in the editor and looking at the dropdown — a five-second +check that nothing prompts anyone to perform, because the code reads as complete. + +**Symptom.** Designers see an empty dropdown, conclude the tags are not ready yet, +and leave the field blank. The feature is then wired to an unset tag, which is the +same failure as TG-01 arriving by a different route. + +**Detect.** Extract every filter string and confirm each resolves: + +```bash +rg -o -N 'Categories\s*=\s*"([^"]+)"' -r '$1' Source/ | sort -u > /tmp/filters +while read -r ns; do + n=$(rg -c "Tag=\"$ns" Config/ Plugins/*/Config/ 2>/dev/null | awk -F: '{s+=$2} END {print s+0}') + [ "$n" = "0" ] && echo "EMPTY FILTER: $ns" +done < /tmp/filters +``` + +On the measured reference one filter named a namespace that appeared nowhere else +in the project, while the real tags lived under a differently-named root. That +filter was one of only fourteen constrained properties in the entire codebase — +one of the few places tag input was validated at all, and it was misconfigured. + +**Guardrail.** Open every constrained property once and confirm the dropdown is +non-empty. Re-check after any namespace rename: `Categories` is a string, and no +rename tool touches it. + +--- + +### TG-04 - Unconstrained tag property + +**Mechanism.** A tag field without `Categories` offers the entire project +vocabulary — thousands of entries, of which a handful are legal here. + +**Why it is silent.** Any tag the designer picks is a valid tag. It serialises, it +loads, it compares. The system that reads it simply never matches, which is the +TG-01 observable again. + +**Why the obvious check misses it.** There is nothing to see. The absence of one +optional metadata specifier does not appear in a diff as a change, does not warn, +and does not look like an omission in review — most properties do not have it +either, so it reads as normal. + +**Symptom.** Wrong-namespace tags assigned in assets; the system ignores them +silently. The designer's mental model — "I set the tag, so it is wired" — is +reasonable and wrong. + +**Detect.** Report the coverage ratio, which is the only number that matters: + +```bash +total=$(rg -c "FGameplayTag(Container)?\s+\w+;" Source/ | awk -F: '{s+=$2} END {print s+0}') +constrained=$(rg -c "meta\s*=\s*\([^)]*Categories" Source/ | awk -F: '{s+=$2} END {print s+0}') +echo "constrained $constrained of $total" +``` + +On the measured reference this was fourteen of sixty-four — under sixteen per cent +— and one of the fourteen was broken per TG-03. A low ratio is not itself a bug; +it is a list of places where a typo cannot be caught. + +**Guardrail.** `meta=(Categories=...)` on every editable tag property and `UPARAM` +on every Blueprint-exposed tag argument. Track the ratio as a number that only +goes up. + +--- + +### TG-05 - A grep pattern that misses a declaration source + +**Mechanism.** Plugin tag config uses `GameplayTagList=` without the leading `+` +and under a different section header than project config. A pattern anchored on +`+` never sees it. + +**Why it is silent.** The audit succeeds. It produces a number, the number is +plausible, and nothing indicates it is partial. A confident wrong answer is +produced by a working tool. + +**Why the obvious check misses it.** The obvious check *is* the grep, and greps +do not report what they could not have matched. This is the general hazard of +proving absence by search: **an empty or small result proves the pattern, not the +codebase.** + +**Symptom.** An audit reports a small tag vocabulary and concludes the project is +safer than it is. Renames and migrations are then planned against an inventory +that is missing most of its entries. + +**Detect.** Count with two patterns and require them to agree: + +```bash +rg -c '^\+GameplayTagList=' Config/ # project form +rg -c 'GameplayTagList=\(Tag=' Config/ Plugins/*/Config/ # both forms +``` + +On the measured reference the first pattern found 99 tags in one file; the second +found 212 across five files, four of which were plugin configs. The difference is +not an edge case — it is the majority of the vocabulary. + +**Guardrail.** Reconcile every count against a documented per-namespace source +list. Never conclude "absent" from one pattern; widen it, then widen the path. + +--- + +### TG-06 - Tag declared in a source no static search can see + +**Mechanism.** DataTable-sourced tags live in binary assets. Runtime registration +leaves no declaration text at all. + +**Why it is silent.** These tags work perfectly. They resolve, match and replicate +like any other. They are invisible only to *tooling*, and tooling is not what +exercises them. + +**Why the obvious check misses it.** Every text-based inventory is structurally +incapable of seeing them. Widening the grep does not help, because there is no +text to find — the failure is in the choice of instrument, not its configuration. + +**Symptom.** Tags exist at runtime that no inventory, document or search reports. +A rename sweep misses them; a "delete unused tags" pass deletes something that is +in use. + +**Detect.** Ask the runtime, not the filesystem: + +```bash +rg -n "GameplayTagTableList|AddNativeGameplayTag|AddTagTableRow" Config/ Source/ +``` + +Any hit means your static inventory is incomplete by construction. Close the gap +by dumping the tag registry from a running build and diffing it against the static +inventory; the difference is exactly the set this entry describes. + +**Guardrail.** Document every namespace's declaration source in one place. Forbid +runtime registration outside module startup, so at least the timing is knowable. + +--- + +## Matching semantics + +### TG-07 - Exact/hierarchical mismatch inside one feature + +**Mechanism.** One code path uses hierarchical matching and another uses exact +comparison, on the same data, sometimes in the same class. + +**Why it is silent.** Both calls are correct in isolation and both are correct for +the tags used during development — typically leaf tags, where hierarchical and +exact matching give identical answers. The two forms only diverge for a tag that +has descendants, which is a case that arrives later. + +**Why the obvious check misses it.** Reading either call site shows a defensible +choice. The defect exists only in the relationship between two call sites that may +be hundreds of lines apart, and no tool reports "these two comparisons on the same +field disagree about semantics". + +**Symptom.** The rule works for the tags used in testing and fails for others — +the failure distribution that is hardest to reproduce, because the reporter's +repro steps look identical to the ones that pass. + +**Detect.** List both forms per file and look for files that contain both: + +```bash +rg -l "MatchesTag\(" Source/ > /tmp/hier +rg -l "MatchesTagExact\(|HasTagExact\(" Source/ > /tmp/exact +comm -12 <(sort /tmp/hier) <(sort /tmp/exact) +``` + +Files in the intersection are candidates. On the measured reference one mapping +class used hierarchical matching in two methods and exact matching in a third, all +against the same tag container. + +**Guardrail.** State matching intent explicitly at each call site — prefer the +`...Exact` form even where `==` would do. Review the read and write paths of a +rule together, never separately. + +--- + +### TG-08 - Unintended enrolment via a new child tag + +**Mechanism.** A system matches on a parent tag. Adding a child anywhere in the +project enrols that child in the system automatically. + +**Why it is silent.** The enrolment is the feature working as designed. Nothing +malfunctions; a new tag simply acquires behaviour its author did not know about, +and that behaviour is usually plausible. + +**Why the obvious check misses it.** The person adding the tag is looking at the +tag list, not at every hierarchical match in the codebase. The person who wrote +the match is not present. There is no declaration anywhere that says "this parent +is load-bearing", so the coupling exists only in the reader's memory. + +**Symptom.** A newly added, unrelated tag activates behaviour nobody wired, in a +subsystem nobody touched, in the same release as dozens of other changes. + +**Detect.** Enumerate the parents that carry behaviour, then their descendants: + +```bash +rg -o -N 'MatchesTag\(([A-Za-z_:]+)\)' -r '$1' Source/ | sort -u # load-bearing parents +# for each parent P, list current descendants: +rg -o -N 'Tag="(P\.[^"]+)"' -r '$1' Config/ Plugins/*/Config/ +``` + +Read the descendant list and confirm every entry belongs in that subsystem. Do +this at the moment you add a hierarchical match, and record the result next to +the parent declaration. + +**Guardrail.** Document each parent's contract at its declaration and treat parent +tags used in hierarchical matches as public API. Reserve a leaf-only convention +for tags that are identities. + +--- + +### TG-09 - Partial string lookup in gameplay code + +**Mechanism.** A helper resolves a tag from a substring, returning the first match +in registry order. + +**Why it is silent.** It returns a real, valid tag. Everything downstream works +normally with it. The answer is only wrong in the sense that it is not the answer +you would have chosen, and there is no way to notice that from the outside. + +**Why the obvious check misses it.** The function has a plausible name and a +`bool` parameter that defaults to the safe behaviour. Reviewing the helper shows a +correct implementation of what it says it does. The hazard is in the *export* — +it crosses a module boundary with a permissive mode available — and export +surfaces are reviewed for compilation, not for misuse potential. + +**Symptom.** The resolved tag changes when an unrelated tag is added elsewhere in +the project. Non-deterministic behaviour, no error, and a repro that depends on +the contents of the tag registry rather than on anything in the feature. + +**Detect.** Find partial-match helpers and check who calls them with matching +enabled: + +```bash +rg -n "PartialMatch|bMatchPartialString|FindTagByString" Source/ +rg -n "FindTagByString\([^)]*true" Source/ # callers opting into partial +``` + +On the measured reference the helper was exported from the game module with the +permissive parameter defaulting to off, and its only current callers were cheat +commands passing it on — safe today, one production caller away from +non-deterministic. + +**Guardrail.** Confine partial matching to cheat and console paths. Do not export +such a helper across a module boundary at all; if you must, do not give it a +permissive mode. + +--- + +## Replication + +### TG-10 - Divergent tag registries under fast replication + +**Mechanism.** Fast replication sends tags by registry index. Client and server +must hold identical registries for an index to mean the same thing on both ends. + +**Why it is silent.** Indices are always valid. A mismatched index does not fail +to deserialise — it resolves to a different, perfectly real tag. The receiving +side then acts correctly on the wrong fact. + +**Why the obvious check misses it.** PIE shares one process and therefore one +registry, so the divergence cannot exist in the environment where testing happens. +Any check performed in the editor is performed on a configuration where the bug is +impossible by construction. + +**Symptom.** Wrong tags received, or replication failures, only in networked +builds with asymmetric plugin loading — and the symptom is a wrong behaviour, not +an error. + +**Detect.** Find every conditional path into tag registration: + +```bash +rg -n "FastReplication" Config/ +rg -n -B4 "AddNativeGameplayTag|RegisterGameplayTag" Source/ \ + | rg "IsRunningDedicatedServer|WITH_EDITOR|#if|IsClient|bIsServer" +``` + +Any conditional wrapping tag registration is a divergence risk. Confirm by +dumping the registry from a client build and a dedicated-server build and diffing +them; that diff is the only authoritative answer. + +**Guardrail.** Register vocabulary unconditionally on both sides. Gate the +*content* a tag refers to, never the vocabulary itself. + +--- + +### TG-11 - Replicated container size cap + +**Mechanism.** A setting fixes the number of bits used for a replicated tag +container's element count, which caps how many tags fit in one container. + +**Why it is silent.** Under the cap, everything is correct. Over it, the excess +entries are simply not transmitted — a truncation, not an error. The receiving +side sees a valid container with fewer tags. + +**Why the obvious check misses it.** The cap lives in a settings file as a bit +count, not as a tag count, so its practical meaning requires a mental exponent. +Nothing connects it to the container that will eventually exceed it, and the +container that exceeds it does so only under gameplay conditions that accumulate +tags. + +**Symptom.** Tags beyond the cap are lost in transit with no compile-time or +runtime signal. Debugging starts at the gameplay logic, which is correct. + +**Detect.** Read the cap, convert it, and compare against your worst case: + +```bash +rg -n "NumBitsForContainerSize|FastReplication" Config/ +``` + +On the measured reference the value was 6 bits — a hard limit of 63 entries per +replicated container. Then find containers that grow without bound: + +```bash +rg -n "AddTag\(|AppendTags\(" Source/ | rg -i "replicat|net" +``` + +**Guardrail.** Know the cap as a tag count, not a bit count, and write it next to +any container that grows dynamically. Assert the size where it grows. + +--- + +## Counters + +### TG-12 - Tag counter swallowing zero or negative deltas + +**Mechanism.** A stack-modifying function guards on `if (Amount > 0)` with no +`else`, so a zero or negative amount returns silently. + +**Why it is silent.** Returning without acting is a legal response to a degenerate +input, and the caller has no return value to check. The write is lost with the +same observable as a write that was never attempted. + +**Why the obvious check misses it.** The function *does* validate — there is a +warning path for an invalid tag right beside the silent path for an invalid +amount. Seeing the tag validation satisfies the reviewer's question "is this +input checked?", and the asymmetry between the two checks is invisible unless you +compare them deliberately. + +**Symptom.** A lost write reported as "the ability did nothing". Because the tag +was valid, the tag-level warning never fires, so the log is clean. + +**Detect.** Compare validation coverage between the two parameters: + +```bash +rg -n -A6 "void .*Stack(Count|Tag)?\(|AddStack|RemoveStack" Source/ \ + | rg "if\s*\(.*Amount|ensure|UE_LOG" +rg -n "//\s*@?TODO" Source/ | rg -i "stack|amount|negative|zero" +``` + +On the measured reference the invalid-tag path carried warnings at both the add +and remove sites while the invalid-amount path returned silently at both, and an +unresolved TODO next to the remove path asked exactly the question this entry +describes. + +**Guardrail.** Validate the amount as strictly as the tag. Define and implement +the contract for removing more than exists — clamp, fail loudly, or allow +negative. Silence is not a contract. + +--- + +### TG-13 - Unguarded accumulation on a stack count + +**Mechanism.** Unchecked `int32` addition on a tag count, with no upper bound. + +**Why it is silent.** Every individual addition is correct. The wrap requires an +accumulation that no test performs, and after it wraps the value is still a valid +integer that every subsequent comparison evaluates without complaint. + +**Why the obvious check misses it.** Overflow guards are absent from most +arithmetic in most codebases, so their absence here does not stand out. The +addition is one line in a function whose logic is otherwise about tags, and the +reviewer's attention is on the tag semantics. + +**Symptom.** A count wraps to negative under a modifier loop or a long session. +Every `>= required` check then fails, so the ability with the most resources +behaves as if it had none. + +**Detect.** Find unbounded accumulation on stack counts: + +```bash +rg -n -B2 -A2 "StackCount\s*\+=|Count\s*\+=\s*\w+" Source/ \ + | rg -v "Clamp|Min\(|check\(|ensure" +``` + +Any addition with no clamp in the surrounding lines is the finding. + +**Guardrail.** Clamp against a documented maximum on every add, and state the +maximum where the container is declared. + +--- + +## Stubs and documentation + +### TG-14 - Tag-shaped property with no reader + +**Mechanism.** A tag container is exposed for authoring while the code that would +consume it is stubbed or deferred. + +**Why it is silent.** An unread property is not an error at any level of the +stack. It serialises, it appears in the editor, and it survives every cook. The +authored value simply goes nowhere. + +**Why the obvious check misses it.** The property exists, is documented by its +name, and appears in the class next to properties that do work. "Is this +implemented?" is answered by the presence of the field. The absent half is the +*reader*, and no default tooling asks who reads a property. + +**Symptom.** Designers configure the field; nothing happens. The feature appears +implemented in the editor and in review, so the investigation starts from "the +data must be wrong". + +**Detect.** For each editable tag property, search for a reader: + +```bash +P='HiddenByTags' # substitute your property +rg -n "\b$P\b" Source/ # declaration plus any uses +rg -n "\b$P\b" Source/ | rg -v "UPROPERTY|^\s*FGameplayTag" # actual reads +``` + +A declaration with no read is a stub. On the measured reference one widget base +class exposed such a container while both of its consumer sites hardcoded the +result of the check to a constant, with the real expression parked in a comment. + +**Guardrail.** Do not ship an editable property whose consumer does not exist. If +it must exist for serialisation reasons, mark it non-editable or log when it is +non-empty. + +--- + +### TG-15 - Restoring a stubbed tag check without restoring its inputs + +**Mechanism.** The stub from TG-14 coexists with other damage: the values the +restored check would select between are themselves overwritten at runtime. + +**Why it is silent.** Both halves are individually invisible. The stub is silent +per TG-14; the overwrite is a normal-looking assignment in an initialisation path. +Neither produces a symptom while the feature is switched off. + +**Why the obvious check misses it.** The obvious check is reading the stub, which +tells you exactly what to implement. It does not tell you that the inputs to the +thing you are about to implement have been clobbered somewhere else, because that +assignment lives in a different method with a different purpose. + +**Symptom.** The obvious one-line fix produces wrong behaviour instead of correct +behaviour, and the fix gets blamed and reverted — burying the original defect +under a failed attempt. + +**Detect.** Before re-enabling any stub, verify every input it reads is still +authored where it appears to be authored: + +```bash +V='ShownVisibility' # a field the stub would select between +rg -n "\b$V\b\s*=" Source/ # who writes it, and when +``` + +On the measured reference the designer-authored visibility fields were overwritten +by a setter called during construction, so implementing the tag check alone would +have selected between two values that no longer held the authored data. + +**Guardrail.** Treat "implement the stub" as a task that starts with tracing the +stub's inputs, not with writing the missing expression. + +--- + +### TG-16 - Vocabulary without documentation + +**Mechanism.** The per-tag comment field is optional and tag names are terse. + +**Why it is silent.** An undocumented tag works exactly as well as a documented +one. The cost is paid by people, later, and never by the build. + +**Why the obvious check misses it.** There is no check. Nothing measures +documentation coverage, and the field is empty by default, so an empty one looks +like every other one. + +**Symptom.** Duplicate tags with different spellings for the same concept; tags +nobody dares delete because nobody knows what reads them. The vocabulary grows +monotonically because deletion is unsafe. + +**Detect.** Measure the ratio: + +```bash +total=$(rg -c 'GameplayTagList=\(Tag=' Config/ | awk -F: '{s+=$2} END {print s+0}') +empty=$(rg -c 'DevComment=""' Config/ | awk -F: '{s+=$2} END {print s+0}') +echo "undocumented $empty of $total" +``` + +On the measured reference this was 68 of 99 in the project config — sixty-nine per +cent undocumented, in a project otherwise notable for its care. + +**Guardrail.** Require a comment in review for every declared tag, and run a +periodic declared-but-unread / read-but-undeclared diff so the vocabulary can +shrink as well as grow. + +--- + +### TG-17 - Vocabulary with no central registry + +**Mechanism.** File-local native tag macros scatter definitions across the module, +each valid, none discoverable from the others. + +**Why it is silent.** Every definition works. Scattering costs nothing at runtime +and nothing at build time; it costs only the ability to answer a question. + +**Why the obvious check misses it.** Looking at the central registry file shows a +well-organised list, and it is genuinely well organised — it is simply a minority +of the vocabulary. The check confirms what it can see, and what it cannot see does +not announce itself. + +**Symptom.** No single place answers "what tags does this project have". Audits +and renames are incomplete by construction, and each incomplete audit produces a +number that looks authoritative. + +**Detect.** Compare the central file against the whole tree, and include plugins: + +```bash +rg -c "UE_DEFINE_GAMEPLAY_TAG" Source/ Plugins/ | awk -F: '{s+=$2} END {print "total", s}' +rg -l "UE_DEFINE_GAMEPLAY_TAG" Source/ Plugins/ | wc -l +rg -c "UE_DEFINE_GAMEPLAY_TAG" Source/*/GameplayTags.cpp 2>/dev/null +``` + +On the measured reference this gave 94 definitions across 31 files, of which 35 +sat in the central registry — so roughly two thirds of the vocabulary lived +outside the file whose name promises to hold it. Note that restricting the search +to the main source directory undercounts by 18, which is TG-05 arriving again in a +different disguise. + +**Guardrail.** A central registry header for shared tags; file-local macros only +for tags used in exactly one file and never authored in data. Measure the central +share and treat a falling ratio as drift. + +--- + +## Migration + +### TG-18 - Rename without a redirect + +**Mechanism.** Tag references in assets and configs are strings resolved at load +time. A rename leaves them unresolvable. + +**Why it is silent.** An unresolvable tag reference loads as an empty tag. The +asset opens, the editor shows a blank field, and the system reads "no tag set" — +which is a legal state that means "not configured yet". + +**Why the obvious check misses it.** The rename itself is a code or config change +that compiles and passes review. The damage is in binary assets that nobody opens +during the change, and it becomes permanent only when one of them is resaved for +an unrelated reason — at which point the old string is gone from disk entirely. + +**Symptom.** Assets silently lose the tag. If they are resaved before anyone +notices, the reference cannot be recovered from the asset; it has to be +reconstructed from memory or version control. + +**Detect.** Count redirects against rename history: + +```bash +rg -n "GameplayTagRedirects" Config/ Plugins/*/Config/ # declared redirects +git log --oneline -S 'GameplayTagList' -- Config/ | head # config tag churn +``` + +On the measured reference the first command returned nothing at all, confirmed +with two independent patterns. Zero redirects means either that no tag has ever +been renamed, or that every rename was destructive — and the second command tells +you which. + +**Guardrail.** Ship the redirect in the same commit as the rename. Never remove a +redirect until every branch has been resaved and verified. Redirects are cheap; +reconstructing a designer's work is not. diff --git a/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/patterns.md b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/patterns.md new file mode 100644 index 0000000..d2995d3 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gameplay-tag-governance/references/patterns.md @@ -0,0 +1,205 @@ +# Patterns: tag vocabulary in a measured reference product + +A worked example of the governance rules from `SKILL.md`, against one specific +reference project read as source. It exists so the failure modes in +[failure-modes.md](failure-modes.md) can be read as *shapes that occur together* +rather than as a checklist of independent risks. + +The project audited below is a careful, well-organised codebase. That is the +point. Every finding here survived review by competent engineers, because none of +them has a runtime signature. + +Markers: **[measured]** — read in source or config; **[derived]** — conclusion +from measured facts; **[open]** — not answerable without the editor or a +networked build. + +--- + +## 1. Six declaration sources, and only one of them is easy to find + +The project declares tags from five distinct sources simultaneously +**[measured]**: + +| Source | Count | Visible to a naive text search | +|---|---:|---| +| Project config, `+`-prefixed | 99 | yes | +| Plugin configs, no `+`, different section | 113 across 4 files | **no** | +| Native macros, central registry file | 35 | by symbol | +| Native macros, elsewhere in the tree | 59 across 30 files | by symbol | +| DataTable rows in binary assets | **[open]** | **no** | + +Two numbers in that table are the lesson. + +**212 versus 99.** A pattern anchored on the project-config form finds 99 tags and +reports them confidently. The real config vocabulary is 212 across five files +**[measured]**. The pattern is not subtly wrong; it misses the majority. Recipe: +TG-05. + +**94 across 31 files, 35 central.** Native definitions split three ways — 24 plain +macros, 35 with comments, 35 file-local statics **[measured]**. The file whose +name promises to be the registry holds roughly a third of them. Restricting the +search to the primary source directory instead of including plugins undercounts by +18 **[measured]** — the same failure as the config case, one directory level up. +Recipe: TG-17. + +**[derived]** There is no single artefact in this project that answers "what tags +exist". Any audit, rename or deletion built on one is incomplete by construction, +and will produce a number that looks authoritative. + +--- + +## 2. Validation coverage: fourteen of sixty-four, one of them broken + +Editable tag properties: 64. Properties carrying a `Categories` constraint: 14 +**[measured]** — under sixteen per cent. + +That ratio is a list of places where a typo cannot be caught. But the sharper +finding is inside the fourteen: one constraint named a namespace that appears +nowhere else in the project, while the tags it was meant to offer live under a +differently-named root **[measured]**. + +**[derived]** The filter does not narrow the picker, it empties it. And it does so +in one of the few places where tag input was validated at all — so the single most +protective mechanism in the file was the one that was misconfigured. + +The general form is worth stating: **a constraint is only a guardrail if something +proves it resolves.** An empty dropdown is indistinguishable from a namespace +nobody has populated yet, so the editor cannot tell you, and neither can the +compiler. Recipes: TG-03, TG-04. + +--- + +## 3. A misspelling that works because every side repeats it + +The token `Primarly` — for `Primarily` — appears at six sites across four files +**[measured]**: two project tag declarations, two platform-specific config files, +and a native tag macro plus its use in a widget class. + +**[derived]** It functions precisely because every side misspells it identically. +Fixing any one site breaks the join, with no error at either end and a diff that +every reviewer would approve. + +This is the cleanest available demonstration of the invariant at the top of the +skill: **when a string literal is the join, correctness means "all sides agree", +not "the string is right".** Recipe: TG-02. + +The project's exposure to this class is total in a specific way: **zero +`GameplayTagRedirects` are declared**, confirmed with two independent search +patterns **[measured]**. Every rename in this project is therefore destructive, +including the corrective rename that fixing `Primarly` would require. Recipe: +TG-18. + +--- + +## 4. Replication settings that carry an unstated limit + +| Setting | Value **[measured]** | What it means in tags | +|---|---|---| +| Fast replication | enabled | Tags travel as registry indices; client and server registries must match exactly | +| Bits for container size | 6 | A hard cap of 63 entries per replicated tag container | + +**[derived]** The first row makes plugin-asymmetric tag loading a correctness +problem rather than a cosmetic one, and PIE cannot reproduce it: one process, one +registry. The second row is a gameplay limit expressed as a bit count, sitting in +a settings file, with nothing connecting it to the containers that will grow. + +**[open]** Whether client and server registries actually match under asymmetric +feature-plugin loading is unanswerable without a dedicated-server build, and is +left open. Recipes: TG-10, TG-11. + +--- + +## 5. Counters: the tag is validated, the amount is not + +A tag-to-count container in the project validates its tag argument at both the add +and the remove site, with a warning on failure. Neither site validates the +*amount*: a guard of the form "if the amount is positive" returns silently +otherwise, and the accumulator has no clamp **[measured]**. An unresolved TODO +next to the remove path asks exactly what should happen when more is removed than +exists **[measured]**. + +**[derived]** This asymmetry is the entire finding, and it is invisible for a +structural reason: seeing the tag validation satisfies the reviewer's question +"are the inputs checked here?". One checked parameter reads as a checked function. + +Worth noting for balance: the same container implements all three replication +callbacks correctly, keeping its array and lookup map consistent **[measured]** — +which is exactly the discipline that a different structure in the same project +fails to apply. Recipes: TG-12, TG-13. + +--- + +## 6. A stub with a second layer under it + +One widget base class exposes an editable tag container for "hide when these tags +are present". Both consumer sites hardcode the check result to a constant, with +the real expression parked in a comment **[measured]**. That alone is an ordinary +stub: the property is authored, read nowhere, and does nothing. + +The second layer is what makes it worth including here. A setter in the same class +overwrites the designer-authored visibility values during construction +**[measured]**. + +**[derived]** So the obvious one-line fix — implement the hardcoded check — +produces *wrong* behaviour rather than correct behaviour, because it would select +between two values that no longer hold what the designer authored. The fix then +gets blamed and reverted, burying the original defect under a failed attempt. + +**A stub can be deeper than its TODO admits.** Before implementing one, trace its +inputs. Recipes: TG-14, TG-15. + +--- + +## 7. Documentation coverage + +68 of 99 project-config tags carry an empty comment field — sixty-nine per cent +undocumented **[measured]**. + +**[derived]** This is why vocabularies only grow. A tag nobody can explain is a tag +nobody can safely delete, so the safe action is always to add another one. Over a +project's life that produces duplicate tags for one concept, which is the +condition that makes hierarchical matching dangerous (TG-08) and renames necessary +(TG-18) — both of which this project is unprotected against. + +One tag's own comment in the project admits that the tag sits in the wrong +namespace and needs to move **[measured]**. With zero redirects declared, that move +cannot be made without breaking every asset that references it. The comment is an +accurate description of a debt that the configuration makes unpayable. + +--- + +## 8. What the reference got right + +1. **A central registry file exists** and is well organised — the problem is its + share, not its quality **[measured]**. +2. **Replication callbacks on the tag-stack container are complete**, keeping the + replicated array and the lookup map consistent in all three **[measured]**. +3. **Tag validation with explicit warnings** at the stack add/remove sites — the + discipline is present, it is simply applied to one parameter of two + **[measured]**. +4. **Constraints exist at all.** Fourteen constrained properties is a low ratio, + and it is fourteen more than most projects have. + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source and +config in a single workspace. Source addresses stay in the research archive that +produced this skill; each `TG-` identifier resolves back to the audited location +there. Numbers without an author are worth less than no numbers, so: this is where +these came from, and they can be produced on request. + +One correction is recorded here rather than hidden: an earlier form of this +material stated that 59 native definitions were spread over "40+ files". The +re-measurement performed while writing this file gave **94 definitions across 31 +files**, and that is the number above. The original figure was never verified; the +recipe in TG-17 is what produced the corrected one. + +## Evidence boundary + +Every claim above refers to one project on one engine version. These are examples +and failure evidence, not guarantees about other engine versions or other samples. +The counts are properties of that project on that day and are quoted only to show +the shape of the contrast a recipe produces — re-run the recipes against your own +tree rather than importing the numbers. diff --git a/plugins/ue-design-skills/skills/ue-gas-architecture/SKILL.md b/plugins/ue-design-skills/skills/ue-gas-architecture/SKILL.md new file mode 100644 index 0000000..465bfb5 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gas-architecture/SKILL.md @@ -0,0 +1,452 @@ +--- +name: ue-gas-architecture +description: >- + Design or review production Gameplay Ability System architecture in Unreal + Engine: ability packages and revocation receipts, buffered ability input on + the ability system component, activation policies and exclusivity groups, + attribute set responsibility boundaries, gameplay effect contexts, tag + relationship policy, match phases as abilities, and the health/death seam. + Use when adding combat, abilities, equipment-granted powers, phases or rounds, + damage and healing, or diagnosing ability activation and lifecycle bugs. +--- + +# UE GAS architecture + +The invariant: + +> GAS owns capabilities and effects; domain components own durable gameplay +> state; data packages compose capabilities; every temporary grant keeps a +> receipt. + +Worked analysis of a production GAS layer, including the parts worth copying +verbatim: [patterns](references/patterns.md). +Detection recipes for silent GAS failures: [failure modes](references/failure-modes.md). + +Related skills: `ue-data-driven-architecture`, `ue-input-architecture`, +`ue-modular-gameplay`, `ue-cosmetics-and-teams`, `ue-multiplayer-authority`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Package grants as ability sets + +Do not scatter calls to grant abilities, apply effects and add attribute sets +across pawns, equipment and modes. Define one immutable capability package: + +```text +AbilitySet +├─ abilities: class + level + optional semantic input tag +├─ gameplay effects: class + level +└─ attribute sets: class +``` + +The same package can be granted by pawn data at spawn, by an equipment +definition while equipped, by a feature action for a mode, by a pickup, or by a +test fixture. + +### Grant order is not cosmetic + +Attributes, then abilities, then effects. Effects in the third group may modify +attributes created by the first; the reverse order silently applies modifiers to +attributes that do not exist yet. + +### Grant only on authority + +The package's grant function must reject non-authority mutation at its first +line. Clients receive replicated specs, effects and attributes; they do not +author grants. + +### Return a grant receipt + +```text +GrantedHandles +├─ AbilitySpecHandles +├─ ActiveGameplayEffectHandles +└─ GrantedAttributeSet pointers +``` + +The revoke function removes exactly those resources. This is what prevents one +equipment item from revoking another owner's grants — and it is the single most +reusable piece of a GAS layer. + +### Decide permanence explicitly + +Passing no receipt means "grant for the lifetime of the ability system +component". That is a legitimate design choice, and it must be a choice rather +than an omission. Review every caller: + +| Caller | Receipt required? | +|---|---| +| pawn baseline from pawn data | optional, permanent by design | +| equipment | yes | +| feature action | yes, per actor and per activation context | +| temporary pickup | yes, or rely on effect duration handles | +| test or cheat command | yes | + +A single reference implementation demonstrates both: the pawn baseline passes +nothing because it is meant to last as long as the player state, and equipment +passes a receipt stored in its own applied-item entry. + +--- + +## 2. Treat input as a buffered concern of the ability system component + +Input events should not call activation directly. Send a semantic input tag to +the ability system component and process once per frame after player input: + +```text +InputAction +→ InputTag pressed/released +→ ASC records spec handles +→ PlayerController.PostProcessInput +→ ASC.ProcessAbilityInput +→ activation policy chooses the activation set +→ TryActivateAbility +``` + +What this buys: + +- simultaneous inputs resolve together instead of in binding order; +- held input is distinguishable from newly pressed input; +- activation can be suppressed globally for a frame; +- activation policy belongs to the ability rather than to each binding site. + +### Required buffers + +Held spec handles, pressed this frame, released this frame. Clear the transient +buffers after processing, and drop stale handles when specs are removed. + +### Order within the frame matters + +Collect held abilities first, then pressed, then activate everything in one pass, +and only then process releases. Activating as you iterate means a held input can +activate an ability and then immediately deliver it a press event it should never +have seen. + +### Add a global input block tag + +When a project-defined "ability input blocked" tag is present on the component: +clear the buffers, activate nothing, and define explicitly whether already-active +held abilities receive a release or a cancel. One tag is cheaper and far more +reviewable than disabling dozens of input actions during UI, death or stun. + +--- + +## 3. Put activation rules on the ability + +### Activation policy — when + +- `OnInputTriggered` — activate once when pressed; +- `WhileInputActive` — activate and retry while held; +- `OnSpawn` — activate when granted and the avatar is ready. + +Do not infer these from the input event type at the binding site. The ability +knows its own lifecycle; the binding does not. + +`OnSpawn` needs two entry points, not one: the ability may be granted to a live +avatar, or the avatar may appear for an already-granted ability. Handle both, and +skip activation when the avatar is being torn down. + +### Activation group — concurrency + +- `Independent` — unaffected by others; +- `Exclusive_Replaceable` — exclusive, but another exclusive may displace it; +- `Exclusive_Blocking` — exclusive and prevents other exclusives from starting. + +The component tracks active counts per group and enforces transitions. This +replaces an N×N web of per-ability blocking tags for broad concurrency rules. + +Rules: + +- an ability may enter a group only if not blocked; +- entering an exclusive group cancels other replaceable exclusives; +- group counts must balance across activate and end; +- an ability in a replaceable group must not be able to refuse cancellation — + that combination is a contradiction and should log rather than obey; +- changing group at runtime is explicit and validated. + +Death is the clearest worked example: on activation it cancels everything not +explicitly marked as surviving death, refuses cancellation, and moves itself into +the blocking group for the duration. + +--- + +## 4. Move tag relationships into a policy asset + +Abilities carry their intrinsic tags. A pawn or archetype selects a relationship +map: + +```text +AbilityTag +├─ tags to block +├─ tags to cancel +├─ activation required tags +└─ activation blocked tags +``` + +This removes repeated policy from N ability defaults and lets the same ability +classes run under different archetype rules. It matters most when abilities +arrive from separate feature plugins that must not know about each other. + +### Relationship validation + +- [ ] the key tag is valid and appropriately scoped; +- [ ] parent-tag matching is intentional, not incidental; +- [ ] block and cancel are not confused with each other; +- [ ] required and blocked sets are not contradictory; +- [ ] cycles are documented and tested; +- [ ] every archetype that needs policy actually selects a map. + +A linear scan is acceptable for a small matrix. Profile before assuming it scales: +in the measured reference the lookup is a linear array walk performed on every +activation and on every block/cancel application, and its authors marked it as +provisional in three separate places. + +--- + +## 5. Split attribute sets by direction of responsibility + +Do not create one character-attributes dumping ground. + +**Receiving set.** Health, max health, the meta attributes damage and healing, +clamping, and the out-of-health signal. Damage and healing are transient inputs +to execution, not durable replicated state: they arrive, convert into a health +change, and reset to zero in the same call. + +**Source set.** Base damage, base heal, outgoing coefficients — the values +captured *from the instigator*. + +Execution calculations capture the source set from the instigator and the +receiving set from the target. The split makes "who contributes this value" +structural rather than a comment. + +### Gates on the receiving set + +- clamp in the correct hook, and in all of them that can change the value; +- guard against duplicate out-of-health events with an explicit flag that resets + when the value recovers; +- record pre-execution values when messages need a before-and-after; +- notify domain components through delegates; +- keep game-flow state — death started, death finished — out of the attribute set + entirely. + +### Mark what modifiers must not touch + +Attributes that may only be changed by an execution should say so in metadata. +This is the difference between "we agreed not to modify health directly" and a +rule the editor enforces. + +--- + +## 6. Extend the gameplay effect context for combat metadata + +When damage needs source or target data beyond the stock context, derive a +project context and implement allocation, script-struct identity, duplication and +network serialization. + +Store only what must survive execution and replication. A context is not a place +to hang a mutable combat object graph. + +Validation: + +- [ ] the context survives duplication with its custom fields intact; +- [ ] custom fields serialize deterministically; +- [ ] object references are network-safe, and fields that are deliberately not + replicated are documented as server-only at their declaration; +- [ ] execution calculations handle absent hit data; +- [ ] script exposure does not permit invalid mutation. + +--- + +## 7. Separate numeric death detection from the death lifecycle + +```text +effect execution changes Health +→ receiving set emits out-of-health +→ health component sends a death gameplay event +→ death ability starts +→ health component enters DeathStarted, then DeathFinished +→ game mode handles respawn +``` + +Responsibilities: GAS owns the *number*; the health component owns the *state* +and its replication; the death ability owns the *sequence*; the game mode owns +the player lifecycle. + +Do not represent death only as "health is zero". Consumers need a monotonic state +they can observe over the network, and a predicted transition needs somewhere to +be corrected. + +### The death trigger must be unambiguous + +Configure it in the ability's class defaults so exactly one path can start it, and +have the ability's end path force the final transition — insurance against a +designer-authored sequence that does not reach the end. + +--- + +## 8. Model match phases as abilities when the lifecycle fits + +A phase ability runs on the ability system component of the game state: + +```text +Game.Playing +├─ Game.Playing.Warmup +├─ Game.Playing.SuddenDeath +└─ Game.PostGame +``` + +Starting a phase grants and activates the ability, registers its tag, cancels +incompatible siblings, and notifies observers. Ending the ability ends the phase. + +Hierarchical tags give the rule for free: **parents and children coexist, siblings +do not.** Starting `Game.Playing.SuddenDeath` while `Game.Playing` is active +leaves the parent running; starting `Game.PostGame` ends the whole subtree. + +### Use this pattern when + +- the phase has a real enter/active/exit lifecycle; +- ability tasks, effects and tags are natural fits for its logic; +- the server owns transitions; +- observers want tag-based matching rather than an enum switch. + +### Do not use it when + +- a replicated enum fully describes the state; +- the phase must never be cancellable; +- clients must query authoritative phase state and no replication path has been + designed for it. This is the trap: if the phase abilities are server-only, a + client-side query of "is this phase active" answers false forever, and it + answers it without an error. + +### Observer handles are mandatory + +Registration returns a handle and supports unsubscribe. An array of callbacks +with no handles grows until the world resets, and every entry pins whatever it +captured. + +--- + +## 9. Failure feedback belongs at the ability system seam + +Activation failure should produce structured tags, not a log line: + +```text +Ability.ActivateFail.Cooldown +Ability.ActivateFail.Cost +Ability.ActivateFail.TagsBlocked +Ability.ActivateFail.IsDead +``` + +For a locally controlled avatar, route them straight to feedback. For a +server-side rejection of a locally predicted action, notify the owning client with +the ability and the failure tags. Map tags to user-facing text and to animations +in data, and deliver them over a message bus so the UI never learns about GAS. + +Never parse log strings in UI. + +### Compose costs rather than accepting one + +A single cost effect is rarely enough. A list of inline cost objects — each with +its own check, its own application, its own failure tag, and an optional +"only charge on hit" rule — turns cost into data. The hit determination should be +computed once and cached, on authority only. + +--- + +## 10. Equipment and feature integration + +Equipment grants ability sets when equipped and retains the receipt in its +applied-item entry. Unequip reverses exactly that receipt. + +A feature action that grants abilities retains, per actor: direct ability handles, +attribute instances, ability-set receipts, and the extension request handle. + +Every integration answers: + +- Who owns the ability system component — the pawn or the player state? +- Which object is the source object for the spec? +- When is avatar info initialized? +- Is the grant permanent or reversible? +- What happens if the actor is destroyed before the feature deactivates? + +See `ue-modular-gameplay` for actor-extension timing and rollback. + +--- + +## 11. Review checklist + +### Ability set + +- [ ] authority-only grant, checked at the top of the function; +- [ ] grant order is attributes, abilities, effects; +- [ ] every class reference non-null and of the expected type; +- [ ] semantic input tag valid and constrained; +- [ ] receipt retained for every temporary grant; +- [ ] revoke removes only owned resources; +- [ ] source object deliberately chosen. + +### Ability system component + +- [ ] input buffers processed once per frame, after player input; +- [ ] blocked-input behaviour explicit for held abilities; +- [ ] activation group counters balance; +- [ ] avatar change activates spawn-policy abilities exactly once; +- [ ] relationship map selected per archetype; +- [ ] failure tags reach local feedback. + +### Attributes and damage + +- [ ] source and target responsibilities separated into different sets; +- [ ] meta attributes are not treated as durable state; +- [ ] clamping occurs in every path that can change the value; +- [ ] team and friendly-fire rules applied in one authoritative calculation; +- [ ] out-of-health cannot fire twice for one death; +- [ ] custom effect context serializes. + +### Phases and death + +- [ ] the server owns transitions; +- [ ] observers can unregister; +- [ ] sibling and parent tag semantics are tested, not assumed; +- [ ] clients have a defined observation path that does not depend on a + server-only subsystem; +- [ ] death start and finish are durable states, not health comparisons. + +--- + +## 12. Decision summary + +Use GAS for a concern when at least one is true: it is a grantable and revocable +capability; prediction or networked activation matters; it is naturally expressed +through effects, tags, cost and cooldown; it benefits from task-based lifecycle +with cancel and end. + +Keep it out of GAS when the concern is durable domain state better served by a +replicated component, presentation only, static definition composition, or a +global lifecycle with no ability semantics. + +A production GAS architecture is not "everything is an ability". It is a clean +seam between capability lifecycle, numeric effects, domain state and data-driven +composition. + +--- + +## Provenance + +The patterns and failure modes in `references/` come from a line-by-line audit of +the GAS layer of Epic's Lyra Starter Game on Unreal Engine 5.6 — roughly 4 900 +lines across 51 files — read as source rather than run. + +Source addresses stay in the research archive that produced this skill; what ships +is the detection recipe. Each entry carries a stable identifier (`GA-01` and up) +that resolves back to the audited location, so any specific claim can be produced +on request. + +## Evidence boundary + +One project, one engine version, one workspace. The architecture described here +is transferable; the specific defects are evidence, not guarantees about other +versions. Re-run the detection recipes against your own tree before acting on any +specific claim. diff --git a/plugins/ue-design-skills/skills/ue-gas-architecture/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-gas-architecture/references/failure-modes.md new file mode 100644 index 0000000..efd0e96 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gas-architecture/references/failure-modes.md @@ -0,0 +1,502 @@ +# Failure modes: GAS architecture + +Fourteen ways a Gameplay Ability System layer misbehaves without reporting +anything. + +The shared property, stated once: **GAS failures present as absence.** An ability +that does not activate, a phase query that answers false, a grant that is never +revoked, a cost that is never charged. The framework is built to tolerate a +refused activation, because a refused activation is a normal gameplay state. So +the diagnostic burden is entirely yours, and the useful question is never "did it +error" but "which side of the seam owns this fact". + +Each entry gives the mechanism, why it stays silent, why the obvious check misses +it, the symptom a human reports, a `Detect` recipe, and the guardrail. +Identifiers (`GA-01` and up) are stable and resolve back to the audited source in +the research archive. + +Recipes use `rg` and run from a project's source root. They were executed against +the audited project while this file was written. + +--- + +## Ownership and receipts + +### GA-01 - Grant with no receipt, permanent by omission + +**Mechanism.** An ability set is granted with the out-parameter for handles left +null. The grant then lasts as long as the ability system component. + +**Why it is silent.** Permanent is a legitimate lifetime, and the API is designed +to offer it. Nothing distinguishes "permanent because we decided" from "permanent +because the argument was easy to omit". + +**Why the obvious check misses it.** The call is correct, the API is used as +documented, and the reviewer sees a grant that works. The defect only exists +relative to an intention that lives in nobody's head at review time. + +**Symptom.** A capability that should have been removed with the feature, the +equipment or the round survives it. Usually noticed after a respawn, when the +player has two of something. + +**Detect.** List every grant site and classify it by whether it keeps a receipt: + +```bash +rg -n -B3 "GiveToAbilitySystem" --glob "*.cpp" . +``` + +Any call passing `nullptr` for the handles parameter is a permanent grant. In the +audited project this is exactly one site — the pawn baseline on the player state — +and the equipment path beside it passes a receipt stored in its applied-item +entry. One permanent grant with a stated reason is a design; several are a leak. + +**Guardrail.** Make permanence explicit at the call site, in a comment naming the +lifetime the grant is tied to. Review new grant sites against the table of +allowed permanent grants rather than against the API signature. + +--- + +### GA-02 - Revocation that removes more than it granted + +**Mechanism.** Teardown clears abilities by class, by tag, or by clearing all +abilities on the component, instead of by the handles the grant returned. + +**Why it is silent.** The abilities the caller wanted removed *are* removed. The +extra removals affect capabilities another owner granted, and that owner is not +watching. + +**Why the obvious check misses it.** The teardown function looks thorough. "Remove +everything matching this class" reads as more robust than "remove these three +handles", not less. + +**Symptom.** Unequipping one item removes an ability granted by a different +feature. Reproduces only when two owners granted overlapping capabilities, which +is rare in a test and common in a shipped game. + +**Detect.** Find removals that are not keyed on stored handles: + +```bash +rg -n "ClearAbility\(|ClearAllAbilities|RemoveActiveGameplayEffect" --glob "*.cpp" . -B5 \ + | rg -n "(GetAssetTags|StaticClass|ForEach|\.Class)" +``` + +Removal keyed on anything other than a handle produced by the matching grant is +the finding. + +**Guardrail.** Every revoke path takes a receipt and removes exactly its contents. +If a receipt was not kept, the correct fix is to keep one, not to widen the +removal. + +--- + +### GA-03 - Grant order that applies effects before their attributes exist + +**Mechanism.** An ability set grants effects before it grants attribute sets, so +an effect modifies an attribute that has not been added yet. + +**Why it is silent.** Applying a modifier to a missing attribute is not an error +in the framework — the modifier simply finds nothing to modify. The effect is +applied, the handle is valid, the log is clean. + +**Why the obvious check misses it.** Each block of the grant function is correct +in isolation, and the whole function reads as a straightforward three-part loop. +Order dependencies between the parts are invisible unless you know the effects +reference the attributes. + +**Symptom.** An attribute sits at its default value despite an effect that should +have initialized it. Investigated as a bad effect asset. + +**Detect.** Read the grant function and confirm the sequence: + +```bash +rg -n -A40 "::GiveToAbilitySystem" --glob "*.cpp" . \ + | rg -n "(AddAttributeSetSubobject|GiveAbility|ApplyGameplayEffect)" +``` + +The line numbers must run attributes, then abilities, then effects. The audited +implementation gets this right and the ordering is worth copying verbatim. + +**Guardrail.** State the order as a comment in the grant function, and add a test +that grants a set whose effect initializes an attribute the same set provides. + +--- + +## State that lies + +### GA-04 - Server-only subsystem queried from clients + +**Mechanism.** Phase abilities are configured server-only, and the subsystem that +tracks them is created on every client anyway. Its map of active phases is +therefore always empty on a client, and the "is this phase active" query always +returns false there. + +**Why it is silent.** False is a valid answer to that question. The client is not +told that it asked a question it cannot answer; it is told "no". + +**Why the obvious check misses it.** The subsystem exists on the client, is +correctly initialized, and answers immediately. Every structural check passes. +The creation gate that would have prevented it exists too — commented out, with +the unconditional `return true` left in place. Reading that function tells you +someone thought about it, not what they concluded. + +**Symptom.** Client-side UI that reacts to match phase never reacts. Blamed on +replication, on the widget, on the tag — rarely on the query itself. + +**Detect.** Find subsystem creation gates that were considered and abandoned: + +```bash +rg -n -A12 "ShouldCreateSubsystem" --glob "*.cpp" . | rg -n "^\s*//.*(return|check|World)" +``` + +Then, for every server-authoritative subsystem, check whether any client code +path queries its state: + +```bash +rg -n "IsPhaseActive|IsRoundActive|GetCurrentPhase" --glob "*.cpp" . +``` + +**Guardrail.** A subsystem whose state is authority-only either refuses to be +created without authority, or its queries return a tri-state that distinguishes +"no" from "I cannot know". Clients observe phase through replicated tags, +replicated state or messages — never through the authoritative tracker. + +--- + +### GA-05 - Observer registration with no handle + +**Mechanism.** A subscription API appends a callback to an array and returns +nothing. There is no way to unsubscribe. + +**Why it is silent.** Subscriptions work. The list grows, and a growing list has +no symptom until it is large or until a captured object needed to die. + +**Why the obvious check misses it.** The API is complete from the caller's +perspective: subscribe, get called back. The missing half is a function that does +not exist, and reviews do not notice absent functions. + +**Symptom.** Callbacks firing on objects that logically ended, and a list that +grows until the world resets. In the audited project the authors documented this +themselves, in two consecutive comments above the function, and shipped it — +including the honest second thought that a handle would not help if callers +ignored it. + +**Detect.** Find subscription functions that return void: + +```bash +rg -n "void \w+::When\w+|void \w+::(Register|Subscribe|Observe)\w*\(" --glob "*.cpp" . +``` + +Any registration whose return type is void is a subscription that cannot be +undone. + +**Guardrail.** Registration returns a handle; the handle unregisters. If callers +cannot be trusted to hold handles, bind weakly *and* sweep expired entries — but +the sweep is an addition to the handle, not a substitute for it. + +--- + +### GA-06 - Write-only field + +**Mechanism.** A member is declared and assigned in the constructor, and read +nowhere. + +**Why it is silent.** It is a correctly initialized member of a working class. It +costs nothing at runtime and breaks nothing. + +**Why the obvious check misses it.** Searching for the name finds two hits — a +declaration and an assignment — which reads as "declared and used". An +unused-variable warning does not fire, because the write *is* a use. + +**Symptom.** No runtime symptom. The cost is comprehension: the field's name +promises behaviour, and a future engineer will implement against a flag that +nothing consumes. In the audited project the field is a cancellation-logging flag +whose own comment describes it as temporary, added while tracking a bug that has +since closed. + +**Detect.** Compare writes against reads for suspicious members: + +```bash +F='bLogCancelation' +rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" . +``` + +Two hits — declaration and assignment — with no comparison, no branch and no +pass-by-value is the finding. This is the mirror image of the missing-writer +check: here the writer exists and the reader does not. + +**Guardrail.** Delete it. A flag with no consumer is a comment that pretends to be +code, and it will eventually be believed. + +--- + +## Performance and correctness on the activation path + +### GA-07 - Function-static containers on a per-activation path + +**Mechanism.** A function called during every ability activation declares +function-local `static` containers as scratch space to avoid reallocation. + +**Why it is silent.** It is correct on a single thread, and every activation +happens on the game thread today. The reuse is invisible because each call clears +before use. + +**Why the obvious check misses it.** The pattern looks like a deliberate +optimization, and it is one. Nothing at the call site says "this function is not +reentrant", and reentrancy through an ability that activates another ability is +plausible but not obvious. + +**Symptom.** Nothing, until the function is called from a second thread or +reentrantly — at which point tag requirements are computed from another call's +scratch data and an ability activates when it should not. + +**Detect.** Find statics in functions on the activation path: + +```bash +rg -n "static F\w+Container|static TArray" --glob "*.cpp" . -B10 \ + | rg -n "(CanActivate|SatisfyTagRequirements|ActivateAbility)" +``` + +The audited project has three such statics in one activation-path function. + +**Guardrail.** Use a member scratch buffer on an instanced object, or accept the +allocation. If the static stays, document the single-thread and +non-reentrancy assumption at the declaration, where the next reader will see it. + +--- + +### GA-08 - Linear scan marked provisional, on the hot path + +**Mechanism.** A relationship or lookup table is walked linearly on every +activation, with a comment acknowledging that a real index would be needed at +scale. + +**Why it is silent.** It is correct at every size. Only the cost changes, and cost +does not report itself. + +**Why the obvious check misses it.** The comment is the check, and it is +addressed to a future that has not arrived. Reviewers read "for now" as "someone +is tracking this". + +**Symptom.** Activation cost grows with the size of a designer-authored data +asset, discovered during a late-project profiling pass when the table has grown +by two orders of magnitude. + +**Detect.** Find provisional comments on lookup paths: + +```bash +rg -n "Simple iteration|for now|O\(n\)|linear" --glob "*.cpp" . -A4 \ + | rg -n "(for\s*\(|ForEach)" +``` + +In the audited project the same provisional comment appears three times in one +mapping class, and the mapping is consulted on every activation and on every +block-and-cancel application. + +**Guardrail.** Record the size at which the structure must change, and add an +assertion or a test that fails when the data crosses it. "For now" without a +number never ends. + +--- + +## Attributes and damage + +### GA-09 - Meta attribute treated as durable state + +**Mechanism.** An incoming-damage or incoming-healing attribute — designed as a +one-shot channel that converts to a health change and resets to zero — is read, +replicated or stored as if it held a value between executions. + +**Why it is silent.** The attribute exists and always has a value. Reading it +outside an execution returns zero, which is a plausible number. + +**Why the obvious check misses it.** Nothing in the type distinguishes a meta +attribute from a state attribute. The distinction lives in a comment and in the +post-execution code that zeroes it. + +**Symptom.** UI or logic that reads "current damage" and always sees zero, or a +replicated attribute that costs bandwidth and carries nothing. + +**Detect.** Confirm each meta attribute is zeroed after conversion and excluded +from replication: + +```bash +rg -n "SetDamage\(0|SetHealing\(0" --glob "*.cpp" . +rg -n "DOREPLIFETIME.*\b(Damage|Healing)\b" --glob "*.cpp" . +``` + +A meta attribute that appears in the replication list, or one that is never +zeroed, is the finding. + +**Guardrail.** Group meta attributes under an explicit comment banner in the +header, keep them out of the replication list, and mark state attributes that +executions alone may change so the editor enforces it rather than the team +remembering it. + +--- + +### GA-10 - Out-of-health fired more than once for one death + +**Mechanism.** The zero-health signal is emitted from a change handler that can +run several times as clamping, max-health changes and replication settle. + +**Why it is silent.** Each individual emission is correct. Listeners are expected +to handle a broadcast; nothing says "at most once per life". + +**Why the obvious check misses it.** The emitting code is a single line in a +single place. The multiplicity comes from how many paths reach it, which is +visible only by enumerating callers of the enclosing handler. + +**Symptom.** Two death sequences, doubled elimination messages, score credited +twice. Often first noticed in the kill feed rather than in gameplay. + +**Detect.** Find the emission and check for a latch: + +```bash +rg -n -B12 "OnOutOfHealth.Broadcast" --glob "*.cpp" . | rg -n "bOutOfHealth|bAlready|if\s*\(!" +``` + +An emission with no guarding flag, or a flag that is never reset when health +recovers, is the finding. The audited implementation has both the latch and the +reset — copy that shape. + +**Guardrail.** Latch the signal, reset the latch when the value recovers, and +assert in development that the death transition runs once per life. + +--- + +### GA-11 - Placeholder attribute initialization + +**Mechanism.** Attributes are initialized by assigning one to another in code — +health set to max health — with a comment saying a data-driven initialization +will replace it. + +**Why it is silent.** The result is a live, plausible character. Nothing about a +full-health spawn looks provisional. + +**Why the obvious check misses it.** The data-driven mechanism usually *exists* +alongside it — an initialization-data field on the grant structure, a curve table +setting — so a reviewer looking for "is initialization data-driven?" finds the +machinery and stops. + +**Symptom.** Designer-authored initialization values have no effect, because the +code path that would consume them is bypassed by the placeholder. + +**Detect.** Find initialization marked as temporary, then check whether the +data-driven path has any consumer: + +```bash +rg -n "TEMP|placeholder|Eventually this will" --glob "*.cpp" . -A3 \ + | rg -n "(SetNumericAttributeBase|InitFromMetaDataTable|InitStats)" +rg -n "GlobalCurveTableName" Config/ +``` + +In the audited project the placeholder is present, the global curve table is set +to none, and the initialization-data field on the grant structure has no traced +consumer. + +**Guardrail.** A placeholder initialization is acceptable only with the +data-driven path deleted or disabled, so nobody can author data that silently +does nothing. + +--- + +## Lifecycle + +### GA-12 - Activation group counters that do not balance + +**Mechanism.** A per-group active count is incremented on activation and +decremented on end. Any path that ends an ability without notifying the component +leaks a count. + +**Why it is silent.** A leaked count does not error. It makes the group look +occupied, so later activations are refused — which is indistinguishable from +correct exclusivity. + +**Why the obvious check misses it.** Both the increment and the decrement exist +and are correctly paired in the normal path. The leak comes from an abnormal end: +a destroyed avatar, a cancelled ability that refuses cancellation, an early return. + +**Symptom.** After some sequence involving death or a cancelled ability, exclusive +abilities stop activating for that pawn. Restarting the match fixes it, which +makes it look like corruption rather than arithmetic. + +**Detect.** Find increments and decrements and confirm they are in paired +notification hooks: + +```bash +rg -n "ActivationGroupCounts" --glob "*.cpp" . -B4 +``` + +Any mutation outside the activated/ended notification pair is a leak candidate. +Then assert the invariant directly: the exclusive counts must never exceed one in +total. + +**Guardrail.** Assert the invariant in development builds at every mutation, and +test the abnormal ends explicitly — avatar destroyed mid-ability, cancel refused, +ability ended from a task callback. + +--- + +### GA-13 - Death represented only as a health comparison + +**Mechanism.** Consumers ask "is health at or below zero" instead of reading a +durable death state. + +**Why it is silent.** The comparison is true at the right moments. It is also true +during the window before the death sequence starts, and true again if health is +restored and re-zeroed within one sequence. + +**Why the obvious check misses it.** The comparison is simple, local and obviously +correct. The state machine it replaces lives in another component, and using it +requires knowing it exists. + +**Symptom.** Systems that disagree about whether a character is dead, especially +across the network, and a predicted death that cannot be corrected because there +is no state to roll back. + +**Detect.** Find health comparisons used as death tests: + +```bash +rg -n "GetHealth\(\)\s*<=?\s*0|Health\s*<=\s*0\.?0?f?" --glob "*.cpp" . +``` + +Every hit outside the attribute set itself should be reading a death state +instead. + +**Guardrail.** One replicated, monotonic death state with an explicit +start and finish, owned by a domain component, with the ability system supplying +only the numeric signal that triggers it. + +--- + +### GA-14 - Failure reason discarded at the seam + +**Mechanism.** Activation fails, the framework records a reason, and the project +does not forward it — or forwards it only on the server, where no UI exists. + +**Why it is silent.** A refused activation is normal. The player sees nothing, +which is exactly what a refused activation looks like when the reason is missing. + +**Why the obvious check misses it.** The failure path exists and is exercised +constantly. What is absent is the delivery of the reason to the machine that can +render it, and absence of delivery has no call site to review. + +**Symptom.** The player presses a button and nothing happens, with no feedback +distinguishing "on cooldown" from "out of ammo" from "you are dead". Designers +compensate with generic click sounds. + +**Detect.** Check that failure tags are both produced and delivered to the owning +client: + +```bash +rg -n "OptionalRelevantTags|NotifyAbilityFailed|ActivateFail" --glob "*.cpp" . -A6 \ + | rg -n "(Client|Broadcast|Message)" +``` + +Failure tags produced with no client-notification path is the finding. The audited +project does this correctly: a client notification carries the tags, the ability +maps them to text and animation in data, and delivery goes over a message bus so +the UI never references GAS. + +**Guardrail.** Every failure reason reaches the locally controlled player as a +tag, never as a log line, and the mapping from tag to feedback lives in data. diff --git a/plugins/ue-design-skills/skills/ue-gas-architecture/references/patterns.md b/plugins/ue-design-skills/skills/ue-gas-architecture/references/patterns.md new file mode 100644 index 0000000..44c081d --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-gas-architecture/references/patterns.md @@ -0,0 +1,340 @@ +# Patterns: a production GAS layer read end to end + +A worked reading of one real Gameplay Ability System layer — roughly 4 900 lines +across 51 files — audited as source. + +This file is different in balance from the other reference documents in this +bundle. Most of them are about what a codebase got wrong. **This layer is mostly +worth copying**, and the useful output is a list of the specific decisions that +earn their keep, each with the reason it was made. The defects are in +[failure-modes.md](failure-modes.md) and are fewer than the patterns. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from measured +facts; **[open]** — not answerable from source, and left open. + +--- + +## 1. The shape of the layer + +Two ability system components, on different owners, for different reasons +**[measured]**: + +- one on the **player state**, with the pawn as avatar — so capabilities survive + the pawn's death; +- one on the **game state**, with itself as avatar — which is what makes match + phases possible at all (§5). + +The only global hooks are three config lines **[measured]**: a replacement +globals class, a cue manager class, and an explicitly empty global curve table. +**[derived]** That is a remarkably small integration surface for a layer this +size, and the reason is that everything else is composed from data assets rather +than registered globally. + +--- + +## 2. The ability set, and why the receipt is the whole idea + +An ability set is a primary data asset holding three arrays: abilities, effects, +attribute sets **[measured]**. Grants happen in the order **attributes, +abilities, effects** **[measured]** — not cosmetic, since effects in the third +group may modify attributes created by the first. + +Two details make it reusable rather than merely tidy: + +**The input tag lives in the grant data, not in the ability** **[measured]**. The +same ability class can be bound to different semantic inputs by different sets, +and the ability itself hides the stock input category entirely **[measured]**. + +**The revocation receipt.** The grant function takes an optional out-parameter — +three parallel arrays of handles — and the revoke function removes exactly those +**[measured]**: + +```text +GrantedHandles +├─ AbilitySpecHandles -> ClearAbility +├─ GameplayEffectHandles -> RemoveActiveGameplayEffect +└─ GrantedAttributeSets -> RemoveSpawnedAttribute +``` + +**[derived]** This is a cloakroom ticket. Whoever granted holds the ticket and can +undo exactly their own grant without touching anyone else's. It is the answer to +the question that breaks most modular gameplay implementations — *who removes +this, and how do they know it was theirs?* + +Three consumers use it, and their differences are the design **[measured]**: + +| Consumer | Receipt | Why | +|---|---|---| +| equipment | stored in the applied-item entry | must reverse on unequip | +| feature action | array per actor | must reverse on deactivation | +| pawn baseline on player state | **passes null** | intended to last as long as the player state | + +**[derived]** The optional out-parameter is therefore an API that encodes a +lifetime decision. Passing null means "permanent, deliberately". One such call in +a project is a design; several are GA-01. + +Authority is gated at the first line of both grant and revoke **[measured]**, and +invalid entries log with the asset name and index rather than crashing or +silently skipping **[measured]** — a small thing that turns a content error into a +searchable message. + +--- + +## 3. Input as a buffered concern + +Enhanced input does not call activation. It sends a semantic tag, the component +records spec handles, and one call per frame from the controller's post-process +input step does the work **[measured]**. + +The ordering inside that call is the part to copy, and the reason is in the +authors' own comment **[measured]**: collect held abilities with the +while-input-active policy, then pressed abilities with the on-triggered policy, +**then activate everything in one pass**, and only then process releases. Activate +as you iterate and a held input first activates an ability and then delivers it a +press event it should never have seen. + +Two more decisions worth taking: + +- **A single global block tag** clears every buffer and returns early + **[measured]**. Suppressing all ability input during UI, death or stun is one + tag rather than dozens of disabled input actions. +- **Input events are replicated as events rather than directly** **[measured]**, + with comments explaining that the wait-for-input ability tasks depend on it. A + subtle choice that is easy to reverse by accident when copying. + +--- + +## 4. Activation policy and exclusivity + +Two small enums replace two large problems. + +**Policy — when to activate** **[measured]**: on input triggered, while input +active, on spawn. **[derived]** Automatic fire, single fire and passives become one +enum in class defaults rather than three different implementations. + +The spawn policy needs two entry points, not one — the ability granted to a live +avatar, and the avatar appearing for an already-granted ability — and the audited +implementation wires both **[measured]**, while skipping activation when the +avatar is being torn down **[measured]**. + +**Group — how to relate to others** **[measured]**: independent, exclusive +replaceable, exclusive blocking. The component keeps per-group counts, cancels +other replaceable exclusives when an exclusive starts, and asserts that no more +than one exclusive is ever active **[measured]**. + +**[derived]** Without this you express "firing cancels reloading, but the finisher +cancels nothing" as an N×N matrix of blocking tags on every ability. With it, one +enum plus one behaviour tag. + +One invariant is enforced rather than documented: an ability in the replaceable +group **cannot** refuse cancellation — the call logs an error and does nothing +**[measured]**. **[derived]** That combination is a contradiction, and catching it +in code rather than in review is the difference between a rule and a convention. + +The death ability is the worked example **[measured]**: on activation it cancels +everything not tagged as surviving death, refuses cancellation, and *moves itself* +into the blocking group for the duration. + +--- + +## 5. Match phases as abilities + +The highest return per line in the layer: about 330 lines of C++ for a +hierarchical match state machine **[measured]**. + +A phase is an ability running on the game state's ability system component. While +the ability is active, the phase is active. The ability's defaults differ from the +base in exactly the two ways that matter **[measured]**: server-initiated +execution and server-only security. A client cannot start a phase. + +**Hierarchy comes from tag nesting, and that is the entire rule** **[measured]**. +On starting a phase, every active phase whose tag does not match the incoming tag +is cancelled: + +| Active | Starting | Result | +|---|---|---| +| `Game.Playing` | `Game.Playing.SuddenDeath` | parent stays, child starts | +| `Game.Playing`, `Game.Playing.SuddenDeath` | `Game.Playing.Overtime` | sibling ends, parent stays | +| `Game.Playing.*` | `Game.PostGame` | whole subtree ends | +| `Game.Playing` | second ability, same tag | both run | + +**Parents and children coexist; siblings do not.** + +What the pattern gets for free, because a phase is an ordinary ability +**[derived]**: ability tasks as phase timers and triggers, effects applied for the +phase's duration, cancellation by the same rules as anything else, and all of the +per-mode logic authored in data rather than in new C++. + +One observer detail is worth copying deliberately: subscribing to a phase that is +**already active** fires the callback immediately **[measured]**. That closes the +classic race where a system subscribes after the phase started and waits forever. +The end-observer path has no equivalent **[measured]**, which is consistent — +there is no "already ended" to catch up on. + +### And two limitations the authors recorded themselves + +- **Observers return no handle** **[measured]**, with two consecutive comments + above the function saying so — including the honest second thought that a handle + would not help if callers ignored it. GA-05. +- **The subsystem's creation gate is commented out** **[measured]**, leaving an + unconditional `return true`. So the subsystem exists on clients, where the + active-phase map is always empty and the phase query always answers false. + **[derived]** Client code must learn about phases through replicated tags, + effects or messages — never through this subsystem. GA-04. + +--- + +## 6. Attribute sets split by direction + +Three classes on a clean axis **[measured]**: + +- a base providing accessor macros and a six-parameter change delegate whose + comment honestly warns that some parameters are null on clients; +- a **receiving** set: health, max health, and the meta attributes damage and + healing; +- a **source** set: base damage, base heal. + +The executions make the split structural rather than nominal **[measured]**: they +capture the source set from the instigator and write into the receiving set on the +target. **[derived]** The same actor usually owns both, but their roles in a +calculation are different, and that difference lives in the class layout instead +of a comment. + +### The meta-attribute pattern + +Incoming damage arrives in a non-replicated attribute, converts to a health change +in post-execution, and resets to zero in the same call **[measured]**. + +**[derived]** Damage and healing become one-shot channels rather than state. They +need no replication, no manual reset, and no "current damage" that a reader might +mistake for something durable. Health and damage are additionally marked as hidden +from modifiers **[measured]** — so only an execution can change them, enforced by +the editor rather than by agreement. + +### Three interception levels, each doing one job + +**[measured]**: a pre-execution veto that can cancel the whole effect (where +damage immunity and a development-only god mode live, both respecting a +self-destruct exception so suicide still works); a post-execution step that +converts meta attributes, broadcasts change delegates and publishes a damage +message; and clamping applied in both change hooks with a shared helper. + +A latch flag prevents a duplicate out-of-health broadcast and resets when health +recovers **[measured]** — the correct shape for GA-10. + +### Where the damage number comes from + +Four multiplications **[measured]**: base damage, distance attenuation, physical +material attenuation, and a team multiplier that is **0 or 1 rather than a +branch**. The two attenuation curves come from an interface implemented by the +damage *source* — the weapon — rather than from the execution. **[derived]** Falloff +therefore lives in weapon data where a designer owns it, and the execution stays +generic. + +--- + +## 7. The death seam: three responsibilities, three classes + +```text +effect execution changes Health +→ receiving set broadcasts out-of-health +→ health component sends a death gameplay event +→ death ability runs the sequence +→ health component enters DeathStarted, then DeathFinished +``` + +**[measured]** throughout. The division is the point: + +- **GAS owns the number** — how much health, and the signal that it reached zero; +- **the health component owns the state** — a replicated enum with its own + replication callback that can roll back a predicted transition and logs invalid + ones; +- **the death ability owns the sequence** — what plays, when control returns. + +**[derived]** None of the three knows the others' internals; they communicate +through delegates and a gameplay event. + +Three details worth taking: + +- the death trigger is configured in the ability's **class defaults** + **[measured]**, so exactly one path can start it; +- the ability's end path **always** calls the finish transition **[measured]** — + insurance against a designer-authored sequence that does not reach the end; +- death state is neither an attribute nor a tag but a replicated enum + **[measured]**; the tags exist too, set loosely, because the state is already + replicated by other means. + +Alongside, a damage message and an elimination message go out over a message bus +**[measured]**, so kill feeds, statistics and accolades never reference GAS. + +--- + +## 8. Two more things worth copying + +**Composable costs.** The stock system offers one cost effect. Here, an ability +holds a list of inline cost objects, each with its own check, its own application, +its own failure tag, and an optional "only charge on hit" rule whose hit +determination is computed once, cached, and only on authority **[measured]**. +**[derived]** Ammunition, inventory items and player resources become three small +classes rather than three special cases. + +**Failure reasons that reach the player.** The stock system swallows why an +activation failed. Here the component notifies the owning client when the avatar +is not locally controlled, the ability maps failure tags to user-facing text and +to animations through two data maps, and delivery goes over a message bus +**[measured]**. **[derived]** "Out of ammo" on the HUD costs no bespoke RPC, and +the UI never learns that GAS exists. GA-14. + +--- + +## 9. Copy carefully + +Two places where the audited implementation is explicitly provisional, and says so: + +- **The relationship mapping is a linear array walk**, consulted on every + activation and on every block-and-cancel application, with the same "simple + iteration for now" comment in three separate functions **[measured]**. + **[derived]** Fine at the size measured; profile before assuming it scales. + GA-08. +- **Three function-static containers** are used as scratch space in a function on + the activation path **[measured]**. **[derived]** Correct on one thread, not + reentrant, and nothing at the call site says so. GA-07. + +And one thing that reads as finished and is not: attribute initialization is a +placeholder that assigns health from max health in code, with a comment saying a +data-driven source will replace it **[measured]**. The global curve table is +configured to none **[measured]**, and the initialization-data field on the grant +structure has no traced consumer **[open]**. GA-11. + +--- + +## 10. What source reading could not settle + +- **Binary assets were not read.** The concrete phase abilities, the ability sets, + the contents of the relationship mapping for each archetype, and the activation + groups configured on real abilities are all **[open]**. +- **Who starts the first phase is unknown** **[open]**. A search for the start + call finds only the subsystem itself, so the caller is in a scripted asset. +- **One field could not be classified from this layer alone**: a + cancellation-logging flag is declared and assigned in the constructor and read + nowhere in the audited directories **[measured]**. Whether something outside them + reads it is **[open]** — though its own comment describes it as temporary, + added while tracking a bug. GA-06. + +--- + +## Provenance + +Measured against the GAS layer of Epic's Lyra Starter Game on Unreal Engine 5.6, +read as source in a single workspace. Source addresses stay in the research +archive that produced this skill; each `GA-` identifier resolves back to the +audited location there, so any specific claim above can be produced on request. + +## Evidence boundary + +One project, one engine version, one workspace. Binary assets were not read, so +every statement about authored content is marked open rather than concluded. The +architecture is transferable; the specific defects are evidence, not guarantees +about other versions. Re-run the recipes in +[failure-modes.md](failure-modes.md) against your own tree before acting on +anything here. diff --git a/plugins/ue-design-skills/skills/ue-input-architecture/SKILL.md b/plugins/ue-design-skills/skills/ue-input-architecture/SKILL.md new file mode 100644 index 0000000..2110007 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-input-architecture/SKILL.md @@ -0,0 +1,387 @@ +--- +name: ue-input-architecture +description: >- + Design or review Enhanced Input architecture in Unreal Engine: physical + mapping contexts, input actions, semantic gameplay tags, input config data, + native versus ability input paths, per-frame buffering on the ability system + component, player-mappable settings, device switching, and runtime injection + and removal of input by feature plugins. Use when adding controls, rebinding, + gamepad or touch support, ability input, mode-specific input, or debugging + actions that bind correctly and never fire. +--- + +# UE input architecture + +The invariant: + +> Physical controls, semantic commands and gameplay implementations are three +> separate layers joined by validated data. + +```text +key / gamepad / touch +→ input mapping context +→ input action +→ input config: input action → input tag +→ native handler OR ability system component +→ ability set: input tag → ability +``` + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes for silent input failures: [failure modes](references/failure-modes.md). + +**The characteristic failure of this architecture is a button that does nothing +and logs nothing.** Read the failure modes before adopting it — most of them are +variations on that one observable. + +Related skills: `ue-gameplay-tag-governance`, `ue-gas-architecture`, +`ue-modular-gameplay`, `ue-ui-architecture`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Keep the three layers independent + +### Physical — the mapping context + +Owns key, button and axis to input action; modifiers and triggers; priority +relative to other contexts; player-mappable metadata. + +Changing a keyboard layout or a gamepad binding should stop here. + +### Semantic — the input config + +An immutable asset owning pairs: + +```text +IA_Move → InputTag.Move +IA_WeaponFire → InputTag.Weapon.Fire +IA_Dash → InputTag.Ability.Dash +``` + +Gameplay depends on tags, not on asset paths or physical keys. + +### Capability — handlers and ability sets + +Native tags bind to explicit handlers. Ability tags travel to the ability system +component as payload. Granted ability specs carry the matching semantic tag. + +Changing an ability implementation should not require touching a mapping context. + +--- + +## 2. Split native and ability actions + +Make the distinction explicit in the config, with two separate lists. + +**Native** — deterministic controller and pawn functions: move, look, crouch, +autorun. Each costs one explicit binding line, and the tag is a **lookup key at +bind time only**. The handler receives an input value; the tag does not exist at +runtime. + +**Ability** — bound through one pressed/released pair with the row's tag passed +as a payload argument. One pair of functions serves any number of abilities, and +the count of new abilities you can add without writing code is unbounded. + +### Classification test + +Native when the action controls continuous motion, must execute immediately and +deterministically, and needs no ability lifecycle. + +Ability when the action activates a grantable capability, can be blocked or +cancelled by tags, has cost or cooldown or prediction, or changes with equipment +or mode. + +### Watch what the ability path fixes at bind time + +If ability binding is hardcoded to two trigger events, then hold, tap and +double-tap variants must be expressed through triggers and modifiers in the +mapping context rather than through trigger-event selection. That is a +reasonable constraint — but it is a constraint, and it is invisible until +someone needs a third event. + +--- + +## 3. Join semantic tags at grant time + +When an ability set grants an ability, add the configured input tag to the +spec's dynamic source tags. Lookup then becomes: + +```text +pressed input tag +→ scan activatable specs +→ exact tag match +→ buffer the spec handle +``` + +Do not store an input action pointer on the ability. That couples a capability to +one content asset and blocks per-mode and per-device remapping. + +### The match is exact, not hierarchical + +A spec tagged `InputTag.Weapon` is **not** matched by a press of +`InputTag.Weapon.Fire`. The tag hierarchy organises the vocabulary and filters +the editor picker; it does not participate in lookup. Assuming otherwise +produces a binding that is correct in every visible respect and never fires. +Recipe: IN-01. + +### Validate the join + +The input config and the ability set are independent assets. Their tag equality +is a foreign key with no database behind it, so validation is the only compiler +it will ever have: + +- every input-driven ability has a non-empty tag; +- every intended ability tag is reachable from at least one active config; +- every config ability tag has a granted consumer in the intended composition; +- duplicates are intentional; +- category constraints hold on both sides. + +--- + +## 4. Buffer ability input, process once per frame + +Forwarding a tag must not activate immediately. Record held handles, pressed this +frame, released this frame; process from the controller's post-input step in +three passes — held with the while-active policy, then pressed with the +on-triggered policy, then activate everything in one batch, then releases. + +This removes event-order dependence and makes simultaneous input resolvable. +See `ue-gas-architecture` for the activation policies and the global block tag. + +**Accept the consequence:** the ability path is one input-processing step behind +the native path within the same frame. For most games that is invisible. It is +still a real difference between the two paths, and it belongs in the +classification decision rather than being discovered later. + +--- + +## 5. Initialize input at a named readiness boundary + +Binding depends on several independently arriving pieces: a local controller, the +input subsystem, the pawn data and its config, the project's input component +subclass, and any contexts added by features. + +Do not treat `BeginPlay` as ready. Use init states and emit a semantic event — +"bind inputs now" — once the dependencies hold. + +Feature actions must respond to **both** the generic extension-added event and +the semantic ready event, because either can arrive first. Set the readiness flag +*before* broadcasting, or a handler that checks it during the broadcast will see +false. + +### Base initialization checklist + +- [ ] the component is the expected input subclass, checked with a diagnostic + that names the fix; +- [ ] clearing existing mappings is deliberate and its consequences are handled; +- [ ] default contexts added to the local player subsystem; +- [ ] settings registration separated from runtime activation (§6); +- [ ] native actions bound explicitly; +- [ ] ability actions bound **and their handles retained** (§7); +- [ ] the ready event is sent exactly once per initialization; +- [ ] late feature handlers can bind safely. + +### Clearing all mappings is a broadcast dependency + +If initialization clears every mapping on the subsystem, it also removes contexts +added by already-active features. Recovery then depends on those features hearing +the ready event and re-adding their contexts — which makes correctness a property +of event ordering rather than of stored state. It works; it is fragile; and it +must be a documented decision rather than a side effect. Recipe: IN-08. + +--- + +## 6. Separate "register with settings" from "activate now" + +These are two different operations: + +```text +RegisterInputMappingContext(IMC) // discoverable and rebindable +AddMappingContext(IMC, priority) // actually active +``` + +A context may be active but not remappable, or registered but not currently +active. All four combinations are meaningful. + +**One flag must not gate both.** If it does, unchecking "expose this for +rebinding" also silently removes the input — a designer changes a settings +checkbox and the character stops responding, with nothing in the log. Recipe: +IN-06. + +--- + +## 7. Modular input needs ownership receipts + +### Adding a mapping context + +Data: context, priority, register-with-settings flag. Retain per activation +context: the extension and delegate handles, the exact players touched, and the +exact contexts and priorities activated. + +### Adding ability bindings + +Data: a list of input configs. Retain per pawn: **every returned bind handle**, +which config produced it, and the readiness subscription. + +### Symmetric removal + +```text +remove the exact bind handles this config created +remove the exact contexts this action added +unregister from settings only if this action registered them +release extension and delegate handles +``` + +**Never implement removal by clearing all bindings or all mappings on the pawn.** +That destroys other features' ownership, and it will look like it works. + +The measured reference composes this well and tears it down badly: bind handles +are declared as locals and discarded, so the removal function *cannot* be +implemented without changing the component that binds. Copy the composition; +write the teardown yourself. Recipes: IN-03, IN-04. + +--- + +## 8. Device switching belongs above gameplay semantics + +Gameplay should not branch on keyboard versus gamepad for the same command. The +UI layer owns the active input type, glyphs, controller brand and style, +change notifications, and disconnected-device handling. + +Gameplay may still use **distinct actions and tags** where the semantics genuinely +differ — mouse delta and analog stick look need different modifiers and +sensitivity, so: + +```text +InputTag.Look.Mouse +InputTag.Look.Stick +``` + +That is semantic distinction, not device sniffing. The test is whether the two +paths do different arithmetic; if they do, they are different commands. + +Per-player sensitivity, dead zones and inversion belong in input modifiers that +read settings — not in handler code. Note that a modifier which resolves a setting +by property **name** through reflection will break silently on a rename, with no +compiler help. Recipe: IN-10. + +--- + +## 9. Context priority and consumption + +A handler can be perfectly bound and never fire because an earlier context +consumes the key. + +Before adding a mapping: list every active context in priority order; find all +mappings for the physical key; inspect triggers, modifiers and consume-input +behaviour; establish whether the contexts can be active simultaneously; and +confirm no raw key path bypasses the system. + +Priority is meaningful only between concurrently active contexts. Document why a +feature needs to outrank the baseline. + +### Avoid raw key fallback + +Raw key handling bypasses rebinding, device parity, triggers and modifiers, +context priority, and input profiles. Use it only for editor and debug controls, +with a stated reason. + +--- + +## 10. The config-file boundary + +With Enhanced Input, the project input config file should contain the +player-input and input-component classes, user-settings class registration, debug +bindings, UI input defaults, and legacy axis properties only where the engine +still requires them. + +Gameplay mappings live in mapping context assets; semantic mappings live in input +config assets. **Legacy axis configuration that sits beside Enhanced Input is +actively misleading** — it looks like the place sensitivity is set, and it is not. +Recipe: IN-09. + +--- + +## 11. Test matrix + +### Composition + +Baseline pawn data and config; a mode action set adding a second config; +equipment granting and removing matching abilities; a feature activating both +before and after pawn readiness; two features adding distinct configs at once. + +### Lifecycle + +```text +spawn → bind → activate feature → press / hold / release +→ deactivate feature → verify no response +→ reactivate → verify exactly one response +``` + +### Devices + +Keyboard and mouse; gamepad; touch; hot-swap while UI is open; a saved and +reloaded remap; two local players with separate profiles. + +### Network and abilities + +Owning-client predicted activation; server rejection with failure feedback; input +held across a possession or avatar change; the block tag clearing buffered state; +an ability removed while its input is held. + +--- + +## 12. Review checklist + +- [ ] Physical key knowledge stops at the mapping context. +- [ ] Gameplay uses semantic tags, not asset references. +- [ ] Native versus ability classification is intentional. +- [ ] Ability input is buffered, not activated inside the event callback. +- [ ] Config and ability-set tags are validated as a join, in both directions. +- [ ] Tag matching semantics are known to be exact, and the data suits that. +- [ ] Binding waits for a named readiness state, not for begin play. +- [ ] Base and feature bindings both retain handles. +- [ ] Runtime activation and settings registration are independent. +- [ ] Deactivation removes only what this owner added. +- [ ] Priority and consumption conflicts have been audited for shared keys. +- [ ] Device icons and method switching stay in the UI layer. +- [ ] Activate, deactivate, reactivate produces no duplicate callbacks. + +--- + +## 13. When this architecture is too much + +For a single-mode project with one device family, no dynamic abilities and no +rebinding, direct bindings are honest and this indirection is overhead. + +Introduce semantic tags when at least one becomes true: abilities are granted +dynamically; several pawn archetypes reuse controls; features add input at +runtime; rebinding and device parity matter; gameplay must be independent of +content asset paths. + +The architecture is paid for by **modularity of content** and **many archetypes +with different action sets**. Without either, the cost is real and the benefit is +not: to answer "what does the left mouse button do?" a reader must traverse +mapping context, action, config, tag, pawn data, ability set and spec tags — +seven assets instead of one call stack. + +--- + +## Provenance + +The findings behind the failure modes come from a line-by-line source audit of +the input layer of Epic's Lyra Starter Game on Unreal Engine 5.6 — about 870 +lines across seven file pairs, plus its integration points in the hero component +and two feature actions — read as source rather than run. + +Source addresses stay in the research archive that produced this skill; what +ships is the detection recipe. Each entry carries a stable identifier (`IN-06` +and up) that resolves back to the audited location. + +## Evidence boundary + +One project, one engine version, one workspace. Mapping contexts and input +actions are binary assets and were not read, so every claim about which keys map +to which actions is out of scope rather than concluded. Re-run the recipes +against your own tree before acting on any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-input-architecture/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-input-architecture/references/failure-modes.md new file mode 100644 index 0000000..005dc62 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-input-architecture/references/failure-modes.md @@ -0,0 +1,405 @@ +# Failure modes: input architecture + +Eleven ways a tag-addressed input layer binds correctly and never fires. + +The shared property is unusually literal here: **the ability input path has no +error case.** A press resolves to a tag, the tag is compared against granted +ability specs, and no match means no activation. That is the same code path as +"you do not currently have this ability", which is a normal state. Nothing logs, +because from the framework's view nothing went wrong. + +The native path is louder — an unresolved tag produces an error at bind time — +which is worth knowing when triaging: **if the button is native and silent, the +problem is upstream of the config; if it is an ability and silent, the problem +could be anywhere along seven assets.** + +Recipes use `rg` from a project's source root, and were executed against the +audited project while this file was written. + +--- + +## The join + +### IN-01 - Tag lookup is exact, and the hierarchy suggests otherwise + +**Mechanism.** The ability spec lookup compares the pressed tag against the +spec's dynamic source tags with an exact match. A spec tagged with a parent is +not found by a press of its child. + +**Why it is silent.** No match means no activation, which is indistinguishable +from not having been granted the ability. There is no "close match" concept to +warn about. + +**Why the obvious check misses it.** Everything else about the tag layer is +hierarchical — the editor picker filters by category, the vocabulary is a tree, +and neighbouring systems in the same codebase do use hierarchical matching. The +reasonable assumption is wrong here, and nothing at either the config or the +ability-set end states the matching semantics. + +**Symptom.** An ability that is granted, bound, and correct in every inspectable +respect simply does not activate. Designers re-author the tag, re-save the +assets, and eventually route around it by duplicating the leaf. + +**Detect.** Read the comparison and confirm the semantics, then check whether any +authored tag relies on the other one: + +```bash +rg -n "HasTagExact|MatchesTag" --glob "*.cpp" . | rg -i "input|spec" +``` + +Then list every input tag used on an ability spec and every tag used in a config, +and confirm the sets are equal rather than merely overlapping. In the audited +project both lookup sites use exact matching, on adjacent lines. + +**Guardrail.** State the matching semantics in a comment at the lookup and in the +tag namespace's own documentation. If hierarchical matching is wanted, implement +it deliberately — do not let readers infer it from the tag shape. + +--- + +### IN-02 - Two tag sources for one namespace + +**Mechanism.** Native input tags are declared in C++ as compile-time symbols; +ability input tags are declared in a config file as strings. Both live under the +same namespace root. + +**Why it is silent.** Both work. The engine resolves both to the same registry, +and a tag from either source behaves identically at runtime. + +**Why the obvious check misses it.** Looking at either source shows a complete, +well-organised list. There is no artefact anywhere that shows both, so the +vocabulary appears smaller than it is and its split appears not to exist. + +**Symptom.** An audit of the input vocabulary misses half of it. A rename sweep +covers one source. A developer looking for a tag symbol in C++ concludes a tag +does not exist when it is declared in config. + +**Detect.** Count both sources and compare against the namespace: + +```bash +rg -n 'UE_DEFINE_GAMEPLAY_TAG\w*\([^,]+,\s*"InputTag\.' --glob "*.cpp" . +rg -n 'GameplayTagList=\(Tag="InputTag\.' Config/ Plugins/*/Config/ +``` + +Both non-empty is the finding. In the audited project the native movement tags +come from C++ and the ability tags from config, and the two sets do not overlap +at all — which is a defensible split, undocumented. + +**Guardrail.** One namespace, one declaration source, documented. If a split is +deliberate, draw the line at a sub-namespace boundary and write it down where +both halves are visible. + +--- + +## Ownership and teardown + +### IN-03 - Bind handles discarded, making removal impossible + +**Mechanism.** The binding helper returns handles through an out-parameter. Both +callers declare that array as a local and let it fall out of scope. + +**Why it is silent.** The bindings work perfectly. Handles are needed only to +undo, and nothing undoes during a normal session — pawns are recreated between +matches, which disposes of the input component and hides the leak. + +**Why the obvious check misses it.** The out-parameter is passed correctly, the +call is idiomatic, and a removal function exists elsewhere in the class. The +defect is the *storage duration* of a local variable, which no review of API +usage examines. + +**Symptom.** Deactivating a feature leaves its ability bindings on a live pawn. +Harmless while the ability is also revoked — the tag reaches the component and +finds no spec — but a re-grant resurrects phantom input. + +**Detect.** Find handle arrays declared at the call site rather than as members: + +```bash +rg -n -B3 "BindAbilityActions|BindAction\(" --glob "*.cpp" . \ + | rg "TArray\s+\w+;|TArray<\w*Handle>\s+\w+;" +``` + +In the audited project this appears at two sites in one component: the initial +binding and the feature-driven addition. + +**Guardrail.** Handles live on the object whose lifetime governs the binding, +stored in the same change that creates them. If there is nowhere to store them, +the input is not removable and that must be stated before it ships. + +--- + +### IN-04 - Removal API that exists and does nothing + +**Mechanism.** The class exposes a removal function whose body is a comment. A +feature action calls it during teardown. + +**Why it is silent.** The call compiles and returns. Nothing distinguishes +"removed the bindings" from "did nothing", because neither returns a result. + +**Why the obvious check misses it.** The teardown path *is* implemented in the +feature action — it iterates its tracked pawns and calls removal on each — so a +review of the action finds a complete, symmetric implementation. The empty body +is one call away in a different subsystem. + +**Symptom.** As IN-03, from the other end. A team scopes "implement the missing +removal" from the empty function, and finds the estimate was an order of +magnitude low because the handles no longer exist to remove. + +**Detect.** Find removal functions with no statements, and dead removal helpers: + +```bash +rg -n -A4 "void \w+::Remove\w+\(" --glob "*.cpp" . | rg -B2 "@?TODO" +rg -c "RemoveBinds|RemoveBindingByHandle" --glob "*.cpp" --glob "*.h" . +``` + +The second command matters: in the audited project a working removal helper +exists on the input component and has **zero callers** — the mechanism is +present, and the data it needs was thrown away by IN-03. + +**Guardrail.** A removal function with no body must not compile quietly: make it +pure virtual, assert, or delete the caller. Then check that its helper has at +least one caller — a correct utility nobody can call is the same as no utility. + +--- + +### IN-05 - Tracking container that is never written + +**Mechanism.** The context-adding action keeps a list of the controllers it +touched, removes from it, and iterates it during reset — but the add path takes +the per-context data as a parameter and never writes to it. + +**Why it is silent.** The reset loop finds the list empty and completes. Contexts +still come off, but only reactively — when the receiver is reported as going +away — so behaviour depends on teardown order. + +**Why the obvious check misses it.** Four of five operations on the container +exist, including an assertion at activation that it starts empty. Any check of +"is this container used?" answers yes emphatically. Only the producer is missing. + +**Symptom.** Mapping contexts survive feature deactivation when the receiving +controller outlives the feature. Intermittent, ordering-dependent. + +**Detect.** Compare reads against writes per container: + +```bash +C='ControllersAddedTo' +rg -n "\b$C\b" --glob "*.cpp" --glob "*.h" . +rg -n "\b$C\b\s*\.\s*(Add|AddUnique|Emplace|Push)" --glob "*.cpp" . +``` + +In the audited project the second search is empty, the authors left a comment on +the reset path noting exactly that, and a sibling action in the same directory +does the equivalent correctly — so the two files diff against each other. + +**Guardrail.** A container whose only mutations are removals has no producer. +Assert non-empty after the add path runs. + +--- + +## Configuration traps + +### IN-06 - One flag secretly gating two operations + +**Mechanism.** Registering a context for rebinding and activating it are separate +operations. In one code path the activation call is nested **inside** the +`if` that tests the register-with-settings flag. + +**Why it is silent.** Both operations are legitimate and both are performed when +the flag is set — which it is by default. Nothing about the default case is +wrong, so the coupling never manifests during development. + +**Why the obvious check misses it.** The flag's name describes exactly one of the +two things it controls. A designer reading the property, and a reviewer reading +the property, both understand it correctly and neither is looking at the brace +structure twenty lines away. Worse: the *same* project has a second +implementation of the same feature where the flag correctly controls only +registration, so a reader who checks one path and generalises is wrong. + +**Symptom.** Unchecking "expose this for rebinding" makes the character stop +responding entirely, with nothing in the log. Diagnosed as an asset problem. + +**Detect.** Find activation calls nested inside a settings-registration branch: + +```bash +rg -n -B8 "AddMappingContext\(" --glob "*.cpp" . | rg "bRegisterWithSettings|RegisterInputMappingContext" +``` + +Any hit is the finding. Then check every other implementation of the same +operation in the project and compare — the discrepancy between two paths is +stronger evidence than either path alone. + +**Guardrail.** Registration and activation are separate statements at the same +brace level, each guarded by its own condition. Test all four combinations of the +two flags, because all four are meaningful. + +--- + +### IN-07 - Nested condition that disables an unrelated feature + +**Mechanism.** The loop that applies default mapping contexts sits inside an +`if` that tests for a semantic config. A pawn with no config therefore gets no +contexts either. + +**Why it is silent.** Every shipped pawn has a config, so the guard is always +true and the nesting is never exercised. The two things are unrelated in concept +and adjacent in code. + +**Why the obvious check misses it.** The guard is correct for the block it was +written for. The extra scope it captures is a brace position, not a statement, +and nothing names the coupling. + +**Symptom.** A new pawn archetype authored without a config is completely +unresponsive — not "no abilities", but no movement, no look, nothing. Read as a +much deeper problem than a missing asset reference. + +**Detect.** Read the initialization function's brace structure: + +```bash +rg -n -A25 "::InitializePlayerInput" --glob "*.cpp" . | rg -n "if \(|for \(|AddMappingContext" +``` + +A mapping-context loop indented inside a config guard is the finding. + +**Guardrail.** Contexts and configs are independent inputs to initialization and +belong in sibling blocks. Validate that a pawn without a config is a deliberate +configuration rather than a broken one. + +--- + +### IN-08 - Clearing all mappings, then relying on a broadcast to restore them + +**Mechanism.** Pawn input initialization clears every mapping on the subsystem +before adding its own — removing contexts that active features added earlier. A +readiness broadcast afterwards prompts those features to re-add theirs. + +**Why it is silent.** It works, because the broadcast happens and the features +listen. Correctness is real; it is just a property of event ordering rather than +of stored state. + +**Why the obvious check misses it.** The clear is a defensible line with an +obvious intent — start from a known state. The features' re-add is in a different +file, reached by an event. Nothing at the clear site says what it is destroying, +and nothing at the re-add site says why it is necessary. + +**Symptom.** A feature's input disappears when a pawn re-initializes, in any +ordering where that feature does not receive or does not act on the ready event — +a feature activated during the broadcast, one whose handler early-returns, or one +whose receiver is not yet registered. + +**Detect.** Find the clear and confirm the recovery path: + +```bash +rg -n -A3 "ClearAllMappings" --glob "*.cpp" . +rg -n "SendGameFrameworkComponentExtensionEvent" --glob "*.cpp" . -B4 +``` + +A clear with no explicit re-application, relying on an event broadcast elsewhere, +is the finding. In the audited project the flag that gates readiness is set +*before* the broadcast — the correct order, and worth verifying rather than +assuming. + +**Guardrail.** Either do not clear what you do not own, or re-apply from a stored +list rather than from a broadcast. If the broadcast is the mechanism, comment it +at the clear site so the dependency is visible from both ends. + +--- + +## Cost and misleading surfaces + +### IN-09 - Legacy axis configuration beside Enhanced Input + +**Mechanism.** The project input config retains legacy axis entries — dead zones, +sensitivities — while actual sensitivity and dead zone are applied by input +modifiers. + +**Why it is silent.** The legacy values are read by the legacy path, which is not +in use. Changing them does nothing, and doing nothing produces no error. + +**Why the obvious check misses it.** They are exactly where a developer would +look for sensitivity, with exactly the names they would search for. The file +answers the question asked; the answer is stale. + +**Symptom.** Time lost tuning values that have no effect, and a plausible but +wrong belief about where input tuning lives — which then propagates into +documentation. + +**Detect.** Compare the legacy entries against the modifiers that really apply: + +```bash +rg -n "AxisConfig|Sensitivity=|DeadZone" Config/*.ini +rg -n "class \w*InputModifier" --glob "*.h" . +``` + +Both non-empty is the finding. In the audited project eleven legacy axis entries +coexist with four modifiers that read player settings, and the modifiers win. + +**Guardrail.** Delete legacy entries the engine does not require, or annotate them +in place as inert with a pointer to the modifiers. A misleading config is more +expensive than a missing one. + +--- + +### IN-10 - Setting resolved by property name through reflection + +**Mechanism.** An input modifier scales a value by looking up a settings property +**by name** through reflection, with a cache. + +**Why it is silent.** A rename produces no compile error, because the name is a +string in data. The lookup fails at runtime and the modifier returns unscaled +input — a plausible value. + +**Why the obvious check misses it.** Renaming a property is exactly the operation +IDEs make safe. The developer performs a rename refactor, everything compiles, +every reference updates, and this one does not because it is not a reference. + +**Symptom.** Sensitivity, inversion or dead zone silently stops respecting the +player's setting after an unrelated refactor. Reported as "the setting does +nothing" long after the rename. + +**Detect.** Find name-based property resolution: + +```bash +rg -n "FindPropertyByName|FindFProperty|GetPropertyByName" --glob "*.cpp" . +``` + +Each hit is a link the compiler does not check. Add a startup assertion for each. + +**Guardrail.** Resolve settings through typed accessors. Where reflection is +unavoidable, assert at startup that every name resolves, so a rename fails loudly +at launch rather than quietly in gameplay. + +--- + +## Loading + +### IN-11 - Blocking load on the pawn initialization path + +**Mechanism.** A mapping context soft reference is resolved with a synchronous +load during pawn input initialization. + +**Why it is silent.** It succeeds. The asset arrives, input works, and the cost is +a stall measured in milliseconds on a warm cache. + +**Why the obvious check misses it.** A synchronous load is the simplest correct +code, and correctness is what review checks. The cost appears only on a cold +cache, on slower storage, or at a moment — player spawn — that nobody profiles +because it happens before the part they were measuring. + +**Symptom.** A hitch at spawn that is attributed to the level, the character mesh +or the ability system, because the input layer is not where anyone looks for a +frame spike. + +**Detect.** Find synchronous loads on initialization paths, and contrast them with +the bare accessors elsewhere: + +```bash +rg -n "LoadSynchronous\(\)" --glob "*.cpp" . -B6 | rg "Init|BeginPlay|Setup" +rg -n "\.Get\(\)" --glob "*.cpp" . -A2 | rg -v "ensure|check|if\s*\(|UE_LOG" +``` + +The pairing is the interesting part: the audited project loads synchronously on +the base path and uses bare, unguarded accessors on the feature paths — so one +route hitches and the other silently skips when a bundle did not load. + +**Guardrail.** Preload input assets with the bundle that owns them and resolve +with a guarded accessor that logs on failure. Neither blocking nor silently +skipping is acceptable on a path this visible. diff --git a/plugins/ue-design-skills/skills/ue-input-architecture/references/patterns.md b/plugins/ue-design-skills/skills/ue-input-architecture/references/patterns.md new file mode 100644 index 0000000..bc097a4 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-input-architecture/references/patterns.md @@ -0,0 +1,225 @@ +# Patterns: tag-addressed input in a measured reference product + +A worked example of the three-layer architecture from `SKILL.md`, measured in one +reference project. It exists to answer the question the norms cannot: **what does +this architecture actually cost, and what does it actually buy?** + +The short version, and it is the reason this file is worth reading: the +architecture works, the composition benefits are real and measurable, and the +teardown half was never finished. Those three facts are not in tension. They are +what a reference project looks like. + +Markers: **[measured]** — read in source or config; **[derived]** — conclusion +from measured facts; **[open]** — not answerable from source alone. + +--- + +## 1. The shape, in one chain + +```text +key +→ mapping context +→ input action +→ [payload: input tag] handler on the hero component +→ ability system component buffers the spec handle +~~~ end of the input-processing step ~~~ +→ controller post-process input +→ activation policy decides +→ activate +``` + +Seven hops, five of them across assets **[measured]**. That count is the whole +trade: it is why a feature plugin can add controls without touching the base +module, and it is why a broken binding takes an afternoon to diagnose. + +--- + +## 2. What the split actually buys, with numbers + +| Benefit | Measured evidence | +|---|---| +| Code cost per new ability control | **zero lines** — one loop binds every row of the config **[measured]** | +| Code cost per new native control | one explicit line; the project has exactly five and they have not changed **[measured]** | +| Feature-plugin controls without base-module edits | one plugin ships 11 input actions and its own config, and the base module contains no reference to them **[measured]** | +| One switch to suppress all ability input | a single tag on the ability system component clears the buffers and returns early **[measured]** | +| Same action, different ability per character | the join is a string in two assets; no code path is per-character **[derived]** | + +The fourth row is the one most often underestimated. Without a semantic layer, +suppressing input during a menu means either a flag consulted in every handler or +unbinding and rebinding at runtime. Here it is one tag, checked once per frame, +in one place. + +--- + +## 3. What it costs, also with numbers + +**Compile-time errors become runtime silence.** The native path logs an error +when a tag cannot be resolved at bind time. The ability path has **no error case +at all** — the pressed tag is compared against granted specs, and no match is +indistinguishable from not having the ability **[measured]**. This asymmetry is +worth internalising as a triage rule: a silent native control is a config +problem; a silent ability control could be anywhere along the chain in §1. + +**The lookup is exact, and everything around it looks hierarchical.** Both lookup +sites use exact tag matching **[measured]**, while the editor picker filters by +category and the vocabulary is a tree. A spec tagged with a parent is never found +by a press of its child. Nothing at either end of the join states this. Recipe: +IN-01. + +**Two declaration sources for one namespace.** Native tags are C++ symbols; +ability tags are config strings; the two sets do not overlap **[measured]**. Both +are defensible individually. Together they mean no single artefact lists the input +vocabulary. Recipe: IN-02. + +**Linear scans on every press.** The lookup iterates all activatable specs +**[measured]**. At the scale of this project that is free; it is proportional to +ability count times presses per frame, and neither number is bounded by the +design. + +--- + +## 4. Teardown: the half that was not finished + +This is the cluster worth studying, because the four findings compound into one +outcome and each looks minor alone. + +| Finding **[measured]** | Alone | In combination | +|---|---|---| +| Bind handles declared as locals at two call sites and discarded | a missed `TArray` member | removal has no data to work with | +| The removal function's body is a single TODO comment | an unimplemented method | the feature action's teardown calls into it and returns successfully | +| A working `RemoveBinds` helper exists with **zero callers** | dead code | the mechanism was built and could never be reached | +| The context action's tracking list is read, iterated and removed from — and never written | an empty loop | contexts come off only reactively, so behaviour depends on teardown order | + +**[derived]** Deactivating a feature leaves its ability bindings on a live pawn. +The practical damage is limited because pawns are recreated between matches and +because an orphan binding sends a tag that matches no spec — but the leak is real, +and a later re-grant of the same ability resurrects phantom input. + +The instructive part is the third row. Someone wrote the removal helper. It is +correct. It has never been called, because the handles it needs were thrown away +one function earlier. **A correct utility that nothing can call is +indistinguishable from a utility that does not exist** — and it is worse, because +its presence answers "is removal supported?" with a yes. + +The original authors had also left a comment on the broken tracking line noting +that the container is never modified **[measured]**. That is a known defect that +shipped, which is the most honest possible evidence for the general rule: *finding +a defect and fixing a defect are different projects, and reference code shows you +the first.* + +Recipes: IN-03, IN-04, IN-05. + +--- + +## 5. One flag, two operations — the sharpest single finding + +In the base initialization path, the call that makes a mapping context **active** +sits *inside* the branch that checks whether the context should be registered with +the player's settings **[measured]**. + +In the feature-plugin path for the same data structure, the two are separate: the +settings registration is skipped by the flag, and the activation call runs +unconditionally **[measured]**. + +**[derived]** So the same checkbox means "expose this for rebinding" in one place +and "expose this for rebinding **and also make it work at all**" in another. Clear +the flag on a base mapping and the controls go silent, with no error and no +obvious connection between cause and effect. + +Two properties make this the best example in the file: + +1. **It is verifiable in ten seconds** by reading the two files side by side, and + invisible in any amount of reading of either file alone. +2. **The correct version exists in the same codebase**, which removes any argument + about intent. This is not a design decision applied inconsistently; one of the + two is wrong. + +Recipe: IN-06. The same initialization path carries a second nesting defect: the +loop over default mapping contexts is nested inside a check for the pawn's input +config, so a pawn with no config gets no contexts either and is completely mute +**[measured]** — two unrelated features disabled by one condition. Recipe: IN-07. + +--- + +## 6. Initialization order as a broadcast dependency + +Initialization clears **every** mapping on the subsystem before adding its own +**[measured]**. That includes contexts added by features that activated earlier. + +Recovery works, and it works by broadcast: the component sets its readiness flag +and then sends a semantic "bind inputs now" event, which both feature actions +listen for alongside the generic extension event **[measured]**. Features re-add +what was just removed. + +**[derived]** This is correct and it is fragile in a specific way: correctness is a +property of *event ordering* rather than of stored state. Nothing reconciles what +was cleared against what was restored, and a feature that is active but not +listening for the ready event loses its contexts permanently. + +Worth copying from this: the readiness flag is set **before** the broadcast, not +after **[measured]**. A handler that checks readiness during the broadcast sees +true. That ordering is one line and it is the difference between the pattern +working and not. Recipe: IN-08. + +--- + +## 7. Sediment: things that exist and do nothing + +A reference project accumulates these, and each one costs a reader time: + +- two mapping-management functions whose bodies contain only a check and a comment, + called from the initialization path, doing nothing **[measured]** — the remains of + a superseded settings API; +- a user-settings class and a key-profile class registered in config, both of which + are empty overrides **[measured]** — extension points presented as features; +- a legacy axis-configuration block sitting beside the modern input system, whose + deadzone and sensitivity values are overridden by modifiers that read player + settings **[measured]**. A reader tuning the legacy values changes nothing. + Recipe: IN-09; +- a settings-driven modifier that resolves a property **by name through + reflection** **[measured]**, so renaming the underlying setting breaks it + silently with no compiler involvement. Recipe: IN-10. + +**[derived]** None of these is a bug. Collectively they are the reason "read the +reference and copy what it does" is a bad adoption strategy: a meaningful share of +what it does is nothing. + +--- + +## 8. What the reference got right + +1. **The payload trick.** Passing the semantic tag as a bound-delegate payload is + what makes one function pair serve unlimited abilities **[measured]**. It is + the single most portable idea in this file. +2. **Buffer then activate in one batch.** Collecting handles and activating after + all input is processed removes event-order dependence, and the authors left a + comment explaining precisely why — held input would otherwise activate an + ability and then deliver a press event to it **[measured]**. +3. **Setting readiness before broadcasting it** (§6). +4. **Separate actions for mouse look and stick look**, with different tags and + different handlers, because the stick needs frame-time scaling and the mouse + does not **[measured]**. This is semantic distinction done correctly — not + device sniffing in gameplay code, but two genuinely different commands. +5. **Device switching kept entirely out of the input layer.** The input module + contains no reference to input-device types at all **[measured]**; device + changes affect icons, haptics and UI style, and never gameplay bindings. + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source and +config in a single workspace: roughly 870 lines across the input module, plus the +hero component, two feature actions and the input configuration file. Blueprint +assets — the mapping contexts and input actions themselves — are binary and were +not read, so every statement here about which keys map to what is **[open]**. + +Source addresses stay in the research archive that produced this skill; each +`IN-` identifier resolves back to the audited location there. + +## Evidence boundary + +These are examples and failure evidence from one project on one engine version, +not guarantees about others. The counts are properties of that project on that +day and are quoted to show the shape of a contrast, not to be imported. Re-run the +recipes against your own tree. diff --git a/plugins/ue-design-skills/skills/ue-modular-gameplay/SKILL.md b/plugins/ue-design-skills/skills/ue-modular-gameplay/SKILL.md new file mode 100644 index 0000000..8508502 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-modular-gameplay/SKILL.md @@ -0,0 +1,425 @@ +--- +name: ue-modular-gameplay +description: >- + Design or review modular gameplay in Unreal Engine: experience definitions, + feature plugins, reusable action sets, feature actions, actor extension + handlers, component init-state chains, primary asset bundles, and runtime + activation and deactivation. Use when adding game modes, seasonal or + downloadable features, modular components, abilities, input or UI, or when + fixing initialization-order and feature teardown bugs. +--- + +# UE modular gameplay + +The invariant: + +> A feature may select and attach behaviour without modifying the receiving base +> class, but it owns every resource it adds and must remove exactly that +> resource. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Eighteen detection recipes for silent modularity failures: +[failure modes](references/failure-modes.md). + +Read the failure modes before adopting this architecture. Nearly all of them are +teardown defects, and teardown defects share one property: **activation is what +gets tested, so the half that is broken is the half nobody exercises.** + +Related skills: `ue-data-driven-architecture`, `ue-input-architecture`, +`ue-ui-architecture`, `ue-gas-architecture`, `ue-multiplayer-authority`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Separate the four layers + +```text +Playlist / session entry +└─ Experience definition + ├─ feature plugins to activate + ├─ reusable action sets + ├─ mode-specific actions + └─ default pawn data + +Feature data asset +└─ lifecycle and asset scanning of one plugin + +Action set +└─ reusable composition slice + +Feature action +└─ one apply/rollback operation +``` + +Do not merge them: + +- the playlist owns map, session and menu metadata, not gameplay composition; +- the experience owns one runtime composition, not plugin lifecycle policy; +- the feature data asset owns plugin registration and scanning, not a game mode; +- an action set expresses reuse; it is not an inherited base experience; +- an action performs one focused mutation and owns its rollback state. + +### Prefer composition over experience inheritance + +A mode should read as a sum: + +```text +SharedInput + StandardComponents + StandardHUD + ModeDelta +``` + +not as a hidden defaults chain: + +```text +Base -> Shooter -> Team -> TeamDeathmatch +``` + +Flat action sets keep dependencies, asset loading and teardown auditable. A +reference product measured for this skill goes further and *forbids* the second +form in validation, pointing authors at composition instead. + +--- + +## 2. Experience loading is a state machine + +```text +Unloaded +→ LoadingAssets +→ LoadingGameFeatures +→ ExecutingActions +→ Loaded +→ Deactivating +→ Unloaded +``` + +Each transition needs an explicit completion criterion. Never track this with a +handful of booleans — `bAssetsReady && bPluginsReady && bActionsDone` cannot +express "failed", and failure is the state you will need most. + +### The server selects; every peer realises locally + +Replicate the selected definition, never the result of loading it: + +```text +server selects experience ID +→ replicated definition reaches the client +→ server and client run the same load pipeline +→ role-specific bundles and actions decide their local work +``` + +This is "replicate intent, realise presentation locally" applied to composition. + +### Document selection precedence + +If an experience can come from matchmaking, a URL, developer settings, the +command line, world settings and a default, define the exact precedence, log the +winning source, validate the asset ID before loading, and provide an explicit +fallback. + +### Delay only for a named dependency + +Deferring selection by one frame so that startup settings can initialise is +legitimate **because the dependency has a name**. A next-tick timer with no +identified prerequisite is initialization-order debt with a timer wrapped around +it. Recipe: MG-12. + +--- + +## 3. Bundle and role boundaries + +Collect the experience plus every action set as primary assets, then request +named bundles: a shared bundle always, the client bundle when a client exists, +the server bundle when authority exists. + +Mark soft references at their declaration: + +```cpp +meta=(AssetBundles="Client") +meta=(AssetBundles="Client,Server") +``` + +| Action target | Expected role | +|---|---| +| HUD, layout, input icons | Client | +| scoring authority, bots, team assignment | Server | +| replicated scoring component | Client + Server | +| ability classes and effects needed for prediction | usually both | + +**A client/server flag on an action does not guarantee correct cook bundles.** +Validate both, separately. + +### Editor bundle loading is not shipping + +A policy that loads client data when not a dedicated server and server data when +not a client-only build loads **both** in the editor. Every memory and load-time +figure taken in-editor therefore describes a configuration that ships to nobody. +Use it for direction of change only. Recipe: MG-18. + +--- + +## 4. Extend actors through receivers and extension handlers + +Do not scan the world for actors and patch them. Use the component manager's +extension protocol: + +```text +receiver actor registers itself +action registers an extension handler for a target class +manager invokes the handler for existing AND future receivers +returned request handle owns the subscription +``` + +The receiver lifecycle is symmetric and belongs in the base class: + +```text +PreInitializeComponents : add receiver +BeginPlay : send ready events +EndPlay : remove receiver +``` + +Note what this buys: base gameplay classes register as receivers and **know +nothing about features**. If a class does not inherit the modular base, those +three calls are written by hand — and forgetting the third is a leak with no +symptom until the second activation. + +### Use semantic readiness events + +A generic "extension added" event means only that the actor is known to the +manager. It may not yet have its pawn data, ability system, controller or input. +Emit named events at the real dependency boundary — abilities ready, bind inputs +now, actor ready — and have handlers accept both the generic event and the +semantic one, so activation order stops mattering. Recipe: MG-13. + +### Retain request handles + +Handler registration and component requests return ownership handles that are +reference counted. **Dropping the handle is the unsubscribe operation**; there is +no explicit removal call. Storing them per activation context is therefore not +bookkeeping, it is the only teardown mechanism that exists. + +--- + +## 5. Use an init-state chain + +Networked actor composition has non-deterministic arrival order. Replace +scattered checks in `BeginPlay`, possession callbacks and replication callbacks +with a monotonic chain: + +```text +Spawned → DataAvailable → DataInitialized → GameplayReady +``` + +- **Spawned** — the actor and component exist. +- **DataAvailable** — required replicated and config data exist. +- **DataInitialized** — every participating feature reached DataAvailable and + cross-feature initialisation ran. +- **GameplayReady** — safe for active gameplay. + +The load-bearing barrier is "have all features on this actor reached +DataAvailable", which synchronises components without any of them naming the +others. + +### Rules + +- states only advance; they reset with the actor's lifecycle, not by backtracking; +- every transition has an observable prerequisite; +- replication callbacks re-enter the same check rather than duplicating init; +- feature-specific work runs once, at a named transition; +- test server, owning client and simulated proxy separately — a proxy never + acquires a controller, and a gate that waits for one deadlocks it while working + perfectly in a single-process test. + +--- + +## 6. Action design contract + +An action class implements behaviour; its instance stores targets and +configuration. A world-aware base action should filter to game worlds, handle +worlds that already exist at activation, subscribe for future ones, and release +its world delegates on deactivation — leaving resource tracking to the concrete +action. + +### State must be keyed by activation context + +One action object can be active in several contexts at once — client and server +in a single-process test are the common case. **State in plain member fields is a +cross-context leak.** Key it: + +```cpp +struct FPerContextData { /* ... */ }; +TMap ContextData; +``` + +Recipe: MG-11. + +### One responsibility per action + +Good: add components; add abilities; add input binding; add input mapping +context; add widgets; add cue path. + +Bad: "set up shooter mode". + +A focused action has a testable rollback and can be reused in an action set. + +### The mandatory ownership ledger + +Write this table **before** implementing any action: + +| Added resource | Stored ownership | Removal | +|---|---|---| +| extension callback | request handle | release the handle | +| component | component request | remove the owned component | +| ability, effect, attribute set | granted handles | revoke the handles | +| input bind | bind handles per pawn | remove those exact binds | +| input mapping context | player + context | remove from the same subsystem | +| widget extension | extension handle | unregister | +| layout | widget or layout handle | deactivate and remove | +| global delegate | delegate handle | unbind | + +**If the middle column is blank, reject the action.** Every teardown failure in +the measured reference is a blank middle column that shipped. Recipes: MG-01 +through MG-05. + +--- + +## 7. Deactivation is not optional + +"Modular" is not proven by successful activation. The round trip is the test: + +```text +inactive baseline +→ activate → verify additions exactly once +→ deactivate → verify exact baseline +→ reactivate → verify no duplicates and no stale state +``` + +Run it with: actors that existed before activation; actors spawned after it; an +actor destroyed before the feature deactivates; two worlds in one process; client +and server roles; a partial load failure; and a plugin required by two +experiences at once. + +### Removal must be idempotent + +Resources come off by two independent paths: the explicit reset at deactivation, +and reactively when the actor dies first and the manager reports it. Both must be +safe to run, in either order, and must operate on "is there a record?" rather than +assuming one. Recipe: MG-08. + +### Reference-count shared plugin activation + +If two experiences need one plugin, count activations. One owner deactivating +must not unload a plugin the other still requires. **Check that the count is not +editor-only** — a reference count compiled out of shipping builds is a reference +count that protects development and nothing else. Recipe: MG-16. + +### Diff requirements on experience change + +```text +old - new = deactivate and unload +new - old = load and activate +intersection = keep active +``` + +Do not unload everything and reload it immediately. If this diff is not +implemented, define experience changes as map-travel-only and say so in writing — +that is a legitimate scope decision, and leaving it undefined is not. + +--- + +## 8. Failure handling + +A production loader must not assert its way through configuration failure. +Represent failure as a state: + +```text +Loaded +Failed(reason, failing asset / plugin / action) +``` + +Handle: an invalid asset ID; a missing plugin URL; a cancelled load; plugin +activation failure; action validation failure; timeout or partial completion; and +a deactivation pauser that never completes. + +The loading screen consumes loader state and reason. **A permanent loading screen +is not error handling** — it is the absence of it, rendered. + +--- + +## 9. Validation checklist + +### Experience + +- [ ] Asset ID registered and resolves. +- [ ] Default pawn data valid, or explicitly no pawn. +- [ ] Plugin names resolve to URLs. +- [ ] Action sets non-null and duplicate-free. +- [ ] Reusable behaviour in action sets; only the mode delta inline. +- [ ] Client and server bundles cover every soft reference. + +### Action + +- [ ] Target actor class valid; addition list non-empty. +- [ ] Existing and future actors both handled. +- [ ] Semantic readiness event used where the generic one is too early. +- [ ] Every request, delegate, grant, widget and input handle retained. +- [ ] Deactivation reverses activation exactly. +- [ ] Repeated activate/deactivate is idempotent. +- [ ] Per-context state cannot leak into another world. +- [ ] Index or key passed to a deferred callback matches the entry it names. + +### Init states + +- [ ] Every prerequisite is explicit. +- [ ] Every participating feature can reach every state. +- [ ] No feature waits on a signal produced only by a path it disabled. +- [ ] Authority and local-control branches tested. +- [ ] Late replication re-enters the same check. + +--- + +## 10. Review questions + +Before approving a modular feature: + +1. Which definition selects it? +2. Which plugin owns its code and content? +3. Which action applies each effect? +4. Which readiness state makes that effect safe? +5. Which handle proves ownership? +6. What is the exact inverse operation? +7. What is loaded on client, on server, in the editor? +8. Can another feature share the same plugin or resource? +9. What happens on partial failure? +10. Does activate → deactivate → activate return to the same observable state? + +**If questions 4 to 6 have no precise answer, the feature is dynamically attached +but not modular.** That distinction is the entire subject of this skill: attaching +is easy and demonstrable, detaching is neither, and only one of them gets +demonstrated. + +--- + +## Provenance + +The findings behind the failure modes come from a line-by-line source audit of +the feature-action layer of Epic's Lyra Starter Game on Unreal Engine 5.6, read +rather than run. The engine-side component manager and feature subsystem are not +part of that project, so every statement about them is inferred from call sites +in the audited code and is marked as such rather than presented as measured. + +That audit found the architecture sound and its teardown half incomplete in five +separate actions — which is why this skill treats the ownership ledger and the +round-trip test as mandatory rather than advisory. Several of the gaps are +recorded by the original authors in their own comments, which is the strongest +confirmation a source audit can produce. + +Source addresses stay in the research archive; what ships is the detection +recipe. Each entry carries a stable identifier (`MG-01`, `MG-02`, …) that +resolves back to the audited location there. + +## Evidence boundary + +One project, one engine version, one workspace. Binary assets were not read, so +which actions the shipped features actually configure is unknown and is recorded +as open rather than guessed. Re-run the recipes against your own tree before +trusting any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-modular-gameplay/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-modular-gameplay/references/failure-modes.md new file mode 100644 index 0000000..6386261 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-modular-gameplay/references/failure-modes.md @@ -0,0 +1,701 @@ +# Failure modes: modular gameplay + +Nineteen ways a feature attaches correctly and detaches incompletely. + +The shared property is stated once, because it explains the whole list: +**activation is what gets demonstrated.** A feature is switched on, someone +checks the ability appeared, the widget rendered, the input worked, and the +review is done. Deactivation is exercised by nobody during development, because +in development you restart the editor instead. So the broken half is +systematically the half nobody looks at, and its symptoms — a duplicate on the +second activation, a bind that survives its feature — appear later and are +attributed elsewhere. + +A second property applies to most entries: the reference implementations of these +actions **look symmetric**. There is an `Add` and there is a `Remove`, a reset +loop and a reactive removal path. The defect is almost never a missing function; +it is a function that removes something other than what was added, or that reads +a container nobody fills. + +Recipes use `rg` and are written to be run from a project's source root. Most are +a pair: one search finds the addition, the other looks for the matching removal. + +--- + +## Teardown + +### MG-01 - Rollback function with an empty body + +**Mechanism.** An action's removal path calls a method on a game component, and +that method's body is a comment saying it is not implemented yet. + +**Why it is silent.** Every caller succeeds. The call compiles, executes and +returns; nothing distinguishes "removed the bindings" from "did nothing" at the +call site, because neither returns a result. + +**Why the obvious check misses it.** The action's teardown *is* implemented and +reads correctly — it iterates its tracked pawns and calls a removal method on +each. Reviewing the action finds a complete, symmetric implementation. The empty +body is one call away, in a different file owned by a different subsystem, and it +is the game component's problem rather than the feature's. + +**Symptom.** Ability input bindings survive the deactivation of the feature that +added them. The second activation adds a second set. Nothing errors. + +**Detect.** Find removal functions whose body is only a comment: + +```bash +rg -n -A4 "void \w+::Remove\w+\(" --glob "*.cpp" . \ + | rg -B2 "^\s*//\s*@?TODO|^\s*\}" +``` + +Read each hit: a `Remove` whose body contains no statements is the finding. In +the measured reference the removal of an additional input config was exactly this +— a three-line function containing one TODO comment. + +**Guardrail.** A removal function with no body must not compile silently. Make it +pure virtual, assert inside it, or delete the caller. An unimplemented rollback +that is *called* is worse than one that does not exist, because the call site +proves the intent was there. + +--- + +### MG-02 - Bind handles discarded at the moment they are created + +**Mechanism.** A binding call takes an out-parameter array of handles. The caller +declares that array as a local variable and lets it go out of scope. + +**Why it is silent.** The bindings work. Handles are only needed to undo, and +nothing undoes during a normal session. + +**Why the obvious check misses it.** The out-parameter is passed correctly and +the call is idiomatic — some codebases even name the argument in a comment at the +call site, which makes it look deliberate. The defect is the *storage duration* +of a local variable, which no search for "is the API used correctly?" examines. +This is also the root cause of MG-01: the removal function cannot be implemented, +because the information it would need was thrown away one function earlier. + +**Symptom.** Rollback is not merely unimplemented but *impossible* without +changing the component that binds. An estimate of "implement the missing removal" +comes out an order of magnitude low. + +**Detect.** Find handle arrays declared as locals at a binding call: + +```bash +rg -n -B2 "BindAbilityActions|BindAction\(" --glob "*.cpp" . \ + | rg "TArray\s+\w+;|TArray\s+\w+;" +``` + +Any handle container declared in the same scope as the call, rather than as a +member, is discarded. In the measured reference this appeared at two separate +call sites in one component. + +**Guardrail.** Handles returned by a binding API are stored by the object whose +lifetime governs the binding, in the same change that creates the binding. If +there is nowhere to store them, the feature is not removable and that must be +stated before it ships. + +--- + +### MG-03 - Tracking container that is read but never written + +**Mechanism.** An action declares a container of the actors it has affected, +removes from it in the reactive path, iterates it during reset — and never adds +to it. The function that would add takes the per-context data as a parameter and +ignores it. + +**Why it is silent.** The reset loop runs, finds the container empty, and +completes successfully. An empty loop is a legal outcome; it looks exactly like a +feature that affected nothing because no receivers existed. + +**Why the obvious check misses it.** Every symptom of correctness is present: the +container is declared, an `ensure` at activation asserts it starts empty, the +reset loop pops from it correctly, and the reactive path removes from it. Four of +the five operations exist. Only the write is missing, and "is this container +used?" answers yes at four sites. + +**Symptom.** Input mapping contexts are never removed at deactivation. They come +off only if the receiving actor happens to die first and the reactive path fires. +The result depends on teardown order, which makes it intermittent. + +**Detect.** For every tracking container, compare reads against writes: + +```bash +C='ControllersAddedTo' +rg -n "\b$C\b" --glob "*.cpp" --glob "*.h" . # all uses +rg -n "\b$C\b\s*\.\s*(Add|AddUnique|Emplace|Push)" --glob "*.cpp" . # writes only +``` + +Uses without writes is the finding. The measured reference is unusually clear +here: a sibling action in the same directory tracks its pawns with +`AddUnique` in exactly the place the broken one omits it, so the two files can be +diffed against each other. The original authors had also left a comment on the +broken line noting that the container is never modified — a known defect that +shipped. + +**Guardrail.** A container whose only writes are removals is a container with no +producer. Assert non-empty after the add path runs, or derive the collection from +the resource itself rather than maintaining a parallel list. + +--- + +### MG-04 - Reset path narrower than the reactive removal path + +**Mechanism.** An action has two removal routes: a reset invoked at deactivation, +and a per-actor removal invoked when the manager reports a receiver going away. +The per-actor route cleans up more than the reset does. + +**Why it is silent.** The reset completes and clears its own bookkeeping, so the +action's internal state is consistent afterwards. What is left behind lives in +another subsystem — a pushed layout widget still on its layer — and that +subsystem was never told. + +**Why the obvious check misses it.** Both functions exist and both are correct +for what they enumerate. Comparing them requires reading two loops in different +parts of a file and noticing that one iterates two collections and the other +iterates one. Nothing names the missing collection. + +**Symptom.** A layout widget pushed by a feature stays on screen after the +feature is deactivated. On reactivation there are two. + +**Detect.** Enumerate what each removal path touches and diff the sets: + +```bash +rg -n -A20 "::Reset\(" --glob "*.cpp" . | rg "Empty\(|Unregister|Deactivate|Remove" +rg -n -A20 "::Remove\w+\(" --glob "*.cpp" . | rg "Empty\(|Unregister|Deactivate|Remove" +``` + +A collection handled in the second output and absent from the first is the +finding. + +**Guardrail.** One removal implementation, called by both routes. If they must +differ, the reset calls the per-actor removal in a loop rather than reimplementing +a subset of it. + +--- + +### MG-05 - Two removal semantics for one logical resource + +**Mechanism.** The same kind of resource is revoked differently depending on which +field granted it: one path clears immediately, the other defers removal until the +resource finishes naturally. + +**Why it is silent.** Both are legitimate APIs with legitimate uses. Deferred +removal is correct when an ability is mid-activation; immediate removal is +correct when tearing down. Neither errors. + +**Why the obvious check misses it.** Each call site is individually defensible, +and they are in different functions. The inconsistency is only visible if you ask +"what happens to a granted ability at deactivation?" and discover the answer is +"it depends which field you used", which is not a question the code invites. + +**Symptom.** Deactivation behaviour differs between two configurations that a +designer considers equivalent. Timing-dependent, so it reproduces inconsistently. + +**Detect.** List every revocation call for one resource kind and compare: + +```bash +rg -n "ClearAbility|SetRemoveAbilityOnEnd|RemoveActiveGameplayEffect|RemoveSpawnedAttribute" \ + --glob "*.cpp" . +``` + +Two different revocation calls for the same resource kind in one teardown path is +the finding. + +**Guardrail.** Pick one semantic per resource kind and document it at the grant +site. If both are genuinely needed, the choice belongs in the data, named, not in +which code path happened to grant it. + +--- + +### MG-06 - Removal that is not idempotent + +**Mechanism.** Resources come off by two independent routes — explicit reset at +deactivation, and reactive removal when the actor dies first. If removal assumes a +record exists, running it twice or in the wrong order misbehaves. + +**Why it is silent.** In the common ordering — feature deactivates while actors +are alive — only one route runs. The double-run needs an actor destroyed during +teardown, which is a shutdown-order accident. + +**Why the obvious check misses it.** Each route is correct in isolation. Nothing +in either function says "this may also be reached from the other path"; the +coupling is a property of the manager's event dispatch, which is engine code. + +**Symptom.** An assert or a stale entry during shutdown, in an ordering nobody can +reproduce on demand. + +**Detect.** Check that removal is guarded by record existence, and that reset +loops tolerate mutation: + +```bash +rg -n -A8 "::Remove\w+\(" --glob "*.cpp" . | rg "Find\(|Contains\(|if\s*\(" +rg -n -A6 "while\s*\(!\w+\.IsEmpty\(\)\)" --glob "*.cpp" . +``` + +The second pattern is the correct shape: a reset written as "while not empty, +take the top and remove it" survives a removal that mutates the same collection. +A plain range-for over a collection that removal mutates does not. + +**Guardrail.** Removal operates on the presence of a record, never on the +assumption of one. Write reset loops as drain loops. + +--- + +## Correctness of what gets attached + +### MG-07 - Index incremented inside the guard it is guarding + +**Mechanism.** A loop builds one deferred callback per configuration entry, +passing the entry's index as payload. The increment sits inside the branch that +skips invalid entries, so after any skip the payload index no longer matches the +list. + +**Why it is silent.** Every index remains in range and resolves to a real entry. +The callback applies a valid, well-formed configuration — just the wrong one. +There is no out-of-bounds access and nothing to assert on. + +**Why the obvious check misses it.** Reading the loop shows an increment that +appears to advance per iteration; it takes reading the brace structure to see it +is inside the guard. Editor-time validation flags the invalid entry that triggers +it, so the condition is believed impossible — but that validation is an editor +check, not a runtime guarantee, and it does not run on cooked data. + +**Symptom.** Actors receive another entry's abilities. Because both sets are +valid, this presents as a design or data error and gets investigated in the +content. + +**Detect.** Find counters incremented inside a conditional within a loop: + +```bash +rg -n -B8 "\b\w*Index\+\+|\+\+\w*Index" --glob "*.cpp" . | rg "if\s*\(|continue" +``` + +Read each hit's brace structure. In the measured reference the increment sat +inside a null-class guard, six lines below the `if` that opened it. + +**Guardrail.** Derive the index from the loop itself rather than maintaining a +parallel counter. If a counter is needed, increment it in the loop header where +its relationship to the iteration is structural. + +--- + +### MG-08 - Soft reference resolved with a bare accessor and no fallback + +**Mechanism.** An action resolves a soft class or asset pointer with a +non-loading accessor. If the bundle did not load it, the accessor returns null +and the entry is skipped. + +**Why it is silent.** The skip is unconditional and unlogged. A feature that +grants nothing looks identical to a feature whose grant list is empty, and both +are legal. + +**Why the obvious check misses it.** Every call site is correct given the +precondition "bundles are loaded", which is true in the editor, where everything +is loaded. The precondition is established in a completely different subsystem — +asset bundle configuration — and the code that depends on it does not state that +it does. + +**Symptom.** A packaged build silently omits abilities, widgets or mappings that +work in the editor. Bisecting finds nothing, because no code changed. + +**Detect.** Compare loading resolutions against non-loading ones, and check for a +fallback: + +```bash +rg -c "LoadSynchronous" --glob "*.cpp" . +rg -n "\.Get\(\)" --glob "*.cpp" . -A2 | rg -v "ensure|check|if\s*\(|UE_LOG" +``` + +In the measured reference the feature-action directory used three synchronous +loads against ten bare accessors, none of the latter guarded or logged. Both +numbers matter: the loads are a runtime hitch, the bare accessors are a silent +omission, and they appear in the same files. + +**Guardrail.** Either load, or assert, or log — never skip silently. A soft +reference that is expected to be bundle-loaded should assert that it was. + +--- + +### MG-09 - Unreachable code after a return + +**Mechanism.** A validation function returns its result and is followed by a +second, unreachable return. + +**Why it is silent.** It is genuinely harmless. The compiler may not even warn. + +**Why the obvious check misses it.** It is invisible to behaviour-driven review +because it has no behaviour. It is included here for one reason: **it is evidence +about the file rather than a defect in it.** Dead code after a return means +someone edited this validation and did not read what followed. Treat it as a +marker for where attention lapsed, and read the surrounding function more +carefully than you otherwise would. + +**Symptom.** None. Its value is diagnostic. + +**Detect.** + +```bash +rg -n -A3 "^\s*return \w+;\s*$" --glob "*.cpp" . | rg -A1 "^\s*$" | rg "return" +``` + +**Guardrail.** Enable and honour unreachable-code warnings. Where one appears, +re-read the function rather than only deleting the line. + +--- + +## Context and scope + +### MG-10 - Action state in plain members rather than per-context + +**Mechanism.** One action object can be active in several state-change contexts +simultaneously — client and server in a single process is the ordinary case. +State kept in plain member fields is shared between them. + +**Why it is silent.** With one context, which is every standalone run, the +behaviour is correct. The corruption requires two contexts, and the second +context's writes look like legitimate updates. + +**Why the obvious check misses it.** Member fields on an action object are the +obvious place to keep state, and nothing in the base class signature suggests +otherwise. The requirement is transmitted only by convention — a context +parameter that every lifecycle method receives and that a single-context +implementation can safely ignore. + +**Symptom.** One process's client and server tread on each other's tracking data. +Teardown removes the wrong context's resources, or removes them twice. + +**Detect.** Find action state that is not keyed by context: + +```bash +rg -n -A12 "class \w*GameFeatureAction\w*" --glob "*.h" . \ + | rg "^\s*(TArray|TMap|TSet|TSharedPtr|bool|int32)\s+\w+" \ + | rg -v "FGameFeatureStateChangeContext" +``` + +Any mutable member that is not inside a per-context map is a candidate. The +correct shape — a per-context struct in a map keyed by the change context — +appears in four separate actions in the measured reference and is worth copying +verbatim. + +**Guardrail.** All action state lives in a per-context map. Assert at activation +that the context's slot is empty, and reset it if not — that assertion is what +turns a silent leak into a visible one. + +--- + +### MG-11 - Activation reference count compiled out of shipping + +**Mechanism.** A count that prevents one owner from deactivating a plugin another +still needs is wrapped in an editor-only guard, and additionally checked against +an editor runtime flag inside. + +**Why it is silent.** In the editor, where multiple worlds share a process, the +count is present and correct. In a shipping build there is one world per process, +so the case it protects against does not arise — until it does. + +**Why the obvious check misses it.** The mechanism exists, is correct, and is +tested every time anyone uses multi-world play-in-editor. Discovering that it does +not ship means reading the preprocessor guard around it, which is at the top of a +file nobody opens while reasoning about deactivation. + +**Symptom.** No symptom in a single-world build. A latent hazard the moment any +build hosts two experiences, which is exactly when a project starts caring about +modularity. + +**Detect.** + +```bash +rg -n -B4 -A10 "RequestToDeactivate|NotifyOfPluginActivation|ActivationCount" \ + --glob "*.cpp" . | rg "#if|GIsEditor|WITH_EDITOR" +``` + +A reference count inside an editor guard, in code that ships, is the finding. + +**Guardrail.** Reference counting is a correctness mechanism, not a development +convenience. If it only exists in the editor, say so where it is declared and +state what protects shipping instead. + +--- + +### MG-12 - Plugins leak enabled between experiences + +**Mechanism.** Experience teardown deactivates the plugins it listed, but a switch +between experiences is not implemented as a diff — so plugins required by the old +experience and not the new one may stay enabled. + +**Why it is silent.** An enabled plugin that nothing uses has no symptom. Its +content stays resident and its actions have already been deactivated, so the +observable state is correct. + +**Why the obvious check misses it.** Deactivation exists and works for the case it +was written for: ending play. Experience-to-experience switching is a different +lifecycle that reuses the same teardown, and nothing in that teardown announces +its scope. + +**Symptom.** Memory that does not return to baseline across mode changes, and a +second experience that inherits capabilities from the first. + +**Detect.** Look for the diff, and for the authors' own admissions: + +```bash +rg -n -i "//\s*@?TODO.*(leak|deactivat|unload|experience)" --glob "*.cpp" . +rg -n "GameFeaturePluginURLs" --glob "*.cpp" . -B4 -A8 | rg "Difference|Intersect|Diff" +``` + +The first search matters more than the second. In the measured reference the +experience manager carries an eight-item block of the authors' own TODOs, one of +which states plainly that plugins are leaked enabled and that diffing requirements +is the intended fix. **A defect the authors documented is the cheapest one you +will ever find** — read those blocks first in any codebase you are adopting. + +**Guardrail.** Implement the requirement diff, or define experience changes as +map-travel-only and write that decision down. An undefined middle is what leaks. + +--- + +### MG-13 - Editor loads both role bundles + +**Mechanism.** The loading policy loads client data when not a dedicated server +and server data when not a client-only build. In the editor neither condition +excludes, so both load. + +**Why it is silent.** It is intentional and it is correct — the editor must be +able to run either role. Nothing is wrong; the number it produces simply describes +no shipping configuration. + +**Why the obvious check misses it.** Profiling produces a figure, and figures from +a profiler carry an authority their provenance does not. The condition that +inflates it is two boolean expressions in a policy class, in a different +subsystem from the one being measured. + +**Symptom.** Memory and load-time budgets built on editor measurements, wrong in a +direction that feels safe until a platform limit is real. + +**Detect.** + +```bash +rg -n -A6 "GetGameFeatureLoadingMode|bLoadClientData|bLoadServerData" --glob "*.cpp" . +``` + +If neither flag excludes in the editor, both bundle sets load. In the measured +reference the authors noted the resulting editor hitching in a comment beside it. + +**Guardrail.** Editor asset measurements are directional only. Any absolute budget +requires a packaged build of the target role, and any budget quoted without one +carries that caveat in writing. + +--- + +## Timing and readiness + +### MG-14 - A delay with no named dependency + +**Mechanism.** Selection or initialisation is deferred by a tick to let something +else finish first. + +**Why it is silent.** It works. One frame is enough today, on this machine, for +this content set. + +**Why the obvious check misses it.** A next-tick deferral is a common, accepted +idiom, and the code around it is correct. Whether it is legitimate depends +entirely on whether the thing being waited for is *named* — and an unnamed wait +looks identical to a named one. + +**Symptom.** Initialisation order breaks on a slower machine, a larger content +set, or after an unrelated subsystem changes its own timing. The failure appears +far from the deferral. + +**Detect.** Find deferrals and check for a stated prerequisite: + +```bash +rg -n -B6 "SetTimerForNextTick|GetTimerManager\(\)\.SetTimer" --glob "*.cpp" . \ + | rg -i "//|comment|wait|until|so that" +``` + +A deferral whose surrounding comment names what it waits for is debt with a +receipt. One with no comment is debt with a timer wrapped around it. In the +measured reference the one-frame delay before experience selection *is* +documented — it waits for startup settings — which is the correct form. + +**Guardrail.** Every delay names its prerequisite in a comment at the call site, +and is replaced by an explicit readiness signal when one becomes available. + +--- + +### MG-15 - Handler attached to the generic extension event + +**Mechanism.** An action registers for "extension added", which fires when the +manager first learns of the actor — before its pawn data, ability system, +controller or input exist. + +**Why it is silent.** The handler runs, finds what it needs missing, and returns +without acting. Returning early on a missing prerequisite is defensive +programming, and looks like it. + +**Why the obvious check misses it.** The registration is correct, the handler is +correct, and the early return is correct. Nothing is wrong at any single point. +The feature simply never applies, because the one event that would have carried it +arrived too early and no later event re-triggers it. + +**Symptom.** A feature that works when activated after the actor exists and does +nothing when activated before, or the reverse — order-dependent, and the order +differs between the editor and a packaged run. + +**Detect.** List which events each handler responds to: + +```bash +rg -n -A12 "HandleActorExtension|FExtensionHandlerDelegate" --glob "*.cpp" . \ + | rg "EventName ==|NAME_" +``` + +A handler that responds only to the generic added/removed events, with no +semantic ready event, is the finding. The correct shape in the measured reference +dispatches on both: the generic event *and* a named readiness event such as +"abilities ready" or "bind inputs now", so whichever arrives second does the work. + +**Guardrail.** Handlers accept both the generic event and a semantic readiness +event, making activation order irrelevant. Emitting a new readiness event costs a +name and one call — it requires no change to the feature infrastructure. + +--- + +### MG-16 - Receiver registration written by hand and left incomplete + +**Mechanism.** Actors register with the component manager in three places: add +receiver early, send a ready event at begin play, remove receiver at end play. +Classes inheriting a modular base get all three. Classes that do not inherit it +have them written by hand. + +**Why it is silent.** Missing the removal leaks a registration whose owner is +gone. Weak references prevent a crash; the record remains. Missing the ready +event means features silently never attach to that actor class. + +**Why the obvious check misses it.** The class works. It renders, it ticks, it +does its job. The two or three lines that make it visible to features are +infrastructure boilerplate, and boilerplate absent from a file looks like +boilerplate that was not needed there. + +**Symptom.** A feature attaches to every actor class except one, with no error — +or registrations that accumulate across level transitions. + +**Detect.** Compare the three calls per class: + +```bash +rg -l "AddGameFrameworkComponentReceiver" --glob "*.cpp" . > /tmp/add +rg -l "RemoveGameFrameworkComponentReceiver" --glob "*.cpp" . > /tmp/rem +comm -23 <(sort /tmp/add) <(sort /tmp/rem) +``` + +Any file that adds and never removes is the finding. In the measured reference +one HUD class inherits a non-modular engine base and therefore carries all three +calls by hand — correctly, but by hand, which is the fragile arrangement. + +**Guardrail.** Put the three calls in a base class and inherit it. Where that is +impossible, add a test that asserts the receiver count returns to baseline after a +level transition. + +--- + +## Composition + +### MG-17 - Handle dropped without realising it was the unsubscribe + +**Mechanism.** Extension handler and component requests return reference-counted +handles. There is no explicit removal API — releasing the last reference *is* the +removal. + +**Why it is silent.** Keeping a handle costs nothing visible, and dropping one +produces no event. Both directions of the mistake are quiet: holding too long +leaks a subscription, releasing too early silently detaches a live feature. + +**Why the obvious check misses it.** Nothing in the calling code says +"unsubscribe". Reviewing teardown for removal calls finds none — correctly, because +there are none — and it takes knowing the ownership model to recognise that +emptying an array of handles is the teardown. + +**Symptom.** Either features that cannot be removed, or features that detach when +an unrelated container is cleared. + +**Detect.** Confirm every handle-producing call has a stored destination: + +```bash +rg -n "AddExtensionHandler|AddComponentRequest" --glob "*.cpp" . -A3 \ + | rg -v "\.Add\(|=\s*\w+\.Add" +``` + +A call whose result is not stored is a subscription that ends immediately. + +**Guardrail.** Name the container for what it does — a handle array is an +ownership ledger, not a cache — and comment at its declaration that clearing it is +the unsubscribe operation. + +--- + +### MG-18 - Plugin activation order assumed + +**Mechanism.** Plugin URLs from an experience and from each of its action sets +merge into one list and activate in parallel. Action *execution* order is +deterministic; plugin *activation* order is not. + +**Why it is silent.** Parallel activation usually completes before anything +depends on the result, and when it does not, the dependent code has its own +readiness check that eventually passes. + +**Why the obvious check misses it.** Action ordering *is* deterministic and +documented — experience actions first, then action sets in order. Having verified +that, it is natural to assume the same of plugins, and nothing contradicts it. + +**Symptom.** A feature that depends on another plugin's registration works +consistently until content or load timing changes. + +**Detect.** Check whether the plugin list is loaded as a batch: + +```bash +rg -n -A10 "GameFeaturePluginURLs" --glob "*.cpp" . | rg "for\s*\(|LoadGameFeaturePlugin" +``` + +A loop issuing loads without sequencing between them means order is not +guaranteed. + +**Guardrail.** Express cross-plugin dependencies as readiness states rather than +activation order. If ordering is genuinely required, sequence the loads and say +why in the code. + +--- + +## Failure handling + +### MG-19 - Asynchronous deactivation counted but not honoured + +**Mechanism.** The deactivation context supports pausers for asynchronous +teardown. The implementation counts them and, if any are outstanding, logs an +error and proceeds. + +**Why it is silent.** It is not entirely silent — it logs. But it logs and +*continues*, so teardown completes and the sequence appears successful. A log line +during shutdown competes with everything else logged during shutdown. + +**Why the obvious check misses it.** The support looks present: the context is +constructed with a pauser callback, the count is tracked, the branch exists. +Reviewing for "is async deactivation handled?" finds all the machinery. Reading +what the branch *does* is a separate step. + +**Symptom.** An action with asynchronous teardown is cut short. Whatever it was +waiting to release stays held, and the failure is attributed to whichever +subsystem later notices the leak. + +**Detect.** Find pauser handling and read the branch: + +```bash +rg -n -A8 "NumExpectedPausers|DeactivatingContext" --glob "*.cpp" . \ + | rg "UE_LOG|Error|return|check" +``` + +An error log with no wait and no abort is the finding. The measured reference +states its own limitation in the log text — that asynchronous deactivation is not +fully supported — which makes it a documented gap rather than an unknown one. + +**Guardrail.** Either wait for pausers before completing teardown, or reject +actions that request asynchronous deactivation at validation time. Logging and +continuing converts a design gap into a runtime one. diff --git a/plugins/ue-design-skills/skills/ue-modular-gameplay/references/patterns.md b/plugins/ue-design-skills/skills/ue-modular-gameplay/references/patterns.md new file mode 100644 index 0000000..9d8bde2 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-modular-gameplay/references/patterns.md @@ -0,0 +1,302 @@ +# Patterns: a feature-action layer read end to end + +A worked reading of one real modular gameplay implementation: how a feature +attaches behaviour to actors it does not own, and what happens when it detaches. + +Individual defects are in [failure-modes.md](failure-modes.md), one entry each, +with a recipe. This file is about the shapes worth copying, and about one +finding that only becomes visible when you count. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from measured +facts; **[open]** — not answerable from the available source, and left open. + +**A scope caveat that shapes everything below.** The engine-side component +manager, request handle, feature action base and feature subsystem are not part +of the audited project — they live in engine plugins that were not in the tree. +Every statement about them here is inferred from call sites **[derived]**, and is +marked so. This is the honest version of a common situation: you can audit how a +project *uses* an engine subsystem far more cheaply than you can audit the +subsystem, and the usage is where your defects will be anyway. + +--- + +## 1. The core idea: the action does not look for actors + +The mechanism that makes this architecture work is one inversion. + +A feature action does **not** scan the world for actors to modify. It registers +an extension handler for a target *class* with a component manager, and the +manager invokes that handler for every matching receiver — the ones that already +exist and the ones that appear later **[measured]**. + +```text +receiver actor registers itself with the manager +action registers an extension handler for a class +manager invokes the handler for existing AND future receivers +returned request handle owns the subscription +``` + +The receiver side is three calls, and in the audited project they live in modular +base classes rather than in gameplay code **[measured]**: add receiver in +pre-initialize, send a ready event at begin play, remove receiver at end play. + +**[derived]** The consequence is the whole point of the architecture: base +gameplay classes register as receivers and **know nothing about features at all**. +Nothing in the character, player state or game mode references the feature system. +A feature plugin can be added to the project without touching any of them. + +The audited project contains one instructive exception. A HUD class inherits a +non-modular engine base, so those three calls are written out by hand in it +**[measured]**. Correct — and fragile in a way the inherited version is not, +because the third call is easy to omit and its omission has no symptom until +something counts registrations. Recipe: MG-16. + +--- + +## 2. Ownership is a reference count, and nothing says so + +Handler registration and component requests return shared handles **[measured]**. +There is no explicit unregister call anywhere in the audited code +**[measured]** — the only way to detach is to release the last reference. + +Every reset path therefore begins by emptying an array of handles **[measured]**, +and that line *is* the unsubscribe operation even though it reads as ordinary +container cleanup. + +**[derived]** Two failure directions follow, and both are quiet. Hold a handle too +long and the subscription leaks. Clear the array early and a live feature detaches +with no event. Reviewing teardown for "removal calls" finds none — correctly, +because there are none — so the reviewer must already know the ownership model to +see that teardown is present. Recipe: MG-17. + +One detail in the audited code confirms the model from the inside: when a +component already exists, the action still issues a request for it in some cases, +with a comment explaining that requests are reference counted **[measured]**. +**[derived]** That is a project telling you, in its own code, how the engine +subsystem it depends on behaves — which is the best evidence available when the +subsystem's source is not in the tree. + +--- + +## 3. Per-context state, and why it is not optional + +One action object can be active in several state-change contexts at once — client +and server in a single process being the ordinary case. Every lifecycle method +receives a context, and four separate actions in the audited project key their +state by it **[measured]**: + +```cpp +struct FPerContextData { /* tracked resources */ }; +TMap ContextData; +``` + +This is the pattern to copy verbatim. **[derived]** State in plain member fields +works perfectly with one context and corrupts silently with two, and the second +context's writes look like legitimate updates. + +Two supporting habits appear consistently and are worth copying with it +**[measured]**: + +- at activation, `ensure` that this context's slot is empty and force a reset if + not, **before** calling the base implementation — an assertion that converts a + leak into a visible failure; +- at deactivation, call the base implementation **first**, then reset the + context's data. + +The base class itself keeps its delegate handles in a map keyed the same way +**[measured]**, which is how it supports being active in two contexts at once. + +Recipe: MG-10. + +--- + +## 4. Semantic readiness events + +A generic "extension added" event means only that the manager knows about the +actor. **[derived]** It does not mean the actor has its pawn data, ability system, +controller, or initialised input — so a handler attached to that event alone will +find its prerequisites missing, return early, and never run again. + +The audited project solves this by emitting its own named events at the real +dependency boundaries — abilities ready, bind inputs now **[measured]** — and by +having every handler dispatch on **both** the generic event and the semantic one +**[measured]**: + +```text +extension removed / receiver removed -> remove +extension added / -> add +``` + +**[derived]** Whichever arrives second does the work, so activation order stops +mattering. This is the single most transferable idea in the file, and it costs a +`static const FName` plus one call to emit — no change to the feature +infrastructure is required. Recipe: MG-15. + +--- + +## 5. The finding that only appears when you count + +Eight actions were audited. Five have incomplete teardown **[measured]**: + +| Action | Attaches | Detaches | +|---|---|---| +| add abilities | grants abilities, attribute sets, ability sets; stores every handle | mostly correct — but two revocation semantics for one resource kind, depending on which field granted it | +| add widgets | pushes layouts, registers extensions, stores handles | reset unregisters extensions and does **not** deactivate layouts; the per-actor path does both | +| add input binding | delegates to a game component | the component's removal method has an empty body | +| add input mapping context | adds contexts to the input subsystem | tracks affected controllers in a container **nothing ever writes to** | +| splitscreen config | votes to disable | correct — decrements the vote | + +**[derived]** Read one at a time, each looks like an isolated oversight. Read +together, they are one pattern: **the add path is complete in all eight, and the +remove path is complete in three.** That ratio is the argument for the ownership +ledger in the skill — not a rule someone invented, but the shape the evidence +takes. + +Two of these deserve their detail, because the detail is what makes them +recognisable elsewhere. + +### The container nobody fills + +An action declares a container of the controllers it has affected. An `ensure` at +activation asserts it starts empty. The reset loop drains it correctly. The +reactive path removes from it. **Nothing adds to it** — the function that would +takes the per-context data as a parameter and ignores it **[measured]**. + +Four of five operations exist, so any search for "is this container used?" +answers yes emphatically. The one missing operation is the producer. + +Two things make this the clearest entry in the file. First, a **sibling action in +the same directory** does it correctly, with `AddUnique` in exactly the place the +broken one omits **[measured]** — so the two files diff against each other. +Second, the original authors left a comment on the broken line stating that the +container is never modified **[measured]**. A known defect that shipped. Recipe: +MG-03. + +### The rollback that cannot be written + +The input-binding action's teardown iterates its tracked pawns and calls a removal +method on the game component. That method's body is a single TODO comment +**[measured]**. + +But the cause is one level further down: the function that creates the bindings +declares the handle array as a **local variable** and lets it go out of scope +**[measured]**, at two separate call sites. + +**[derived]** So the removal is not merely unimplemented, it is impossible to +implement without changing the component — the information it would need was +discarded at creation. An estimate of "implement the missing rollback" made from +reading the empty function comes out an order of magnitude low. Recipes: MG-01, +MG-02. + +--- + +## 6. What the reference got right + +Stated deliberately, because the section above is not an argument against the +architecture: + +1. **Actions do not search for actors.** The extension handler protocol handles + existing and future receivers uniformly **[measured]**. +2. **Per-context state in all four attaching actions** **[measured]**, with the + activation-time emptiness assertion. +3. **Reset loops written as drain loops** — "while not empty, take the top, + remove it" **[measured]** — which survives a removal that mutates the same + collection. A range-for would not. +4. **Two removal routes, both idempotent**: explicit reset at deactivation, and + reactive removal when the actor dies first, keyed on whether a record exists + **[measured]**. +5. **Composition enforced over inheritance.** Validation rejects deriving one + scripted experience from another and points the author at action sets + **[measured]**. That is a design rule with a compiler behind it. +6. **Action sets are separate primary assets** with their own bundle data + **[measured]**, which is what allows them to be loaded alongside the experience + rather than through it. +7. **A one-frame deferral with a named reason** **[measured]** — it waits for + startup settings. A delay whose prerequisite is named is debt with a receipt; + an unnamed one is debt with a timer around it. Recipe: MG-14. +8. **A policy class as the place for pre-activation work.** Logic that must run + before activation, is process-global rather than per-world, or needs to see all + of a plugin's actions at once, belongs in a lifecycle observer rather than in an + action **[measured]**. One action in the audited set is pure data with a getter, + executed entirely by such an observer **[measured]** — a clean separation worth + copying. + +--- + +## 7. The authors' own TODO block + +The experience manager carries a block of eight TODOs from its authors +**[measured]**, covering: asynchronous experience loading; explicit error +handling instead of assertions; running actions in phases rather than all at +once; support for deactivating an experience; and — stated plainly — that plugins +are **leaked enabled** between experiences, with diffing requirements named as the +intended fix. + +**[derived]** This is the cheapest finding in any audit, and the most reliable. +The authors know; they wrote it down; nobody adopting the code reads it, because +adoption reads the class list rather than the comments. + +Two practical consequences: + +- **Read the TODO blocks first** in any codebase you are adopting. They are a + defect list written by the people best placed to write one. +- Treat "the authors documented this gap" as *stronger* evidence than a static + finding of your own, not weaker. It removes the possibility that you have + misread the intent. + +Recipe: MG-12. The same reading also surfaces asynchronous deactivation: pausers +are counted, and when any are outstanding the code logs an error and proceeds +**[measured]**. The limitation is stated in the log text itself. Recipe: MG-19. + +--- + +## 8. Editor loading is not shipping loading + +The loading policy loads client data when not a dedicated server, and server data +when not a client-only build **[measured]**. **[derived]** In the editor neither +condition excludes, so both bundle sets load — and the authors noted the resulting +hitching in a comment beside it **[measured]**. + +Every in-editor memory and load-time figure therefore describes a configuration +that ships to nobody. Directionally useful, absolutely wrong. Recipe: MG-13. + +--- + +## 9. What source reading could not settle + +- **The engine subsystems were not in the tree.** Component manager, request + handle, feature action base, feature subsystem, and the add-components action + **[open]**. Everything above about them is inferred from call sites. +- **Binary assets were not read.** Which actions the shipped feature plugins + actually configure, and with what data, is **[open]**. What *is* measured: those + plugins contain no feature-action C++ of their own, so they use only actions + from the base module and the engine **[measured]**. +- **One teardown question is genuinely undecidable from this tree.** Whether + clearing the handle array causes the manager to dispatch removal events — which + would make the narrower reset path in MG-04 harmless — depends on engine code + that was not available **[open]**. It is recorded as an open question rather than + resolved in either direction, because a plausible answer either way would be a + guess. + +That last item is worth its space. It would have been easy to state MG-04 as a +confirmed defect; the honest form is "this is a defect unless an engine behaviour +we could not read compensates for it", and a reader adopting the code can settle +it in ten minutes with the engine source that they have and the audit did not. + +--- + +## Provenance + +Measured against the feature-action layer of Epic's Lyra Starter Game on Unreal +Engine 5.6, read as source in a single workspace. Source addresses stay in the +research archive that produced this skill; each `MG-` identifier resolves back to +the audited location there, so any specific claim above can be produced on +request. + +## Evidence boundary + +One project, one engine version, one workspace, with the engine-side subsystems +outside the readable tree. These are examples and failure evidence, not +guarantees about other versions. Claims marked open stay open. Re-run the recipes +in [failure-modes.md](failure-modes.md) against your own tree before acting on +anything here. diff --git a/plugins/ue-design-skills/skills/ue-multiplayer-authority/SKILL.md b/plugins/ue-design-skills/skills/ue-multiplayer-authority/SKILL.md new file mode 100644 index 0000000..d45c5bd --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-multiplayer-authority/SKILL.md @@ -0,0 +1,390 @@ +--- +name: ue-multiplayer-authority +description: >- + Design and review the authority model of a networked Unreal Engine product: + which side owns each fact, what is predicted versus replicated versus + validated, where responsibility boundaries run between GameState, PlayerState, + Controller, Pawn and components, RPC and validation discipline, hit + registration and lag compensation, replication modes, relevancy and cost, and + the cheat surface. Use when starting a multiplayer feature, porting + single-player code to network, reviewing trust boundaries, or investigating + desync, rubber-banding, "works in PIE but not on dedicated server", or + suspected cheating. +--- + +# UE multiplayer authority + +The invariant: + +> Every fact in a networked game has exactly one owner. Authority is not a +> property of code that says `HasAuthority` — it is the guarantee that no other +> machine can produce that fact. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Twenty detection recipes for silent trust failures: +[failure modes](references/failure-modes.md). + +Read the failure modes before adopting or reviewing a networked design. Each +entry names the mechanism, why it stays silent, why the obvious check misses it, +and a command you can run against your own tree. + +Related skills: `ue-gas-architecture`, `ue-gameplay-messaging`, +`ue-modular-gameplay`, `ue-cosmetics-and-teams`, `ue-architecture-guardrails`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Start from one question + +Before any code: **for each fact in the game, who is allowed to be wrong about +it?** + +- If a client being wrong is a cosmetic glitch → the client may own it. +- If a client being wrong is an advantage → the server must own it, and must be + able to produce the value independently. +- If a client being wrong is unrecoverable → the server must own it *and* + persist it. + +This question, answered per fact and written down, is the authority model. +Everything else in this document is enforcement of that table. + +### The three states of a networked fact + +| State | Meaning | Who can be wrong | +|---|---|---| +| **Predicted** | Client computes it immediately for responsiveness; server may overrule | Client, temporarily, visibly | +| **Replicated** | Server computes it, clients receive it | Nobody — clients only observe | +| **Validated** | Client proposes it, server independently verifies before accepting | Client may lie; server must detect | + +**The third is the expensive one, and it is the one that is usually missing.** +"Replicated" is often mistaken for "validated": a value that travelled from the +client to the server and was then replicated back out to everyone is *client +authored*, no matter how authoritative it looks on arrival. + +--- + +## 2. What actually enforces authority + +Unreal offers several mechanisms that look similar and enforce very different +amounts. Know which is which before relying on one. + +| Mechanism | Enforces | Does not enforce | +|---|---|---| +| `HasAuthority()` guard | that this code path runs on the server | anything about callers that skip the guard | +| `UFUNCTION(Server)` | routing to the server | that the arguments are truthful | +| `WithValidation` | that a `_Validate` function runs and can disconnect | anything, if `_Validate` returns `true` unconditionally | +| `BlueprintAuthorityOnly` | that **Blueprint** cannot call this | nothing at all for C++ callers | +| replicated property | one-way flow server → client | that the server computed the value itself | +| `GetLifetimeReplicatedProps` conditions | who receives it | who may change it | + +### Rules + +1. **A public mutator on a replicated fact must guard itself**, not rely on its + callers. `if (!HasAuthority()) return;` at the top of the function, not at the + call site — call sites multiply. +2. **`WithValidation` with an empty `_Validate` is worse than none**: it passes + review as validation while enforcing nothing. +3. **`BlueprintAuthorityOnly` is a Blueprint editor affordance.** Never cite it as + a security property. +4. **Write the guard even when the only current caller is the server.** The guard + is documentation the compiler enforces; the caller list is not stable. + +--- + +## 3. Responsibility boundaries by layer + +Put each fact in exactly one place. The most common source of networked bugs is a +fact that lives in two. + +| Layer | Exists on | Owns | Must not own | +|---|---|---|---| +| **GameMode** | server only | match rules, spawning, scoring decisions | anything a client must read | +| **GameState** | server + all clients | match-wide replicated facts: phase, timer, team scores | per-player secrets | +| **PlayerState** | server + all clients | per-player facts everyone may see: name, team, score, abilities | input, camera, UI state | +| **PlayerController** | server + owning client | input, camera, client-directed commands, UI ownership | facts other players need | +| **Pawn / Character** | server + all clients | embodiment: transform, movement, physical state | player identity, persistent progression | +| **Components** | follow their owner | one cohesive slice of the owner's facts | facts belonging to a different layer | + +### Rules + +1. **`GameMode` exists only on the server.** Any client code path referencing it + is a bug that PIE will hide. +2. **Persistent player facts belong to `PlayerState`, not the Pawn.** Pawns die. +3. **Input and camera belong to the Controller.** A Pawn reading input directly + cannot be possessed by AI or spectated. +4. **A component's authority is its owner's authority.** A component attached to a + client-owned actor cannot be authoritative regardless of its internal guards. +5. **Simulated proxies have no controller.** Any logic gated on "has a controller" + silently excludes them; any logic that assumes one will null-deref on them. + +--- + +## 4. RPC discipline + +### Choosing the direction + +```text +Client -> Server UFUNCTION(Server, Reliable/Unreliable, WithValidation) +Server -> owner UFUNCTION(Client, ...) +Server -> everyone UFUNCTION(NetMulticast, ...) +``` + +### Rules + +1. **Every `Server` RPC carries `WithValidation` and a `_Validate` that actually + checks.** Range, ownership, state legality, rate. An RPC without validation is + a client-side API into your server. +2. **Reliable is a budget, not a default.** Reliable RPCs are ordered and queued; + overflowing the queue disconnects the client. Anything cosmetic, frequent or + idempotent should be unreliable. +3. **Never multicast what a replicated property already carries.** The property + handles joiners and relevancy; a multicast does not. +4. **Multicasts do not reach late joiners or newly relevant clients.** Anything a + late joiner must know is state, not an event. +5. **Validate on the server that the sender owns the thing it is acting on.** + Routing to the server proves nothing about the *subject* of the call. +6. **Rate-limit anything a client can call in a loop.** + +### Debug and cheat commands + +Cheat entry points reachable from a client are the highest-value target in the +codebase. They must be compiled out of shipping builds — not merely hidden, +disabled by a flag, or guarded by UI. A `Server` RPC that executes an arbitrary +cheat string, present in a shipping binary, is a remote console. + +--- + +## 5. Prediction + +Prediction exists to hide latency, not to save server work. The server must still +compute everything it accepts. + +### Safe to predict + +- movement through the engine's own prediction (correction path already exists); +- animation, audio, VFX, camera; +- UI affordances (cooldown wheels, ammo counters) that a correction can restore; +- ability activation under a prediction key, where the server can reject. + +### Never predict + +- damage application; +- score, currency, progression; +- inventory grants; +- death and elimination; +- anything persisted. + +### Rules + +1. **Every predicted value needs a defined correction path** — what happens + visually when the server disagrees. Undefined correction means the client + keeps a wrong value forever. +2. **Prediction is a client-side optimisation of a server-side truth.** If the + server cannot independently produce the value, the feature is not predicted, + it is client-authoritative. +3. **A prediction key is not validation.** It reconciles ordering; it does not + check whether the client was entitled to act. + +--- + +## 6. Hit registration — the decision that shapes the whole product + +This is where authority is usually lost, and it is a design decision that must be +made explicitly and early, because retrofitting is expensive. + +There are three positions: + +| Model | Server does | Cost | Cheat exposure | +|---|---|---|---| +| **Server-authoritative trace** | traces from its own state at the current tick | cheap | none, but shots feel like they miss under latency | +| **Lag-compensated rewind** | stores per-tick historical positions, rewinds to the client's timestamp, re-traces | expensive: memory, CPU, complexity | low; the standard for competitive shooters | +| **Client-reported hit** | accepts the client's hit result | free | total | + +### Rules + +1. **Choose consciously and write the choice down.** Most projects arrive at the + third by default, because the first feels bad and the second is work. +2. **If you accept client-reported hits, know exactly what a malicious client + gains** — typically arbitrary damage, arbitrary range, arbitrary target, + arbitrary surface multipliers. +3. **The server cannot validate what it cannot reproduce.** Without stored + historical positions there is nothing to validate a client's hit *against*, + so bolting on "validation" later is not a patch, it is building the rewind + subsystem. +4. **Sanity checks are not validation, but they are not worthless.** Maximum + range, line-of-sight from the shooter's replicated position, fire-rate limits, + team check, and "was the target ever near there recently" raise the cost of + trivial cheats even without full rewind. +5. **Never let the client choose damage-modifying inputs.** The surface/material + that multiplies damage, the distance used for falloff, and the trace origin + must all come from server-known state, even in a client-reported-hit model. + These are the cheapest things to move server-side and the most valuable. + +### Recognising a stub + +Watch for parameters threaded through many function signatures but never branched +on, and for booleans initialised to a constant with a `//@TODO`. These are the +attachment points of a rewind subsystem that was designed and never built. They +make the code *look* like it validates. Recipes: NA-06, NA-07 in the failure +modes. + +--- + +## 7. Replication cost and relevancy + +Authority decisions determine correctness; relevancy decisions determine whether +the product ships. + +### Rules + +1. **Pick an ability-system replication mode deliberately.** `Full` replicates all + gameplay effects to all clients; `Mixed` replicates effects to the owner and + tags/cues to everyone; `Minimal` replicates no effects to anyone. + Player-controlled actors want `Mixed`; AI want `Minimal`. +2. **Net cull distance is a gameplay decision, not a performance knob.** Setting + it beyond the map size disables relevancy culling entirely — every actor + replicates to every client, which is correct for a small arena and fatal for a + large map. +3. **Prefer replicated state over event spam.** State is idempotent, survives + packet loss, and reaches late joiners. +4. **FastArraySerializer for collections that change incrementally**, with the + `PreReplicatedRemove` / `PostReplicatedAdd` / `PostReplicatedChange` callbacks + all implemented — an empty callback is a client-side desync with no symptom on + the server. +5. **Replication graph before you need it, or never.** Enabling it late changes + which actors clients see, and every relevancy assumption in gameplay code has + to be re-verified. If it ships disabled, treat routing configuration as + untested. +6. **Cosmetic state should not replicate.** Replicate intent (which skin, which + team), realise presentation locally. + +--- + +## 8. Lifecycle across the network + +1. **Readiness is a state chain, not a timing assumption.** Actor spawned, + controller assigned, PlayerState replicated, data initialised, gameplay ready — + each an explicit state with explicit transitions. +2. **Client and server reach each state in a different order.** Any code that + assumes an order will work for the listen-server host and fail for remote + clients. +3. **A simulated proxy passes through the chain without ever acquiring a + controller.** Gates must handle that case explicitly rather than by omission. +4. **Every dynamic addition needs its exact removal**: granted abilities, applied + effects, registered listeners, spawned cues, added tag paths. Test the removal + on both server and client, in the same session, more than once. +5. **Death is not destruction.** Define the full sequence — death started, effects + removed, input released, cues cleaned, actor destroyed or respawned — and give + each step one owner. + +--- + +## 9. The cheat surface review + +Enumerate, for the shipping build, every path a malicious client can reach: + +```text +1. Every Server RPC -> does _Validate independently verify? +2. Every RPC argument -> is any of it trusted without recomputation? +3. Every cheat/debug entry point -> is it compiled out, not just disabled? +4. Every damage input -> which come from the client? +5. Every fail-open error path -> what does the code do when a check errors out? +6. Every client-side gate -> is there a matching server-side one? +``` + +Item 5 deserves attention: a permission check that returns "allowed" when it +cannot determine the answer converts an unknown state into an exploit. Failure to +determine must deny. + +--- + +## 10. Test matrix + +**PIE is not a network test.** A listen-server host shares memory with its client; +half of the failure modes here cannot occur in it. + +### Mandatory configurations + +- dedicated server + two remote clients; +- one client with 150 ms latency and 2% packet loss; +- a client joining mid-match; +- a client disconnecting mid-action; +- an actor becoming relevant and then irrelevant. + +### Per feature + +- does it work for a remote client, not just the host? +- does a late joiner reach the correct state? +- what does a simulated proxy see? +- what happens if the client sends the RPC twice, out of order, or with absurd + arguments? +- is the state correct after death and respawn, twice? + +### Authority regression + +- for each fact in the authority table, assert on the server that the client + cannot change it — with a test client that tries. + +--- + +## 11. Review checklist + +- [ ] Every fact has a named owner in a written authority table. +- [ ] Each fact is classified predicted / replicated / validated. +- [ ] Public mutators of replicated state guard themselves. +- [ ] Every `Server` RPC has `WithValidation` with a real `_Validate`. +- [ ] No trust in client-supplied damage inputs: origin, distance, surface, target. +- [ ] Hit registration model is chosen and documented. +- [ ] Cheat entry points are compiled out of shipping. +- [ ] Error paths deny rather than allow. +- [ ] Replication mode and net cull distance chosen per class, not inherited. +- [ ] FastArray callbacks all implemented. +- [ ] Every addition has a verified removal. +- [ ] Feature tested on a dedicated server with a remote client under latency. +- [ ] `BlueprintAuthorityOnly` is not cited as enforcement anywhere. + +--- + +## 12. When client authority is the right answer + +Deliberate client authority is legitimate when: + +- the fact is purely presentational; +- the game is cooperative or single-player-with-friends and the threat model is + "no adversary"; +- the cost of server validation exceeds the value of the protected fact; +- an anti-cheat layer outside the game handles the threat. + +In every one of these cases, the decision is written down with its threat model. +The failure worth avoiding is not "we trusted the client" — it is **"we did not +know we trusted the client."** + +That distinction is what this skill exists to preserve. A codebase can be an +excellent architectural reference and simultaneously an unsafe network reference; +those are separate axes, and copying the first without noticing the second is the +common accident. + +--- + +## Provenance + +The findings behind the failure modes come from a line-by-line source audit of +Epic's Lyra Starter Game on Unreal Engine 5.6, read rather than run. That project +is an outstanding architectural reference and a poor network-security reference; +those are independent axes, and the audit exists to keep them apart. Source +addresses stay in the research archive that produced this skill; what ships is the +detection recipe, because an address in someone else's tree is not something you +can act on and a recipe is. + +Each entry carries a stable identifier (`NA-01`, `NA-02`, …) that resolves back to +the audited location in that archive. If you need the original address to settle a +dispute, it exists and can be produced. + +## Evidence boundary + +The measured material refers to one specific reference project on one engine +version in one workspace. It is evidence of failure modes, not a guarantee about +other engine versions or other samples. Several claims in it are explicitly marked +as unanswerable without a dedicated-server build, and they stay unanswered. Re-run +the recipes against your own tree before trusting any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/failure-modes.md new file mode 100644 index 0000000..f29c779 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/failure-modes.md @@ -0,0 +1,764 @@ +# Failure modes: multiplayer authority + +Twenty ways a networked Unreal project loses authority without anything going +wrong on screen. Each entry names the mechanism, why it stays silent, why the +obvious check misses it, the symptom, a recipe you can run against your own tree, +and the guardrail. + +Two properties recur across this list and are worth stating once, because they +explain why these survive review: + +- **PIE cannot reproduce most of them.** A listen-server host shares memory with + its client, so the client/server split that produces the bug does not exist + during the test that would have caught it. Entries NA-01, NA-04, NA-11, NA-13, + NA-17 and NA-19 are invisible in PIE by construction. +- **Several of them look correct in review.** A `WithValidation` marker, a + `BlueprintAuthorityOnly` specifier, a `bIsSimulated` parameter threaded through + signatures — each reads as evidence of a server-side path that is not there. + The specifier is real; the enforcement is not. + +Recipes below use `rg`. They are written to be run from the root of a project's +source tree and to be read as a *comparison of two numbers*, not as a single +search: the whole class of defect here is "the thing is present, and the thing +that would make it mean something is absent". + +--- + +## Trust boundaries + +### NA-01 - Replicated is mistaken for validated + +**Mechanism.** A value travels client to server and is then replicated back out to +everyone. On arrival it lives in an authoritative-looking replicated property, and +every later reader treats it as a server fact. + +**Why it is silent.** Replication works perfectly. The value is consistent on every +machine, arrives in order, survives relevancy changes and late joins. Consistency +is not authorship, but consistency is what everyone observes and tests. + +**Why the obvious check misses it.** Review asks "is this replicated?" and the +answer is yes. There is no compiler concept, specifier or naming convention that +distinguishes a server-computed replicated value from a server-relayed one. Both +compile to the same declaration. + +**Symptom.** Nothing at all, until someone modifies a client. Then the value is +whatever that client chose, and it is authoritative across the whole session. + +**Detect.** This one is a table, not a grep. List every replicated property and +name, for each, the server-side computation that produced it: + +```bash +rg -n "UPROPERTY\([^)]*Replicated" Source/ | wc -l # facts to account for +rg -n "GetLifetimeReplicatedProps" Source/ # where they are declared +``` + +For each result, answer in writing: *what did the server compute to get this?* If +the answer is "the client sent it", the fact is client-authored. A property whose +only writer is an RPC handler that assigns the parameter directly is the shape to +look for. + +**Guardrail.** Classify every fact as predicted / replicated / validated in a +written table before the feature is built. "Validated" requires the server to +independently reproduce the value, not merely to receive it. + +--- + +### NA-02 - `WithValidation` with an empty `_Validate` + +**Mechanism.** The RPC carries the `WithValidation` specifier and the paired +`_Validate` function exists, so the engine calls it. Its body returns `true` +unconditionally. + +**Why it is silent.** The mechanism is fully wired and working. The engine really +does call the validator and really would disconnect a client that failed it. There +is simply no input it rejects, and a validator that accepts everything is +indistinguishable at runtime from one that has nothing to reject yet. + +**Why the obvious check misses it.** The specifier is what reviewers grep for, and +it is present. Declarations and implementations live in different files, so a +header review — the natural place to audit an RPC surface — never sees the body. + +**Symptom.** The RPC accepts anything. Review signs off on the subsystem as +validated, and the estimate for "add validation" is later scoped as a patch when +it is a design. + +**Detect.** Read every validator body, not the declarations: + +```bash +rg -n --stats "WithValidation" Source/ # declared validators +rg -n -A3 "_Validate\(" Source/ | rg -B1 "^\s*return true;" # bodies that accept all +``` + +Compare the counts. On the measured reference the second command returned two +sites, both belonging to cheat RPCs. Any validator whose body is a bare +`return true;` is an unimplemented validator regardless of the specifier. + +**Guardrail.** Require every `_Validate` body to check range, ownership, state +legality and rate. Treat a one-line `return true;` in a validator as a review +blocker, the same way an empty `catch` block is. + +--- + +### NA-03 - `BlueprintAuthorityOnly` cited as enforcement + +**Mechanism.** A function is marked `BlueprintAuthorityOnly`. The specifier stops +Blueprint graphs from calling it on a non-authoritative actor. C++ callers are +entirely unaffected. + +**Why it is silent.** In a project where the function is only ever called from +Blueprint, the specifier does exactly what everyone believes it does. The gap +opens the first time a C++ caller appears, and that caller works, so nothing draws +attention to it. + +**Why the obvious check misses it.** The word "authority" is in the specifier. +Grepping for authority-related terms finds it and it looks like a guard. It is an +editor affordance whose name describes intent rather than enforcement. + +**Symptom.** A function documented and reviewed as authority-guarded is reachable +from any C++ path, including client code, and mutates replicated state from there. + +**Detect.** Enumerate the sites and check each for a real guard in the body: + +```bash +rg -n "BlueprintAuthorityOnly" Source/ # claimed guards +rg -n -A6 "BlueprintAuthorityOnly" Source/ \ + | rg "HasAuthority|GetLocalRole" # actual guards +``` + +A large gap between the two counts is the finding. On the measured reference the +specifier appeared at five sites across three subsystems, with no `HasAuthority` +in the guarded bodies. + +**Guardrail.** Never cite `BlueprintAuthorityOnly` as a security property in a +review, a comment or a design document. Add an explicit `HasAuthority()` guard +inside the function body; keep the specifier only as an editor hint. + +--- + +### NA-04 - Guard placed at the call site rather than in the mutator + +**Mechanism.** A public mutator of replicated state relies on its callers to check +authority. Today every caller does. + +**Why it is silent.** The invariant holds for as long as the caller list is what it +was when the code was written. Nothing degrades gradually; the code is correct +until one day it is not, and the transition is a new caller, not an edit to the +function. + +**Why the obvious check misses it.** Reading the function shows nothing wrong — +there is no missing line, because the check was never meant to be there. Reading +the call sites shows a check at each one. Both halves pass review independently. +The defect only exists in the relationship between them, which no single file +review inspects. + +**Symptom.** Works until a new caller appears — commonly a subclass in a feature +plugin, a Blueprint node, or a second caller added by someone who reasonably +assumed the mutator guards itself. Then a client writes replicated state directly. + +**Detect.** Find public or virtual writers of replicated state whose own body has +no guard: + +```bash +# 1. names of replicated properties +rg -n -o "UPROPERTY\([^)]*Replicated[^)]*\)\s*\n\s*\w+\s+(\w+)" -r '$1' Source/ +# 2. for each name, find assignments and check the enclosing function +rg -n -B12 "\bDeathState\s*=" Source/ | rg "HasAuthority|void |virtual" +``` + +Substitute your own property name in step 2. The finding is an assignment whose +preceding function signature is public or virtual and whose body contains no +authority check. + +**Guardrail.** `if (!HasAuthority()) return;` as the first line of the mutator. +Write it even when the only current caller is the server: the guard is +documentation the compiler enforces, and the caller list is not a stable property. + +--- + +### NA-05 - Client-supplied damage inputs + +**Mechanism.** Distance, trace origin, surface material or target identity arrive +in the RPC payload and feed the damage calculation directly. + +**Why it is silent.** Every value is in a plausible range, because a normal client +sends plausible values. The damage numbers that result look exactly like correct +damage numbers. There is no error state and no log line, because nothing failed. + +**Why the obvious check misses it.** The damage code is usually reviewed for +correctness of the formula, and the formula is correct. The question that matters +is not "does this compute the right number?" but "who chose the inputs?", and the +inputs arrive as ordinary function parameters several layers away from where they +entered the process. + +**Symptom.** Arbitrary damage multipliers from a modified client, invisible in +logs. In a competitive game this surfaces as player reports that a specific +weapon or surface "feels wrong", long before anyone suspects the trust boundary. + +**Detect.** Trace each damage-modifying input back to its origin: + +```bash +rg -n "PhysicalMaterial|PhysMat|TraceStart|HitLocation|Distance" \ + Source/ --glob "*Damage*" --glob "*Execution*" +``` + +For each hit, answer: did the server compute this, or did it arrive from the +client? On the measured reference the falloff distance and the surface multiplier +both came from client-supplied fields, while the team check beside them was +genuinely server-side — the model was incomplete, not absent, which is the harder +case to spot. + +**Guardrail.** Recompute every damage-modifying input from server-known state. The +shooter's replicated position and a server-side collision query cover most of it +and are cheap. Client-supplied values may inform presentation, never arithmetic +that changes an outcome. + +--- + +### NA-06 - Validation flag initialised to a constant + +**Mechanism.** A boolean local is initialised to a literal — `const bool bIsValid += true;` — and then consumed by branches that read as validation. There is no +computation between the initialisation and the use. + +**Why it is silent.** A permanently-valid result is a legal runtime state. Every +branch takes the path it would take on a successful validation, which is also the +path taken during all normal play, so the code behaves exactly as intended for +every non-malicious input. + +**Why the obvious check misses it.** The flag has a meaningful name, appears in +conditionals, and is used more than once. Any search for "is validation present?" +finds a variable called `bIsTargetDataValid` being consumed by a gate and stops +there. The tell is the *initialiser*, which is three lines above and looks like +boilerplate. + +**Symptom.** The code path reads as validated in review and in estimates. Nothing +is observable at runtime until a modified client sends implausible data and it is +accepted. + +**Detect.** Look for boolean locals initialised to a literal and later consumed as +gates, restricted to network-relevant code: + +```bash +rg -n "const bool b\w+\s*=\s*(true|false)\s*;" Source/ +``` + +Then, for each hit, check whether the name promises a computation +(`bIsValid`, `bCanApply`, `bIsAuthorized`) and whether the variable is read by a +branch. On the measured reference this returned eight sites; most were legitimate +call-parameter constants, and exactly one — a target-data validity flag inside a +weapon ability — was a validation stub. The recipe's value is that it reduces the +whole codebase to a list short enough to read. + +**Guardrail.** Treat any boolean local in a network path whose initialiser is a +literal and whose name asserts a property as unimplemented until proven otherwise. +If validation is not written yet, name the variable for what it is. + +--- + +### NA-07 - Dead parameters that imply a server path + +**Mechanism.** A parameter such as `bIsSimulated` is threaded through many function +signatures, passed a literal at the only call site, and never branched on +anywhere. + +**Why it is silent.** The parameter is passed and received correctly at every +level. Nothing is unused in the compiler's view, so no warning fires. The code +does what it does today with complete consistency. + +**Why the obvious check misses it.** Signature review is exactly the wrong +instrument here, and it is the instrument reviewers reach for when auditing a +subsystem's shape. A parameter appearing in ten signatures is strong evidence to a +human that a feature exists. Confirming it requires following the value to a +branch, which no IDE affordance does for you. + +**Symptom.** A reviewer concludes a simulated or authoritative path exists. +Estimates for "just add server validation" come out an order of magnitude low, +because the attachment points were mistaken for the subsystem. + +**Detect.** Count mentions, then count branches: + +```bash +P='bIsSimulated' +rg -n "\b$P\b" Source/ | wc -l # mentions: many +rg -n "if\s*\(\s*!?\s*$P\b|\?\s*.*:\s*.*\b$P\b" Source/ # branches: ? +``` + +On the measured reference the parameter appeared at ten sites and was branched on +at zero. A high first number with an empty second is a mounting point, not a +feature. + +**Guardrail.** Trace every network-relevant parameter to a branch. No branch means +no feature: either implement it or delete the parameter, because its presence is +an active lie to the next reader. + +--- + +### NA-08 - Retrofitting hit validation without rewind + +**Mechanism.** The server holds no historical positions, so there is nothing to +validate a timestamped client hit against. Validation is scheduled as a task on +top of the existing weapon code. + +**Why it is silent.** The absence is an absence. Nothing in the codebase points at +the missing subsystem; there is no stub, no TODO, no failing test. The weapon +system works, feels good, and passes playtests, because every client is honest +during a playtest. + +**Why the obvious check misses it.** The question "do we validate hits?" gets +answered by looking at the hit path, which is full of plausible machinery. The +right question is one level down and is about a subsystem nobody expects to find +by reading weapon code: does the server store per-tick positions at all? + +**Symptom.** "Add server validation" is scoped as a patch and turns out to be a +subsystem. Partial attempts produce shots that visibly miss under latency, which +is then treated as a regression and reverted. + +**Detect.** Ask whether the machinery exists at all, before reading any weapon +code: + +```bash +rg -n "FSavedMove|ServerMove|CompressedFlags|SavedMoves|PositionHistory" Source/ +``` + +Empty output means there is no lag compensation and no custom movement +prediction, and therefore no basis for hit validation in principle. On the +measured reference this returned nothing at all, which explains every other +finding in this file about client trust: they are consequences of one absent +subsystem, not independent oversights. + +**Guardrail.** Decide the hit-registration model — server trace, lag-compensated +rewind, or client-reported — before the weapon system exists, and record the +decision with its threat model. Interim sanity checks (max range, line of sight +from the replicated position, fire rate, team) raise cheat cost without full +rewind and are worth having in every model. + +--- + +### NA-09 - Fail-open permission check + +**Mechanism.** A check returns "allowed" on an indeterminate or error result. The +error branch was written to avoid blocking legitimate play during development. + +**Why it is silent.** The indeterminate case does not arise in normal play. Both +participants are always fully initialised, always have a PlayerState, always have +a resolved team. The permissive branch is dead code for the entire duration of +every test. + +**Why the obvious check misses it.** The function has a check, the check is +correct in the determinate cases, and unit tests cover those cases. Reviewing the +error branch requires asking "what if this returns the error value?", and the +answer during development is "it doesn't". + +**Symptom.** Rules bypassed exactly in the edge cases nobody tested — early +initialisation, missing PlayerState, a disconnecting client, an actor spawned +before its owner. For a friendly-fire check this means damage applied where team +rules should have prevented it. + +**Detect.** Find the error branches of permission checks and read their return +value: + +```bash +rg -n -B3 -A3 "InvalidArgument|Indeterminate|Unknown" Source/ \ + | rg "return true|return .*Allow|bCanDamage\s*=\s*true" +rg -n "//\s*@?TODO.*temporar" -i Source/ +``` + +The second command is worth running on its own: a permissive branch marked +temporary is the highest-confidence instance of this class. + +**Guardrail.** Failure to determine must deny. Make the error branch explicit and +restrictive, and if a permissive default is genuinely needed during development, +gate it on a build configuration so it cannot ship. + +--- + +### NA-10 - Cheat RPC surface in shipping + +**Mechanism.** `Server` RPCs taking arbitrary command strings are declared outside +any build-configuration guard. The implementation may be neutralised; the surface +is not. + +**Why it is silent.** The commands are useful, work as intended, and are reached +through a UI that only appears in development. Hiding the entry point in the UI +feels like removing it, and the difference never shows up in play. + +**Why the obvious check misses it.** The guard that reviewers look for is usually +present *somewhere* — around the cheat manager, around the console UI, around the +implementation body. Confirming the RPC declaration itself is unguarded means +checking a header for the absence of a preprocessor directive, which is not +something a code search naturally expresses. + +**Symptom.** A remote console in the shipped binary. Any client that can encode +the RPC can execute arbitrary cheat strings on the server. + +**Detect.** Check whether the declarations sit inside a build guard: + +```bash +rg -n "UFUNCTION\([^)]*Server[^)]*\)" -A2 Source/ | rg -i "cheat|exec|command" +rg -n -B8 "void Server(Cheat|Exec)" Source/ | rg "#if|UE_BUILD_SHIPPING" +``` + +An empty second result with a non-empty first is the finding. Note the +distinction the recipe forces: whether the *implementation* is neutralised in a +shipping configuration is a separate question that only a packaged build answers, +and it should stay marked as open until someone answers it. + +**Guardrail.** Wrap declaration *and* implementation in `#if !UE_BUILD_SHIPPING` +or an equivalent cheat-manager guard. Hiding the UI is not removal. + +--- + +## Replication shape + +### NA-11 - Multicast used where state belongs + +**Mechanism.** A one-shot multicast conveys a fact that a late joiner needs. It +reaches everyone connected and relevant at the moment it fires. + +**Why it is silent.** For everyone present, it works. The fact arrives, the visual +plays, the state is right. The failure is scoped precisely to players who were not +there, and those players see a plausible world — just the wrong one. + +**Why the obvious check misses it.** Testing is done by players who join before +the match starts, which is the configuration in which the bug cannot occur. +Joining mid-match is a separate deliberate test that has to be designed. + +**Symptom.** Players who join, or become relevant, after the event see the wrong +state permanently. It never self-corrects, because nothing re-sends an event. + +**Detect.** Enumerate multicasts and ask what each one carries: + +```bash +rg -n "UFUNCTION\([^)]*NetMulticast" -A3 Source/ +``` + +For each, answer: if a player joined one second after this fired, what would they +see? Then run the test — join mid-round and inspect the state of every feature. +This is a runtime check by nature; the grep only produces the list to test. + +**Guardrail.** Anything a late joiner must know is replicated state. Multicasts +carry transient presentation only — a sound, an impact, a one-frame effect whose +absence is invisible a second later. + +--- + +### NA-12 - Reliable RPC overflow + +**Mechanism.** Frequent or per-frame reliable RPCs fill the reliable queue. + +**Why it is silent.** The queue is generous. At two players on a local network it +never fills. The cost scales with player count, latency and packet loss +simultaneously, so it stays invisible across the entire range of configurations +used during development. + +**Why the obvious check misses it.** Each individual call site is defensible — +one reliable RPC is nothing. The failure is a sum over call sites and time, and no +review reads all of them together. + +**Symptom.** Clients disconnect under load, typically first observed in a +playtest at full player count, and usually attributed to the network rather than +to a specific call site. + +**Detect.** Count reliable RPCs and look for the ones a client can trigger +repeatedly: + +```bash +rg -n "UFUNCTION\([^)]*\bReliable\b" Source/ | wc -l +rg -n "UFUNCTION\([^)]*\bReliable\b" -A3 Source/ \ + | rg -i "tick|update|move|aim|look|input|cosmetic" +``` + +The second command finds reliable calls on names that imply frequency. Confirm +with a load test at maximum player count and 2% packet loss, watching for +queue-overflow disconnects. + +**Guardrail.** Reliable is a budget, not a default. Cosmetic, frequent or +idempotent calls are unreliable; rate-limit anything a client can call in a loop. + +--- + +### NA-13 - Incomplete FastArray callbacks + +**Mechanism.** `PreReplicatedRemove`, `PostReplicatedAdd` and +`PostReplicatedChange` are partially implemented. A derived cache or local +listener list is updated in some of them and not others. + +**Why it is silent.** The server is always right, because the server never runs +the callbacks that are missing — it mutates the array directly. Only clients +apply deltas, so only clients drift, and they drift into a state that is +internally consistent. + +**Why the obvious check misses it.** All three callbacks exist as overrides, so a +structural review finds a complete-looking implementation. The empty one has a +valid signature and compiles. Server-side tests pass because the defect cannot +exist on the server. + +**Symptom.** Client-side desync with no symptom on the server — the hardest class +of bug to attribute, because every server-side log and dump shows correct state. + +**Detect.** Compare the three callbacks per FastArray type: + +```bash +rg -n "FFastArraySerializer" -l Source/ | while read -r f; do + echo "== $f" + rg -n -A4 "Pre(ReplicatedRemove)|Post(ReplicatedAdd|ReplicatedChange)" "$f" +done +``` + +Read the bodies. An empty body next to two implemented ones is the finding. +On the measured reference one replicated structure had an empty removal hook and +no callers at all — harmless while unused, a silent desync the moment it is +adopted, which is the worst possible time to discover it. + +**Guardrail.** Implement all three or none. Assert cache/array consistency in +development builds so the drift becomes an error rather than a mystery. + +--- + +### NA-14 - Relevancy disabled by an inherited constant + +**Mechanism.** A net cull distance chosen for a small arena is inherited by a +large map. Beyond the map's own size, the setting disables relevancy culling +entirely. + +**Why it is silent.** It is not a bug on the map it was chosen for — it is the +correct setting there, and it makes gameplay better. The value carries forward +because it is in a base class, and base-class defaults are not re-examined when a +new map is added. + +**Why the obvious check misses it.** The number is a squared distance, which +makes it unreadable at a glance: nine hundred million does not announce itself as +three hundred metres. Reviewers see a tuned-looking constant and move on. + +**Symptom.** Bandwidth and server CPU scale with total actor count instead of +local density. Discovered at player-count scale-up, late, when the fix requires +re-verifying every gameplay assumption about who can see whom. + +**Detect.** Find the value and convert it: + +```bash +rg -n "SetNetCullDistanceSquared|NetCullDistanceSquared\s*=" Source/ Config/ +``` + +Take the square root of each result and compare it to the largest map's diagonal. +On the measured reference the value corresponded to a 300 m radius on arena maps — +correct there, fatal if inherited into an open world. + +**Guardrail.** Treat net cull distance as a per-class gameplay decision with a +stated map-size assumption written next to it, not as an inherited performance +knob. + +--- + +### NA-15 - Replication graph enabled late + +**Mechanism.** The graph implementation and its routing configuration exist, and +the graph ships disabled. The routing is therefore never exercised. + +**Why it is silent.** Disabled means the default replication path runs, and the +default path is correct. The configuration file looks maintained. Nothing +degrades, because nothing is using it. + +**Why the obvious check misses it.** The presence of a configured routing table is +read as evidence that the feature is in use. Confirming otherwise means finding +the one boolean that disables the whole thing, which lives in a settings class +rather than next to the routing. + +**Symptom.** Enabling it changes which actors clients see, breaking gameplay code +that had silently assumed universal relevancy. The breakage appears far from the +change, in features nobody associated with replication. + +**Detect.** Check the enable flag before trusting the routing: + +```bash +rg -n "bDisableReplicationGraph|bEnableReplicationGraph" Source/ Config/ +rg -n "ClassNodeMapping|RepGraph" Config/ +``` + +A disabled flag beside a populated routing table means the routing is untested +configuration, not working configuration. + +**Guardrail.** Enable it early or treat it as unavailable. Re-verify every +relevancy assumption at the moment it is turned on, and treat that as a project +milestone rather than a settings change. + +--- + +### NA-16 - Wrong ability-system replication mode for AI + +**Mechanism.** `Mixed` is inherited for AI-controlled actors. `Mixed` replicates +gameplay effects to the owning client — and an AI has no owning client to +receive them. + +**Why it is silent.** It is correct for player-controlled actors, which is where +it was chosen, and it is not wrong for AI either — merely wasteful. Waste has no +symptom until bandwidth is the constraint. + +**Why the obvious check misses it.** All three modes are one-line calls in +constructors spread across different classes. There is no place where the choice +is expressed as a policy, so nobody reviews the set of choices together. + +**Symptom.** Bandwidth spent replicating gameplay effects nobody reads. Shows up +as an unexplained baseline in profiling, attributed to player count. + +**Detect.** List every mode choice in one place: + +```bash +rg -n "SetReplicationMode|EGameplayEffectReplicationMode::" Source/ \ + | sed 's/.*EGameplayEffectReplicationMode:://' | sort | uniq -c +``` + +On the measured reference all three ability-system components used `Mixed`, with +`Minimal` and `Full` absent from the module entirely — a default applied +uniformly rather than a decision made three times. + +**Guardrail.** `Mixed` for player-controlled, `Minimal` for AI, `Full` only with a +written reason. Put the three choices in one comment block so they are reviewed +as a policy. + +--- + +## Lifecycle + +### NA-17 - Readiness gate that assumes a controller + +**Mechanism.** An initialisation chain waits for a controller. Simulated proxies +on remote clients never receive one. + +**Why it is silent.** The host is fine. The listen-server's own pawn has a +controller, and so does every pawn in a single-process test. The stall happens +only on machines that are not running the authority, which is the configuration +least often exercised. + +**Why the obvious check misses it.** The gate reads as obviously correct — of +course initialisation waits for the controller. The missing knowledge is not in +the code, it is that simulated proxies exist and never satisfy the condition. A +reviewer who knows that spots it instantly; one who does not cannot find it by +reading more carefully. + +**Symptom.** Remote clients see uninitialised characters forever. Reproduces only +with a dedicated server, so it is often first reported by QA and not reproducible +by the developer. + +**Detect.** Find readiness conditions that mention a controller: + +```bash +rg -n -B4 -A4 "GetController\(\)|Controller\s*!=\s*nullptr|IsPlayerControlled" \ + Source/ --glob "*Init*" --glob "*Extension*" --glob "*Ready*" +``` + +For each, ask what a simulated proxy does at that line. The measured reference +handles this correctly — its init-state gate permits progression without a +controller — and is worth reading precisely because the naive version of the same +gate deadlocks every simulated proxy while working perfectly in PIE. + +**Guardrail.** An explicit named state chain, with the no-controller case handled +deliberately rather than by omission. Write down which states a simulated proxy +reaches. + +--- + +### NA-18 - Addition without verified removal + +**Mechanism.** Granted abilities, applied effects, registered listeners and +spawned cues accumulate across death/respawn or feature activation cycles, +because the addition path was built and the removal path was assumed. + +**Why it is silent.** One cycle is fine. Two cycles produce a duplicate, and a +duplicate of an idempotent-looking effect is usually invisible. The cost grows +linearly with session length, and development sessions are short. + +**Why the obvious check misses it.** Both halves exist in most cases — there is a +removal function, it is called, and it removes *something*. Whether it removes +exactly what was added requires holding the handle, and a review that sees +`Remove` next to `Add` reasonably concludes the pair is complete. + +**Symptom.** Duplicated effects after the second respawn; leaks that only appear +in long sessions; a second feature activation that adds a second copy of every +widget. + +**Detect.** Compare addition and removal sites per resource type, then test: + +```bash +rg -n "GiveAbility|ApplyGameplayEffect|AddLoadedGameplayCuePath|RegisterListener" \ + Source/ | wc -l +rg -n "ClearAbility|RemoveActiveGameplayEffect|RemoveLoadedGameplayCuePath|Unregister" \ + Source/ | wc -l +``` + +Unequal counts are a starting point, not a verdict. The verdict is runtime: +baseline, activate, deactivate, compare to baseline, reactivate — on both server +and client, twice in one session. + +**Guardrail.** Every dynamic addition retains an ownership handle and an exact +rollback, tested on both server and client, twice in one session. Assert on the +removal count where the API returns one. + +--- + +### NA-19 - Prediction without a correction path + +**Mechanism.** A client-predicted value has no defined behaviour when the server +disagrees. The prediction key reconciles ordering; nothing reconciles the value. + +**Why it is silent.** The server rarely disagrees on a local network with an +honest client. Disagreement requires latency, packet loss or a rejected action, +and the development environment supplies none of them. + +**Why the obvious check misses it.** The prediction machinery is present and +correct, and prediction keys are a real feature that really works. Reviewing +"is this predicted properly?" finds a correct implementation of prediction. The +missing piece is a design question — what does the player see when we were +wrong? — that has no code location to be missing from. + +**Symptom.** The client keeps a wrong value indefinitely. UI and gameplay diverge +with no error anywhere, and the player's experience is that the game lied to them +once and never corrected. + +**Detect.** Force disagreement and watch: + +```bash +rg -n "ScopedPredictionWindow|PredictionKey|GetPredictionKeyForNewAction" Source/ +``` + +For each predicted value, run the paid test: introduce 150 ms latency, make the +server reject the action, and observe the client for ten seconds. A value that +does not converge has no correction path. This one cannot be settled statically. + +**Guardrail.** Define, per predicted value, what the correction looks like to the +player. A prediction key reconciles ordering, not entitlement. + +--- + +### NA-20 - PIE treated as a network test + +**Mechanism.** The listen-server host shares memory with its client, so the +serialisation, ordering and relevancy boundaries that produce networked defects do +not exist during the test. + +**Why it is silent.** PIE multiplayer looks like multiplayer. Two windows, two +players, two views of one world. Everything that works there works convincingly, +including the things that only work because the boundary is absent. + +**Why the obvious check misses it.** The check *is* the test, and the test passes. +There is no artefact to inspect and no code to review; the defect is in the +choice of configuration, which is made once, early, by whoever set up the +workflow, and never revisited. + +**Symptom.** Features ship having never run in the configuration they will run in. +Entries NA-01, NA-04, NA-11, NA-13, NA-17 and NA-19 above cannot occur in PIE at +all. + +**Detect.** Audit the test configuration rather than the code: + +```bash +rg -n "PlayNetMode|RunUnderOneProcess|bLaunchSeparateServer" Config/ Saved/Config/ +``` + +If the editor is configured for a listen server in one process, no networked +feature in the project has been tested. Confirm by running one known-networked +feature under a dedicated server with a remote client and comparing behaviour. + +**Guardrail.** Dedicated server plus two remote clients, one with 150 ms latency +and 2% packet loss, as the minimum configuration for accepting a networked +feature — including a mid-match joiner. diff --git a/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/patterns.md b/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/patterns.md new file mode 100644 index 0000000..21d481d --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-multiplayer-authority/references/patterns.md @@ -0,0 +1,178 @@ +# Patterns: authority in a measured reference product + +A worked example of the authority model from `SKILL.md`, against one specific +reference project read as source. It is here for one reason: the failure modes in +[failure-modes.md](failure-modes.md) are easier to believe, and much easier to +apply, when you can see how they cluster in a real codebase. + +**Read this first.** The project audited below is an outstanding architectural +reference and a poor network-security reference. Those are independent axes. None +of what follows is a criticism of the sample — a learning project that ships no +product has no reason to build lag compensation — but copying its structure +without noticing its trust boundaries is a real and common accident, and it is the +accident this skill exists to prevent. + +Markers: **[measured]** — read in source; **[derived]** — conclusion from measured +facts; **[open]** — not answerable without a dedicated-server build, and left +open rather than guessed. + +--- + +## 1. The headline: one absence explains almost everything + +No custom saved-move type, no custom server-move path, no compressed-flags +extension, and no historical position storage anywhere in the project +**[measured]**. + +Without stored per-tick positions the server has nothing to rewind to, and +therefore **cannot** validate a client's hit claim even in principle +**[derived]**. The client trust documented below is not carelessness at individual +call sites; it is the necessary consequence of a subsystem that was never built. + +This is the single most valuable thing to carry away from the audit, and it +generalises: **before reading a weapon system, ask whether the machinery that +would make validation possible exists at all.** If it does not, every "missing +check" you find downstream is one symptom of one decision, not a list of +independent bugs. Recipe: NA-08. + +Practical implication for anyone adopting such code: "add server validation to the +weapon system" is not a patch. It is building the rewind subsystem first. + +--- + +## 2. Hit registration was client-authored end to end + +Four findings, each of which reads as an isolated defect and is not: + +| What was measured | Consequence | +|---|---| +| The targeting trace early-returns unless locally controlled — the server never traces **[measured]** | The client decides what was hit | +| A target-data validity flag is a `const bool` initialised to `true`, consumed twice, with no computation between **[measured]** | The code path reads as validated in review; NA-06 | +| Distance falloff is computed from a client-supplied trace origin; the damage multiplier is taken from a client-supplied surface material **[measured]** | A modified client controls falloff and multiplier directly **[derived]**; NA-05 | +| A `bIsSimulated` parameter appears at ten sites, is passed a literal at its single call site, and is branched on nowhere; a hit-replacement helper is never called **[measured]** | Attachment points for the rewind subsystem that was never built **[derived]**; NA-07 | + +One check in that path was genuinely server-side: the team comparison guarding +whether damage may be caused at all **[measured]**. That detail matters more than +it looks. **The model was incomplete, not absent** — which is the harder case, +because the presence of one real server-side check makes the surrounding code read +as server-side too. + +--- + +## 3. Three mechanisms that look equivalent and are not + +One equipment component demonstrated all three in a single file **[measured]**: + +| Declaration | Actually enforces | +|---|---| +| `UFUNCTION(Server, Reliable, BlueprintCallable)` | routing only — no `WithValidation` | +| `UFUNCTION(BlueprintAuthorityOnly, ...)` | blocks Blueprint callers; no effect on C++ | +| plain C++ writes elsewhere in the same class | nothing | + +`BlueprintAuthorityOnly` is an editor affordance. It appears in review as an +authority guard and is not one **[derived]**. Across the whole project the +specifier appeared at five sites in three subsystems, none of which carried an +authority check in the guarded body **[measured]**. Recipe: NA-03. + +Unguarded state transitions followed the same shape: the death-state writers were +public, virtual, and unguarded, with the only barrier being a server-initiated +activation policy on one ability **[measured]**. The guard belongs in the mutator, +not in one of its callers **[derived]** — the call-site list is not a stable +property, and virtual functions mean a subclass in a feature plugin inherits the +hole. Recipe: NA-04. + +--- + +## 4. Cheat surface and error handling + +Two arbitrary-string cheat RPCs were declared outside any build-configuration +guard **[measured]**. Their validators return `true` unconditionally +**[measured]** — which is defensible for a cheat RPC and is also exactly the shape +NA-02 looks for, so the recipe finds them first and you must read what you found. +Whether the implementations are neutralised in a shipping configuration is +**[open]** without a packaged build; the RPC surface itself is unconditional. + +A team-comparison helper returned a permissive result on an invalid-argument +outcome, carrying a comment marking it temporary **[measured]**. Since that +function gates friendly fire, the failure mode is damage applied where team rules +should have prevented it **[derived]**. + +Rule extracted, and it is worth more than the finding: **failure to determine must +deny.** Recipe: NA-09. + +--- + +## 5. Replication configuration + +| Setting | Measured | Reading | +|---|---|---| +| Ability-system replication mode | `Mixed` on all three ability-system components; `Minimal` and `Full` absent from the module **[measured]** | Correct for player-controlled actors; wasteful when inherited by AI **[derived]**; NA-16 | +| Net cull distance | A squared value corresponding to a 300 m radius, set in the character base class **[measured]** | Every character relevant to every client at all times **[derived]** — correct for a small arena, fatal if inherited into a large map; NA-14 | +| Replication graph | Ships disabled, with class routing configured **[measured]** | The routing is untested configuration, not working configuration **[derived]**; NA-15 | +| A replicated fast-array | Empty removal callback, no callers in the project **[measured]** | Harmless while unused; a silent client-side desync the moment it is adopted **[derived]**; NA-13 | + +The last row is the one worth pausing on. An unused structure with an incomplete +callback set is invisible to every kind of testing, because nothing exercises it. +It becomes a defect at adoption time — the moment when the person adopting it has +the least context to recognise what they inherited. + +--- + +## 6. What the reference got right + +For balance, and because these are the parts worth copying directly: + +1. **An explicit init-state chain** instead of timing assumptions. Its gate + permits progression for actors without a controller, which is what allows + simulated proxies to initialise at all **[measured]**. The naive version of the + same gate deadlocks every simulated proxy on remote clients while working + perfectly in PIE **[derived]** — study the correct one before writing your own. + Context: NA-17. +2. **A replication mode chosen deliberately** for player ability-system components + rather than left at the engine default **[measured]**. +3. **Replicate intent, realise presentation locally** — the cosmetic system sends + which parts to wear, not the assembled result **[measured]**. +4. **Correct teardown where it exists**: one pawn-extension path checks avatar + identity before cancelling abilities, clearing input and removing cues + **[measured]**. Note the shape — identity check first, then reverse-order + removal. +5. **The team check in the damage execution is genuinely server-side** + **[measured]**. The model is not absent, it is incomplete. + +--- + +## 7. Questions the audit could not answer + +Left open deliberately. Each requires a dedicated-server or packaged build, which +was out of scope: + +1. Whether the cheat RPC implementations are neutralised in a shipping + configuration. +2. Actual bandwidth per client, and whether the relevancy radius holds at full + player count. +3. Behaviour of the replication graph when enabled — routing completeness, + dormancy usage. +4. Whether client and server tag registries match when feature plugins load + asymmetrically. +5. Real reconciliation behaviour of ability prediction under packet loss. + +An audit that answers questions it cannot answer is worth less than one that +enumerates them. If you adopt this material, these five are yours to close. + +--- + +## Provenance + +Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source in +a single workspace. Source addresses stay in the research archive that produced +this skill; each `NA-` identifier above resolves back to the audited location +there. Numbers without an author are worth less than no numbers, so: this is where +these came from, and they can be produced on request. + +## Evidence boundary + +Every claim above refers to one project on one engine version. These are examples +and failure evidence, not guarantees about other engine versions or other samples. +Verify version-sensitive claims against your target engine before acting on them. +In particular: PIE cannot reproduce most of the failure modes described here, +because a listen-server host shares memory with its client. diff --git a/plugins/ue-design-skills/skills/ue-reference-project-adoption/SKILL.md b/plugins/ue-design-skills/skills/ue-reference-project-adoption/SKILL.md new file mode 100644 index 0000000..00370f5 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-reference-project-adoption/SKILL.md @@ -0,0 +1,343 @@ +--- +name: ue-reference-project-adoption +description: >- + Adopt architecture from an Unreal Engine reference or sample project (a + first-party starter game, an engine template, an internal legacy title) + without inheriting its shortcuts: classify each borrowed element as defect, + showcase stub, deliberate sample-scope narrowing, architecture tax or + foreign-product legacy, then decide copy/close/skip per element. Use when + starting a project from a template, lifting a subsystem out of a sample, + reviewing code that was copied from one, or explaining why a setting copied + from the reference does nothing. +--- + +# UE reference project adoption + +The invariant: + +> A reference project shows you a **shape**, not a **contract**. Anything you copy +> without finding its consumer, its writer and its owner is a shape you now have to +> finish yourself. + +Worked classification, tax analysis and vocabulary hygiene: +[patterns](references/patterns.md). +Nineteen detection recipes for silent failures: +[failure modes](references/failure-modes.md). + +Read the failure modes before adopting or reviewing borrowed architecture. Each +entry names the mechanism, why it stays silent, why the obvious check misses it, +and a command you can run against your own tree. + +Related skills: `ue-architecture-guardrails`, `ue-modular-gameplay`, +`ue-multiplayer-authority`, `ue-gameplay-tag-governance`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. The one thing that makes this hard + +Broken code announces itself. It crashes, it logs, it fails a test. + +An unfinished mechanism announces nothing. It returns a neutral value forever, and +that neutral value is indistinguishable from "the feature is off right now". + +So the ranking that matters during adoption is **not severity**. It is *probability +of copying it without noticing*. Those are close to inverted: the most dangerous +elements in a mature sample project are the ones that never misbehave. + +### Severity is not copyability + +| | Pattern is safe to copy | Pattern is unsafe to copy | +|---|---|---| +| **Fails loudly** | plain defect — fix the line, keep the design | — | +| **Never fails** | foreign legacy — rename and keep | **stubs and scope narrowings** | +| **Works as intended** | tax you can afford | tax sized for someone else | + +Only the bottom-right cell costs a subsystem rewrite instead of a line edit. + +--- + +## 2. Five categories, and the decision each one implies + +Classify every element you intend to borrow. The category, not the symptom, +determines what you do. + +### C1 — Defect + +A mistake the reference's authors would fix if they noticed. + +**Decision:** fix the line; keep the surrounding design. A defect is not evidence +against the pattern that contains it. + +### C2 — Showcase stub + +Declared, editable, sometimes even read — but with no writer, no consumer, or a +hardcoded neutral result. Usually marked with a TODO, often not. + +**Decision:** either implement it before exposing it, or delete the field. Never +ship an editable property whose consumer is a literal constant. Designers will +spend days tuning it. + +### C3 — Deliberate sample-scope narrowing + +The reference intentionally did not build a subsystem, because a sample does not +need it. Nothing is broken. The reference is internally consistent. + +**Decision:** this is the expensive one. Decide explicitly, at adoption time, +whether your product needs the missing subsystem — and budget it. The failure mode +is not "we copied a bug", it is "we shipped without noticing there was a hole". + +### C4 — Architecture tax + +Correct, working, and sized for the reference author's team, release cadence and +platform matrix. + +**Decision:** adopt by *need*, not by *completeness*. Copying an entire modular +stack because it is there means paying for an answer to a question you do not have. + +### C5 — Foreign-product legacy + +Constants, names and comments carried over from a different title. + +**Decision:** rename on entry. Harmless as code, corrosive as vocabulary — in a +year, half the team will believe the foreign term is an engine concept. + +--- + +## 3. Four tests, run before you copy anything + +Cheapest first. Most stubs die at test 3.1. + +### 3.1 Find the writer, not the reader + +For any tunable field, search for **assignment**, not usage. + +```bash +rg -n "\bMyField\b" Source/ # declaration plus two reads: looks alive +rg -n "\bMyField\b\s*(=|\+=|-=)" Source/ # one hit, and it is the declaration +``` + +Watch the second line. A member declared as `float MyField = 0.f;` matches every +assignment pattern you can write, so the naive search answers "there is a writer" +in exactly the case where there is none. Subtract the declaration before you +conclude anything: + +```bash +rg -n "\bMyField\b\s*(=|\+=|-=)" Source/ \ + | rg -v "\b(bool|u?int\d+|float|double|F[A-Z]\w+|T\w+<)\s+MyField\b" +``` + +A field that is read and compared but never written is worse than an unused field: +reading the code convinces you it is wired up. This is the single highest-yield +check in this skill — and, as the two commands above show, the one whose first +draft most often lies to you. + +### 3.2 Find the comparison, not the plumbing + +For any field named `Priority`, `Order`, `Weight`, `Rank`, `SortKey`: search for its +use inside a sort, predicate or comparison. Being passed through six functions and +stored in a struct means nothing. + +If there is no comparison, the effective ordering is *registration order*, which in +a plugin-based project means *feature activation order* — non-deterministic from +the designer's point of view. + +### 3.3 Ask who fills this in + +Any of these is a stub or a narrowing until proven otherwise: + +- an empty container next to a TODO; +- `const bool bSomething = false;` with the real expression in a comment; +- `const bool bIsValid = true;` consumed by a validation branch; +- a parameter threaded through several call layers that is never branched on; +- a function that exists, is public, and has zero callers in the codebase. + +The distinction between C2 and C3 is only whether the author ever intended to fill +it in. Both cost you the same. + +### 3.4 Ask what the adversary does + +For anything on a trust boundary, ask: *what stops the other machine from lying +about this fact?* If the answer is "nothing", you are looking at C3, and the cost +of closing it is yours to price. + +### 3.5 The paid test + +When static reading is ambiguous, change the setting and measure whether anything +moves. Expensive, and the only test that cannot be fooled. Use a written protocol +with a restore point, one knob at a time. + +--- + +## 4. Sample-scope narrowings a product must close + +These are the holes that reference projects most commonly leave, in rough order of +how much they cost to close later: + +| Narrowing | Why a sample omits it | What it costs a product | +|---|---|---| +| No lag compensation / server rewind | needs a real server, real latency, real test rig | Cannot be retrofitted cheaply — it dictates how abilities, tracing and damage are structured | +| Replication graph present but disabled | tuning it requires production-scale player counts | Traffic and CPU limits discovered in beta | +| No tag/asset redirect discipline | a sample never renames anything after release | Every rename becomes a breaking change with no migration path | +| Authority checks missing on state writers | single-process PIE never exercises the boundary | Silent state divergence, cheat surface | +| Fail-open defaults on unknown input | sample never produces unknown input | Security and gameplay decisions default to "allowed" | +| Preload/streaming machinery bypassed by config | sample content fits in memory | Hitching that looks like an engine bug | +| No teardown for dynamic additions | sample never unloads a feature | Leaks and duplicate registrations on the second load | + +None of these is a bug report against the reference. Each is a line item in your +adoption budget. + +### The rule + +> Every C3 narrowing you accept must be written down as an accepted risk with an +> owner, or scheduled. "We did not notice" is the only unacceptable outcome. + +--- + +## 5. Architecture tax: adopt need, not completeness + +A reference project's modularity is an answer to a specific question — usually +"how do several teams ship seasonal content into one binary without merge +conflicts". If you do not have that question, you pay the cost and receive nothing. + +Before adopting a modular stack, state the question it answers for *you*: + +- more than one team shipping into the same binary? +- content added after launch without a client patch? +- game modes swapped at runtime by data? +- platform-specific feature sets? + +Zero yes answers means the stack is tax. One or two means adopt the layers that +answer them and stop there. + +### The C++/Blueprint boundary is also tax + +A reference project's C++/BP split reflects its author's staffing. A codebase where +C++ exposes overwhelmingly *calls* rather than *override points* assumes a strong +C++ team is always available to extend it. Copy that boundary without that team and +designers hit a wall on their first non-trivial request. + +Measure both sides before copying the ratio: how many `BlueprintCallable` entry +points versus how many `BlueprintImplementableEvent` and `BlueprintNativeEvent` +extension points, and who on your team can add to the second column. Measured +counts from one reference project are in [patterns](references/patterns.md). + +### Measure the whole mount surface, not the primary content root + +In a plugin-based project, plugins mount their own content roots. Any census, +budget or memory metric computed over the primary game root alone is wrong by +however much the plugins hold — in one measured reference, by roughly a factor of +six. Enumerate all mount roots, not just the obvious one. + +--- + +## 6. Foreign legacy hygiene + +On entry, rename: + +- constants and type names carrying another product's brand; +- asset type identifiers whose names describe another game's content; +- comments copied from a different class; +- tags whose comment admits they are in the wrong namespace. + +Do this at import time. After the first sprint it becomes archaeology, and the +foreign vocabulary starts appearing in design documents. + +--- + +## 7. Fork and drift discipline + +If you copy code rather than depend on a plugin: + +1. Record the reference version and commit at import time, in a file next to the + copy. +2. Keep imported code in a directory that is obviously imported. +3. Record every intentional divergence with a one-line reason. +4. Re-diff against the upstream on each engine upgrade; an unrecorded divergence is + indistinguishable from an upstream fix you are about to lose. + +Without this, engine upgrades turn every borrowed subsystem into a manual merge +with no ground truth. + +--- + +## 8. Adoption procedure + +For each subsystem you intend to borrow: + +1. **Name the question** it answers for your product. If you cannot, stop. +2. **Read the consumer, not the declaration.** Confirm each field has a writer and + each ordering field has a comparison. +3. **List the trust boundaries** it crosses and what validates each one. +4. **Classify every surprising element** as C1–C5. +5. **Decide per category:** fix / implement or delete / budget and schedule / trim + to need / rename. +6. **Write down accepted C3 risks** with an owner. +7. **Record the import version** and directory. +8. **Add a test that would fail** if the narrowing you accepted became a bug. + +--- + +## 9. Review checklist + +- [ ] Every borrowed tunable field has a verified writer. +- [ ] Every ordering field participates in an actual comparison. +- [ ] No editable property is consumed by a hardcoded literal. +- [ ] Every parameter threaded through layers is branched on somewhere. +- [ ] Trust boundaries are enumerated; each has a named validator or an accepted + risk. +- [ ] Missing subsystems (C3) are written down, not discovered later. +- [ ] Modular layers adopted map to a stated product need. +- [ ] C++/BP boundary matches the team that has to live with it. +- [ ] Metrics computed over all mount roots. +- [ ] Foreign-product vocabulary renamed at import. +- [ ] Import version and divergences recorded. +- [ ] Teardown exists for every dynamic addition adopted. + +--- + +## 10. When copying wholesale is right + +Adopt as-is, without this ceremony, when **all** of the following hold: + +- the element is self-contained and has no trust boundary; +- you can name its consumer within one file; +- it has no editable tuning surface; +- replacing it later is a local edit, not a migration. + +Small utilities, math helpers, container types and single-purpose components +qualify. Subsystems, lifecycle machinery and anything that touches replication, +persistence or the tag registry do not. + +--- + +## 11. What this skill does not say + +It does not say reference projects are bad guidance. A mature sample is the best +available evidence for *composition* — how definition and instance separate, how +readiness becomes an explicit chain, how presentation decouples from gameplay. + +It says: the sample is evidence for the shape and silent about the finish. The +references exist so the silence is enumerated instead of discovered in production. + +--- + +## Provenance + +The classification, the four tests and every entry under `references/` come from a +line-by-line audit of Epic's Lyra Starter Game on Unreal Engine 5.6, read as source +rather than run. Source addresses stay in the research archive that produced this +skill; what ships is the detection recipe, because an address in someone else's +tree is not something you can act on and a recipe is. + +Each entry carries a stable identifier (`RA-01`, `RA-02`, …) that resolves back to +the audited location in that archive. If you need the original address to settle a +dispute, it exists and can be produced. + +## Evidence boundary + +The measured material refers to one specific reference project on one engine +version in one workspace. It is a worked example of the classification, not a +defect list to carry into other engine versions or other samples. Re-run the four +tests against your own reference before trusting any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/failure-modes.md new file mode 100644 index 0000000..f1c63d7 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/failure-modes.md @@ -0,0 +1,739 @@ +# Failure modes when adopting a reference project + +Nineteen ways a borrowed subsystem stays quiet while doing nothing. + +They share one property: **each is discovered late by construction.** Every entry +here compiles, passes review, produces no warning and no log line. What is absent +is a writer, a comparison, a validator or an owner — and absence does not raise. + +Each entry gives you a `Detect` you can run against your own tree today. The +recipes are deliberately crude: a grep that returns nothing where reads exist is +worth more than a subtle static analysis nobody runs. + +`RA-nn` identifiers are stable and resolve to the audited location in the research +archive behind this skill. + +Categories: **C1** defect · **C2** showcase stub · **C3** deliberate sample-scope +narrowing. Architecture tax (C4) and foreign vocabulary (C5) are not failure modes +and live in [patterns](patterns.md). + +--- + +## C1 — Defects + +Fix the line, keep the design. A defect is not evidence against the pattern that +contains it. + +### RA-01 - Tag picker filtered to a namespace that does not exist + +**Mechanism.** An editable gameplay-tag property carries a +`meta = (Categories = "Some.Root")` filter, and no tag in the project is declared +under that root. The real tags live under a different prefix. + +**Why it is silent.** The filter's only job is to narrow an editor dropdown. A +filter that matches nothing produces an empty dropdown, which is visually identical +to "no tags of this kind have been authored yet". Nothing compiles differently, +nothing logs, and the property simply stays unset forever. + +**Why the obvious check misses it.** The filter string is syntactically valid and +the property is genuinely declared, read and used — every "is this wired up?" search +answers yes. Nobody searches for the *filter value* as a tag prefix, because it +reads like a category label rather than a key that has to resolve. In the measured +reference this was one of only fourteen places in the entire project where tag input +was constrained at all, so the one guardrail present was also the one misconfigured. + +**Symptom.** A designer opens the dropdown, sees nothing, concludes the feature is +unimplemented, and leaves the field empty. The system then behaves as if the +designer chose "none". + +**Detect.** Extract every filter value and confirm each resolves to declared tags: + +```bash +rg -o 'Categories\s*=\s*"([^"]*)"' -r '$1' Source/ Plugins/ | sort -u +# then, per root R: +rg -c "\"R\." Config/*.ini Source/ +``` + +A root with zero declarations is a dead filter. + +**Guardrail.** A `Categories` filter is a guardrail only if something proves it +resolves. Add a startup or CI assertion that every filter value matches at least one +declared tag; an empty picker must be an error, not a shrug. + +--- + +### RA-02 - A misspelling that works because every side repeats it + +**Mechanism.** A tag or string key is misspelled in its declaration and in every +consumer. The string is the join between config, code and platform overrides, so +identical errors on both sides join correctly. + +**Why it is silent.** Correctness for a string join means *all sides agree*, not +*the string is a word*. Six identical misspellings across four files function +exactly as six correct ones would. There is no side that could disagree. + +**Why the obvious check misses it.** Searching for the correct spelling returns +nothing, and an empty result reads as "this feature is not implemented here" rather +than "it is spelled differently". The compiler never sees the literal as a name, and +spell-checkers do not run over identifiers. The bug only becomes visible at the +moment someone *fixes* it at one site. + +**Symptom.** Nothing at all — until a cleanup pass corrects the spelling in one +file. The join then breaks with no error, and the regression is attributed to +whatever else shipped that week. + +**Detect.** Find tag literals repeated as raw strings instead of referenced through +one constant: + +```bash +rg -oN '"[A-Za-z][A-Za-z0-9]*(\.[A-Za-z][A-Za-z0-9]*){2,}"' Source/ Config/ \ + | sort | uniq -c | sort -rn | head -40 +``` + +Any literal appearing more than once is a join with no single source of truth. + +**Guardrail.** One declaration site per tag; every other site references the +constant. Then a rename is a compile error instead of a silent disconnection. + +--- + +### RA-03 - A loop over a container that has no producer + +**Mechanism.** A local container is declared next to a TODO explaining how it will +one day be filled, then iterated immediately. Nothing ever fills it. + +**Why it is silent.** Zero iterations is a legal, successful outcome. The function +returns normally and reports success, because doing nothing to an empty set is +indistinguishable from doing the work correctly. + +**Why the obvious check misses it.** The code reads as a complete mechanism: +declaration, loop, and a body that does real work. Review attention goes to whether +the body is correct. The missing piece is one level up — the container has a +consumer and no producer, and reviewers check consumers. + +**Symptom.** A documented behaviour ("these entries are always loaded") reports zero +at runtime, and is diagnosed as a content or configuration problem for days. + +**Detect.** For every container that is iterated, look for a writer: + +```bash +rg -n "for\s*\(.*:\s*(\w+)\)" -r '$1' Source/ | sort -u > /tmp/iterated +# per name N: +rg -n "\bN\b\s*\.\s*(Add|Append|Emplace|Push|Insert|Reserve)" Source/ +``` + +Reads without writes mean the loop is decoration. + +**Guardrail.** Do not ship a loop over a container that has no producer in the same +translation unit or an obvious injection point. If the producer is future work, the +loop is future work too. + +--- + +### RA-04 - Progress arithmetic with inverted operands + +**Mechanism.** A startup progress fraction computes each sub-step's contribution +with the operands the wrong way round, so the reported value does not advance while +a single long step runs. + +**Why it is silent.** The result stays inside the valid range and still increases +across step boundaries. It is a wrong number, not an invalid one, so no assertion +and no clamp fires. + +**Why the obvious check misses it.** On a developer machine with a warm cache the +whole sequence finishes in under a second, and nobody watches one step long enough +to see it stall. The code is exercised on every launch and observed on none. + +**Symptom.** On a cold shader cache or a slow disk the bar freezes for a long time +and players report a hang. The engineering response is to look for a deadlock, +because the bar is trusted. + +**Detect.** Two options. Force the slow path with whatever artificial-delay cvar the +loading system provides and watch whether the value moves within a step. Or unit-test +the progress function directly: feed it a fixed step count and assert the value is +strictly increasing at every sub-step, not just at step boundaries. + +**Guardrail.** Progress arithmetic gets a test. It is the canonical example of code +that runs constantly and is observed only under conditions developers do not have. + +--- + +### RA-05 - Listener removed using the wrong channel key + +**Mechanism.** A tag-addressed pub/sub subsystem walks parent tags when +broadcasting. On finding an expired listener it removes the handle using the +*original broadcast channel* rather than the ancestor tag currently being visited. +Handle identifiers are unique only within a channel. + +**Why it is silent.** Both outcomes are legal operations on a map. Either the key is +absent and removal is a no-op — the stale listener survives in the parent bucket — or +a different listener happens to hold the same channel-local id and is removed from +the wrong bucket. Neither path raises. + +**Why the obvious check misses it.** The removal call is present, correctly typed, +and looks right. The defect is in *which key* is passed, and both candidate keys are +in scope, same-typed, and similarly named. Type checking cannot separate them, and +review reads the statement as "remove the expired listener". + +**Symptom.** A listener that was unregistered keeps receiving messages, or an +unrelated listener silently stops receiving them. The failure surfaces in whichever +system owned the collateral listener, arbitrarily far from the subsystem at fault. + +**Detect.** Inside any parent-tag or hierarchy traversal, confirm every mutation +keys on the loop cursor rather than the function parameter: + +```bash +rg -n -B12 "(Unregister|Remove\w*Listener|RemoveAt)" Source/ \ + | rg -n "for\s*\(|ParentTag|Ancestor" +``` + +Then read each hit and check the key. + +**Guardrail.** Make listener identity an explicit pair type — channel plus id — so +passing the wrong channel fails to compile instead of silently mis-keying. + +--- + +## C2 — Showcase stubs + +The defining property: **the declaration and the consumer both exist; the writer +does not** — or the consumer is a literal. Neither a compiler warning nor an +"unused symbol" search finds these. + +### RA-06 - Editable tunable with a live reader and no writer + +**Mechanism.** A designer-facing cooldown is compared against a "last event" +timestamp field. The timestamp is declared and read, and is never assigned anywhere +in the codebase. + +**Why it is silent.** Time elapsed since an unset timestamp is time since world +start, which always exceeds any sane threshold. The comparison is permanently true, +so the gate the setting was meant to impose never rejects anything. Permanently +allowing is a valid runtime state, so nothing errors. + +**Why the obvious check misses it.** The field is read twice and participates in a +comparison, so every "is this used?" heuristic — grep, IDE find-usages, unused-symbol +warnings — answers yes. What is missing is the **writer**, not the reader, and no +default tool asks that question. This is the highest-yield check in the whole skill +precisely because it inverts the usual direction of the search. + +**Symptom.** A designer tunes the value for days, reports it has no effect, and is +told to check their data. The setting is then either abandoned or "fixed" by +changing something else that happens to correlate. + +**Detect.** For every `EditAnywhere` or `EditDefaultsOnly` numeric property and for +every timestamp it is compared against, search for an assignment rather than a +mention: + +```bash +rg -n "\bLastFireTime\b" Source/ # reads: several +rg -n "\bLastFireTime\b\s*(=|\+=|-=)" Source/ # one hit - and it is the declaration +``` + +The second search does **not** come back empty, and that is the trap. A member +declared with an initializer (`double LastFireTime = 0.0;`) matches every +assignment pattern you can write, so the naive search reports a writer that does +not exist. Subtract the declaration before concluding: + +```bash +rg -n "\bLastFireTime\b\s*(=|\+=|-=)" Source/ \ + | rg -v "\b(bool|u?int\d+|float|double|F[A-Z]\w+|T\w+<)\s+LastFireTime\b" +``` + +Empty after that subtraction, with live reads elsewhere, is a stub. Verified on +the reference: the unsubtracted search returns one line, the subtracted search +returns none. + +**Guardrail.** Never ship an editor-exposed property whose value has no writer. +Every tunable needs one owning assignment site and one test that moves the value and +asserts an observable delta. + +--- + +### RA-07 - Ordering field that is stored, copied and never compared + +**Mechanism.** A `Priority` integer is accepted by eight registration overloads, +stored on the entry, copied into the request struct, and never appears in a sort, a +predicate or a comparison anywhere. + +**Why it is silent.** The field has a defined value at all times and travels +correctly through the entire API. Ordering still happens — it is just registration +order. A plausible order is produced on every run, so nothing looks wrong. + +**Why the obvious check misses it.** The value is written, read, copied and passed +across a large public surface. Usage counts are high. Any review that asks "is +`Priority` used?" finds a dozen sites and stops. The right question is narrower: does +it appear inside a `Sort`, `<`, `>` or predicate? That question is not one anyone +thinks to ask about a field that is obviously plumbed. + +**Symptom.** Widget or handler order changes between runs and between machines. In a +plugin-based project, registration order is feature activation order, so the symptom +reads as a race condition and gets chased in the wrong subsystem. + +**Detect.** Search for the comparison, not the plumbing: + +```bash +F='(Priority|Order|Weight|Rank|SortKey)' +rg -n "\b$F\b" Source/ Plugins/ | wc -l # plumbing: many +rg -n "(Sort|StableSort|Algo::)" Source/ Plugins/ | rg "\b$F\b" # ordering: ? +rg -n "\b$F\s*[<>]=?\s*\w|\w\s*[<>]=?\s*\w*$F\b" Source/ Plugins/ \ + | rg -v -- "->" # comparisons: ? +``` + +The `-v -- "->"` matters. Without it, every `Entry->Priority = X` is counted as a +comparison because the arrow contains `>`, and a field that is only ever assigned +looks like a field that is ordered. On the measured reference the naive form +reported two "comparisons"; both were assignments. + +A large first number with empty second and third lines is an ordering field that +does not order. + +**Guardrail.** An ordering field ships with the comparator that consumes it, in the +same change. If ordering is not implemented yet, do not expose the field. + +--- + +### RA-08 - Tag-driven visibility with the check hardcoded off + +**Mechanism.** A widget base class exposes an editable tag container for +"hide when these tags are present". Both consumer sites read +`const bool bHasHiddenTags = false;` with the real expression parked in a comment +next to a TODO. The listener registration that would supply the tags is also a +comment. + +**Why it is silent.** `false` is a legitimate value: it means "no hiding tags are +active right now". The widget is visible, which is a normal state, so no test and no +reviewer can tell the difference between "correctly not hidden" and "never able to +hide". + +**Why the obvious check misses it.** The TODO admits *one* missing piece and thereby +misdirects. The obvious fix — implement `bHasHiddenTags` — does not restore the +intended behaviour, because a nearby visibility setter overwrites the +designer-authored shown and hidden visibility values on its first call, which arrives +during construction. A stub can be deeper than its own TODO says it is, and a reviewer +who trusts the TODO stops one layer too early. + +**Symptom.** A designer fills the tag container, observes nothing, and the class is +recorded in team lore as "the tag hiding does not work" without anyone establishing +why. + +**Detect.** Find neutral literals feeding decision branches: + +```bash +rg -n "const bool b\w+ = (false|true);\s*//" Source/ Plugins/ +``` + +For each hit, check whether the property that should feed it is `EditAnywhere`. Then +check that nothing else overwrites the same state before the branch is reached — the +second step is the one people skip. + +**Guardrail.** An editable property whose consumer is a literal constant must not +ship. When you do implement one, re-derive the whole path rather than the single line +the TODO names. + +--- + +### RA-09 - Parameters threaded through for a subsystem that was never built + +**Mechanism.** A `bIsSimulated` flag appears in three method signatures and is +passed through six call sites. Its sole caller passes a literal `false`. A companion +"replace this hit" method exists and is never called; the boolean recording its result +is only ever read. + +**Why it is silent.** A flag that is always `false` produces one consistent code +path, which is the path everything is tested on. A method with no callers has no +behaviour to be wrong. Both are inert rather than broken. + +**Why the obvious check misses it.** The threading through multiple layers is exactly +what a real, wired-up feature looks like — that shape is itself the disguise. Usage +searches return many hits. Only two narrower questions expose it: is the parameter +ever branched on, and does the sole call site pass a variable or a literal? + +**Symptom.** An engineer reads the signatures, concludes server-side hit rewind +exists, and budgets zero for it. The gap is discovered when the game meets real +latency, at which point the structure of abilities, tracing and damage is already +committed. + +**Detect.** Find parameters that are carried but never branched on: + +```bash +rg -n "\bbIsSimulated\b" Source/ # many hits: plumbed +rg -n "if\s*\(\s*!?bIsSimulated" Source/ # none: never decides anything +rg -n "\w+\(.*\bfalse\b.*\)" Source/ | rg "bIsSimulated|Simulated" +``` + +Also list public functions with zero callers: `rg -n "FunctionName"` returning only +the declaration and definition. + +**Guardrail.** Do not ship attachment points for a subsystem that does not exist. +An unused parameter is a claim about the architecture; if the claim is false, delete +the parameter or write the subsystem. + +--- + +### RA-10 - Field silently dropped in both directions of a conversion + +**Mechanism.** Helper functions convert between a gameplay message and cue +parameters. A context tag container present on both sides is not copied in either +direction. Both sites carry a TODO. + +**Why it is silent.** The conversion succeeds and produces a fully valid object. The +dropped field simply arrives empty, and empty is the same value a caller who did not +set it would produce. There is no partial-failure signal because nothing failed. + +**Why the obvious check misses it.** Round-tripping is the natural test, and a +round trip through both directions loses the same field consistently — so a +comparison of "before" against "after" on the fields anyone thought to compare +passes. Structural equality is the test that would catch it, and structural equality +is what nobody writes for conversion helpers. + +**Symptom.** Effects lose their contextual tags when routed through the conversion, +so downstream selection by context silently picks the default variant. + +**Detect.** For every conversion helper, compare field counts on both sides: + +```bash +rg -n "^\s*(FGameplayTagContainer|F\w+|TArray<\w+>)\s+\w+;" Source/**/Struct.h +rg -n -A30 "ToOtherForm|FromOtherForm|Convert\w+" Source/ | rg "\bField\b" +``` + +Any field declared on both types and mentioned in neither direction is dropped. + +**Guardrail.** Conversion helpers get a structural round-trip test that enumerates +fields by reflection rather than by hand. A hand-written comparison tests the fields +the author remembered, which is the same set they remembered to copy. + +--- + +### RA-11 - Replicated structure with no callers and an empty removal hook + +**Mechanism.** A replicated fast-array serializer type is fully declared, with the +removal callback implemented as an empty body. Nothing in the codebase constructs or +uses it. + +**Why it is silent.** Unused replicated types cost nothing at runtime and generate no +warning. The empty callback is a valid override — many are legitimately empty. + +**Why the obvious check misses it.** The type is complete and idiomatic, so it reads +as infrastructure that some subsystem depends on. Deleting it feels risky, so it +survives every cleanup. Meanwhile a reader treats its existence as evidence that +replicated messaging is solved. + +**Symptom.** A team builds on top of it, assuming it works, and discovers on the +first multi-client test that no path ever populated it. + +**Detect.** Find declared-but-unconstructed types: + +```bash +rg -ln "struct F\w+ : public FFastArraySerializer" Source/ Plugins/ +# per type T: +rg -n "\bT\b" Source/ Plugins/ | rg -v "\.h:|struct|USTRUCT|template" +``` + +Hits only in the header mean the type has no user. + +**Guardrail.** Replication infrastructure ships with at least one caller and one +test that observes a replicated change. Untested replication is not infrastructure, +it is a plan. + +--- + +## C3 — Deliberate sample-scope narrowings + +Nothing here is broken. The reference is internally consistent with every one of +these. They become expensive only when the sample becomes the base of a product. + +### RA-12 - No lag compensation, and the whole hit chain follows from it + +**Mechanism.** The project contains no custom saved-move, server-move or compressed +flags implementation, and therefore no historical positions on the server. Local +targeting early-returns when the pawn is not locally controlled, so the server never +traces. A `bIsTargetDataValid` local is initialised to `true` and consumed as if it +were a validation result. Damage falloff is computed from the client's supplied trace +start, and the material multiplier from the client's supplied physical material. + +**Why it is silent.** Every step is the only available behaviour given the missing +subsystem. The server cannot validate positions it never recorded, so trusting the +client is not a slip — it is the sole option. Correct-looking code all the way down, +with the gap one level below the code. + +**Why the obvious check misses it.** Reviewing the damage path finds an authority +check and a validity boolean, which together look like a validation layer. The +boolean is a literal and the authority check answers a different question ("may this +instigator damage this target?") than the one that matters ("did this shot happen?"). +Reading for the presence of checks finds checks; reading for what the adversary can +assert finds nothing stopping them. + +**Symptom.** In production, clients report impossible hits and the diagnosis becomes +"someone forgot a validation call" — which sends engineers to patch individual sites +instead of pricing the missing subsystem. + +**Detect.** Ask what the adversary supplies, then trace each such value: + +```bash +rg -n "class \w+ : public FSavedMove_Character|ServerMove|CompressedFlags" Source/ +rg -n "const bool b\w*Valid\w* = true" Source/ +rg -n "TargetData|HitResult" Source/ | rg -i "damage|falloff|physmat" +``` + +No saved-move type plus client-supplied geometry in the damage math means full +client trust. + +**Guardrail.** Choose a hit-registration model *before* borrowing a weapon stack: +server-authoritative with rewind, client-claim with plausibility bounds, or explicit +full trust. Write the choice down. Each has a different cost and a different cheat +surface, and the choice dictates structure you cannot cheaply change later. + +--- + +### RA-13 - Replication graph implemented, configured, and shipped disabled + +**Mechanism.** A replication graph implementation and around a dozen tuning cvars +exist. Project config disables it and the class routing table is essentially empty, +with the few entries present marked not-routed. + +**Why it is silent.** The default replication path works fine at sample scale. +Disabled infrastructure produces no error and no measurable difference until player +counts and actor counts grow. + +**Why the obvious check misses it.** The presence of the implementation, the cvars +and the config section reads as "this is configured". Nobody diffs the routing table +against the actor classes that actually exist, because a populated-looking config +section satisfies the eye. + +**Symptom.** Bandwidth and server CPU limits are discovered in beta, at the point +where the actor class layout is fixed and re-routing is a large change. + +**Detect.** Compare routing entries against replicated classes: + +```bash +rg -n "bDisableReplicationGraph|ClassNodeMapping" Config/ +rg -c "GetLifetimeReplicatedProps" Source/ Plugins/ +``` + +A handful of routing entries against dozens of replicated classes means the graph +has never carried traffic. + +**Guardrail.** Copying a disabled subsystem's configuration copies untested +configuration. Either enable it and measure under load, or delete the config and +record that you owe the work. + +--- + +### RA-14 - Preload machinery made unreachable by a load-mode constant + +**Mechanism.** A cue manager supports delayed loading with garbage-collect and +map-load hooks. A load-mode constant is set to "load upfront", which short-circuits +every switch site — including the one that installs the three delegates. + +**Why it is silent.** Loading everything upfront is correct behaviour. The +machinery is not broken; it is bypassed. Memory is higher and nothing reports it, +because nothing is measuring against a budget. + +**Why the obvious check misses it.** Searching for the delegates finds them bound in +source, so the wiring appears present. The early return that makes the binding +unreachable is in a different function, guarded by a constant that reads like a +development convenience. The project's own diagnostic command reports zeros, which is +then read as "no cues are preloaded yet" rather than "this counter is dead". + +**Symptom.** A team diagnosing memory or hitching concludes the engine's preload +system is broken and descends into engine code, when the sample simply opted out. + +**Detect.** Find switch sites gated by a constant, and confirm delegate binding is +reachable: + +```bash +rg -n "LoadMode|ELoadMode|EEditorLoadMode" Source/ Plugins/ +rg -n -B6 "AddUObject|AddRaw|BindUObject" Source/ | rg "return;|LoadUpfront" +``` + +**Guardrail.** Distinguish "the mechanism is broken" from "the sample configured it +off" before you spend a day in engine code. Record which subsystems the reference +bypasses; that list is part of what you are adopting. + +--- + +### RA-15 - Public virtual state writers with no authority check + +**Mechanism.** Death-state transitions are written by two public virtual methods on +a health component, neither of which checks for network authority. + +**Why it is silent.** Single-process play-in-editor never separates authority from +simulation, so the code path is exercised constantly and always on the authority. The +missing check has no observable consequence in the environment where the sample runs. + +**Why the obvious check misses it.** Authority checks are usually reviewed at the +RPC boundary, and these methods are not RPCs — they are ordinary virtual functions +that a caller in the right context invokes correctly. Being public and virtual is +what makes it dangerous: the class invites overriding and calling from anywhere, and +the invitation carries no precondition. + +**Symptom.** A client-side subclass or a Blueprint calls the method during +prediction, the client's death state diverges from the server's, and the resulting +desync is chased as a replication bug. + +**Detect.** List state writers and check each for an authority guard: + +```bash +rg -n "void (Start|Finish|Set)\w*(Death|State|Team)\w*\(" Source/ +rg -n -A6 "void (Start|Finish|Set)\w*(Death|State|Team)\w*\(" Source/ | rg -c "HasAuthority" +``` + +A count below the number of writers names the unguarded ones. + +**Guardrail.** Every writer of replicated state either checks authority or is +private with a single guarded caller. Public plus virtual plus unguarded is an +invitation. + +--- + +### RA-16 - Faction query fails open on invalid input + +**Mechanism.** A team comparison returns an "invalid argument" result for unknown +membership, and the damage path treats anything that is not an explicit "same team" +as damageable. The site carries a TODO calling itself temporary. + +**Why it is silent.** Unknown membership does not occur in a sample where every pawn +is assigned a team at spawn. The fail-open branch is never taken, so it is never +observed to be wrong. + +**Why the obvious check misses it.** The function returns a three-valued result, +which reads as careful design. The defect is in the *consumer's* collapse of three +values into two, in a different file. Reviewing the query finds correct code; +reviewing the damage path finds a plausible boolean. + +**Symptom.** Neutral, spectating or mid-transition actors take or deal damage. The +bug appears only in modes the sample does not have, which is to say, in yours. + +**Detect.** Find three-valued results collapsed to booleans: + +```bash +rg -n "Invalid|Unknown|Indeterminate" Source/ | rg -i "team|faction|relation" +rg -n -B4 -A4 "CanCauseDamage|CanDamage|IsHostile" Source/ +``` + +Read the branch: does the invalid case land with "allowed" or "denied"? + +**Guardrail.** Unknown must be denied, never allowed, on any decision that costs +something. If denial is not acceptable, the unknown state itself is the bug. + +--- + +### RA-17 - No tag or asset redirects declared at all + +**Mechanism.** Config declares zero gameplay tag redirects. Every tag name is a hard +string with no migration path. + +**Why it is silent.** A project that has never renamed a tag needs no redirects. The +absence is invisible right up to the first rename, and samples do not rename after +release. + +**Why the obvious check misses it.** This is an absence with no consumer to inspect +— there is no line of code to review. It is only visible if you go looking for a +capability you have not needed yet, which is exactly the thing nobody does during +adoption. + +**Symptom.** The first tag rename after content exists breaks every asset that +referenced it. Since references live in binary assets, the breakage is silent data +loss rather than a compile error, discovered per-asset over weeks. + +**Detect.** + +```bash +rg -c "GameplayTagRedirects|\+ActiveGameRedirects|ClassRedirects" Config/ +``` + +Zero, in a project that has shipped content, means renames have not yet happened — +not that they are safe. + +**Guardrail.** Establish redirect discipline before the first rename, not after. +Add a CI check that a removed tag name has a corresponding redirect entry. + +--- + +### RA-18 - Server RPC declared without validation + +**Mechanism.** A quick-bar style component exposes a `Server, Reliable, +BlueprintCallable` function with no `WithValidation`, accepting an index from the +client. + +**Why it is silent.** Well-behaved clients send valid indices, and the sample only +ever has well-behaved clients. Out-of-range access either clamps harmlessly or hits a +path no test covers. + +**Why the obvious check misses it.** Nearby functions are marked +`BlueprintAuthorityOnly`, which reads as an access restriction and satisfies a +reviewer scanning for guards. That specifier constrains Blueprint execution context +only — it does not validate parameters, and it does not constrain C++ callers at all. +A guard that answers a different question is worse than no guard, because it stops +the search. + +**Symptom.** A modified client sends an arbitrary index. Depending on the container, +this is a crash, a read of unrelated memory, or equipping something the player never +earned. + +**Detect.** + +```bash +rg -n "UFUNCTION\([^)]*\bServer\b[^)]*\)" Source/ Plugins/ | rg -v "WithValidation" +``` + +Every hit takes client-controlled input on trust. + +**Guardrail.** Every server RPC validates its parameters, including range and +ownership, and `BlueprintAuthorityOnly` is never counted as validation. + +--- + +### RA-19 - Readiness gate admits simulated proxies with no controller + +**Mechanism.** An initialisation-state transition treats "has a controller" as a +precondition, but admits simulated proxies that have none, so readiness is reached +on different grounds depending on net role. + +**Why it is silent.** Both branches produce "ready", and ready is the state everything +downstream waits for. No consumer asks *why* readiness was granted. + +**Why the obvious check misses it.** The gate is short and reads as a careful +role-aware special case — which is what it is. The problem is that "ready" then means +two different things, and the divergence lives in every consumer rather than in the +gate. Reviewing the gate finds nothing wrong. + +**Symptom.** Systems that assume a controller exists after readiness crash or +no-op on simulated proxies. The failure appears only with a remote client, so it +survives all local testing. + +**Detect.** Find role-dependent readiness and enumerate its consumers: + +```bash +rg -n -B4 -A10 "CanChangeInitState|HasReachedInitState" Source/ | rg "IsLocallyControlled|GetController|ROLE_SimulatedProxy" +rg -n "HasReachedInitState|GameplayReady" Source/ | wc -l +``` + +**Guardrail.** One readiness state means one set of guarantees. If proxies reach it +by a different route, they need a different state name, so consumers must choose. + +--- + +## How to use this list + +Do not read it as a defect report about someone else's project. Read it as a list of +**question shapes**: + +1. Which of my editable properties has no writer? +2. Which of my ordering fields never reaches a comparison? +3. Which of my decision branches reads a literal? +4. Which of my parameters is threaded but never branched on? +5. Which of my unknown-input paths fails open? +6. Which of my public state writers has no authority check? +7. Which of my capabilities exists only as configuration that is switched off? + +Each has a one-line `Detect` above. Run all seven against your own tree before you +run any of them against the reference. + +## Evidence boundary + +Every entry was read in source in one reference project on one engine version. They +are worked examples of the classification, not a defect list to carry into other +engine versions or other samples. An entry that does not reproduce under its own +`Detect` in your tree does not apply to you. diff --git a/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/patterns.md b/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/patterns.md new file mode 100644 index 0000000..1dfcb37 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-reference-project-adoption/references/patterns.md @@ -0,0 +1,294 @@ +# Adoption patterns: a worked classification + +This is the C1-C5 classification from `SKILL.md` applied end to end against one +real reference project, plus the tax analysis and vocabulary hygiene that follow +from it. + +Individual silent failures are not repeated here. They live in +[failure modes](failure-modes.md), one entry each, with a detection recipe you +can run against your own tree. This file is about *how the categories behave* and +what each one costs, so that when you meet a new element you can place it. + +--- + +## 1. What each category feels like from the inside + +The categories are not severity levels. They are answers to a different question: +**what does the author of the reference believe about this element?** + +| | Author's belief | Your evidence | Your move | +|---|---|---|---| +| **C1 Defect** | "this works" | it demonstrably does not | fix the line, keep the design | +| **C2 Showcase stub** | "this is not finished" | reads as finished | implement or delete the field | +| **C3 Scope narrowing** | "a sample does not need this" | nothing is wrong at all | budget it or accept the risk in writing | +| **C4 Architecture tax** | "this is worth it for us" | correct, working, expensive | adopt by need, not completeness | +| **C5 Foreign legacy** | nobody believes anything | a name from another product | rename on entry | + +The dangerous boundary is C2 against C3, because from the outside they look +identical: a mechanism that produces a neutral result forever. The distinction is +whether the author intended to come back. It changes nothing about your cost, and +everything about how you talk to your team: a C2 is a hole in the reference, a C3 +is a hole in your understanding of what you bought. + +--- + +## 2. C1 - Defects + +A defect in a reference project is the cheapest category and the least +interesting. Fix the line, keep the surrounding design, and do not let it +discredit the pattern that contains it. + +Three shapes recur often enough to be worth naming: + +**The filter that filters nothing.** A field constrains designer input to a +namespace, and the namespace does not exist in this project. The picker comes up +empty, which reads as "no matching entries yet" rather than "this constraint is +misconfigured". Where such filters are rare in a codebase, each one is load-bearing +and worth checking individually. + +**The symmetric typo.** A misspelled identifier used as a string join across +config, code and platform overrides works perfectly, because every side misspells +it identically. It is only discovered by someone fixing the spelling in one place. +The lesson is not "check spelling"; it is that when a string literal is the join, +correctness means *all sides agree*, and the way to get that is one declaration +site referenced everywhere else. + +**The loop over an empty container.** A local collection is declared next to a +TODO, never populated, and then iterated. The loop reads as a mechanism. It is a +placeholder for one. + +None of these survives a careful read of the consumer. They survive a read of the +declaration, which is why they get copied. + +--- + +## 3. C2 - Showcase stubs + +The defining property is precise, and worth memorizing: + +> The declaration exists. The consumer exists. The **writer** does not - or the +> consumer is a literal. + +Neither a compiler warning nor an "unused symbol" search finds these, because the +symbol is used. Reading the code convinces you it is wired up. This is why the +highest-yield check in the whole skill is searching for assignment rather than +mention. + +Sub-shapes, each with an entry in the failure-mode index: + +- a tunable compared against a timestamp that is never assigned (RA-06); +- an ordering field stored, copied through the entire API surface, and never + compared (RA-07); +- an editable tag container whose consumer is a hardcoded `false` (RA-08); +- a parameter threaded through several call layers, whose only call site passes a + literal, waiting for a subsystem that was never written (RA-09); +- a field dropped silently in both directions of a conversion helper (RA-10); +- a replicated structure with no callers at all (RA-11). + +### The trap inside the trap + +A stub can be deeper than its TODO admits. In one measured case, implementing the +obvious missing condition would still not restore the intended behaviour, because +a second method on the same class overwrites the designer-authored values on its +first call. The TODO points at one line; the repair is two. + +Treat a TODO as evidence that the author knew about *something*, not as a +specification of what is missing. Read the whole class before pricing the fix. + +### What to do + +Either implement it before exposing it, or delete the field. There is no third +option that is honest. An editable property whose consumer is a literal constant +is a trap for designers, who will tune it for days and then report that the build +is broken. + +--- + +## 4. C3 - Deliberate sample-scope narrowings + +Nothing here is broken. The reference is internally consistent with every one of +these. They become expensive only when the sample becomes the base of a product. + +This is the category that costs a subsystem rather than a line, and the one that +adoption discussions habitually skip, because there is nothing to point at. + +### The largest one: no lag compensation + +When a reference project has no server-side rewind - no historical position +storage, no custom saved-move pipeline - everything downstream follows +mechanically: + +- the server does not trace, because tracing is gated on being locally + controlled, and the server controls nobody; +- the client's target data is consumed as authoritative, because there is nothing + to compare it against; +- damage falloff is computed from the client's own reported trace origin; +- surface-material multipliers are taken from the client's reported hit. + +Read as a defect list, this is four bugs. Read correctly, it is one design +decision with four consequences. "The reference forgot to validate hits" is the +wrong diagnosis, and it leads to four local patches that do not add up to server +authority. + +**Adoption decision:** choose a hit-registration model *before* borrowing a weapon +stack. Server-authoritative with rewind, client-claim with plausibility checks, or +full trust. A sample usually implements the third. Each has a different cost and a +different cheat surface, and the choice dictates how abilities, tracing and damage +are structured - which is why it cannot be retrofitted cheaply. + +### Configuration that ships disabled + +A subsystem can be present, implemented, documented by its own console variables, +and switched off in shipped configuration with an essentially empty routing table. +The implementation being there tells you nothing about whether it was ever +exercised. Copying that configuration copies untested settings, and the settings +are the part that needs production-scale traffic to tune. + +### Machinery made unreachable by a mode switch + +A load-mode enum set to "load everything upfront" can short-circuit every branch +that would install a delay-load subsystem's delegates. The inherited machinery is +functional; the sample bypasses it. The diagnostic tools built for that subsystem +then report zeros forever, which reads as "nothing to report" rather than "this +code never runs". + +Important qualification: this is a *configuration choice of the sample*, not an +engine defect. Diagnosing it as a broken engine feature sends you into engine +source for no reason. Check the mode before you check the mechanism. + +### The rest, in one table + +| Narrowing | What it costs a product | +|---|---| +| State transitions on a health or death component written without an authority check, on public virtual methods | Death becomes a state with no owner; every caller is an authority | +| Team or faction lookup that fails open on an invalid argument | Unknown membership resolves to "damage allowed" | +| No tag or asset redirects configured anywhere | Every rename is a breaking change with no migration path | +| A reliable server RPC declared without validation | Unvalidated input from the network; a Blueprint-authority marker does not constrain C++ callers | +| An init-state gate that admits simulated proxies with no controller | "Ready" means something different per net role | + +Each row is a line item in your adoption budget, not a bug report against the +reference. + +### The rule + +> Every C3 narrowing you accept must be written down as an accepted risk with an +> owner, or scheduled. "We did not notice" is the only unacceptable outcome. + +--- + +## 5. C4 - Architecture tax + +Correct, working, and sized for the reference author's team, release cadence and +platform matrix. The question is never "is this good architecture". It is "does my +product have the problem this solves". + +### Measured shape of one reference project + +| Element | Measured | When it stops paying | +|---|---|---| +| Definition-to-instance stack: experience selects feature plugins, which contribute action sets, which grant abilities and data | 5 gameplay-feature plugins out of 81 plugins total | One game mode and no post-launch content: you buy distributed control flow and receive no modularity in return | +| C++ / Blueprint boundary | 493 Blueprints scanned, median 5 nodes each; 1256 `BlueprintCallable` entry points against 45 `BlueprintImplementableEvent` plus `BlueprintNativeEvent` override points | The 28:1 ratio means C++ exposes *calls*, not *extension points*. That assumes a permanently available C++ team. Without one, designers stall on their first non-trivial request | +| Plugin-mounted content roots | primary game root holds 2 838 assets; all 105 mount roots hold 17 256; project-owned content is 3 876 | Any census, budget or memory number computed over the primary root alone understates by roughly six times | + +The second row is the one people copy without noticing they copied it. A C++ to +Blueprint ratio is a staffing decision wearing an architecture costume. Measure +both columns before adopting the boundary, and name the person who will add to the +second one. + +The third row is a measurement discipline, not an architecture choice, and it +generalizes: in any plugin-based project, enumerate all mount roots. Tooling that +defaults to the primary game root will quietly report a fraction of reality. + +### How to decide + +State the question the modular stack answers *for you*: + +- more than one team shipping into the same binary? +- content added after launch without a client patch? +- game modes swapped at runtime by data? +- platform-specific feature sets? + +Zero yes answers means the whole stack is tax. One or two means adopt the layers +that answer them and stop there. Adopting for completeness is paying for an answer +to a question you do not have. + +--- + +## 6. C5 - Foreign-product legacy + +Three recurring shapes: + +- global constants and type identifiers carrying another title's brand, usually + because the subsystem was lifted from that title; +- a tag whose own descriptive comment admits it is in the wrong namespace and + needs to move; +- a class comment copy-pasted from a different class, describing something the + class is not. + +Harmless as code. Corrosive as vocabulary. Within a year nobody remembers why a +primary asset type is named after another game's content, and part of the team +assumes it is an engine term. The comment case is worse than the constant case: a +wrong name is eventually questioned, a wrong explanation is believed. + +Rename at import time. After the first sprint it becomes archaeology. + +--- + +## 7. The counterexample worth copying + +A register-and-unregister pair, in which the unregister path removes exactly the +paths that registration added and then asserts on the removal count: + +```cpp +// on unregister +const int32 NumRemoved = Registry.RemovePathsAddedBy(Feature); +ensure(NumRemoved == PathsAddedBy(Feature).Num()); +``` + +In an audit whose dominant finding was "add works, remove is incomplete", this was +the one place where teardown was both implemented and asserted. Copy the shape: +paired add/remove, plus an assertion that the counts match. The assertion is what +turns a silent leak into a test failure. + +Note the asymmetry that remained even there: registration went through the +project's own manager subclass while unregistration reached for the engine base +class, and a refresh call made on the way in had no counterpart on the way out. +Even the good example is worth reading twice - copy the shape, fix the asymmetry. + +--- + +## 8. What source reading cannot settle + +Three limits are worth stating, because they bound every claim in this skill: + +1. **Blueprint call sites are invisible to source search.** A function with zero + C++ callers may be called from a Blueprint graph. Absence of C++ callers is not + absence of callers; settling it requires opening the graphs. +2. **Data tables and data assets are binary.** How many entries a table + contributes to a registry cannot be read from text. It requires the editor. +3. **An empty search result proves the pattern, not the absence.** Widen the + pattern, check the path, and check what your search tool excludes by default + before concluding that something does not exist. + +The third has caught real errors in this material more than once. Treat it as a +standing rule rather than a caveat. + +--- + +## Provenance + +The classification and every measured number above come from a line-by-line audit +of Epic's Lyra Starter Game on Unreal Engine 5.6, read as source rather than run, +during a research project whose product was this skill set. + +Source addresses stay in that archive. What ships is the recipe, because an +address in a tree you do not have is not actionable and a recipe is. Entry +identifiers (`RA-01` and up) in [failure modes](failure-modes.md) resolve back to +the audited locations, so any specific claim can be produced on request. + +## Evidence boundary + +One reference project, one engine version, one workspace. This is a worked example +of the classification, not a defect list to carry into other engine versions or +other samples. Re-run the tests from `SKILL.md` against your own reference before +trusting any specific claim here. diff --git a/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/SKILL.md b/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/SKILL.md new file mode 100644 index 0000000..b0ab22c --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/SKILL.md @@ -0,0 +1,265 @@ +--- +name: ue-runtime-allocation-and-caching +description: >- + Design or audit runtime allocation, pooling, caching and replication cost in + Unreal Engine: per-frame and per-event churn, object pools, bounded caches and + ring buffers, pointers into container storage, strong versus weak retention of + spawned components, fast-array and subobject replication, dormancy and + relevancy, and dedicated-server exclusions. Use when profiling hitches, + investigating memory that grows during a session, effect or UI churn, or + deciding how to reduce replicated object count. +--- + +# UE runtime allocation and caching + +The invariant: + +> Every cache needs an owner, a bound and an eviction rule. Every pooled object +> needs a stable handle. Every replication optimization must name which cost it +> reduces — memory, CPU or bandwidth — because they are not the same. + +Measured patterns from a reference product: [patterns](references/patterns.md). +Detection recipes: [failure modes](references/failure-modes.md). + +Related skills: `ue-asset-loading-and-memory`, `ue-cosmetics-and-teams`, +`ue-gameplay-messaging`, `ue-streaming-and-platform-budgets`. + +Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`. + +--- + +## 1. Never store a raw pointer into a container's storage + +This is the highest-severity defect class in this skill, and it is the one that +looks most harmless: + +```cpp +FPooledList& Pool = PooledMap.FindOrAdd(Key); // address of a map value +LiveEntries.Emplace(Component, &Pool, ReleaseTime); // held across frames +``` + +Maps and arrays reallocate on insertion. A later insert invalidates every stored +address. The failure is **data-dependent**: it needs a second key inserted while +an earlier entry is still live, so it survives casual testing and appears under +load. + +A null check does not help. The pointer is not null; it is non-null garbage. + +### Rules + +- store a **key or an index**, and resolve at the point of use; +- if a stable address is genuinely required, use a container that guarantees + stable element addresses, and say so in a comment at the declaration; +- treat "pointer obtained from a container, stored beyond the current scope" as a + review blocker, not a style note; +- when you suspect one, the allocator's stomp mode converts a probabilistic + corruption into a deterministic crash at the right line. + +The same shape appears with lazily-growing statistics maps whose entries are +cached by a widget and dereferenced every frame — the map grows a key when a new +data source appears, and the widget's pointer dies mid-session. Recipes: RC-01, +RC-02. + +--- + +## 2. Bound every cache and name its eviction + +A cache without a bound is a leak with good intentions. For each one, write down: + +```text +owner who invalidates it +key space bounded by what? +capacity a number, or "unbounded" with a justification +eviction time, count, or explicit removal +retention strong or weak +teardown what empties it, and when +``` + +Three shapes recur and each is silent: + +**Read-modify-write growth.** Fetch a collection, append one element, write the +whole thing back. Per-event cost grows linearly and total work is quadratic over +a session — and if nothing ever clears it, the collection is also a leak. Recipe: +RC-03. + +**Re-append without pruning.** A component that re-adds every previously spawned +effect on each trigger and never removes finished ones. When the arrays are +strong references, every one-shot effect is pinned for the whole match. Recipe: +RC-04. + +**Prefilled buffers reported as data.** A fixed-capacity sample cache filled with +zeroes reports a misleading minimum and average until it wraps. Track the fill +count, or exclude the prefix. Recipe: RC-08. + +--- + +## 3. Pools need a lifecycle, not just reuse + +A pool is correct when acquire and release are symmetric and idempotent; released +objects are reset to a known state; there is a high-water policy — grow, cap, or +trim on idle; pooled objects remain reachable by the collector while pooled; the +pool empties on teardown; and a leaked acquire is detectable rather than silent. + +Fixed-size pools must define what happens on exhaustion: block, drop, or grow. +Silent drop is acceptable for cosmetic effects only, and it should be counted. + +**A pool that only grows is not a leak in the technical sense and behaves like one +in practice.** One burst leaves its peak allocated for the rest of the session. + +--- + +## 4. Churn: diff, do not rebuild + +A periodic scan that destroys and recreates its descriptor objects every interval +allocates, invalidates observers, and hides real changes in noise. + +```text +compute the desired set +diff against the current set +add new, remove missing, leave the rest alone +``` + +Reserve capacity once when the size is predictable. In hot paths prefer the +reset-without-freeing operation over the one that releases the buffer — the +latter guarantees a reallocation on the next frame. Do not allocate in paint or +per-frame render-thread paths at all. Recipes: RC-05, RC-06, RC-07. + +--- + +## 5. Protocol identifiers need width and uniqueness discipline + +A batch identifier assigned as *the current in-flight count* rather than a +monotonic sequence produces reuse: A gets 0, B gets 1, A confirms and is removed, +C gets 1 again — and a handler that matches the first plausible entry applies the +confirmation to the wrong batch. + +Compounding it, the same value can be one width at the source, another in the +message and a third in storage — which turns a wraparound into a permanent +mismatch, and an unbounded pending list into an unbounded leak. + +### Rules + +- identifiers are sequences, never counts; +- one width for the whole path, asserted at every boundary; +- define wrap behaviour explicitly; +- handlers match on identity, not on "the first entry that could be it"; +- an unbounded pending list needs a timeout and a drain policy. + +Recipe: RC-09. + +--- + +## 6. Replication: name the axis you are optimizing + +Three different costs, routinely conflated: + +| Mechanism | Memory | CPU | Bandwidth | +|---|---|---|---| +| Dedicated-server presentation exclusion | ↓ | ↓ | **unchanged** | +| Fast-array delta serialization | **↑** per-item key and per-connection state | **↑** delta walk | ↓ | +| Subobject replication | ↓ | ↓ | ↓ | +| Dormancy | — | ↓↓ | ↓ | +| Relevancy and replication graph | — | ↓↓ | ↓ | + +Two counterintuitive facts worth internalizing: + +- excluding cosmetics on a dedicated server saves memory and CPU and changes + bandwidth by **zero**, because the intent list still replicates; +- fast arrays are a **bandwidth** optimization that costs memory and CPU. + +The genuine object-count reducer is subobject replication: N owned items add zero +actors and zero channels. + +### Before claiming a replication optimization + +- [ ] state which axis improves and which regresses; +- [ ] measure that axis specifically, not "performance"; +- [ ] confirm the mechanism is actually **enabled** rather than merely present — + see `ue-multiplayer-authority` for a measured case where the replication + graph ships disabled, dormancy is unused, and a force-update list is + declared, read, and never populated; +- [ ] verify net-mode gating is consistent across presentation systems; +- [ ] check replicated history structures for a removal path. + +--- + +## 7. Dedicated-server exclusions must be systematic + +List every client-only construction and guard it in one reviewable place: +cosmetic actors and meshes, widgets and layouts, effect and audio components, +preview and capture systems, input and settings objects. + +**An inconsistent set is worse than none**, because server memory profiles become +unpredictable — one system guarded while a sibling spawning real actors is not. +Audit by searching for the net-mode guard and diffing against the list of +presentation systems. + +--- + +## 8. Review checklist + +### Allocation + +- [ ] no stored pointers into container storage; +- [ ] hot paths reset rather than free, with capacity reserved; +- [ ] no allocation in paint or render-thread paths; +- [ ] periodic work diffs rather than rebuilds. + +### Caches and pools + +- [ ] owner, bound, eviction and teardown written down for each; +- [ ] key space provably bounded; +- [ ] retention strength intentional; +- [ ] acquire and release symmetric; exhaustion policy defined; +- [ ] prefilled buffers do not distort statistics. + +### Protocol + +- [ ] identifiers are sequences with one consistent width; +- [ ] pending lists have a timeout and a drain; +- [ ] handlers match by identity. + +### Replication + +- [ ] the optimization names its axis; +- [ ] the mechanism is verified enabled; +- [ ] net-mode gating consistent across presentation systems; +- [ ] replicated history bounded. + +--- + +## 9. Measurement plan + +- **long-session soak**: fifty to a hundred spawn, despawn and transition cycles, + looking for monotonic growth rather than an absolute number; +- per-frame allocation counters in the hot scenes; +- the allocator's stomp mode **first** when a dangling-pointer class is + suspected — it is a binary answer, and it costs one run; +- profiler counters for boundedness curves over time, which is the cheapest way + to turn "is this cache bounded?" into a graph; +- network statistics and a network profiler for the replication axes, measured + separately from each other; +- server and client measured independently; never generalize one to the other. + +**Report growth curves, not single samples.** A cache that is 4 MB after one +minute and 40 MB after ten is a defect regardless of where it eventually stops. + +--- + +## Provenance + +The patterns and failure modes in `references/` come from a source audit of the +runtime, feedback, UI and weapon systems of Epic's Lyra Starter Game on Unreal +Engine 5.6, read as source rather than profiled. No profiler was run, so there +are no timing numbers here by construction — the findings are structural, and the +recipes tell you how to measure them yourself. + +Source addresses stay in the research archive that produced this skill. Each +entry carries a stable identifier (`RC-01` and up) that resolves back to the +audited location, so any specific claim can be produced on request. + +## Evidence boundary + +One project, one engine version, one workspace. The defect classes are general; +the specific instances are evidence, not guarantees about other versions. Re-run +the detection recipes against your own tree before acting on any specific claim. diff --git a/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/failure-modes.md b/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/failure-modes.md new file mode 100644 index 0000000..2192518 --- /dev/null +++ b/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/failure-modes.md @@ -0,0 +1,384 @@ +# Failure modes: runtime allocation and caching + +Ten ways a running game allocates more each minute, keeps what it should drop, or +dereferences memory that moved. + +The shared property splits this list in two, and the split is worth stating +because it changes how you look for each half. + +**The dangling-pointer half is data-dependent.** It needs a second key inserted +while a first entry is still live, or a new statistic to appear mid-session. +Neither happens in a short test, both happen in a match, and the crash lands far +from the cause. A null check never catches it, because the pointer is not null. + +**The growth half is time-dependent.** Nothing is wrong at any instant; the +defect is a slope. A single memory sample cannot show it and a single frame +cannot contain it, which is why every recipe below asks for two measurements +rather than one. + +Recipes use `rg` from a project source root, and were executed against the +audited project while this file was written. + +--- + +## Pointers into moving storage + +### RC-01 - A raw pointer into a map value, held across frames + +**Mechanism.** A reference to a map value is taken, its address is stored in a +long-lived entry, and the map later gains a key. The map reallocates; the stored +address points into freed memory. + +**Why it is silent.** Until a second key arrives, the address stays valid and +everything works. The window is narrow and specific — a live entry plus an insert +— so ordinary play produces it and ordinary testing does not. + +**Why the obvious check misses it.** The code often *has* a check, and the check +passes: the pointer is not null, it is non-null garbage. Reading the storing line +shows a reference to a pool, which is exactly what the entry needs. The defect is +a property of the container's reallocation behaviour, which is nowhere in view. + +**Symptom.** A crash inside the pooled object's own method, under load, with a +stack that looks like the pool is corrupt. In the audited project the window is a +one-second component lifespan plus a second visual style — trivially produced by +mixed normal and critical damage in the same second. + +**Detect.** Find addresses taken from containers and stored: + +```bash +rg -n -B2 -A4 "FindOrAdd\(|\.Find\(" --glob "*.cpp" . | rg "&\w+|Emplace\(|= &" +rg -n "\w+\*\s+\w+ = nullptr;" --glob "*.h" . | rg -i "pool|cache|list|entry" +``` + +The second search finds the *declaration* side — a bare pointer member whose name +suggests it points at a collection — which is often easier to spot than the +assignment. + +**Guardrail.** Store a key or an index and resolve at use. If a stable address is +genuinely required, use a container that guarantees stable element addresses and +say so at the declaration. When you suspect one, run under the allocator's stomp +mode: it converts a probabilistic corruption into a deterministic crash on the +right line, and costs one run. + +--- + +### RC-02 - A cache pointer handed out and stored by a widget + +**Mechanism.** A subsystem returns a pointer into its own map so a widget can +read samples cheaply. The widget stores it. The map grows a key when a new data +source appears — a network connection, an optional module — and reallocates. + +**Why it is silent.** The optimization is real and the pointer is valid at hand- +out. Growth is lazy and event-driven: the map gains keys minutes into a session, +when a connection is established or a feature is enabled, long after the widget +cached its pointer. + +**Why the obvious check misses it.** Both sides are idiomatic. Returning a +pointer to avoid copying a sample buffer is good practice; caching it to avoid a +lookup per frame is also good practice. Neither side knows the map is still +growing, and nothing in either signature says the pointer has a lifetime. + +**Symptom.** A crash in the widget's paint path, appearing only after the session +has been running long enough to acquire a new statistic — and therefore never in +a short repro. Compounded when the subsystem's teardown resets the tracker while +the widget still holds the pointer. + +**Detect.** Find getters returning interior pointers, then find who stores them: + +```bash +rg -n "const \w+\* \w+::Get\w*Data\(" --glob "*.cpp" . -A4 | rg "\.Find\(" +rg -n "const \w+\* \w+ = nullptr;" --glob "*.h" . | rg -i "widget|graph|view" +``` + +A getter returning `Find` on a member map, plus a member pointer of that type in +a widget, is the pair. + +**Guardrail.** Return a copy of the small value, or a handle the subsystem can +invalidate, or require the caller to re-fetch each frame. If a pointer must +escape, give the subsystem an invalidation broadcast and make subscribing +mandatory. + +--- + +## Unbounded growth + +### RC-03 - Read, append, write back — with no clear + +**Mechanism.** Each event fetches an entire collection from another system, +appends one element, and writes the whole thing back. Nothing ever clears it. + +**Why it is silent.** Every individual call is correct and fast. The array is a +legitimate data channel to the effects system, and its contents are consumed +visually, so nobody looks for a removal path. + +**Why the obvious check misses it.** The two lines read as an idiomatic +accessor pair. The cost — two full copies per event — is invisible at small sizes, +and the growth is invisible without a second measurement. Reviewing the function +shows three lines of correct code. + +**Symptom.** Cost per event rising through a match, total work quadratic in event +count, and memory that never returns. In a damage-number path this means the +hundredth hit costs measurably more than the first. + +**Detect.** Find get-append-set triples on the same key: + +```bash +rg -n -A3 "= \w*::Get\w*Array\w*\(" --glob "*.cpp" . | rg "Add\(|SetNiagara|Set\w*Array" +``` + +Then, for each hit, search the file for any clear: + +```bash +rg -n "Empty\(\)|Reset\(\)|SetNum\(0\)|Clear" +``` + +An append path with no clear anywhere in the file is the finding. + +**Guardrail.** Give the collection an owner and an eviction rule — a cap, a +time-out, or a clear on the event that logically ends its contents. If the +consuming system owns the lifetime, ask it for a removal API rather than +assuming one. + +--- + +### RC-04 - Re-append every previous element on every trigger + +**Mechanism.** A component that spawns audio and effect components copies all +previously active ones into temporary arrays, appends the new ones, and writes +the union back — with no removal of finished components. + +**Why it is silent.** Effects play and stop correctly, because stopping is the +component's own business. What accumulates is the *references*, and those are +invisible until something counts them. + +**Why the obvious check misses it.** The function reads as careful state +management: gather, combine, store. The absent operation is a filter for finished +components, and absence in a function that is otherwise thorough is the hardest +kind to see. + +**Symptom.** Thousands of retained effect components after a match. Because the +arrays are strong references, none of them can be collected, so this is +simultaneously a memory leak and a quadratic allocation profile in the hottest +cosmetic path — footsteps, at a few per second. + +**Detect.** Find union-rebuild patterns and check for a prune: + +```bash +rg -n -B4 -A6 "\.Append\(" --glob "*.cpp" . | rg "Empty\(\)|Reset\(\)" +rg -n "UPROPERTY\(Transient\)" -A2 --glob "*.h" . | rg "TArray