feat(skills): ship ue-design-skills bundle, licensing and delivery gate
Phase 0 of the handoff plan, as a marketplace rather than a flat skills/ directory. Content moved out of the LyraResearch archive and depersonalised: addresses stay in the archive, recipes ship. - plugins/ue-design-skills: 17 skills, 232 failure-mode entries, each with the six required fields; catalog.json as the harness-neutral source of truth and .claude-plugin/ as one adapter over it. - _gate: 16 rules, one poisoned fixture per rule, plus surface coverage so a declared file cannot silently miss the line rules. - ADR-0002 (harness-neutral bundle behind a marketplace) and ADR-0003 (split licensing: CC BY-ND 4.0 prose, Apache-2.0 code and metadata). - LICENSE files at both levels, CONTRIBUTING.md, docs/licensing-options.md as the material the licence decision grew from. Verified: gate.py 0 violations; test_gate.py 16/16 rules redden on their fixtures with a clean baseline and 2 root files reaching the line rules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,319 @@
|
||||
---
|
||||
name: ue-architecture-guardrails
|
||||
description: >-
|
||||
Audit an Unreal Engine architecture for systemic failure modes across
|
||||
data-driven definitions, feature plugins, ability systems, input, UI,
|
||||
messaging, cosmetics and teams, settings and editor workflows. Use before
|
||||
adopting a modular architecture, at design review, before enabling runtime
|
||||
feature switching, during engine upgrades, or when indirect systems have
|
||||
become hard to debug. Provides ownership, lifecycle, context, validation,
|
||||
observability and integration-test gates.
|
||||
---
|
||||
|
||||
# UE architecture guardrails
|
||||
|
||||
This skill is for **cross-system review**. It does not cover the implementation
|
||||
of any one subsystem — each of those has its own skill and its own detection
|
||||
recipes, and this one tells you which of them to run and in what order.
|
||||
|
||||
Systemic patterns and the go/no-go review: [patterns](references/patterns.md).
|
||||
Fifteen cross-boundary failure modes: [failure modes](references/failure-modes.md).
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
The invariant behind all of it:
|
||||
|
||||
> Every systemic failure in a modular architecture lives **between** two
|
||||
> subsystems that are each individually correct. That is why no subsystem's
|
||||
> tests find it, and why review by subsystem cannot either.
|
||||
|
||||
---
|
||||
|
||||
## The seven invariants
|
||||
|
||||
A modular, data-driven architecture is healthy only when all seven hold:
|
||||
|
||||
1. **Justified abstraction** — every layer pays for real, named variability.
|
||||
2. **Explicit ownership** — every dynamic resource has an owner, a receipt and
|
||||
an inverse.
|
||||
3. **Symmetric lifecycle** — activate/deactivate and create/destroy are round
|
||||
trips, tested as round trips.
|
||||
4. **Explicit context** — world, local player, authority and activation context
|
||||
are never inferred from globals.
|
||||
5. **Compiled data graph** — semantic joins, references and bundles are
|
||||
validated by something that can fail a build.
|
||||
6. **Observable indirection** — composition, blockers and listeners can be
|
||||
dumped at runtime.
|
||||
7. **Cross-boundary tests** — lifecycle is tested across subsystem seams, not
|
||||
inside one manager.
|
||||
|
||||
**If a design review cannot show evidence for an invariant, treat it as absent.**
|
||||
Not "probably fine" — absent. Every finding in the failure modes was in a
|
||||
codebase where the invariant was assumed rather than demonstrated.
|
||||
|
||||
---
|
||||
|
||||
## 1. Abstraction pressure test
|
||||
|
||||
For each proposed layer, fill this in before writing code:
|
||||
|
||||
| Layer | Concrete variability #1 | #2 | Operations it removes | New lifecycle and tests it adds |
|
||||
|---|---|---|---|---|
|
||||
|
||||
Reject or defer a layer when it has fewer than two real variants, when direct
|
||||
code would be smaller and equally safe, when the project will not use dynamic
|
||||
activation or a packaging boundary, or when the team cannot maintain the
|
||||
validation and observability the layer requires.
|
||||
|
||||
**Do not adopt a large-studio architecture as an identity marker.** The cost is
|
||||
paid per layer, per release, by whoever debugs it. Recipe: AG-01.
|
||||
|
||||
---
|
||||
|
||||
## 2. Ownership audit
|
||||
|
||||
Build a ledger for every dynamically added resource:
|
||||
|
||||
```text
|
||||
resource · owner · scope · receipt · apply · revoke · sharing policy · rollback
|
||||
```
|
||||
|
||||
Red flags, each of which has been observed in shipping code:
|
||||
|
||||
- a weak pointer presented as ownership;
|
||||
- a global subsystem that "owns everything";
|
||||
- cleanup that clears **all** resources rather than owned ones;
|
||||
- a handle discarded immediately after registration;
|
||||
- an owner that can die before the resource;
|
||||
- two owners that can both remove one shared resource.
|
||||
|
||||
### The kill-owner test
|
||||
|
||||
Destroy each owner at the worst moment: during an async load, after apply but
|
||||
before deactivation, after the target actor is gone, during map travel. Every
|
||||
owned resource must disappear or transfer explicitly.
|
||||
|
||||
Subsystem recipes: `ue-modular-gameplay` MG-01 to MG-03,
|
||||
`ue-input-architecture` IN-03, `ue-gas-architecture` GA-01. Recipe: AG-02.
|
||||
|
||||
---
|
||||
|
||||
## 3. Lifecycle as a transaction
|
||||
|
||||
```text
|
||||
validate → load → activate 1..N → committed
|
||||
failure at K → reverse K-1..1 → failed, with a reason
|
||||
```
|
||||
|
||||
Three tests, all mandatory:
|
||||
|
||||
```text
|
||||
baseline → activate → deactivate → baseline
|
||||
baseline → activate → fail each step → baseline
|
||||
baseline → activate → deactivate → reactivate → exactly one copy
|
||||
```
|
||||
|
||||
The third is the one that finds duplicates, and it is the one nobody runs during
|
||||
development, because in development you restart the editor instead. Recipe:
|
||||
AG-03.
|
||||
|
||||
---
|
||||
|
||||
## 4. Context audit
|
||||
|
||||
For every mutable map and every subscription, ask which key is actually
|
||||
required: process, game instance, world handle, activation context, local
|
||||
player, actor, or net role. **A global key is valid only when the behaviour is
|
||||
genuinely global.**
|
||||
|
||||
Required scenario matrix:
|
||||
|
||||
| Scenario | What it catches |
|
||||
|---|---|
|
||||
| multi-world editor play | process-global leaks |
|
||||
| dedicated server + client | local-broadcast and presentation assumptions |
|
||||
| listen server | double execution of local paths |
|
||||
| two local players | UI, input and settings context leaks |
|
||||
| map travel | stale listeners and state on long-lived scopes |
|
||||
|
||||
Recipes: AG-04, plus `ue-game-settings-architecture` GS-03 and
|
||||
`ue-gameplay-messaging` MB-08.
|
||||
|
||||
---
|
||||
|
||||
## 5. Compile the data graph
|
||||
|
||||
Generate and validate the whole chain: production roots → identifiers →
|
||||
references → plugin boundaries → bundles → semantic joins → leaf assets.
|
||||
|
||||
Failures that must fail a build, not a log line: unresolved identifiers, null
|
||||
required fields, plugin dependency cycles, role-mismatched bundles, orphaned
|
||||
producers or consumers of a semantic tag, prototype content under a production
|
||||
root, an input tag with no granted consumer, a UI contribution with no matching
|
||||
point, an incompatible payload family on a parent channel.
|
||||
|
||||
**Editor-only validation is feedback, not a gate.** Recipes: AG-05, plus
|
||||
`ue-data-driven-architecture` DD-09 and DD-10.
|
||||
|
||||
---
|
||||
|
||||
## 6. Temporal dependency audit
|
||||
|
||||
Draw initialization prerequisites as a directed graph and reject: any cycle; a
|
||||
prerequisite produced only by an optional or disabled path; an anonymous
|
||||
next-tick delay; a generic "extension added" event used where semantic readiness
|
||||
is required; a transition that can block with no diagnostic reason.
|
||||
|
||||
Then permute: feature activation before and after actor spawn, data before and
|
||||
after possession, replication arrival order, contribution before and after its
|
||||
point, listener before and after event. **Correct systems converge to the same
|
||||
state.** Recipes: AG-06, plus `ue-modular-gameplay` MG-14 and MG-15.
|
||||
|
||||
---
|
||||
|
||||
## 7. Observability requirements
|
||||
|
||||
At minimum, expose read-only dumps for: the selected composition and its load
|
||||
state; pending assets, plugins and actions with reasons; active feature plugins
|
||||
and who requires them; per-actor readiness state and blockers; active input
|
||||
contexts and semantic joins; granted capabilities by owner receipt; UI layers,
|
||||
points and contributions with owners; message channels and listener counts;
|
||||
authoritative identity and its mirrors; requested versus effective settings.
|
||||
|
||||
Logs need correlation fields: world, local player, actor, composition, plugin,
|
||||
action, channel, owner.
|
||||
|
||||
**If diagnosis requires stepping through many assets with no state dump, the
|
||||
indirection is operationally incomplete** — the architecture works and cannot be
|
||||
supported. Recipe: AG-07.
|
||||
|
||||
---
|
||||
|
||||
## 8. The false-genericity test
|
||||
|
||||
For every public field or hook that implies behaviour — an ordering value, a
|
||||
viewer parameter, an automatic mode, a removal function, a policy suffix:
|
||||
|
||||
1. find its consumer;
|
||||
2. mutate the value in a fixture;
|
||||
3. assert an observable delta;
|
||||
4. test both the native and the scripting API where both are exposed.
|
||||
|
||||
No consumer and no delta means: implement it, remove it, or mark it
|
||||
unsupported. **Never let a configurable no-op survive as stable API.**
|
||||
|
||||
This single test has the highest yield of anything in this skill. In the audited
|
||||
material it independently found an ordering field that never sorts, a viewer
|
||||
parameter that is ignored, an automatic quality pipeline with no caller, a
|
||||
removal function with an empty body, and a composition part with no writer.
|
||||
Recipe: AG-08, plus `ue-ui-architecture` UI-01, `ue-cosmetics-and-teams` CT-10,
|
||||
`ue-game-settings-architecture` GS-06 to GS-08.
|
||||
|
||||
---
|
||||
|
||||
## 9. Event, state and command
|
||||
|
||||
Classify every interaction:
|
||||
|
||||
- **state** — a queryable authoritative model;
|
||||
- **command** — one accountable service with a result and a failure;
|
||||
- **event** — zero or many optional observers, a past-tense fact.
|
||||
|
||||
Red flags: UI reconstructing state from transient messages; a message asking an
|
||||
unknown listener to perform required work; a local bus assumed to replicate;
|
||||
presentation treated as authority. Recipes: AG-09, plus
|
||||
`ue-gameplay-messaging` MB-12 to MB-14.
|
||||
|
||||
---
|
||||
|
||||
## 10. Version and adoption audit
|
||||
|
||||
Before copying any framework or sample, classify each element: architectural
|
||||
invariant, sample scaffolding, project-specific behaviour, incomplete path,
|
||||
confirmed defect, prototype content, version-sensitive API. Then pin the engine
|
||||
version, the sample version, plugin versions and any serialized schema version,
|
||||
and re-run behaviour probes after every upgrade.
|
||||
|
||||
**A current documentation page does not prove the content you copied is
|
||||
current.** See `ue-reference-project-adoption` for the full classification and
|
||||
its four tests. Recipe: AG-10.
|
||||
|
||||
---
|
||||
|
||||
## 11. Risk-based integration matrix
|
||||
|
||||
Do not attempt a Cartesian product. Require this spine:
|
||||
|
||||
1. standalone baseline;
|
||||
2. dedicated server + remote client;
|
||||
3. listen server;
|
||||
4. two local players;
|
||||
5. feature activates **before** receiver readiness;
|
||||
6. feature activates **after** receiver readiness;
|
||||
7. deactivate and reactivate;
|
||||
8. actor destroyed before feature teardown;
|
||||
9. late join;
|
||||
10. map travel or composition change;
|
||||
11. packaged clean client;
|
||||
12. engine or plugin upgrade fixture.
|
||||
|
||||
**Every critical invariant needs at least one test crossing two subsystem
|
||||
boundaries.** Unit tests inside one manager cannot find anything in this skill.
|
||||
|
||||
---
|
||||
|
||||
## 12. Severity triage
|
||||
|
||||
**Critical — stop feature work.** Ownership or rollback unknown; authority
|
||||
duplicated; world or local-player context leaked; production data graph
|
||||
unvalidated; persistent state reconstructed from messages; client presentation
|
||||
affecting authority.
|
||||
|
||||
**High — fix before runtime switching or shipping.** Initialization cycle or
|
||||
timing dependency; missing bundle with no fallback; native/scripting semantic
|
||||
mismatch; undocumented ordering dependency; insufficient observability; version
|
||||
mismatch.
|
||||
|
||||
**Medium — track with a bound.** Abstraction tax; linear scans acceptable at
|
||||
current data size; editor-only gaps; duplicated scaffolding.
|
||||
|
||||
---
|
||||
|
||||
## 13. Go/no-go
|
||||
|
||||
- [ ] Every layer has two named variability pressures.
|
||||
- [ ] Ownership ledger complete, kill-owner test passes.
|
||||
- [ ] Lifecycle rollback and reactivation tests pass.
|
||||
- [ ] Context matrix passes.
|
||||
- [ ] Data graph validation runs where it can fail a build.
|
||||
- [ ] Prerequisite graph acyclic and observable.
|
||||
- [ ] State dumps and log correlation exist.
|
||||
- [ ] Every public generic hook has a consumer test.
|
||||
- [ ] Event, state and command classifications are correct.
|
||||
- [ ] Replicated intent has a packaged-client realization test.
|
||||
- [ ] Version set pinned.
|
||||
- [ ] Cross-boundary matrix automated or scheduled.
|
||||
|
||||
**No-go rule: three unchecked Critical gates mean stop adding abstraction.**
|
||||
Repair ownership, context and validation first — further indirection multiplies
|
||||
the failure surface rather than the capability.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The systemic patterns and their severities come from a cross-system audit of
|
||||
Epic's Lyra Starter Game on Unreal Engine 5.6 — the same audit that produced the
|
||||
subsystem skills in this bundle. This skill exists because the same shapes
|
||||
appeared in every subsystem: the individual findings are in those skills, and
|
||||
what is here is the structure they share.
|
||||
|
||||
Source addresses stay in the research archive. Each entry carries a stable
|
||||
identifier (`AG-01`, `AG-02`, …) resolving back to the audited material there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Cross-system conclusions from one project at one engine version, read as source
|
||||
rather than run. Where a finding is a general principle rather than a measured
|
||||
instance, the failure modes say so explicitly. The subsystem recipes referenced
|
||||
above are the measured half; this skill is the ordering.
|
||||
+379
@@ -0,0 +1,379 @@
|
||||
# Failure modes: cross-system architecture
|
||||
|
||||
Ten failures that belong to **no single subsystem**, which is why each one
|
||||
survived review of every subsystem it passes through.
|
||||
|
||||
The shared property here is different from the other documents in this bundle.
|
||||
Elsewhere the mechanism is "a check exists and cannot fire". Here it is: **every
|
||||
individual component is correct, and the defect lives in the seam.** Nobody owns
|
||||
a seam. A reviewer of the messaging layer sees a correct bus; a reviewer of the
|
||||
UI sees a correct widget; the failure is that the UI reconstructs state from the
|
||||
bus, and that fact is written down nowhere.
|
||||
|
||||
That has a practical consequence for how these are detected. Most recipes below
|
||||
are not a single search — they are a **join between two searches**, and the
|
||||
finding is in the difference between the two result sets.
|
||||
|
||||
Where a mechanism is fully described by another skill, the entry says so and does
|
||||
not restate it. Identifiers (`AG-01` and up) are stable and resolve back to the
|
||||
audited material in the research archive.
|
||||
|
||||
---
|
||||
|
||||
## Abstraction
|
||||
|
||||
### AG-01 - A layer adopted for completeness rather than pressure
|
||||
|
||||
**Mechanism.** A modular architecture is copied whole from a reference project
|
||||
because it is the reference's architecture, without asking which variability each
|
||||
layer absorbs. Every layer is correctly implemented and paid for in full.
|
||||
|
||||
**Why it is silent.** Nothing fails. The cost is not a defect but a permanent tax:
|
||||
every new mechanic now needs a class, a definition asset, a tag, an action, a
|
||||
plugin, a validation rule and a UI extension. The team experiences this as "the
|
||||
engine is like this".
|
||||
|
||||
**Why the obvious check misses it.** Review asks "is this implemented correctly?"
|
||||
and the answer is yes at every layer. The question that finds it — "what are the
|
||||
two concrete variants this layer separates?" — is not part of any code review,
|
||||
because it is not about code.
|
||||
|
||||
**Symptom.** Feature velocity that falls as the project matures, with no single
|
||||
slow component. Estimates that are consistently wrong in the same direction.
|
||||
|
||||
**Detect.** Count the axes against the variants they serve:
|
||||
|
||||
```bash
|
||||
rg -c "class \w+ : public UGameFeatureAction" --glob "*.h" . # action types
|
||||
fd -e uasset -p "Experiences|GameFeature" | wc -l # variants shipped
|
||||
```
|
||||
|
||||
The finding is arithmetic, not textual: **planned variants fewer than
|
||||
architectural axes.** In the audited reference, five feature plugins out of
|
||||
eighty-one total carried the modular machinery — which pays for that project and
|
||||
would not pay for a single-mode title.
|
||||
|
||||
**Guardrail.** For each layer, name two real variabilities it separates before
|
||||
adopting it. If you cannot, defer the layer. See `ue-reference-project-adoption`
|
||||
for the full classification and the adoption budget.
|
||||
|
||||
---
|
||||
|
||||
### AG-02 - Distributed control flow with no trace
|
||||
|
||||
**Mechanism.** A request travels definition → feature → action → extension event →
|
||||
component → tag → message → widget. Each hop is decoupled by design.
|
||||
|
||||
**Why it is silent.** Decoupling is working exactly as intended. The absence of a
|
||||
call stack is the feature, not a defect.
|
||||
|
||||
**Why the obvious check misses it.** A debugger shows one hop. Each hop's owner
|
||||
can explain their hop. Nobody can explain the path, because reconstructing it
|
||||
requires reading assets, config and code in three modules — and it has to be
|
||||
reconstructed again next time.
|
||||
|
||||
**Symptom.** "Who triggered this?" costs an afternoon. Bugs reproduce reliably
|
||||
while no class looks responsible. New engineers take months rather than weeks.
|
||||
|
||||
**Detect.** Test the property directly rather than searching for it: take a
|
||||
recent bug and ask an engineer who did not write the feature to reconstruct the
|
||||
path **from logs and dumps alone**, without opening assets. Then measure what
|
||||
they needed and did not have.
|
||||
|
||||
Structurally, the precondition is visible:
|
||||
|
||||
```bash
|
||||
rg -c "GetSubsystem<|BroadcastMessage|SendGameFrameworkComponentExtensionEvent" --glob "*.cpp" .
|
||||
rg -c "UE_LOG.*Verbose.*(World|LocalPlayer|Experience|Feature)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
A large first number with a small second is the finding: heavy indirection, no
|
||||
correlated logging.
|
||||
|
||||
**Guardrail.** Structured lifecycle logs carrying world, player, experience,
|
||||
plugin, action and owner identifiers; a generated composition graph; one index of
|
||||
semantic tags and their consumers. Indirection without observability is not
|
||||
architecture, it is a maze.
|
||||
|
||||
---
|
||||
|
||||
## Ownership and lifecycle
|
||||
|
||||
### AG-03 - Activation reviewed, deactivation assumed
|
||||
|
||||
**Mechanism.** Adding a component, a binding, a widget or a grant is easy and
|
||||
visible. The inverse operation is written from memory, or not at all.
|
||||
|
||||
**Why it is silent.** Development restarts the editor instead of deactivating.
|
||||
The happy path is exercised hundreds of times a day; the reverse path is
|
||||
exercised by nobody until a player switches modes.
|
||||
|
||||
**Why the obvious check misses it.** Both functions usually exist and look
|
||||
symmetric. The asymmetry is in what each enumerates — see `ue-modular-gameplay`
|
||||
recipes MG-01 through MG-06 for the six distinct shapes this takes, each with its
|
||||
own recipe.
|
||||
|
||||
**Symptom.** The second activation produces duplicates. Reported long after the
|
||||
first, and usually attributed to the feature that was activated second.
|
||||
|
||||
**Detect.** The behavioural test is stronger than any search, and it is the one
|
||||
gate this whole skill exists to insist on:
|
||||
|
||||
```text
|
||||
baseline → activate → verify additions → deactivate → verify baseline
|
||||
→ reactivate → verify exactly one copy
|
||||
```
|
||||
|
||||
Run it with the actor existing before activation, spawned after activation, and
|
||||
destroyed before deactivation.
|
||||
|
||||
**Guardrail.** An ownership ledger — resource, receipt, inverse — filled in
|
||||
**before** implementation. A blank middle column is a rejected design, not a
|
||||
follow-up task.
|
||||
|
||||
---
|
||||
|
||||
### AG-04 - A subsystem's lifetime mistaken for resource ownership
|
||||
|
||||
**Mechanism.** A long-lived subsystem holds registrations for resources owned by
|
||||
short-lived objects. Weak references prevent crashes, so nothing appears wrong.
|
||||
|
||||
**Why it is silent.** Weak pointers do their job: the dead object is not called.
|
||||
The **record** remains, and records are not visible in any profiler view that
|
||||
answers "is this leaking?".
|
||||
|
||||
**Why the obvious check misses it.** The code is defensively written and looks
|
||||
careful. Weak references read as evidence that ownership was considered — when
|
||||
they are precisely the mechanism that lets the bookkeeping rot silently.
|
||||
|
||||
**Symptom.** Registration lists that grow across a session; cleanup that resorts
|
||||
to clearing everything because per-owner removal was never possible; the same
|
||||
resource removed twice by two owners.
|
||||
|
||||
**Detect.** For every registry, compare adds against removes and check the key:
|
||||
|
||||
```bash
|
||||
rg -n "\.Add\(|\.Emplace\(|\.FindOrAdd\(" --glob "*.cpp" . | rg -i "listener|extension|handle|request"
|
||||
rg -n "\.Remove\(|\.RemoveSwap\(|Unregister" --glob "*.cpp" . | rg -i "listener|extension|handle|request"
|
||||
```
|
||||
|
||||
A registry with adds and no owner-keyed removal is the finding. A registry whose
|
||||
only removal is a full clear is the same finding, one step later.
|
||||
|
||||
**Guardrail.** Every dynamic resource has exactly one named owner and one receipt.
|
||||
Shared resources use reference counts or leases. "The subsystem owns it" is not an
|
||||
answer; subsystems outlive the things they track.
|
||||
|
||||
---
|
||||
|
||||
## Context
|
||||
|
||||
### AG-05 - Global state where the scope is world or player
|
||||
|
||||
**Mechanism.** A registration, cache or setting is keyed globally, while the
|
||||
things it describes belong to a world, a local player or an activation context.
|
||||
|
||||
**Why it is silent.** With one world and one player — the configuration in which
|
||||
almost all testing happens — global and scoped are indistinguishable. The code is
|
||||
correct in the case you run.
|
||||
|
||||
**Why the obvious check misses it.** The accessor reads naturally: a player asking
|
||||
for its own settings, a subsystem holding its own registry. The scope error is one
|
||||
line inside an accessor, or a missing key in a map declaration.
|
||||
|
||||
**Symptom.** Multi-world editor sessions cross-contaminate; split-screen players
|
||||
share what should be per-player; a listen server processes an event twice. Each
|
||||
appears as an unrelated bug in a different subsystem.
|
||||
|
||||
**Detect.** For every mutable registry, ask what the minimum sufficient key is,
|
||||
then check what it actually is:
|
||||
|
||||
```bash
|
||||
rg -n "TMap<.*>\s+\w+;" --glob "*.h" . | rg -v "FObjectKey|FGameFeatureStateChangeContext|ULocalPlayer"
|
||||
rg -n -A4 "::Get\w*Settings\(\)" --glob "*.cpp" . | rg "::Get\(\)|GEngine->"
|
||||
```
|
||||
|
||||
The second search is the specific case documented in
|
||||
`ue-game-settings-architecture` GS-03: a per-player accessor returning a global
|
||||
singleton, whose tell is a proliferation of "primary player only" conditions
|
||||
elsewhere.
|
||||
|
||||
**Guardrail.** Key every mutable record by the minimum context that makes it
|
||||
correct — activation context, world handle, local player where applicable. Then
|
||||
run the required matrix: multi-world editor, dedicated server plus client, listen
|
||||
server, two local players, map travel.
|
||||
|
||||
---
|
||||
|
||||
### AG-06 - Editor behaviour that differs from the shipped configuration
|
||||
|
||||
**Mechanism.** A branch keyed on running in the editor loads more, validates less,
|
||||
or guesses identifiers that the packaged build resolves strictly.
|
||||
|
||||
**Why it is silent.** Both branches are correct for their environment. The editor
|
||||
branch is usually more permissive, so everything works better where you are
|
||||
looking.
|
||||
|
||||
**Why the obvious check misses it.** The branch is a single condition in a
|
||||
subsystem nobody reads while working on a feature. Its consequences appear in
|
||||
memory profiles, cook results and packaged-only failures — three places that are
|
||||
each somebody else's job.
|
||||
|
||||
**Symptom.** Memory numbers that describe no shipping configuration. Features that
|
||||
work in the editor and silently do nothing when packaged. Asset identifiers that
|
||||
resolve in one and not the other.
|
||||
|
||||
**Detect.** Enumerate every editor divergence and judge each one deliberately:
|
||||
|
||||
```bash
|
||||
rg -n "GIsEditor|WITH_EDITOR|IsRunningCommandlet|GIsPlayInEditorWorld" --glob "*.cpp" . -A3
|
||||
rg -n "bShouldGuessTypeAndNameInEditor|PreloadInEditor|bOnlyCookProduction" Config/
|
||||
```
|
||||
|
||||
In the audited reference this finds an editor branch that loads **both** role
|
||||
bundles — which alone invalidates in-editor residency measurement — and a
|
||||
configuration that guesses asset identifiers in the editor and not in the build.
|
||||
|
||||
**Guardrail.** Keep a written list of editor divergences and their justification.
|
||||
Any measurement taken in the editor states which divergences apply to it.
|
||||
See `ue-asset-loading-and-memory` AL-10.
|
||||
|
||||
---
|
||||
|
||||
## Data and validation
|
||||
|
||||
### AG-07 - A data graph with no compiler
|
||||
|
||||
**Mechanism.** Null references, wrong identifiers, cross-plugin cycles, prototype
|
||||
content and semantic mismatches are all valid data. They load, they cook, they run.
|
||||
|
||||
**Why it is silent.** Data does not compile. There is no stage that can reject it
|
||||
except one somebody chose to write — and that validation is typically
|
||||
editor-only, so it does not run where it would matter.
|
||||
|
||||
**Why the obvious check misses it.** The validation *exists*, which satisfies the
|
||||
question "is the data validated?". What it does not do is run in the build. In the
|
||||
audited reference, eleven of twelve validation implementations were compiled out
|
||||
of non-editor builds.
|
||||
|
||||
**Symptom.** A production playlist pointing at a test mode; a health pickup
|
||||
granting a weapon definition; a missing bundle discovered only in a packaged
|
||||
build.
|
||||
|
||||
**Detect.** Count validation, then count how much of it survives the build:
|
||||
|
||||
```bash
|
||||
rg -c "IsDataValid" --glob "*.cpp" .
|
||||
rg -n -B6 "IsDataValid" --glob "*.cpp" . | rg -c "WITH_EDITOR"
|
||||
rg -n "class \w*ValidationCommandlet|UEditorValidatorBase" --glob "*.h" .
|
||||
```
|
||||
|
||||
A ratio close to one, with no commandlet or automation path, means the data graph
|
||||
is unvalidated where it ships. See `ue-data-driven-architecture` DD-09 and DD-10.
|
||||
|
||||
**Guardrail.** Run the same validation in automation that you run in the editor.
|
||||
Add a production-root allow list so prototype paths cannot reach a shipped
|
||||
playlist.
|
||||
|
||||
---
|
||||
|
||||
### AG-08 - A generic hook with no consumer
|
||||
|
||||
**Mechanism.** A public field or extension point is stored, copied and threaded
|
||||
through an API — and never read by anything that changes behaviour.
|
||||
|
||||
**Why it is silent.** Every "is this used?" check answers yes, because the value
|
||||
*is* used: passed, assigned, copied. What is missing is the comparison, the
|
||||
branch, or the sort.
|
||||
|
||||
**Why the obvious check misses it.** This is the single most repeated shape in
|
||||
this whole bundle, and it earns its own cross-system entry because it recurs in
|
||||
every subsystem independently: an ordering field that never sorts, a viewer
|
||||
identity that is ignored, a benchmark decision with no caller, a profile suffix
|
||||
that is never populated, a removal function with an empty body, a flag written in
|
||||
a constructor and never read.
|
||||
|
||||
**Symptom.** A designer configures a documented setting and observes no effect,
|
||||
concludes their data is wrong, and works around it.
|
||||
|
||||
**Detect.** The general form — mentions minus comparisons:
|
||||
|
||||
```bash
|
||||
F='Priority'
|
||||
rg -c "\b$F\b" --glob "*.cpp" --glob "*.h" . # mentions
|
||||
rg -n "\b$F\b\s*(<|>|<=|>=|==)|Sort.*\b$F\b|\b$F\b.*Sort" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Mentions without comparisons is the finding. Per-subsystem instances have their
|
||||
own recipes: `ue-ui-architecture` UI-01, `ue-cosmetics-and-teams` CT-10,
|
||||
`ue-game-settings-architecture` GS-06 and GS-07, `ue-gas-architecture` GA-06,
|
||||
`ue-modular-gameplay` MG-01.
|
||||
|
||||
**Guardrail.** Every public setting gets a consumer test: mutate it in a fixture,
|
||||
assert an observable delta. Anything without one is implemented, removed, or
|
||||
marked unsupported — never left as configurable decoration.
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
### AG-09 - Single-process testing of a distributed property
|
||||
|
||||
**Mechanism.** A listen-server host shares memory with its client, so the
|
||||
client/server split that produces a whole class of defects does not exist during
|
||||
the test that would have caught it.
|
||||
|
||||
**Why it is silent.** The tests pass. They are real tests exercising real code;
|
||||
they simply cannot express the failure.
|
||||
|
||||
**Why the obvious check misses it.** Coverage looks good and the feature demonstrably
|
||||
works. The missing dimension is a **configuration**, not a code path, so no
|
||||
coverage tool reports it.
|
||||
|
||||
**Symptom.** Features that ship having never run in the configuration they will
|
||||
run in. Bugs that appear at first playtest and are attributed to the network layer.
|
||||
|
||||
**Detect.** This one is answered by an inventory rather than a search: list the
|
||||
failure modes that require a separate process, and confirm each has a test in a
|
||||
configuration that has one. From this bundle, the ones that cannot occur in a
|
||||
single process include replicated-versus-validated confusion, call-site authority
|
||||
guards, multicast used where state belongs, incomplete replicated-array callbacks,
|
||||
readiness gates that assume a controller, and prediction with no correction path
|
||||
(`ue-multiplayer-authority` NA-01, NA-04, NA-11, NA-13, NA-17, NA-19).
|
||||
|
||||
**Guardrail.** A dedicated server with two remote clients, one under latency and
|
||||
packet loss, as the **minimum** configuration for accepting a networked feature.
|
||||
Plus a mid-match joiner.
|
||||
|
||||
---
|
||||
|
||||
### AG-10 - A version set that drifts silently
|
||||
|
||||
**Mechanism.** Engine, reference sample and plugins evolve independently. Code
|
||||
copied from a sample at one version keeps running against another.
|
||||
|
||||
**Why it is silent.** Compilation succeeds. Serialized fields still load. A stub
|
||||
that changed behaviour between versions still returns something plausible.
|
||||
|
||||
**Why the obvious check misses it.** Documentation for the current version is
|
||||
easy to find and describes the current version — not the vendored copy in the
|
||||
project. Reading the docs actively produces false confidence.
|
||||
|
||||
**Symptom.** A method that "works differently now"; a structure size assertion
|
||||
that fails after an upgrade; content that disagrees with the plugin that reads it.
|
||||
|
||||
**Detect.** Pin and verify rather than search. Where code manually enumerates the
|
||||
members of an engine structure, a size assertion is the correct tripwire:
|
||||
|
||||
```bash
|
||||
rg -n "static_assert\(sizeof\(" --glob "*.cpp" --glob "*.h" .
|
||||
```
|
||||
|
||||
In the audited reference this appears five times against one engine structure. A
|
||||
failing assertion after an upgrade is a **feature**: it means a new field would
|
||||
otherwise have been silently ignored by five hand-written functions.
|
||||
|
||||
**Guardrail.** Pin the engine, sample and plugin versions. Keep compile-time
|
||||
guards where you enumerate engine structures by hand. Re-run structural and
|
||||
behavioural probes after every upgrade, and never treat a documentation page as
|
||||
evidence about the code in your tree.
|
||||
@@ -0,0 +1,330 @@
|
||||
# Patterns: reviewing an architecture across its seams
|
||||
|
||||
This is the cross-system half of the bundle. The other fifteen skills each audit
|
||||
one subsystem; this one exists because the most expensive failures are not inside
|
||||
a subsystem at all.
|
||||
|
||||
[Failure modes](failure-modes.md) carries the detection recipes. This file
|
||||
carries the review procedure — what to ask, in what order, and what evidence to
|
||||
demand before believing an answer.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why seams are where the cost is
|
||||
|
||||
A subsystem has an owner. Someone can be asked whether the ability system is
|
||||
correct, and they can answer.
|
||||
|
||||
A seam has no owner. "The UI reconstructs state from the message bus" is not a
|
||||
fact about the UI or about the bus; it is a fact about their relationship, and
|
||||
relationships do not appear in any file. Every finding in this skill was reviewed
|
||||
— twice, once from each side — and survived.
|
||||
|
||||
That produces the review technique this whole document is built on:
|
||||
|
||||
> **Ask each question of the pair, not of the component.** Not "is this registry
|
||||
> correct?" but "which context is this registry keyed by, and which context do its
|
||||
> entries belong to?" The defect lives in the difference.
|
||||
|
||||
---
|
||||
|
||||
## 2. The seven invariants
|
||||
|
||||
An architecture of this kind is healthy only when all seven hold:
|
||||
|
||||
1. **Justified abstraction** — every layer pays for real variability.
|
||||
2. **Explicit ownership** — every dynamic resource has an owner, a receipt and an
|
||||
inverse.
|
||||
3. **Symmetric lifecycle** — activate/deactivate and create/destroy are round
|
||||
trips, tested as round trips.
|
||||
4. **Explicit context** — world, local player, authority and activation context
|
||||
are never inferred from globals.
|
||||
5. **Compiled data graph** — semantic joins, references and bundles are validated
|
||||
somewhere that runs in the build.
|
||||
6. **Observable indirection** — composition, blockers and listeners can be dumped.
|
||||
7. **Cross-boundary tests** — lifecycle is tested across subsystem seams.
|
||||
|
||||
**If a review cannot produce evidence for an invariant, record it as absent.** Not
|
||||
"probably fine" — absent. Every finding in the audit that produced this skill was
|
||||
in a system whose author would have said it was fine.
|
||||
|
||||
---
|
||||
|
||||
## 3. Abstraction pressure test
|
||||
|
||||
Fill this before adopting a layer, not after:
|
||||
|
||||
| Layer | Variability #1 | Variability #2 | Operations removed | Lifecycle and tests added |
|
||||
|---|---|---|---|---|
|
||||
| feature plugin | | | | |
|
||||
| experience / mode definition | | | | |
|
||||
| semantic input | | | | |
|
||||
| UI extension points | | | | |
|
||||
| message bus | | | | |
|
||||
|
||||
Defer a layer when it has fewer than two real variants or owners; when direct code
|
||||
would be smaller and equally safe; when the project will not use dynamic
|
||||
activation or a packaging boundary; or when the team cannot maintain the
|
||||
validation and observability the layer requires.
|
||||
|
||||
**The arithmetic that decides it:** planned variants versus architectural axes. If
|
||||
axes exceed variants, the architecture is aspirational. Recipe: AG-01.
|
||||
|
||||
---
|
||||
|
||||
## 4. The ownership ledger
|
||||
|
||||
For every dynamically added resource:
|
||||
|
||||
```text
|
||||
resource
|
||||
owner
|
||||
scope process / game instance / world / local player / actor
|
||||
receipt the handle that proves this owner added it
|
||||
apply the call that adds
|
||||
revoke the call that removes exactly this
|
||||
sharing what happens when two owners want the same resource
|
||||
rollback what happens if apply fails halfway
|
||||
```
|
||||
|
||||
Red flags, each seen in the audited reference:
|
||||
|
||||
- a weak reference presented as ownership;
|
||||
- a global subsystem that "owns everything";
|
||||
- cleanup that clears all resources rather than owned ones;
|
||||
- a handle discarded at the moment it is returned;
|
||||
- an owner that can die before its resource;
|
||||
- two owners able to remove the same shared resource.
|
||||
|
||||
### The kill-owner test
|
||||
|
||||
For each owner, destroy it at the worst moment: during an async load, after apply
|
||||
but before normal deactivation, after the target actor is destroyed, and during
|
||||
map travel. Every owned resource must disappear or transfer explicitly.
|
||||
|
||||
This is a cheap test that finds expensive defects, and almost nobody runs it.
|
||||
|
||||
---
|
||||
|
||||
## 5. Lifecycle as a transaction
|
||||
|
||||
```text
|
||||
validate → load → activate 1..N → committed
|
||||
|
||||
failure at step K
|
||||
→ reverse K-1 .. 1
|
||||
→ failed, with a reason that names the step
|
||||
```
|
||||
|
||||
Three mandatory sequences:
|
||||
|
||||
```text
|
||||
baseline → activate → deactivate → baseline
|
||||
baseline → activate → fail each step in turn → baseline
|
||||
baseline → activate → deactivate → reactivate → exactly one copy
|
||||
```
|
||||
|
||||
The third is the one that finds duplicates, and the one development never runs
|
||||
because development restarts the editor instead. Recipe: AG-03.
|
||||
|
||||
---
|
||||
|
||||
## 6. Context audit
|
||||
|
||||
For every mutable map or subscription, ask which key is **required**: process,
|
||||
game instance, world handle, activation context, local player, actor, or net role.
|
||||
|
||||
A global key is correct only when the behaviour is genuinely global. The
|
||||
temptation is that a global key is simpler and, with one world and one player,
|
||||
indistinguishable.
|
||||
|
||||
### The required matrix
|
||||
|
||||
| Scenario | What it catches |
|
||||
|---|---|
|
||||
| multi-world editor session | process-global leaks |
|
||||
| dedicated server plus remote client | local-broadcast and presentation assumptions |
|
||||
| listen server | doubled local paths |
|
||||
| two local players | UI, input and settings context leaks |
|
||||
| map travel | stale listeners surviving a world |
|
||||
|
||||
Recipes: AG-05, AG-06.
|
||||
|
||||
---
|
||||
|
||||
## 7. The data graph needs a compiler
|
||||
|
||||
Generate and validate the path from production roots through identifiers,
|
||||
references, plugin boundaries, bundles and semantic joins to leaf assets.
|
||||
|
||||
Fail the build on: unresolved identifiers or soft references; null required
|
||||
fields; plugin dependency cycles; role-bundle mismatches; orphan producers or
|
||||
consumers of a semantic tag; prototype or test paths reachable from a production
|
||||
root; an input tag with no granted consumer; a UI extension with no point or
|
||||
contract; an incompatible payload family under a parent channel.
|
||||
|
||||
**The crucial property is where it runs.** Editor-only validation is useful
|
||||
feedback and is not a safety boundary. Recipe: AG-07.
|
||||
|
||||
---
|
||||
|
||||
## 8. Temporal dependencies
|
||||
|
||||
Draw initialization prerequisites as a directed graph — each feature state and
|
||||
what it requires. Reject: cycles; a prerequisite produced only by an optional or
|
||||
disabled path; an anonymous next-tick delay; a generic "extension added" event
|
||||
used where semantic readiness is meant; a blocked transition with no diagnostic.
|
||||
|
||||
### Permutation test
|
||||
|
||||
Vary activation before and after actor spawn, archetype data before and after
|
||||
possession, replication arrival order, UI point before and after extension, and
|
||||
listener before and after event. **Correct systems converge to the same state
|
||||
from every order.** Systems that pass by luck have one order that works.
|
||||
|
||||
---
|
||||
|
||||
## 9. Observability is a requirement, not tooling polish
|
||||
|
||||
Expose read-only dumps for: the selected mode and its load state; pending assets,
|
||||
plugins and actions **with reasons**; active feature plugins and who requires
|
||||
them; per-feature actor readiness and blockers; active input contexts and their
|
||||
semantic joins; granted capabilities by owner receipt; UI layers, points,
|
||||
extensions and owners; message channels with listener counts and owners; the
|
||||
authoritative team identifier and its mirrors; requested versus effective
|
||||
settings with clamps.
|
||||
|
||||
Logs carry correlation fields: world, local player, actor, experience, plugin,
|
||||
action, channel, owner.
|
||||
|
||||
**The acceptance test:** an engineer who did not build the feature diagnoses an
|
||||
injected failure from logs and dumps alone, without opening assets, within a
|
||||
bounded time. If that is impossible, the indirection is operationally incomplete
|
||||
regardless of how clean the code is. Recipe: AG-02.
|
||||
|
||||
One thing worth stealing outright: the audited reference ships console variables
|
||||
that inject artificial delay into mode loading. A reproducible way to exercise
|
||||
slow-load, timeout and cancellation paths is worth more than it costs, and almost
|
||||
nobody builds one.
|
||||
|
||||
---
|
||||
|
||||
## 10. The false-genericity sweep
|
||||
|
||||
For every public field or hook implying behaviour — an ordering value, a viewer
|
||||
identity, an automatic-benchmark decision, a removal function, a profile suffix,
|
||||
a policy flag:
|
||||
|
||||
1. find its consumer;
|
||||
2. mutate the value in a fixture;
|
||||
3. assert an observable delta;
|
||||
4. test both native and visual-script surfaces where exposed.
|
||||
|
||||
No consumer or no delta means: implement it, remove it, or mark it explicitly
|
||||
unsupported. Never let a configurable no-op survive as stable API.
|
||||
|
||||
This sweep is worth running as a scheduled activity rather than a review step,
|
||||
because the shape recurs independently in every subsystem. Recipe: AG-08.
|
||||
|
||||
---
|
||||
|
||||
## 11. Event, state and command
|
||||
|
||||
Classify every interaction:
|
||||
|
||||
- **state** — a queryable authoritative model;
|
||||
- **command** — one accountable handler, with a result and a failure;
|
||||
- **event** — a past-tense fact with zero or many optional observers.
|
||||
|
||||
Red flags: UI reconstructing state from transient messages; a message asking an
|
||||
unknown listener to perform required work; a local bus assumed to replicate;
|
||||
presentation treated as authority.
|
||||
|
||||
The test that settles it: **if a listener subscribes one second late, must it know
|
||||
the current value?** Yes means state. No means an event is acceptable.
|
||||
|
||||
---
|
||||
|
||||
## 12. Version and adoption
|
||||
|
||||
Before copying from a reference, classify each element: architectural invariant
|
||||
(adopt), sample scaffolding (replace), project-specific behaviour (evaluate),
|
||||
unfinished path (do not claim supported), confirmed defect (fix with a regression
|
||||
test), prototype content (exclude from production), version-sensitive API (verify
|
||||
live).
|
||||
|
||||
Pin the engine version and changelist, the sample version, plugin versions, and
|
||||
any serialized schema version. Re-run probes after every upgrade.
|
||||
|
||||
**A current documentation page is not evidence about the code in your tree.**
|
||||
Recipe: AG-10. The full classification is `ue-reference-project-adoption`.
|
||||
|
||||
---
|
||||
|
||||
## 13. The integration matrix
|
||||
|
||||
Do not attempt the full product of configurations. Require this spine:
|
||||
|
||||
1. standalone baseline;
|
||||
2. dedicated server plus remote client;
|
||||
3. listen server;
|
||||
4. two local players;
|
||||
5. feature activates **before** receiver readiness;
|
||||
6. feature activates **after** receiver readiness;
|
||||
7. deactivate and reactivate;
|
||||
8. actor destroyed before feature teardown;
|
||||
9. late join;
|
||||
10. map travel or mode change;
|
||||
11. packaged clean client;
|
||||
12. engine or plugin upgrade fixture.
|
||||
|
||||
Every critical invariant needs at least one test crossing **two** subsystem
|
||||
boundaries. Unit tests inside one manager cannot express these failures.
|
||||
Recipe: AG-09.
|
||||
|
||||
---
|
||||
|
||||
## 14. Triage
|
||||
|
||||
**Critical — stop feature work.** Resource ownership or rollback unknown;
|
||||
authority duplicated; world or player context leaking; production data graph
|
||||
unvalidated; persistent state reconstructed from messages; client presentation
|
||||
affecting authority.
|
||||
|
||||
**High — fix before runtime switching or shipping.** Initialization cycles or
|
||||
timing dependencies; missing packaged-asset fallback; native and visual-script
|
||||
semantics diverging; undocumented ordering dependencies; insufficient
|
||||
observability; version mismatch.
|
||||
|
||||
**Medium — track with a bound.** Abstraction tax; linear scans acceptable at
|
||||
current data size; editor-only polish gaps; duplicated sample scaffolding.
|
||||
|
||||
### The no-go rule
|
||||
|
||||
Three unchecked Critical gates mean **stop adding abstraction and features**.
|
||||
Repair ownership, context and validation first. More indirection on an
|
||||
unvalidated base multiplies the failure surface rather than organising it.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The invariants and gates here are generalized from a full-project audit of Epic's
|
||||
Lyra Starter Game on Unreal Engine 5.6 — its data-driven layer, feature plugins,
|
||||
ability system, input, UI, messaging, cosmetics and teams, settings, loading and
|
||||
streaming — read as source and configuration rather than run.
|
||||
|
||||
Every cross-system failure mode in [failure modes](failure-modes.md) was observed
|
||||
in that project, in a subsystem whose local implementation was correct. That is
|
||||
the reason this skill exists as a separate document rather than as a section in
|
||||
each of the others.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill. Entry
|
||||
identifiers resolve back to the audited locations, so any specific claim can be
|
||||
produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The invariants are transferable;
|
||||
the specific instances are evidence, not guarantees about other versions. Re-run
|
||||
the detection recipes against your own tree before acting on any specific claim.
|
||||
@@ -0,0 +1,298 @@
|
||||
---
|
||||
name: ue-asset-loading-and-memory
|
||||
description: >-
|
||||
Design or audit asset loading and CPU memory residency in Unreal Engine: the
|
||||
asset manager and primary asset identifiers, asset bundles and cook rules,
|
||||
hard versus soft versus weak reference semantics, streamable handle ownership,
|
||||
synchronous load hitches, startup job sequencing, garbage-collection
|
||||
reachability, retention leaks and unload paths. Use when reducing load times
|
||||
or memory, choosing reference types, adding content-pulling definitions,
|
||||
diagnosing hitches, or investigating objects that are never collected.
|
||||
---
|
||||
|
||||
# UE asset loading and CPU memory
|
||||
|
||||
The invariant:
|
||||
|
||||
> An asset stays in memory because something reachable **owns** it. Loading is a
|
||||
> policy decision; residency is an ownership fact. They are set by different code
|
||||
> and must be reviewed separately.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes: [failure modes](references/failure-modes.md).
|
||||
|
||||
Related skills: `ue-data-driven-architecture`, `ue-modular-gameplay`,
|
||||
`ue-streaming-and-platform-budgets`, `ue-runtime-allocation-and-caching`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Separate the five budgets before optimizing anything
|
||||
|
||||
| Budget | What it is | Measured by |
|
||||
|---|---|---|
|
||||
| A. Disk / cook | package bytes shipped | build size, chunk sizes |
|
||||
| B. CPU memory | object graph plus bulk data | memory report, object list |
|
||||
| C. GPU memory | live render resources | platform GPU profiler |
|
||||
| D. Streaming pool | a **limit**, not consumption | streaming stats |
|
||||
| E. Net objects | actors, channels, bandwidth | network profiler |
|
||||
|
||||
This skill owns A and B. C and D belong to
|
||||
`ue-streaming-and-platform-budgets`; E to `ue-runtime-allocation-and-caching`.
|
||||
|
||||
**The most common analytical error is treating D as B.** A pool setting grants
|
||||
permission to use memory. It never reports use.
|
||||
|
||||
---
|
||||
|
||||
## 2. Reference semantics decide load cost
|
||||
|
||||
Three mechanisms, routinely conflated:
|
||||
|
||||
| Mechanism | Effect |
|
||||
|---|---|
|
||||
| reflected property, manual reference collection, a retained streamable handle | **owns** — keeps resident |
|
||||
| weak pointer, object key | **observes** — keeps nothing |
|
||||
| soft object or class pointer | **a path, not a reference** |
|
||||
|
||||
The third line is where most reasoning fails. **A soft pointer saves nothing by
|
||||
itself.** Residency is decided by where the loaded result is stored — a soft
|
||||
reference resolved into a reflected property is a hard reference with extra
|
||||
steps.
|
||||
|
||||
### Measure the transitive closure, not the direct count
|
||||
|
||||
Direct reference counts are misleading. What matters is how many packages load
|
||||
when this one loads. In the audited reference the same architectural layer
|
||||
spanned three orders of magnitude depending on reference type alone: an archetype
|
||||
holding hard references pulled over a thousand packages, while a catalog entry
|
||||
addressing by identifier pulled two.
|
||||
|
||||
That single contrast is the argument for treating reference type as
|
||||
architecture rather than style. Recipe: AL-01.
|
||||
|
||||
### Choosing
|
||||
|
||||
| Need | Reference |
|
||||
|---|---|
|
||||
| address before load; cook and bundle identity | primary asset identifier |
|
||||
| optional, feature-scoped or async content | soft object or class pointer |
|
||||
| required whenever the owner is loaded, and small | hard pointer |
|
||||
| data selects an implementation class | class reference |
|
||||
| owner-exclusive polymorphic fragment | instanced object |
|
||||
|
||||
Rule of thumb: a definition that is a **catalog entry** references by identifier
|
||||
or softly; a definition that is an **archetype** may reference hard — but its
|
||||
closure must be measured and budgeted, not assumed.
|
||||
|
||||
---
|
||||
|
||||
## 3. Bundles express role, not grouping
|
||||
|
||||
Declare the load role beside the reference, then request bundles per runtime role
|
||||
at load time.
|
||||
|
||||
Gates:
|
||||
|
||||
- every soft reference carries an intentional bundle tag;
|
||||
- server-irrelevant presentation is client-only;
|
||||
- prediction-relevant gameplay classes are client and server;
|
||||
- **a missing bundle tag means "always loaded with the owner"** — decide it, do
|
||||
not default into it;
|
||||
- cook rules control budget **A**, never budget **B**.
|
||||
|
||||
A label asset with thousands of soft references and no hard references determines
|
||||
what ships and retains nothing at runtime. Do not read cook rules as residency.
|
||||
|
||||
Watch for a declared bundle name that no property ever uses: it is requested at
|
||||
load time, matches nothing, and looks like a working tier. Recipe: AL-02.
|
||||
|
||||
---
|
||||
|
||||
## 4. Startup discipline
|
||||
|
||||
### Never synchronously load a graph you have not measured
|
||||
|
||||
The dominant boot hitch in a data-driven project is one synchronous load whose
|
||||
hard closure is large. An elegant async bundle pipeline that runs *after* it is
|
||||
decoration.
|
||||
|
||||
- [ ] every synchronous load on the boot path is justified;
|
||||
- [ ] its transitive closure is measured, not assumed;
|
||||
- [ ] an async alternative was considered, with an explicit wait state;
|
||||
- [ ] the loading screen carries a reason string;
|
||||
- [ ] cancellation is distinguishable from completion.
|
||||
|
||||
### Startup jobs must own handles and report truthfully
|
||||
|
||||
A progress bar that never moves is a bug, not a cosmetic issue — it hides which
|
||||
job is slow. Three independent layers can each break it: a macro that never
|
||||
fills the handle, a throttle comparison with swapped operands, and an empty
|
||||
progress receiver. Any one of them is enough. Recipe: AL-03.
|
||||
|
||||
### Global data assets need bundles too
|
||||
|
||||
Soft class pointers on a globally-loaded data asset with **no** bundle tags load
|
||||
synchronously at first use — typically in the frame of first damage. Tag them and
|
||||
preload them with the mode. Recipe: AL-04.
|
||||
|
||||
---
|
||||
|
||||
## 5. Retention: what actually keeps objects alive
|
||||
|
||||
Audit six owners explicitly: reflected containers on long-lived subsystems;
|
||||
manual reference collection in native code; retained streamable handles; delegate
|
||||
bindings capturing strong references; component and actor ownership chains; class
|
||||
default objects reached through class references.
|
||||
|
||||
### A convenience loader must not default to permanent
|
||||
|
||||
The worst shape in this area is a helper that loads an asset and adds it to a
|
||||
reflected set, with the keep-in-memory parameter **defaulting to true** and no
|
||||
release API anywhere. Every casual call then synchronously loads *and permanently
|
||||
roots* the asset. The soft pointer was chosen for deferred loading; the API turns
|
||||
it into a process-lifetime hard reference.
|
||||
|
||||
Required design: the flag is explicit at every call site or defaults to false; a
|
||||
paired release exists and is used; the retained set is inspectable from a console
|
||||
command; retention has an owner and a scope. Recipe: AL-05.
|
||||
|
||||
### Strong keys pin object graphs — and raw pointers lose them
|
||||
|
||||
```text
|
||||
map<strong actor pointer, strong value> one missed unregister pins the actor,
|
||||
its meshes, materials, sounds, effects
|
||||
|
||||
map<raw actor pointer, non-reflected struct of object pointers>
|
||||
garbage collection cannot see the
|
||||
values; the key can dangle
|
||||
```
|
||||
|
||||
Both errors can coexist in one codebase. **Audit direction as well as presence.**
|
||||
Prefer object keys with explicitly-owned values plus a periodic prune. Recipe:
|
||||
AL-06.
|
||||
|
||||
---
|
||||
|
||||
## 6. Unloading must exist before you claim modularity
|
||||
|
||||
If a project cannot answer "what is released when a mode ends", its memory
|
||||
profile only grows. Either implement release, or state plainly that mode changes
|
||||
require map travel — that is a legitimate scope decision, and leaving it
|
||||
undefined is not.
|
||||
|
||||
The measurement is blunt and takes one command: count unload calls, bundle
|
||||
releases and explicit collection requests across the whole project. In the
|
||||
audited reference every one of those counts was **zero**. Recipe: AL-07.
|
||||
|
||||
### Avoid collection ping-pong
|
||||
|
||||
A forced collection at a transition boundary must not collect what the next
|
||||
transition immediately needs. Options: keep transition-critical classes rooted,
|
||||
defer collection past the next load, or do not force it and let the budget drive
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
## 7. Failure handling changes memory behaviour
|
||||
|
||||
- a cancelled load must not run the success path;
|
||||
- asynchronous results must be read, not discarded;
|
||||
- partially loaded state must be releasable;
|
||||
- waiting without a timeout converts a stall into a hang;
|
||||
- a permanent loading screen with no reason string is not error handling.
|
||||
|
||||
The sharpest instance: binding the cancel delegate to the **same callback** as
|
||||
completion. A cancelled load then reports success, downstream code dereferences
|
||||
unloaded assets, and each of those dereferences is individually silent. Recipes:
|
||||
AL-08, AL-09.
|
||||
|
||||
---
|
||||
|
||||
## 8. Review checklist
|
||||
|
||||
### References
|
||||
|
||||
- [ ] reference type chosen per lifecycle and documented in metadata;
|
||||
- [ ] transitive hard closure measured for every archetype-level definition;
|
||||
- [ ] no soft reference resolved into an unbounded reflected cache;
|
||||
- [ ] bundle tags on all soft references, and every requested bundle is used;
|
||||
- [ ] cross-plugin edges intentional and acyclic.
|
||||
|
||||
### Loading
|
||||
|
||||
- [ ] no unmeasured synchronous load on boot or in a hot path;
|
||||
- [ ] async loads own their handles and can be cancelled;
|
||||
- [ ] progress reporting is verified to move, not merely implemented;
|
||||
- [ ] cancellation and failure paths are distinct from success;
|
||||
- [ ] global data assets are bundled, not lazily loaded at first use.
|
||||
|
||||
### Retention
|
||||
|
||||
- [ ] every long-lived container has an owner, a bound and a prune policy;
|
||||
- [ ] strong keys justified; object keys used where identity suffices;
|
||||
- [ ] non-reflected object pointers eliminated or documented as weak by design;
|
||||
- [ ] register and unregister pairs verified across teardown paths;
|
||||
- [ ] a release API exists for anything deliberately kept resident.
|
||||
|
||||
---
|
||||
|
||||
## 9. Measurement plan
|
||||
|
||||
Static first — it is cheap and finds structural problems:
|
||||
|
||||
1. build the dependency graph from the asset registry: direct hard and soft
|
||||
counts, transitive hard closure per definition, cross-plugin edges;
|
||||
2. rank definitions by closure and compare against their architectural layer;
|
||||
3. search for synchronous loads on boot and hot paths;
|
||||
4. inventory retention owners by type.
|
||||
|
||||
Runtime, once a packaged build exists: memory reports before and after a mode
|
||||
cycle; object listings on suspected roots — root-graph evidence, not process
|
||||
memory; fifty to a hundred spawn and travel cycles looking for monotonic growth;
|
||||
an allocation profiler for attribution.
|
||||
|
||||
### Do not measure residency in the editor
|
||||
|
||||
Editor builds commonly load **both** client and server bundles, so an in-editor
|
||||
memory profile represents neither shipping role. Residency numbers require a
|
||||
packaged build; use the editor for relative structural comparison only. Recipe:
|
||||
AL-10.
|
||||
|
||||
---
|
||||
|
||||
## 10. When to stop
|
||||
|
||||
Optimize loading and residency when boot or transition time is a product problem,
|
||||
a platform ceiling is real and near, growth is monotonic across a session, or a
|
||||
specific closure is measurably oversized.
|
||||
|
||||
Do not restructure references on aesthetics. A fourteen-package closure is not a
|
||||
problem; a thousand-package synchronous closure on the boot path is. Measure
|
||||
first, and **keep the measurement** so the next change can be compared against
|
||||
it.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings come from a source and dependency-graph audit of Epic's Lyra Starter
|
||||
Game on Unreal Engine 5.6. The engine's own asset manager and streamable manager
|
||||
are not part of that project, so every claim about their internals is inferred
|
||||
from call sites and marked as such rather than presented as measured.
|
||||
|
||||
No profiler was run and no timing was captured: this material is structural by
|
||||
construction, and it deliberately contains no performance numbers. What it
|
||||
contains is counts — of call sites, of unload paths, of packages in a closure —
|
||||
which is a different kind of evidence and a more portable one.
|
||||
|
||||
Source addresses stay in the research archive. Each entry carries a stable
|
||||
identifier (`AL-01` and up) resolving back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version. Binary assets were not read, so the contents of
|
||||
any specific closure are structural rather than byte-measured. Re-run the recipes
|
||||
against your own tree; the counts quoted show the shape of a contrast, not a
|
||||
target.
|
||||
+389
@@ -0,0 +1,389 @@
|
||||
# Failure modes: asset loading and CPU memory
|
||||
|
||||
Ten ways a project loads more than it meant to, keeps it forever, and reports
|
||||
nothing.
|
||||
|
||||
The shared property here is different from the other skills in this bundle, and
|
||||
it is worth stating because it changes what the recipes look like: **memory
|
||||
defects are not events.** There is no moment at which something goes wrong. An
|
||||
asset is loaded because a reference exists; it stays because a reference still
|
||||
exists. Both are correct behaviour at every instant, and the failure is a
|
||||
property of an aggregate that no single frame contains.
|
||||
|
||||
That has two consequences for detection. First, **counting is the primary
|
||||
instrument** — call sites, unload paths, packages in a closure. Second, a count of
|
||||
**zero** is often the finding: no release API, no unload call, no bundle tag.
|
||||
Several recipes below are therefore searches you expect to come back empty, and
|
||||
the empty result is the evidence rather than the absence of it.
|
||||
|
||||
Recipes use `rg` from a project source root, and were executed against the
|
||||
audited project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
### AL-01 - A reference type chosen by convenience, costing three orders of magnitude
|
||||
|
||||
**Mechanism.** A definition holds hard references to the things it selects — a
|
||||
class, a data asset, a package of capabilities — because hard references are
|
||||
simpler and always resolve. Loading the definition therefore loads its entire
|
||||
transitive closure.
|
||||
|
||||
**Why it is silent.** Everything works, immediately and correctly. Hard
|
||||
references never fail, never arrive late and never need a fallback. The cost is
|
||||
paid once, at load, in a place nobody attributes to this asset.
|
||||
|
||||
**Why the obvious check misses it.** Reviewing the definition shows a handful of
|
||||
properties — five or six references, all necessary. The cost is not the direct
|
||||
count but the closure, and the closure lives in assets this file never mentions.
|
||||
Nothing at the declaration hints at its size.
|
||||
|
||||
**Symptom.** A boot or transition hitch attributed to "the level" or "shaders".
|
||||
In the audited reference the same architectural family ranged from **two**
|
||||
packages when addressing by identifier to **over a thousand** when holding a hard
|
||||
chain — a difference produced entirely by reference type.
|
||||
|
||||
**Detect.** Direct counts are not the measurement; build the closure. From
|
||||
source you can only find the candidates:
|
||||
|
||||
```bash
|
||||
rg -n "TObjectPtr<|TSubclassOf<" --glob "*.h" . | rg -i "definition|data|archetype"
|
||||
rg -n "TSoftObjectPtr<|TSoftClassPtr<" --glob "*.h" . | rg -i "definition|data"
|
||||
```
|
||||
|
||||
Then, from the editor or a registry dump, compute the transitive hard closure per
|
||||
definition and rank by size. **Compare each against its architectural layer**: a
|
||||
catalog entry with a large closure is the finding; an archetype with one may be
|
||||
correct and budgeted.
|
||||
|
||||
**Guardrail.** Write the intended reference type into the property metadata and
|
||||
validate it. Re-measure the closure of every archetype-level definition when it
|
||||
changes; a closure without a recorded baseline cannot be reviewed.
|
||||
|
||||
---
|
||||
|
||||
### AL-02 - A bundle name that no property uses
|
||||
|
||||
**Mechanism.** A bundle tier is declared as a constant and requested at load time
|
||||
alongside the role bundles. No property in the codebase is annotated with it.
|
||||
|
||||
**Why it is silent.** Requesting a bundle that matches nothing is not an error.
|
||||
The load succeeds, the other bundles arrive, and the tier appears to be part of
|
||||
the loading strategy.
|
||||
|
||||
**Why the obvious check misses it.** The constant is declared, referenced at the
|
||||
request site, and named meaningfully. "Is this bundle used?" answers yes at two
|
||||
places. The absent half is the **annotation**, and nothing connects a request to
|
||||
the properties it is meant to gather.
|
||||
|
||||
**Symptom.** No runtime symptom at all. The cost is a false model: the team
|
||||
believes there is an equipped-content tier, plans around it, and the loading
|
||||
strategy has one fewer dimension than it appears to.
|
||||
|
||||
**Detect.** Compare declared bundle names against annotated ones:
|
||||
|
||||
```bash
|
||||
rg -n "AssetBundles=\"[^\"]*\"" --glob "*.h" . -o | sed 's/.*AssetBundles=//' | sort -u
|
||||
rg -n "BundlesToLoad.Add|LoadStateClient|LoadStateServer|FName\(TEXT\(\"" --glob "*.cpp" . \
|
||||
| rg -i "bundle"
|
||||
```
|
||||
|
||||
Any name in the second list and not the first is requested and unpopulated. In
|
||||
the audited project one such tier is requested on every load regardless of net
|
||||
mode, and appears in no annotation in the entire source tree.
|
||||
|
||||
**Guardrail.** Assert at startup that every requested bundle name matches at
|
||||
least one annotated property, or remove the tier. A bundle nobody fills is a
|
||||
plan, not a mechanism.
|
||||
|
||||
---
|
||||
|
||||
## Startup
|
||||
|
||||
### AL-03 - Progress scaffolding where every layer is independently broken
|
||||
|
||||
**Mechanism.** A startup framework reports per-job progress. Three things must
|
||||
work: the job must hand back a load handle, a throttle must permit an update, and
|
||||
a receiver must render it.
|
||||
|
||||
**Why it is silent.** A progress bar that does not move looks like a fast load or
|
||||
a coarse-grained one. There is no error state for "progress was never reported",
|
||||
and the framework's logs still print per-job timings, so it looks instrumented.
|
||||
|
||||
**Why the obvious check misses it.** Each layer is individually plausible. The
|
||||
macro compiles and runs the job. The throttle has a sensible-looking comparison.
|
||||
The receiver is a real function with a real name. Only reading all three together
|
||||
shows that the handle is never assigned, the comparison can never be true, and
|
||||
the receiver's body is a comment.
|
||||
|
||||
**Symptom.** No visible progress during startup, so the slowest job is unknown.
|
||||
The team optimizes what it guesses rather than what it measured.
|
||||
|
||||
**Detect.** Check the three layers separately, and expect each to look fine:
|
||||
|
||||
```bash
|
||||
# 1. does anything assign the out-parameter handle?
|
||||
rg -n -B4 -A8 "TSharedPtr<FStreamableHandle>& \w+" --glob "*.h" --glob "*.cpp" .
|
||||
# 2. is the throttle comparison the right way round?
|
||||
rg -n -B2 -A6 "LastUpdate|LastReport" --glob "*.h" --glob "*.cpp" .
|
||||
# 3. does the receiver have a body?
|
||||
rg -n -A4 "UpdateInitial\w*Percent|::UpdateProgress" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
In the audited project the throttle reads `LastUpdate - Now > interval`, where
|
||||
the elapsed value is monotonically increasing and the stored value starts at
|
||||
zero — so the difference is never positive and the branch is unreachable.
|
||||
|
||||
**Guardrail.** Test that progress fires at all before trusting it. A callback
|
||||
that has never been observed to run is indistinguishable from one that cannot.
|
||||
|
||||
---
|
||||
|
||||
### AL-04 - Global data with soft references and no bundle tags
|
||||
|
||||
**Mechanism.** A globally-loaded data asset holds soft class pointers to effects
|
||||
or classes used across the game. The asset is loaded at startup; its soft targets
|
||||
are not, because nothing tags them into a bundle.
|
||||
|
||||
**Why it is silent.** The pointers resolve on demand and everything works. The
|
||||
first resolution is a synchronous load, and it lands in whatever frame first
|
||||
needs it — commonly the first damage event, which is already a busy frame.
|
||||
|
||||
**Why the obvious check misses it.** The asset is loaded at startup, which
|
||||
answers "is the global data preloaded?" with yes. Whether its *contents* are
|
||||
preloaded is a different question, decided by annotations that are absent — and
|
||||
absent annotations look exactly like properties that do not need them.
|
||||
|
||||
**Symptom.** A hitch at first damage, first dynamic tag, or first use of any
|
||||
global class. Reproduces once per session, which makes it easy to dismiss.
|
||||
|
||||
**Detect.** Find globally-loaded data assets and check their soft properties for
|
||||
tags:
|
||||
|
||||
```bash
|
||||
rg -n -A20 "class \w*GameData" --glob "*.h" . | rg "TSoftClassPtr|TSoftObjectPtr|AssetBundles"
|
||||
```
|
||||
|
||||
Soft properties with no bundle annotation on a startup-loaded asset is the
|
||||
finding. Then find where they first resolve:
|
||||
|
||||
```bash
|
||||
rg -n "GetSubclass\(|GetAsset\(" --glob "*.cpp" . | rg -i "gamedata"
|
||||
```
|
||||
|
||||
**Guardrail.** Tag them and load them with the mode. A soft reference on a global
|
||||
asset without a bundle is a deferred synchronous load with an unpredictable
|
||||
trigger.
|
||||
|
||||
---
|
||||
|
||||
## Retention
|
||||
|
||||
### AL-05 - A keep-in-memory pool with no way out
|
||||
|
||||
**Mechanism.** A convenience loader resolves a soft pointer, and adds the result
|
||||
to a reflected set so it stays alive. The keep-in-memory parameter **defaults to
|
||||
true**. No removal API exists.
|
||||
|
||||
**Why it is silent.** It is a cache doing its job. Every call is faster than the
|
||||
last, nothing is ever missing, and the set has no size at which it complains.
|
||||
|
||||
**Why the obvious check misses it.** The parameter is visible in the signature
|
||||
and reads as a considered choice — which it is, at the declaration. At every call
|
||||
site it is invisible, because it is defaulted. Reviewing a call shows a load; the
|
||||
retention is in the default argument of a function in another file.
|
||||
|
||||
**Symptom.** Residency that only grows across a session, unaffected by mode
|
||||
changes, map travel or anything else. Because the growth is monotonic and
|
||||
attributable to nothing, it is usually first noticed as "we are over budget"
|
||||
rather than as a leak.
|
||||
|
||||
**Detect.** Count additions against removals, and read the default:
|
||||
|
||||
```bash
|
||||
rg -n "LoadedAssets|KeptAssets|RetainedAssets" --glob "*.h" --glob "*.cpp" .
|
||||
rg -n "bKeepInMemory" --glob "*.h" . | head
|
||||
```
|
||||
|
||||
In the audited project the set is added to at one site, read at two — both inside
|
||||
a dump command — and **never removed from anywhere**, with the flag defaulting to
|
||||
true and no call site in the project passing false.
|
||||
|
||||
**Guardrail.** Default to false, or make the flag explicit at every call site.
|
||||
Ship a paired release API and a console command that prints the set's size. A
|
||||
cache without an eviction policy is a leak with a nicer name.
|
||||
|
||||
---
|
||||
|
||||
### AL-06 - Strong keys pin graphs; raw keys lose them
|
||||
|
||||
**Mechanism.** Two opposite errors in the same area. A map keyed by a **strong
|
||||
object pointer** keeps its key alive, so one missed unregister pins an actor and
|
||||
everything it owns. A map holding object pointers inside a **non-reflected**
|
||||
struct is invisible to garbage collection, so the values can be collected while
|
||||
still referenced.
|
||||
|
||||
**Why it is silent.** The first has no symptom until memory is measured — the
|
||||
actor is dead in gameplay terms and alive in memory. The second has no symptom
|
||||
until collection happens to run at the wrong moment, which is rare and
|
||||
non-deterministic.
|
||||
|
||||
**Why the obvious check misses it.** Both look like ordinary containers. The
|
||||
difference between a reflected and a non-reflected struct is one macro, and the
|
||||
difference between a strong and a weak key is one wrapper type. Neither reads as
|
||||
a memory decision at the declaration site.
|
||||
|
||||
**Symptom.** Actors that never disappear from a memory report, or — from the
|
||||
other error — a crash on a pointer that was valid a moment ago. Both are usually
|
||||
investigated far from the container.
|
||||
|
||||
**Detect.** Audit direction as well as presence:
|
||||
|
||||
```bash
|
||||
# strong keys
|
||||
rg -n "TMap<TObjectPtr<\w+>|TMap<A\w+\*" --glob "*.h" .
|
||||
# object pointers outside the reflection system
|
||||
rg -n -B6 "TArray<U\w+\*>|U\w+\* \w+;" --glob "*.h" . | rg -v "UPROPERTY"
|
||||
```
|
||||
|
||||
Both patterns appearing in one codebase is common and is worth stating plainly in
|
||||
a review: they are not variants of one mistake, they are opposite mistakes, and a
|
||||
fix for one does not address the other.
|
||||
|
||||
**Guardrail.** Prefer object keys with explicitly-owned values plus a periodic
|
||||
prune. Every object pointer that must survive collection is reflected; every one
|
||||
that must not is weak. There is no third category.
|
||||
|
||||
---
|
||||
|
||||
### AL-07 - No unload path anywhere
|
||||
|
||||
**Mechanism.** Content is loaded per mode through bundles. Nothing releases it —
|
||||
no unload call, no bundle removal, no explicit collection request.
|
||||
|
||||
**Why it is silent.** For a session-based game with a process restart between
|
||||
matches, it is correct and cheap. The absence only becomes a defect when modes
|
||||
change within one process, which is a product decision made later.
|
||||
|
||||
**Why the obvious check misses it.** There is nothing to see. Reviewing the
|
||||
loading code finds a complete, well-built loading path; the absence of a
|
||||
symmetric release path is not a line anyone reads. The authors of the audited
|
||||
project recorded it themselves, in comments, and shipped it.
|
||||
|
||||
**Symptom.** Memory that never returns to baseline across mode changes. The first
|
||||
mode's content is resident for the whole session alongside the second's.
|
||||
|
||||
**Detect.** One command, and the expected answer is a row of zeros:
|
||||
|
||||
```bash
|
||||
for p in UnloadPrimaryAsset bRemoveAllBundles CollectGarbage FlushAsyncLoading TrimMemory; do
|
||||
printf "%-24s %s\n" "$p" "$(rg -c "$p" --glob '*.cpp' --glob '*.h' . | wc -l)"
|
||||
done
|
||||
```
|
||||
|
||||
In the audited project every one of those is **zero**. That table is the finding,
|
||||
and it takes ten seconds to produce in any project.
|
||||
|
||||
**Guardrail.** Either implement release, or write down that mode changes require
|
||||
map travel. An undefined middle is what grows.
|
||||
|
||||
---
|
||||
|
||||
## Failure handling
|
||||
|
||||
### AL-08 - Cancellation bound to the completion callback
|
||||
|
||||
**Mechanism.** An asynchronous load binds a completion delegate, and binds the
|
||||
**same** delegate to cancellation.
|
||||
|
||||
**Why it is silent.** Cancellation is rare — it needs a mode change or a world
|
||||
teardown mid-load. When it does happen, the system reports success and proceeds,
|
||||
and every downstream consumer of a missing asset fails silently in its own way.
|
||||
|
||||
**Why the obvious check misses it.** Both delegates are bound, which is more
|
||||
than most code does. The lambda is short and reads as deliberate. Nothing
|
||||
distinguishes "handle cancellation" from "handle cancellation *correctly*" at the
|
||||
binding site.
|
||||
|
||||
**Symptom.** A mode reports loaded with partially loaded content. Downstream,
|
||||
unloaded classes resolve to null and are skipped without logging, so the result
|
||||
is a mode missing several features and no diagnostic anywhere.
|
||||
|
||||
**Detect.** Compare the two bindings:
|
||||
|
||||
```bash
|
||||
rg -n -B6 -A6 "BindCancelDelegate" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
If the cancel lambda executes the completion delegate, that is this defect. In
|
||||
the audited project it does, in three lines, immediately below the completion
|
||||
binding.
|
||||
|
||||
**Guardrail.** Cancellation enters a distinct state — failed, or cancelled — that
|
||||
the loading screen and the consumers can observe. Success is a claim; make it
|
||||
provable.
|
||||
|
||||
---
|
||||
|
||||
### AL-09 - An asynchronous result that is never read
|
||||
|
||||
**Mechanism.** A plugin or asset load completes with a result parameter carrying
|
||||
success or failure. The callback decrements a counter and ignores the parameter.
|
||||
|
||||
**Why it is silent.** The counter reaches zero either way, so the pipeline
|
||||
completes. A failed load is indistinguishable from a successful one at every
|
||||
point downstream.
|
||||
|
||||
**Why the obvious check misses it.** The callback has the right signature and is
|
||||
bound correctly. An unused parameter in a delegate signature produces no warning,
|
||||
because the signature is imposed by the API rather than chosen.
|
||||
|
||||
**Symptom.** A feature that silently does not activate, in a build where one of
|
||||
its assets failed to load. The failure is attributed to the feature.
|
||||
|
||||
**Detect.** Find completion callbacks whose result parameter is unused:
|
||||
|
||||
```bash
|
||||
rg -n -A8 "::On\w*LoadComplete\(const \w+::\w+& Result\)" --glob "*.cpp" . \
|
||||
| rg -c "Result"
|
||||
```
|
||||
|
||||
A count of one — the signature only — is the finding.
|
||||
|
||||
**Guardrail.** Read the result, log the failure with the identifier that failed,
|
||||
and enter a failed state. An ignored result is a decision to be surprised later.
|
||||
|
||||
---
|
||||
|
||||
## Measurement
|
||||
|
||||
### AL-10 - Measuring residency in the editor
|
||||
|
||||
**Mechanism.** The editor loads both client and server bundles, because it must
|
||||
be able to run either role.
|
||||
|
||||
**Why it is silent.** The measurement completes and produces a number. The number
|
||||
is real; it just describes a configuration that ships to nobody.
|
||||
|
||||
**Why the obvious check misses it.** It is intentional and correct — nothing to
|
||||
find in the memory system. The inflating condition is a boolean expression in the
|
||||
bundle-selection code, and nobody profiling memory reads the loading policy.
|
||||
|
||||
**Symptom.** Memory budgets built on editor numbers, wrong in a direction that
|
||||
feels safe until a platform limit is real.
|
||||
|
||||
**Detect.** Read the bundle-selection condition before trusting any in-editor
|
||||
number:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A6 "bLoadClient|bLoadServer|LoadStateClient" --glob "*.cpp" . | rg "GIsEditor"
|
||||
```
|
||||
|
||||
An editor branch that unions both sets means editor residency is directional
|
||||
only. The audited project does this in two places, and one of them carries a
|
||||
comment from its authors noting the resulting hitching.
|
||||
|
||||
**Guardrail.** Residency numbers require a packaged build of the target role. Use
|
||||
the editor for relative structural comparison, and say so whenever an in-editor
|
||||
number is quoted.
|
||||
@@ -0,0 +1,360 @@
|
||||
# Patterns: asset loading and residency in a measured reference
|
||||
|
||||
A worked reading of one shipped project's loading path, from asset-manager
|
||||
startup to feature-plugin activation, audited as source.
|
||||
|
||||
Individual defects are in [failure modes](failure-modes.md) with detection
|
||||
recipes. This file is about the shape: which decisions determine load cost, which
|
||||
determine residency, and why those are two different reviews.
|
||||
|
||||
Markers: **[measured]** — read in source or config; **[derived]** — conclusion
|
||||
from measured facts; **[open]** — not answerable without a packaged build or
|
||||
engine source, and left open.
|
||||
|
||||
---
|
||||
|
||||
## 1. The five budgets, and the one that is not a budget
|
||||
|
||||
| Budget | What it is | Measured by |
|
||||
|---|---|---|
|
||||
| A. Disk / cook | package bytes shipped | build and chunk sizes |
|
||||
| B. CPU RAM | object graph plus bulk data | memory report, object list |
|
||||
| C. GPU VRAM | live render resources | platform GPU profiler |
|
||||
| D. Streaming pool | **a limit, not consumption** | streaming stats |
|
||||
| E. Net objects | actors, channels, bandwidth | network profiler |
|
||||
|
||||
This skill owns A and B.
|
||||
|
||||
**The most common analytical error in this area is reading D as B.** A pool
|
||||
setting grants permission to use memory. It never reports use, it never causes
|
||||
use, and lowering it does not free anything that is already resident.
|
||||
|
||||
---
|
||||
|
||||
## 2. Reference type is the largest single lever
|
||||
|
||||
Three mechanisms, routinely conflated:
|
||||
|
||||
| Mechanism | Effect |
|
||||
|---|---|
|
||||
| a reflected property, an explicit GC reference, a retained load handle | **owns** — keeps resident |
|
||||
| a weak pointer, an object key | **observes** — keeps nothing |
|
||||
| a soft pointer | **a path, not a reference** |
|
||||
|
||||
The third line is where reasoning usually fails. A soft pointer saves nothing by
|
||||
itself. Residency is decided by **where the loaded result is stored** — and a
|
||||
soft reference resolved into a strong property is a hard reference with extra
|
||||
steps.
|
||||
|
||||
### The measured spread
|
||||
|
||||
From the dependency graph of the audited project — 3 874 packages, 16 639 hard
|
||||
edges, 9 333 soft edges **[measured]**:
|
||||
|
||||
| Asset | Packages pulled by loading it |
|
||||
|---|---:|
|
||||
| pawn archetype with a hard class-and-object chain | **1 015** |
|
||||
| weapon pickup definition | 800 |
|
||||
| ability set | 121 |
|
||||
| input config | 14 |
|
||||
| action set referencing by identifier and soft pointer | **2** |
|
||||
| playlist referencing by primary asset identifier | **2** |
|
||||
|
||||
**[derived]** Three orders of magnitude, from one decision made per property.
|
||||
That is the whole argument for treating reference type as architecture rather
|
||||
than as style — and the reason a "make it soft" refactor has to start from the
|
||||
closure measurement rather than from the property list.
|
||||
|
||||
### Choosing
|
||||
|
||||
| Need | Reference |
|
||||
|---|---|
|
||||
| address before load; cook and bundle identity | primary asset identifier |
|
||||
| optional, feature-scoped or async content | soft object or class pointer |
|
||||
| required whenever the owner is loaded, and small | hard pointer |
|
||||
| data selects an implementation class | class reference |
|
||||
| owner-exclusive polymorphic fragment | instanced object |
|
||||
|
||||
Rule of thumb: a definition that is a **catalog entry** references by identifier
|
||||
or soft pointer; a definition that is an **archetype** may reference hard, but its
|
||||
closure must be measured and budgeted.
|
||||
|
||||
---
|
||||
|
||||
## 3. Bundles declare role, and their absence declares nothing
|
||||
|
||||
Put the load role beside the reference:
|
||||
|
||||
```cpp
|
||||
UPROPERTY(EditAnywhere, meta=(AssetBundles="Client"))
|
||||
TSoftClassPtr<UUserWidget> WidgetClass;
|
||||
|
||||
UPROPERTY(EditAnywhere, meta=(AssetBundles="Client,Server"))
|
||||
TSoftClassPtr<UGameplayAbility> AbilityType;
|
||||
```
|
||||
|
||||
Then request bundles per runtime role at load time.
|
||||
|
||||
**[measured]** In the audited project the bundle vocabulary is small — client
|
||||
and client-plus-server annotations only — and one declared bundle name is
|
||||
requested on every load while appearing on **no** property in source at all. It
|
||||
is either satisfied entirely by content, or it is dead. **[derived]** A bundle
|
||||
name that no property claims is indistinguishable from a typo, and neither the
|
||||
compiler nor the cooker will say which it is.
|
||||
|
||||
Gates worth adopting:
|
||||
|
||||
- every soft reference has an intentional bundle annotation;
|
||||
- server-irrelevant presentation is client-only;
|
||||
- prediction-relevant gameplay classes are client-and-server;
|
||||
- a missing annotation means "loaded whenever the owner loads" — decide it, do
|
||||
not default into it;
|
||||
- cook rules control budget A, not budget B. A label asset with thousands of
|
||||
soft references and no hard ones determines what ships and retains nothing.
|
||||
|
||||
---
|
||||
|
||||
## 4. Startup: the shape that works, and the shape that only looks like it
|
||||
|
||||
The audited project has a well-formed startup-job framework **[measured]**: named
|
||||
jobs, weights, per-job timing logs, a progress delegate, and boot-timing scopes
|
||||
that integrate with the engine's trace.
|
||||
|
||||
**[measured]** It also has: a macro that never fills the load handle it is
|
||||
designed to return, a progress throttle whose comparison can never be true, and a
|
||||
progress receiver whose body is a comment.
|
||||
|
||||
**[derived]** Three independent layers, each individually harmless, combining
|
||||
into a startup with no progress reporting at all — inside a framework built to
|
||||
report it. This is the single most instructive thing in the audit, and the
|
||||
general form is worth memorizing: **scaffolding is not evidence of function.**
|
||||
Recipes: AL-02, AL-03.
|
||||
|
||||
What to take from it anyway:
|
||||
|
||||
- **name every startup job and log its duration.** Per-job timing is nearly free
|
||||
and it is the only thing that answers "what is slow at boot".
|
||||
- **wrap boot phases in the engine's timing scopes**, so the data lands in the
|
||||
same trace as everything else.
|
||||
- **fail fatally on genuinely required global data.** The audited project does
|
||||
this for its global data asset, with a comment explaining that a soft failure
|
||||
here would be harder to diagnose than a hard one **[measured]**. That is the
|
||||
correct call, and it is rare.
|
||||
|
||||
---
|
||||
|
||||
## 5. The synchronous load that dominates everything else
|
||||
|
||||
**[measured]** The root definition of a game mode is resolved with a synchronous
|
||||
load in the game thread, and its own authors marked it for async conversion.
|
||||
That definition holds a hard pointer to a pawn archetype, which holds hard
|
||||
references to the pawn class, its ability sets, its input config, its tag policy
|
||||
and its camera mode — all hard.
|
||||
|
||||
**[derived]** One line therefore pulls the closure measured at over a thousand
|
||||
packages, **before** the elegant asynchronous bundle pipeline underneath it runs
|
||||
at all. An async pipeline that executes after the bulk of the work is decoration.
|
||||
|
||||
The general lesson is a review question rather than a rule: *for every
|
||||
synchronous load on a boot or transition path, what is its transitive closure?*
|
||||
If nobody has measured it, the loading architecture is unverified regardless of
|
||||
how much of it is asynchronous.
|
||||
|
||||
**[measured]** Note the asymmetry that makes this easy to miss: on a client the
|
||||
same definition arrives by replication and is resolved by the net driver, so the
|
||||
synchronous cost exists on one side only, and profiling the wrong side finds
|
||||
nothing.
|
||||
|
||||
---
|
||||
|
||||
## 6. Handle ownership: the difference between loading and controlling
|
||||
|
||||
**[measured]** In the audited experience loader, every load handle is a local
|
||||
variable. The component stores none of them.
|
||||
|
||||
**[derived]** Three consequences, none of which produce an error:
|
||||
|
||||
1. **Cancellation is impossible.** If the mode changes or the world dies mid-load,
|
||||
there is no reference to cancel.
|
||||
2. **Release is impossible.** Residency is held inside the asset manager rather
|
||||
than by the component, and unloading requires the API that nothing in the
|
||||
project calls.
|
||||
3. **The load survives only incidentally** — because the completion delegate is
|
||||
bound to the handle and the streamable manager keeps it alive until it fires.
|
||||
|
||||
**[measured]** The same codebase contains the correct discipline elsewhere: UI
|
||||
async actions cancel their handles on teardown, and an async helper mixin cancels
|
||||
in its destructor. The pattern was known; it was not applied at the place with the
|
||||
largest closure.
|
||||
|
||||
**Take the shape from the UI code, not from the loader.**
|
||||
|
||||
---
|
||||
|
||||
## 7. Retention: six owners, audited by name
|
||||
|
||||
Anything that stays resident is held by one of:
|
||||
|
||||
1. reflected containers on long-lived subsystems;
|
||||
2. explicit GC references added by native code;
|
||||
3. retained load handles;
|
||||
4. delegate bindings capturing strong references;
|
||||
5. component and actor ownership chains;
|
||||
6. class default objects reached through class references.
|
||||
|
||||
### The convenience loader that defaults to permanent
|
||||
|
||||
**[measured]** A helper resolves a soft pointer, and a boolean parameter
|
||||
defaulting to **true** adds the result to a reflected set. There is no removal
|
||||
API, no call passing false, and no clearing anywhere in the project.
|
||||
|
||||
**[derived]** Every casual call therefore loads synchronously *and* roots the
|
||||
asset for the lifetime of the process. The soft pointer was chosen to defer
|
||||
loading; the helper converts it into a permanent hard reference at the call site,
|
||||
invisibly, by default.
|
||||
|
||||
The design that avoids it: make retention explicit at the call site or default it
|
||||
to false, provide a paired release, make the retained set inspectable from a
|
||||
console command, and give retention an owner and a scope rather than assigning it
|
||||
to "the asset manager". Recipe: AL-05.
|
||||
|
||||
### Both directions of the pointer error
|
||||
|
||||
Strong keys pin more than intended — a map keyed by actor pointers keeps every
|
||||
key actor and its whole presentation graph alive on one missed unregister.
|
||||
|
||||
Raw pointers in a non-reflected struct do the opposite: the collector cannot see
|
||||
them, so the values can be collected while referenced and the keys can dangle.
|
||||
|
||||
**[derived]** Both errors appear in real codebases, sometimes in the same one.
|
||||
Audit *direction*, not just presence.
|
||||
|
||||
---
|
||||
|
||||
## 8. Unloading, or the six "never"s
|
||||
|
||||
**[measured]** In the audited project, searched across all source:
|
||||
|
||||
| Mechanism | Occurrences |
|
||||
|---|---:|
|
||||
| primary asset unload | 0 |
|
||||
| bundle removal | 0 |
|
||||
| explicit collection | 0 |
|
||||
| async flush | 0 |
|
||||
| memory trim | 0 |
|
||||
|
||||
The one forced collection in the project is in the loading-screen hide path
|
||||
**[measured]**, and the experience loader's own teardown carries a comment
|
||||
admitting it deactivated without unloading **[measured]**.
|
||||
|
||||
**[derived]** The residency table for this project has "never" in the release
|
||||
column for six of its rows. For a session-based game restarted between matches
|
||||
that is an acceptable engineering position. For a service title that swaps modes
|
||||
in place it is not — and the difference is a product decision that must be made
|
||||
explicitly rather than discovered from a memory graph.
|
||||
|
||||
**The honest options are two:** implement release, or state plainly that mode
|
||||
changes require travel. What does not work is claiming runtime modularity while
|
||||
the release column is empty.
|
||||
|
||||
### Watch for collection ping-pong
|
||||
|
||||
**[measured]** Hiding the loading screen forces a full purge; showing it again
|
||||
synchronously loads the widget class that purge just collected. A forced
|
||||
collection at a transition boundary must not collect what the next transition
|
||||
immediately needs. Context: AL-07.
|
||||
|
||||
---
|
||||
|
||||
## 9. Failure handling is a memory concern
|
||||
|
||||
- a cancelled load must not run the success path;
|
||||
- the result of an asynchronous activation must be read;
|
||||
- partially loaded state must be releasable;
|
||||
- an unbounded wait converts a stall into a hang;
|
||||
- a loading screen with no reason string is not error handling.
|
||||
|
||||
**[measured]** In the audited project the cancel delegate invokes the *same*
|
||||
callback as completion, and the plugin-activation result parameter is never read.
|
||||
**[derived]** A cancelled or failed load therefore transitions to "loaded", and
|
||||
downstream code resolves soft references that are not there — where, by design,
|
||||
a failed resolve is silently skipped. The result is graceful degradation with no
|
||||
diagnostic, which is the most expensive kind.
|
||||
|
||||
### One tool worth copying outright
|
||||
|
||||
**[measured]** The project ships console variables that inject artificial delay
|
||||
into mode loading. That is a reproducible way to exercise slow-load and
|
||||
cancellation paths without a network or a cold cache, and almost nobody builds
|
||||
one.
|
||||
|
||||
---
|
||||
|
||||
## 10. Measurement plan
|
||||
|
||||
Static first, because it is cheap and finds structural problems:
|
||||
|
||||
1. build the dependency graph from the asset registry: direct hard and soft
|
||||
counts, transitive hard closure per definition, cross-plugin edges;
|
||||
2. rank definitions by closure and compare against their architectural layer —
|
||||
a catalog entry with a large closure is the finding;
|
||||
3. search for synchronous loads on boot and hot paths;
|
||||
4. inventory retention owners by type (§7).
|
||||
|
||||
Runtime, once a packaged build exists:
|
||||
|
||||
- memory reports before and after a full mode cycle;
|
||||
- object lists and reference queries on suspected roots — root-graph evidence,
|
||||
not process memory;
|
||||
- fifty to a hundred spawn, despawn and travel cycles, looking for monotonic
|
||||
growth;
|
||||
- allocation attribution from the engine's profiler;
|
||||
- the allocator's dangling-pointer mode when corruption is suspected.
|
||||
|
||||
### Do not measure residency in the editor
|
||||
|
||||
**[measured]** Editor builds can load **both** role bundles, by an explicit
|
||||
editor branch in bundle selection — a fact the project's own comments
|
||||
acknowledge, noting the resulting hitches.
|
||||
|
||||
**[derived]** An in-editor memory profile therefore represents neither shipping
|
||||
role. Use it for relative structural comparison only; absolute residency requires
|
||||
a packaged build.
|
||||
|
||||
---
|
||||
|
||||
## 11. When to stop
|
||||
|
||||
Optimize loading and residency when boot or transition time is a product problem,
|
||||
when a platform memory ceiling is real and near, when growth is monotonic across
|
||||
a session, or when a specific closure is measurably oversized.
|
||||
|
||||
Do not restructure references on aesthetics. A fourteen-package closure is not a
|
||||
problem. A thousand-package synchronous closure on the boot path is. Measure
|
||||
first, and keep the measurement, so the next change has something to be compared
|
||||
against.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The analysis and every number above come from a source and configuration audit of
|
||||
Epic's Lyra Starter Game on Unreal Engine 5.6, plus a dependency graph extracted
|
||||
from its asset registry through the editor.
|
||||
|
||||
Two boundaries applied to that reading and bound these claims: the engine's own
|
||||
asset manager, streamable manager and feature subsystem are not part of the
|
||||
project, so statements about their internals are inference from call sites rather
|
||||
than from source; and no profiler was run, so there are no timing or
|
||||
memory-footprint numbers here by construction — the package counts are static
|
||||
graph measurements.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill. Entry
|
||||
identifiers in [failure modes](failure-modes.md) resolve back to the audited
|
||||
locations, so any specific claim can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The structure is transferable;
|
||||
the specific defects are evidence, not guarantees about other versions. Re-run
|
||||
the detection recipes against your own tree before acting on any specific claim.
|
||||
@@ -0,0 +1,315 @@
|
||||
---
|
||||
name: ue-cosmetics-and-teams
|
||||
description: >-
|
||||
Design or review modular character cosmetics and team systems in Unreal
|
||||
Engine: replicated cosmetic intent with local realization, controller-owned
|
||||
selection versus pawn-owned realization, fast-array entries, dedicated-server
|
||||
exclusion, tag-driven mesh and animation selection, one authoritative team
|
||||
identifier with subscribed mirrors, team relationship and damage policy, and
|
||||
data-driven material presentation. Use when building skins, modular
|
||||
characters, team assignment, friendly fire, hit-marker filtering or
|
||||
team-coloured UI and effects.
|
||||
---
|
||||
|
||||
# UE cosmetics and teams
|
||||
|
||||
Two invariants, one per half:
|
||||
|
||||
> Replicate cosmetic **intent**, never the spawned presentation.
|
||||
|
||||
> Store team membership **once**, authoritatively; every other copy is a mirror
|
||||
> with a subscription.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes: [failure modes](references/failure-modes.md).
|
||||
|
||||
These two systems are in one skill because they meet at exactly one point —
|
||||
presentation — and the most common design error is letting them meet anywhere
|
||||
else.
|
||||
|
||||
Related skills: `ue-data-driven-architecture`, `ue-gas-architecture`,
|
||||
`ue-gameplay-messaging`, `ue-modular-gameplay`, `ue-multiplayer-authority`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
# Part I — Modular cosmetics
|
||||
|
||||
## 1. Split persistent intent from pawn realization
|
||||
|
||||
A controller-level component owns the selection, so it survives pawn death and
|
||||
respawn. A pawn-level component realizes it on the current body.
|
||||
|
||||
```text
|
||||
controller component pawn component
|
||||
├─ desired part list ├─ replicated part list
|
||||
├─ listens for pawn changes ├─ spawns local presentation components
|
||||
└─ applies and removes on the ├─ computes cosmetic tags
|
||||
current pawn, keeping └─ notifies body and animation selection
|
||||
handles
|
||||
```
|
||||
|
||||
If the only record of a player's appearance lives on the pawn, it dies with the
|
||||
pawn. That is the whole reason for the split.
|
||||
|
||||
### Possession transfer
|
||||
|
||||
On a possession change, in this order: remove owned parts from the old pawn;
|
||||
find the component on the new pawn; add owned parts; retain the returned handles.
|
||||
**Each owner removes only its own parts** — cheats, developer settings and real
|
||||
content coexist in one list, distinguished by a source field rather than by
|
||||
separate lists.
|
||||
|
||||
---
|
||||
|
||||
## 2. Define a part as data, and make it incapable of gameplay
|
||||
|
||||
A minimal part: a class to spawn, a socket to attach to, and a collision mode
|
||||
that defaults to none.
|
||||
|
||||
Default presentation safety: collision off unless explicitly requested; no
|
||||
ability system, attributes or authority inside the cosmetic actor; no replicated
|
||||
gameplay state; validated attach target; locally tracked handle.
|
||||
|
||||
### The guarantee is the net-mode guard, not the convention
|
||||
|
||||
Nothing in the type system prevents putting a gameplay-capable actor into a part
|
||||
class. The real isolation comes from one line: **presentation is not spawned at
|
||||
all on a dedicated server.** The server therefore cannot trace against, collide
|
||||
with, or be influenced by cosmetics, because they do not exist there.
|
||||
|
||||
Be precise about the limit of that guarantee: on a listen server or in
|
||||
standalone, the parts *do* exist on the authority. Recipe: CT-01.
|
||||
|
||||
---
|
||||
|
||||
## 3. Replicate intent with a fast array
|
||||
|
||||
The replicated entry holds the intent — class and socket. It must not hold the
|
||||
local handle or the spawned component. Put both of those in the same struct as
|
||||
**non-replicated** fields and let the annotation express the split, rather than
|
||||
maintaining two parallel types.
|
||||
|
||||
Callbacks realize changes locally: post-add spawns, pre-remove destroys,
|
||||
post-change rebuilds.
|
||||
|
||||
### Decide what "change" means before you need it
|
||||
|
||||
A common and defensible implementation makes post-change destroy and respawn the
|
||||
actor, because propagating a diff into an already-spawned actor is hard. That is
|
||||
fine and it is a decision with a cost: every edit is a full rebuild, so a system
|
||||
that edits parts frequently will hitch. Recipe: CT-02.
|
||||
|
||||
---
|
||||
|
||||
## 4. Tag-driven selection: rules plus a default
|
||||
|
||||
Parts contribute tags — body style, animation style, armour set. A selection set
|
||||
maps required tags to a mesh or an animation layer, with a fallback:
|
||||
|
||||
1. gather tags from all active parts;
|
||||
2. walk the rules in order;
|
||||
3. the first rule whose required tags are **all** present wins;
|
||||
4. otherwise use the default.
|
||||
|
||||
This microscopic pattern — an array of (asset, required tags) plus a default,
|
||||
first match wins — is about ten lines and is worth copying wholesale. It is the
|
||||
same shape whether it selects a mesh, a physics asset or an animation layer.
|
||||
|
||||
### Its one sharp edge
|
||||
|
||||
**Priority is array order.** There are no weights and no specificity ranking, so
|
||||
a broad rule placed above a narrow one silently shadows it, and the symptom is a
|
||||
wrong mesh rather than an error. Document the ordering requirement next to the
|
||||
array, and validate that no rule's tag set is a subset of an earlier rule's.
|
||||
Recipe: CT-03.
|
||||
|
||||
### Do not let cosmetic tags become gameplay tags
|
||||
|
||||
They may share a type. They must not share authority. A tag contributed by a
|
||||
cosmetic part is client-realizable data; nothing that decides gameplay may read
|
||||
it without an explicit, reviewed promotion.
|
||||
|
||||
---
|
||||
|
||||
## 5. Cosmetic lifecycle tests
|
||||
|
||||
Add and remove one part; several owners contributing; controller changes pawn;
|
||||
pawn destroyed before controller cleanup; replicated add, remove and change; late
|
||||
join reconstructs everything; **dedicated server spawns zero presentation
|
||||
actors**; listen server does not realize twice; invalid part class logs rather
|
||||
than crashes; incomplete tag set falls back.
|
||||
|
||||
---
|
||||
|
||||
# Part II — Teams
|
||||
|
||||
## 6. One authoritative source, and loud mirrors
|
||||
|
||||
Put the team identifier on the player state: it survives respawn, already
|
||||
replicates, and belongs to player identity rather than to a body.
|
||||
|
||||
Controllers, pawns and local players may mirror it. The rules that make mirroring
|
||||
safe:
|
||||
|
||||
- mirrors **subscribe** to the source and re-subscribe when the source object
|
||||
changes;
|
||||
- a write to a mirror is **rejected loudly**, not silently ignored;
|
||||
- a pawn may hold its own value only while unpossessed, and must say so in the
|
||||
error when it refuses.
|
||||
|
||||
The second rule is the one teams get wrong. A silent no-op setter on a mirror is
|
||||
indistinguishable from a working setter until something depends on it. Log an
|
||||
error naming the real owner. Recipe: CT-06.
|
||||
|
||||
---
|
||||
|
||||
## 7. Resolve team from an arbitrary object, in one place
|
||||
|
||||
Centralize the cascade so damage, UI, AI and targeting cannot each grow their own
|
||||
version:
|
||||
|
||||
1. the object implements the team interface;
|
||||
2. a pawn resolves through its controller or player state;
|
||||
3. an actor resolves through its **instigator** — this is what lets projectiles
|
||||
and effects carry the shooter's team without owning one;
|
||||
4. otherwise unassigned.
|
||||
|
||||
### Return a relationship, not a boolean
|
||||
|
||||
```text
|
||||
same team / different teams / indeterminate
|
||||
```
|
||||
|
||||
A boolean conflates "enemy" with "unknown", and the caller then has to guess.
|
||||
Every caller must handle the third case explicitly — and "indeterminate means
|
||||
allowed" is the failure that ships. See `ue-multiplayer-authority` for why an
|
||||
error path that permits is worse than one that denies.
|
||||
|
||||
---
|
||||
|
||||
## 8. One damage policy, used by both authority and prediction
|
||||
|
||||
One function answers "may this instigator damage this target". The damage
|
||||
calculation multiplies by its result as a scalar rather than branching, so the
|
||||
team rule participates in the formula like any other modifier.
|
||||
|
||||
**The same function decides whether to show a hit marker.** That single choice
|
||||
removes the entire class of "the client showed a hit and the server dealt no
|
||||
damage" bugs, because there is no second implementation to disagree with.
|
||||
|
||||
### Make the policy configurable before you need it
|
||||
|
||||
Friendly fire, self-damage, damage to unassigned actors and per-mode overrides
|
||||
are policy inputs. If they are hard-coded inside the damage calculation, a mode
|
||||
that wants different rules requires editing the calculation.
|
||||
|
||||
Watch for the shape where a header comment promises settings that the body does
|
||||
not implement — the comment is a design that was not built, and it reads as a
|
||||
feature. Recipe: CT-07.
|
||||
|
||||
---
|
||||
|
||||
## 9. Presentation as a named parameter bag
|
||||
|
||||
A team display asset should hold **named material parameters** — scalars,
|
||||
colours, textures, a short name — rather than a set of assets. Consumers ask for
|
||||
a parameter; materials are authored to expect the names. Adding a team then costs
|
||||
one asset instead of a duplicated material tree.
|
||||
|
||||
Apply it to meshes, effects and UI. Applying to an actor should include child
|
||||
actors by default — that is precisely what carries team colour onto cosmetic
|
||||
parts, and it is the only place the two halves of this skill touch.
|
||||
|
||||
### Two traps in the applying code
|
||||
|
||||
**Colour conversion drops alpha.** Converting a linear colour to a three-element
|
||||
vector to feed a material parameter silently discards the fourth channel. If any
|
||||
team parameter encodes opacity, it is gone. Recipe: CT-09.
|
||||
|
||||
**A viewer-relative parameter that is ignored.** If the API accepts a viewer
|
||||
identity — "always show my own team as blue" — then the implementation must use
|
||||
it. An accepted-and-ignored parameter is worse than an absent one: it makes the
|
||||
feature look supported and unimplementable at the same time. Recipe: CT-10.
|
||||
|
||||
---
|
||||
|
||||
## 10. Replicate intent, realize locally — the shared principle
|
||||
|
||||
| Domain | Replicated intent | Local realization |
|
||||
|---|---|---|
|
||||
| cosmetics | part definitions | spawned components, meshes, materials |
|
||||
| teams | team identifier | colours, UI, effects, nameplates |
|
||||
| composition | selected definition | activated plugins and actions |
|
||||
|
||||
Realization must be **idempotent**: applying the same intent twice produces one
|
||||
result. This is what makes late join deterministic from compact state.
|
||||
|
||||
---
|
||||
|
||||
## 11. Validation checklist
|
||||
|
||||
### Cosmetics
|
||||
|
||||
- [ ] part class is presentation-only and valid for the body;
|
||||
- [ ] socket exists, or a fallback is defined;
|
||||
- [ ] collision disabled by default;
|
||||
- [ ] handle and spawned component are non-replicated;
|
||||
- [ ] dedicated server realizes nothing;
|
||||
- [ ] every owner retains removal handles;
|
||||
- [ ] possession transfer removes then applies, exactly once;
|
||||
- [ ] rule ordering documented, fallback present, no shadowed rules;
|
||||
- [ ] no gameplay decision reads a cosmetic tag.
|
||||
|
||||
### Teams
|
||||
|
||||
- [ ] one authoritative source, authority-only writes;
|
||||
- [ ] mirrors subscribe, re-subscribe, and refuse writes loudly;
|
||||
- [ ] resolution is centralized and includes the instigator step;
|
||||
- [ ] the indeterminate relationship is handled explicitly at every caller;
|
||||
- [ ] damage and hit feedback call the same policy function;
|
||||
- [ ] friendly fire and self-damage are configurable and tested;
|
||||
- [ ] "private" team data is actually filtered, or is renamed;
|
||||
- [ ] display parameters exist in the target materials;
|
||||
- [ ] a viewer-relative API uses the viewer.
|
||||
|
||||
---
|
||||
|
||||
## 12. Keep them separate
|
||||
|
||||
They meet at one seam:
|
||||
|
||||
```text
|
||||
team identifier → display asset → material parameters on cosmetic parts
|
||||
```
|
||||
|
||||
Do not let cosmetics own the team identifier, and do not let the team system own
|
||||
part selection. Separation is what allows skins independent of teams, players
|
||||
with cosmetics and no team, team recolour over any skin, and viewer-relative
|
||||
presentation without touching authoritative state.
|
||||
|
||||
Combine them only inside a small presentation adapter — never in authority or
|
||||
replication ownership.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a source audit of Epic's Lyra
|
||||
Starter Game on Unreal Engine 5.6 — roughly 1300 lines of cosmetics and 1800 of
|
||||
teams, read rather than run. Several of the wiring call sites live in Blueprint
|
||||
graphs, which are binary and were not read; those are marked open rather than
|
||||
concluded, and the skill is written so that none of its claims depend on them.
|
||||
|
||||
Source addresses stay in the research archive. Each entry carries a stable
|
||||
identifier (`CT-01`, `CT-02`, …) resolving back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The limits are specific and worth
|
||||
repeating: the components are added to actors from data assets that were not
|
||||
read, the team colour application has no C++ caller, and the animation-layer
|
||||
selection is invoked from a Blueprint. Absence of a C++ caller is **not** absence
|
||||
of a caller — see the failure modes for how to check properly before deleting
|
||||
anything.
|
||||
@@ -0,0 +1,406 @@
|
||||
# Failure modes: cosmetics and teams
|
||||
|
||||
Eleven ways an appearance system or a team system produces the wrong result
|
||||
without producing an error.
|
||||
|
||||
The shared property differs slightly between the two halves, and both are worth
|
||||
stating.
|
||||
|
||||
**Cosmetics fail by looking almost right.** A wrong mesh, a missing part or a
|
||||
duplicated accessory is a visual outcome, and every visual outcome is plausible
|
||||
to code. Nothing can assert that a character looks correct, so these defects are
|
||||
found by people, late, and reported as "the skin is broken".
|
||||
|
||||
**Teams fail by permitting.** Every interesting team defect resolves to
|
||||
"indeterminate", and the branch that handles indeterminate almost always allows
|
||||
the action, because denying it would have blocked something during development.
|
||||
|
||||
Recipes use `rg` from a project source root and were executed against the audited
|
||||
project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## Cosmetics
|
||||
|
||||
### CT-01 - The isolation guarantee is one line, and it is narrower than it looks
|
||||
|
||||
**Mechanism.** Cosmetic presentation is excluded from the dedicated server by a
|
||||
single net-mode check at the spawn site. That check is the entire reason cosmetic
|
||||
content cannot influence authoritative simulation.
|
||||
|
||||
**Why it is silent.** It works. On a dedicated server no presentation actor
|
||||
exists, so none can collide, block a trace or tick. The guarantee holds exactly
|
||||
where it is claimed.
|
||||
|
||||
**Why the obvious check misses it.** The guarantee is usually restated as
|
||||
"cosmetics cannot affect gameplay", which is a stronger claim than the code
|
||||
makes. On a listen server or in standalone the parts **do** exist on the
|
||||
authority, with whatever collision and components their class carries. Nothing in
|
||||
the type system prevents a part class from containing an ability system or
|
||||
enabled collision — the constraint is content discipline plus one `if`.
|
||||
|
||||
**Symptom.** A cosmetic accessory that blocks a shot, or a part actor that ticks
|
||||
expensively, in exactly the configuration used for local playtesting — and never
|
||||
on the dedicated server where the "real" testing happens.
|
||||
|
||||
**Detect.** Find the guard, confirm there is only one, and check what part
|
||||
classes are permitted to contain:
|
||||
|
||||
```bash
|
||||
rg -n "NM_DedicatedServer|IsNetMode" --glob "*.cpp" <cosmetics-module>/
|
||||
rg -n "CollisionMode|SetActorEnableCollision" --glob "*.cpp" --glob "*.h" \
|
||||
<cosmetics-module>/
|
||||
```
|
||||
|
||||
In the audited project the first command returns **exactly one line** for the
|
||||
whole module. That is elegant and it is also the entire safety boundary, which is
|
||||
worth knowing before relying on it.
|
||||
|
||||
**Guardrail.** State the guarantee accurately: "not present on a dedicated
|
||||
server", not "cannot affect gameplay". Validate part classes at author time —
|
||||
reject any that contain an ability system component, replicated properties or
|
||||
enabled collision by default.
|
||||
|
||||
---
|
||||
|
||||
### CT-02 - Every change rebuilds the actor
|
||||
|
||||
**Mechanism.** The replication callback for a changed entry destroys the spawned
|
||||
presentation and spawns a new one, because propagating a diff into a live actor
|
||||
is hard and rarely needed.
|
||||
|
||||
**Why it is silent.** It is correct. The result after a change is exactly the
|
||||
result after an add, which is what the system promises. The cost is a rebuild,
|
||||
and a rebuild looks identical to a first build.
|
||||
|
||||
**Why the obvious check misses it.** The implementation is deliberate and
|
||||
commented as such. Reviewing it shows a decision, not a defect. The consequence
|
||||
only appears at a usage frequency nobody has yet — a system that changes parts
|
||||
rarely and one that changes them per second are the same code.
|
||||
|
||||
**Symptom.** A hitch on every cosmetic edit. Discovered when a feature starts
|
||||
mutating parts at runtime — a preview screen, a colour cycler, a progression
|
||||
system — long after the mechanism was reviewed.
|
||||
|
||||
**Detect.** Read what the change callback does, and compare it with the add path:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "PostReplicatedChange" --glob "*.cpp" . | rg -i "destroy|spawn|update"
|
||||
```
|
||||
|
||||
If change is implemented as destroy plus spawn, treat the mechanism as
|
||||
add/remove only and design callers accordingly.
|
||||
|
||||
**Guardrail.** Document the cost where authors will see it. If frequent edits are
|
||||
planned, add a real update path for the fields that can change in place, and keep
|
||||
the rebuild for the ones that cannot.
|
||||
|
||||
---
|
||||
|
||||
### CT-03 - Rule order is priority, and nothing says so
|
||||
|
||||
**Mechanism.** A selection set is an array of rules, each with a set of required
|
||||
tags, plus a default. The first rule whose tags are all present wins. Rules are
|
||||
not sorted by specificity.
|
||||
|
||||
**Why it is silent.** Every possible input produces a result — either a rule or
|
||||
the default. There is no "no match" state to report, and the result is always a
|
||||
valid asset.
|
||||
|
||||
**Why the obvious check misses it.** The array looks like a set, and sets have no
|
||||
order. A reviewer checking "are the rules correct?" checks each rule in
|
||||
isolation, where each is correct. The defect exists only in the relationship
|
||||
between a broad rule and a narrower one below it, which requires comparing every
|
||||
pair.
|
||||
|
||||
**Symptom.** A character with a specific tag combination gets the generic mesh.
|
||||
Nobody notices for weeks, because it is the *right kind* of mesh.
|
||||
|
||||
**Detect.** Find selection sets and check for shadowing — a rule whose tag set is
|
||||
a subset of an earlier rule's:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A6 "SelectBest|RequiredTags" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Then, for each set, verify the shadowing property in a test rather than by
|
||||
reading: for every pair of rules `i < j`, assert that rule `i`'s tags are not a
|
||||
subset of rule `j`'s.
|
||||
|
||||
**Guardrail.** Write the ordering rule as a comment on the array and enforce the
|
||||
subset check in validation. A rule that can never win is a content bug the
|
||||
compiler cannot see.
|
||||
|
||||
---
|
||||
|
||||
### CT-04 - The realization path re-applies unconditionally
|
||||
|
||||
**Mechanism.** Recomputing appearance sets the mesh on every change broadcast,
|
||||
relying on the setter to be cheap when the value has not changed — while passing
|
||||
a flag that forces a pose reinitialization.
|
||||
|
||||
**Why it is silent.** The result is correct. The mesh is right, the pose is
|
||||
right, and the cost is invisible at the frequency the system is normally used.
|
||||
|
||||
**Why the obvious check misses it.** The code carries a comment explaining that
|
||||
the setter is a no-op when the mesh is unchanged, which answers the question a
|
||||
reviewer would ask. The forced flag is a separate argument on the same call, and
|
||||
it is not what the comment is about.
|
||||
|
||||
**Symptom.** A visible pose pop or an animation reset whenever anything about
|
||||
cosmetics changes, including changes that do not affect the mesh at all.
|
||||
|
||||
**Detect.** Check the arguments of the re-application call, not just the call:
|
||||
|
||||
```bash
|
||||
rg -n -B6 "SetSkeletalMesh\(" --glob "*.cpp" . | rg "bReinit|true|Broadcast"
|
||||
```
|
||||
|
||||
A hardcoded reinitialization flag on a path that runs on every change is the
|
||||
finding.
|
||||
|
||||
**Guardrail.** Compare before applying, and pass the reinitialization flag only
|
||||
when the mesh actually changed. Idempotent realization means "applying twice
|
||||
costs nothing", not merely "applying twice is correct".
|
||||
|
||||
---
|
||||
|
||||
### CT-05 - Everything is a hard reference
|
||||
|
||||
**Mechanism.** Part classes and selection-set assets are hard references, so the
|
||||
entire cosmetic catalogue is loaded with whatever holds the rules.
|
||||
|
||||
**Why it is silent.** Hard references always resolve. Nothing is ever missing,
|
||||
nothing ever fails to load, and every asset is available the moment it is needed.
|
||||
|
||||
**Why the obvious check misses it.** Reference correctness is what reviews check,
|
||||
and hard references are maximally correct. The cost is memory residency, which
|
||||
belongs to a different discipline and a different person, and which does not
|
||||
appear in any test of the cosmetics system.
|
||||
|
||||
**Symptom.** Memory proportional to the size of the cosmetic catalogue rather
|
||||
than to what is worn, discovered during a platform memory pass and attributed to
|
||||
content rather than to a reference-type decision.
|
||||
|
||||
**Detect.** Count hard versus soft references in the cosmetic data types:
|
||||
|
||||
```bash
|
||||
rg -n "TSubclassOf<|TObjectPtr<" --glob "*.h" <cosmetics-module>/ | wc -l
|
||||
rg -n "TSoftClassPtr<|TSoftObjectPtr<" --glob "*.h" <cosmetics-module>/ | wc -l
|
||||
```
|
||||
|
||||
A large first number with a zero second is the finding. See
|
||||
`ue-asset-loading-and-memory` for what this costs and how to measure it.
|
||||
|
||||
**Guardrail.** Cosmetic catalogues are the textbook case for soft references and
|
||||
bundles: large, optional, and mostly unworn. Decide per field and record the
|
||||
decision in metadata.
|
||||
|
||||
---
|
||||
|
||||
## Teams
|
||||
|
||||
### CT-06 - The mirror setter that silently does nothing
|
||||
|
||||
**Mechanism.** Team identity is authoritative on one object and mirrored on
|
||||
several others. The mirrors implement the interface's setter because they must,
|
||||
and the implementation does nothing.
|
||||
|
||||
**Why it is silent.** Doing nothing is correct — the mirror is not the owner. The
|
||||
call succeeds, returns, and the value is unchanged.
|
||||
|
||||
**Why the obvious check misses it.** The setter exists and is implemented. Any
|
||||
review asking "does this type support setting a team?" answers yes. The caller
|
||||
has no return value to check and no reason to suspect the write did not land.
|
||||
|
||||
**Symptom.** Code that sets a team on a controller or a pawn and observes the old
|
||||
value. The investigation goes to replication, then to ordering, then eventually to
|
||||
the setter.
|
||||
|
||||
**Detect.** Read the bodies of every team setter and classify them:
|
||||
|
||||
```bash
|
||||
rg -n -A4 "::SetGenericTeamId|::SetTeamId" --glob "*.cpp" . \
|
||||
| rg -B1 "UE_LOG|ensure|^\s*\}"
|
||||
```
|
||||
|
||||
The audited project does this **correctly** and is worth copying: the mirror
|
||||
setters log an error naming the real owner rather than returning quietly. That
|
||||
one line converts an invisible failure into a searchable one.
|
||||
|
||||
**Guardrail.** A mirror's setter logs an error identifying the authoritative
|
||||
owner. Never implement a no-op setter to satisfy an interface.
|
||||
|
||||
---
|
||||
|
||||
### CT-07 - The header promises a policy the body does not implement
|
||||
|
||||
**Mechanism.** The damage-permission function is documented as taking friendly
|
||||
fire settings into account. The body contains no settings and no lookup.
|
||||
|
||||
**Why it is silent.** The behaviour is coherent — allies are always safe — and
|
||||
that is the right default for most modes. Nothing malfunctions.
|
||||
|
||||
**Why the obvious check misses it.** The comment *is* the documentation. A team
|
||||
adopting the system reads the header, concludes friendly fire is supported, and
|
||||
plans a mode around it. Discovering otherwise requires reading a body that looks
|
||||
finished.
|
||||
|
||||
**Symptom.** A mode that needs friendly fire is scoped as a configuration change
|
||||
and turns out to be a design change, in a function every damage path depends on.
|
||||
|
||||
**Detect.** Compare the promise with the implementation:
|
||||
|
||||
```bash
|
||||
rg -n -i "friendly fire" --glob "*.h" --glob "*.cpp" .
|
||||
rg -n -i "bFriendlyFire|bAllowFriendly|FriendlyFireSetting" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
In the audited project the first command finds the promise in a header comment
|
||||
and the second returns **zero results project-wide**. A promise with no
|
||||
implementing symbol is the finding, and this recipe generalises to any header
|
||||
comment describing configurability.
|
||||
|
||||
**Guardrail.** Treat a comment describing behaviour as a claim requiring a
|
||||
symbol. If the setting does not exist, the comment describes a plan and must say
|
||||
so.
|
||||
|
||||
---
|
||||
|
||||
### CT-08 - Indeterminate resolves to permitted
|
||||
|
||||
**Mechanism.** Team comparison correctly returns three states. The damage rule
|
||||
treats the indeterminate case as allowed when the target satisfies an unrelated
|
||||
condition — in the audited project, having an ability system component.
|
||||
|
||||
**Why it is silent.** It exists to make something work: an unassigned training
|
||||
target must be damageable. It succeeds at that, and the permissive branch is
|
||||
never reached by any assigned actor during normal play.
|
||||
|
||||
**Why the obvious check misses it.** The function has three branches and handles
|
||||
all three, so it passes any review asking whether the indeterminate case is
|
||||
handled. It *is* handled — permissively — and the marker in the code says
|
||||
"temporary".
|
||||
|
||||
**Symptom.** Any actor with an ability system and no team assignment is damageable
|
||||
by anyone. Harmless while the only such actor is a target dummy; a rule violation
|
||||
the moment a neutral faction, a destructible objective or an NPC exists.
|
||||
|
||||
**Detect.** Find every consumer of the relationship enum and read its
|
||||
indeterminate branch:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A10 "InvalidArgument|Indeterminate|NoTeam" --glob "*.cpp" . \
|
||||
| rg "return true|= true|Allow"
|
||||
rg -n -i "//\s*@?TODO.*(temporary|until)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
The second command is the high-yield one: a permissive branch that its own author
|
||||
marked temporary is the strongest possible confirmation that the finding is real
|
||||
and known.
|
||||
|
||||
**Guardrail.** Failure to determine must deny. Give the target dummy a team
|
||||
instead of giving every unassigned actor an exemption — solve the specific
|
||||
problem in data, not the general rule in code.
|
||||
|
||||
---
|
||||
|
||||
### CT-09 - A colour conversion that drops alpha
|
||||
|
||||
**Mechanism.** A four-channel colour is converted to a three-element vector to be
|
||||
passed as a material parameter. The fourth channel is discarded.
|
||||
|
||||
**Why it is silent.** Three of four channels arrive correctly, so the colour is
|
||||
right. Alpha is frequently unused in team colours, making the loss invisible for
|
||||
as long as nobody encodes anything in it.
|
||||
|
||||
**Why the obvious check misses it.** The conversion is one token inside an
|
||||
otherwise correct call. Both types are colours, both are correct, and the review
|
||||
question — "is the parameter set?" — is answered yes.
|
||||
|
||||
**Symptom.** A team parameter that encodes opacity or a blend factor silently
|
||||
reads as its default. In the same file, the effects-system path preserves all four
|
||||
channels, so the same asset behaves differently in two places.
|
||||
|
||||
**Detect.** Find lossy conversions at parameter-setting sites:
|
||||
|
||||
```bash
|
||||
rg -n "SetVectorParameterValue\w*\(.*FVector\(" --glob "*.cpp" .
|
||||
rg -n "SetVariableLinearColor|SetVectorParameterValue" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Two applicators for the same data with different channel counts is the finding.
|
||||
In the audited project the material paths convert and the effects path does not.
|
||||
|
||||
**Guardrail.** Pass four channels where the API accepts them. If a conversion is
|
||||
unavoidable, assert or document that alpha is unused.
|
||||
|
||||
---
|
||||
|
||||
### CT-10 - An accepted parameter that is ignored
|
||||
|
||||
**Mechanism.** The display-asset accessor takes a viewer identity so that
|
||||
presentation can be relative — "my team always looks blue". The implementation
|
||||
does not read it.
|
||||
|
||||
**Why it is silent.** The function returns the correct absolute asset. Every
|
||||
caller gets a valid result, and the feature the parameter implies has simply
|
||||
never been exercised.
|
||||
|
||||
**Why the obvious check misses it.** The signature and the header comment
|
||||
describe the feature completely. Confirming absence requires reading the body,
|
||||
where the parameter is unused — and an unused parameter produces no warning
|
||||
because it is part of a virtual-looking public API.
|
||||
|
||||
**Symptom.** A team implementing viewer-relative colours writes correct calling
|
||||
code and observes absolute colours. The bug appears to be in their code.
|
||||
|
||||
**Detect.** For any parameter that names a feature, check that the body reads it:
|
||||
|
||||
```bash
|
||||
P='ViewerTeamId'
|
||||
rg -n "\b$P\b" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Occurrences only in the declaration and the definition's signature — with none in
|
||||
the body — is the finding. In the audited project the body carries a comment
|
||||
stating the parameter is currently ignored, which is honest and still ships an
|
||||
API that cannot do what it claims.
|
||||
|
||||
**Guardrail.** Do not accept a parameter you do not use. Remove it, or implement
|
||||
it, or make the function name state the limitation. This is the same defect class
|
||||
as an ordering field that never sorts — see `ue-ui-architecture`, UI-01.
|
||||
|
||||
---
|
||||
|
||||
### CT-11 - "Private" that is not filtered
|
||||
|
||||
**Mechanism.** Team data is split into a public and a private actor to express a
|
||||
replication boundary. The private one has no filtering; it is an empty subclass
|
||||
with a note that privacy is not implemented.
|
||||
|
||||
**Why it is silent.** Everything works. Both actors replicate, both carry their
|
||||
data, and the split is structurally correct — it is only the *filtering* that is
|
||||
absent.
|
||||
|
||||
**Why the obvious check misses it.** The architecture is visible and right: two
|
||||
types, two names, a clear intent. Reading the class names answers the question.
|
||||
Confirming means noticing that the private subclass has no body.
|
||||
|
||||
**Symptom.** Data placed in the private actor because it is private is replicated
|
||||
to every client. Nothing indicates this until someone inspects network traffic —
|
||||
or until a competitor does.
|
||||
|
||||
**Detect.** Check that the privacy-named type actually filters:
|
||||
|
||||
```bash
|
||||
rg -n -A15 "class \w*PrivateInfo|class \w*Private\w*" --glob "*.h" .
|
||||
rg -n "IsNetRelevantFor|GetLifetimeReplicatedProps|COND_" --glob "*.cpp" . \
|
||||
| rg -i "private"
|
||||
```
|
||||
|
||||
An empty subclass, or one with no relevancy or condition logic, is the finding.
|
||||
|
||||
**Guardrail.** A name is not a mechanism. Either implement the filtering — via
|
||||
relevancy, replication conditions or a replication graph — or rename the type to
|
||||
what it is. In the meantime, do not put anything in it that matters.
|
||||
@@ -0,0 +1,288 @@
|
||||
# Patterns: cosmetics and teams in a measured reference product
|
||||
|
||||
A worked example of both halves of `SKILL.md` against one reference project read
|
||||
as source. The two systems are presented together because the audit found what
|
||||
the skill claims: they touch at exactly one point, and both are well built on
|
||||
either side of it.
|
||||
|
||||
This file is unusually positive. That is the finding, not a tone: the
|
||||
architecture here is the strongest in the audited project, and the defects are
|
||||
narrow and specific. Separating "this design is worth copying" from "this
|
||||
implementation is finished" is the entire job.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from
|
||||
measured facts; **[open]** — not answerable from source, because the call site is
|
||||
in a binary asset.
|
||||
|
||||
---
|
||||
|
||||
## 1. The two-level split, and what it buys
|
||||
|
||||
```text
|
||||
controller component owns the desired part list, survives respawn
|
||||
│ on possession change: remove from old, apply to new, keep handles
|
||||
▼
|
||||
pawn component replicated list, spawns presentation, computes tags
|
||||
```
|
||||
|
||||
The controller component subscribes to possession changes **only on the
|
||||
authority**, and if a pawn already exists at subscription time it invokes its own
|
||||
handler with a null previous pawn **[measured]**. That second detail is the
|
||||
pattern's whole robustness: the same code path handles "pawn already here" and
|
||||
"pawn arrives later", so ordering stops mattering.
|
||||
|
||||
Transfer removes by handle from the old pawn and re-applies to the new, guarding
|
||||
against double-add by checking whether the handle is already valid **[measured]**.
|
||||
|
||||
**[derived]** This generalises well beyond appearance. Any state that belongs to
|
||||
the *player* and is realized on the *body* — loadouts, perks, temporary buffs —
|
||||
has the same shape and the same failure mode if you skip it: the state dies with
|
||||
the pawn.
|
||||
|
||||
The link from controller to pawn is found dynamically each time, with no cache
|
||||
**[measured]**. If the pawn has no cosmetics component the lookup returns null and
|
||||
nothing happens, silently — a tolerance that is correct for optional cosmetics and
|
||||
would be wrong for anything required.
|
||||
|
||||
---
|
||||
|
||||
## 2. Replicating intent, with the split declared in the type
|
||||
|
||||
One structure carries both server-only and client-only fields, distinguished by
|
||||
annotation rather than by having two types **[measured]**:
|
||||
|
||||
| Field | Replicated | Lives on |
|
||||
|---|---|---|
|
||||
| the part itself: class plus socket | yes | server → clients |
|
||||
| the handle returned to the caller | **no** | server only |
|
||||
| the spawned presentation component | **no** | client only |
|
||||
|
||||
**[derived]** This is the cleanest available demonstration of "replicate intent,
|
||||
realize locally". The wire carries a class reference and a socket name; the
|
||||
receiving machine reconstructs the appearance. Late join is deterministic because
|
||||
the state is compact and complete.
|
||||
|
||||
Change handling is deliberately coarse: the change callback destroys and respawns
|
||||
rather than propagating a diff, and the authors say so in a comment
|
||||
**[measured]**. Recipe: CT-02.
|
||||
|
||||
Part identity is defined as the pair (class, socket) and deliberately **excludes**
|
||||
the collision mode from the comparison **[measured]** — so two requests differing
|
||||
only in collision are the same part. **[derived]** That is a real decision about
|
||||
what identity means, made once, in the comparison function where it belongs.
|
||||
|
||||
---
|
||||
|
||||
## 3. The isolation guarantee, stated precisely
|
||||
|
||||
One line excludes presentation from the dedicated server, and it is the only such
|
||||
line in the module **[measured]**.
|
||||
|
||||
**[derived]** The consequence is strong and worth stating exactly: on a dedicated
|
||||
server the cosmetic actors do not exist, so they cannot collide, block a trace,
|
||||
tick, or influence the authoritative simulation in any way. That is a real
|
||||
architectural guarantee obtained for one `if`.
|
||||
|
||||
It is also narrower than the sentence people repeat. On a listen server or in
|
||||
standalone the parts exist on the authority, with whatever their class contains,
|
||||
and nothing in the type system prevents a part class from carrying an ability
|
||||
system or enabled collision. The other protections — collision defaulting to off,
|
||||
authority-only mutators, no ability-system includes anywhere in the module
|
||||
**[measured]** — are conventions and content discipline. Recipe: CT-01.
|
||||
|
||||
**The honest form:** *not present on a dedicated server*, not *cannot affect
|
||||
gameplay*.
|
||||
|
||||
---
|
||||
|
||||
## 4. The ten-line pattern worth stealing
|
||||
|
||||
An array of rules, each an asset plus a set of required tags, plus a default.
|
||||
Gather tags, walk the rules in order, first rule whose tags are **all** present
|
||||
wins, otherwise the default **[measured]**.
|
||||
|
||||
It appears at least three times in the audited project — body mesh, forced
|
||||
physics asset, and weapon animation layer **[measured]** — and costs roughly ten
|
||||
lines each time.
|
||||
|
||||
**[derived]** Its sharp edge is that priority is array order, with no weights and
|
||||
no specificity ranking. A broad rule above a narrow one silently shadows it, and
|
||||
the failure produces a plausible wrong asset rather than an error. Recipe: CT-03.
|
||||
|
||||
The tag namespace is declared in config and is deliberately small — a body style
|
||||
and two animation styles **[measured]**. **[derived]** Small is the right size:
|
||||
every tag in this namespace multiplies the rule matrix that has to be reasoned
|
||||
about.
|
||||
|
||||
---
|
||||
|
||||
## 5. Where the two systems touch
|
||||
|
||||
Exactly one place, and the code says so. The cosmetics component broadcasts a
|
||||
change with a comment noting that observers may need to apply team colouring
|
||||
**[measured]**; the team display asset applies material parameters to an actor
|
||||
**and its child actors by default** **[measured]** — and cosmetic parts are child
|
||||
actors.
|
||||
|
||||
**[derived]** That default is what carries team colour onto modular parts. It is
|
||||
one boolean default, and it is the seam.
|
||||
|
||||
**[open]** The actual call is made from a Blueprint: the applying function has no
|
||||
C++ caller. Which leads to the most important methodological point in this file.
|
||||
|
||||
---
|
||||
|
||||
## 6. What could not be answered, and why that matters
|
||||
|
||||
Four things in this subsystem have **no C++ call site** **[measured]**:
|
||||
|
||||
- who applies team colours to an actor;
|
||||
- who selects the weapon animation layer from cosmetic tags;
|
||||
- where the cosmetics component is added to a pawn;
|
||||
- where the controller component is added to a controller.
|
||||
|
||||
All four are wired in Blueprint or data assets, which are binary and were not
|
||||
read.
|
||||
|
||||
**[derived]** The correct conclusion is *"the wiring is in a place this audit
|
||||
cannot see"*, and the incorrect one — available, tempting, and wrong — is *"this
|
||||
function is unused"*. Both components are marked as spawnable from Blueprint
|
||||
**[measured]**, which is positive evidence for the first reading.
|
||||
|
||||
This is the general rule from `ue-reference-project-adoption`, met in its natural
|
||||
habitat: **absence of a C++ caller is not absence of a caller.** Anyone deleting
|
||||
"dead" code in a project with a large Blueprint surface needs the editor's
|
||||
reference viewer, not a text search.
|
||||
|
||||
---
|
||||
|
||||
## 7. Teams: one source, mirrors that refuse loudly
|
||||
|
||||
Membership lives on the player state, replicated push-based, authority-only, with
|
||||
a non-authority write logged as an error **[measured]**.
|
||||
|
||||
Every mirror — controller, bot controller, local player, pawn — reads through to
|
||||
the source, and their setters **log an error naming the real owner** rather than
|
||||
silently doing nothing **[measured]**.
|
||||
|
||||
**[derived]** This is the single most copyable line in the team system. A silent
|
||||
no-op setter on a mirror is indistinguishable from a working one until something
|
||||
depends on it; an error line converts an invisible failure into a searchable one,
|
||||
at a cost of one statement. Recipe: CT-06.
|
||||
|
||||
The pawn mirror is more subtle and also correct: it accepts a direct write **only
|
||||
while unpossessed**, and refuses with an explanatory error once a controller owns
|
||||
it **[measured]**. On possession it takes the controller's value and subscribes to
|
||||
its change delegate; on unpossession it unsubscribes and calls an overridable hook
|
||||
whose default is "no team" **[measured]**.
|
||||
|
||||
**[derived]** That hook is an extension point placed exactly where a neutral
|
||||
faction would need it, which is the difference between a design and an
|
||||
implementation that happens to work.
|
||||
|
||||
---
|
||||
|
||||
## 8. Resolution, comparison, and the one permissive branch
|
||||
|
||||
Resolution is a four-step cascade — interface, instigator, team-info actor,
|
||||
associated player state **[measured]**. **[derived]** The instigator step is what
|
||||
lets projectiles and area effects carry the shooter's team without owning one; a
|
||||
design that omits it grows a team field on every projectile.
|
||||
|
||||
Comparison returns three states rather than a boolean **[measured]**, which is
|
||||
correct and is what makes the next finding visible at all.
|
||||
|
||||
The damage rule permits four ways **[measured]**, and the third is the finding:
|
||||
when the relationship is indeterminate, damage is allowed if the target has an
|
||||
ability system component — with an author's comment marking it temporary until a
|
||||
training dummy gets a team assignment.
|
||||
|
||||
**[derived]** It exists to make one specific actor damageable and it generalises
|
||||
to every unassigned actor with an ability system. Harmless while the only such
|
||||
actor is a dummy; a rule violation the moment a neutral faction or a destructible
|
||||
objective exists. **Solve it in data — give the dummy a team — rather than in the
|
||||
rule.** Recipe: CT-08.
|
||||
|
||||
Friendly fire is promised by a header comment and implemented nowhere: a
|
||||
project-wide search for any friendly-fire symbol returns **zero** **[measured]**.
|
||||
Allies are unconditionally safe. Recipe: CT-07.
|
||||
|
||||
---
|
||||
|
||||
## 9. One policy function, two consumers — the pattern to copy
|
||||
|
||||
The same permission function is called by the authoritative damage calculation
|
||||
and by the client-side decision to show a hit marker **[measured]**.
|
||||
|
||||
**[derived]** This removes an entire bug class by construction. "The client showed
|
||||
a hit and the server dealt no damage" requires two implementations of one rule to
|
||||
disagree, and here there is only one implementation. Any project with predicted
|
||||
feedback should copy this arrangement before copying anything else in the team
|
||||
system.
|
||||
|
||||
The damage calculation consumes the result as a **0-or-1 multiplier** in the
|
||||
formula rather than as an early return **[measured]**. **[derived]** The team rule
|
||||
becomes an ordinary participant in the arithmetic, alongside distance and surface
|
||||
attenuation, which keeps the execution pipeline uniform and makes the rule easy to
|
||||
soften later — a scalar between zero and one is already expressible.
|
||||
|
||||
---
|
||||
|
||||
## 10. Presentation as named parameters
|
||||
|
||||
The display asset is a bag of named material parameters — scalars, colours,
|
||||
textures, a short name — not a set of assets **[measured]**. Materials are
|
||||
authored to expect the names, so adding a team costs one asset rather than a
|
||||
duplicated material tree.
|
||||
|
||||
Two defects in the applying code, both narrow:
|
||||
|
||||
- **Alpha is dropped** by the material paths, which convert a four-channel colour
|
||||
to a three-element vector, while the effects path preserves all four
|
||||
**[measured]**. The same asset therefore behaves differently in two consumers.
|
||||
Recipe: CT-09.
|
||||
- **A viewer identity parameter is accepted and ignored**, with a comment saying
|
||||
so **[measured]**. The signature and header describe viewer-relative colouring;
|
||||
the implementation returns the absolute asset. Recipe: CT-10.
|
||||
|
||||
Reactive binding is done well: an observation action broadcasts **immediately on
|
||||
activation** so a late subscriber gets current state rather than waiting for the
|
||||
next change, and a second action re-subscribes correctly when the observed team
|
||||
changes **[measured]**. **[derived]** Both are small patterns that eliminate a
|
||||
race, and both are reusable anywhere state is observed through a key that can
|
||||
itself change.
|
||||
|
||||
---
|
||||
|
||||
## 11. What to copy, in order
|
||||
|
||||
1. **One policy function for authority and prediction** (§9). Highest value,
|
||||
lowest cost, removes a whole bug class.
|
||||
2. **Loud mirrors** (§7). One log line per setter.
|
||||
3. **Controller-owned intent, pawn-owned realization** (§1), including the
|
||||
already-possessed case.
|
||||
4. **Non-replicated fields inside a replicated entry** (§2) — one type instead of
|
||||
two.
|
||||
5. **Rules-plus-default selection** (§4), with a shadowing check added.
|
||||
6. **Immediate first broadcast on subscription** (§10).
|
||||
|
||||
What to fix while copying: the indeterminate branch (CT-08), the friendly-fire
|
||||
gap (CT-07), the alpha conversion (CT-09), the ignored viewer parameter (CT-10),
|
||||
and the privacy that is a name rather than a mechanism (CT-11).
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6: roughly 1300
|
||||
lines of cosmetics across 11 files and 1800 lines of teams across 22, read as
|
||||
source in a single workspace. Source addresses stay in the research archive that
|
||||
produced this skill; each `CT-` identifier resolves back to the audited location
|
||||
there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Four wiring call sites are in Blueprint or data assets and were not read; they are
|
||||
marked open above and nothing here depends on them. Everything else describes one
|
||||
project on one engine version. Re-run the recipes against your own tree before
|
||||
trusting any specific claim.
|
||||
@@ -0,0 +1,413 @@
|
||||
---
|
||||
name: ue-data-driven-architecture
|
||||
description: >-
|
||||
Design or review data-driven architecture in Unreal Engine: definition layers,
|
||||
data asset versus zero-node Blueprint class defaults versus instanced
|
||||
fragments and actions, primary asset identifiers, soft and hard and class
|
||||
references, asset bundles, modular composition, validation, and the boundary
|
||||
between code and data. Use when adding game modes, pawn archetypes, ability
|
||||
packages, item and equipment definitions, playlists or feature plugins, or
|
||||
when a manager is growing per-variant branching.
|
||||
---
|
||||
|
||||
# UE data-driven architecture
|
||||
|
||||
Data-driven does **not** mean "put fields in a data asset". It means:
|
||||
|
||||
> Code owns lifecycle, algorithms, replication and invariants. Data selects and
|
||||
> composes implementations at explicit architectural layers.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes for silent data-graph failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
**The characteristic failure of a data graph is that it compiles.** A wrong
|
||||
reference, an unregistered identifier, a missing bundle and a copy-pasted
|
||||
prototype are all valid data. Validation is the only compiler this layer will
|
||||
ever have, and building it is part of the architecture rather than a follow-up
|
||||
task.
|
||||
|
||||
Related skills: `ue-modular-gameplay`, `ue-gas-architecture`,
|
||||
`ue-input-architecture`, `ue-ui-architecture`, `ue-cosmetics-and-teams`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Start with layers, not asset classes
|
||||
|
||||
Before creating an asset, name the level of composition it owns. A robust
|
||||
product usually needs several narrow definitions, not one omniscient
|
||||
configuration:
|
||||
|
||||
```text
|
||||
catalog / playlist
|
||||
├─ where and how to launch (map, session, display metadata)
|
||||
└─ gameplay definition
|
||||
├─ feature plugins to enable
|
||||
├─ reusable action sets
|
||||
├─ mode-specific delta
|
||||
└─ pawn archetype
|
||||
├─ pawn class
|
||||
├─ capability packages
|
||||
├─ input semantics
|
||||
├─ policy matrix
|
||||
└─ camera
|
||||
```
|
||||
|
||||
Items form a separate pipeline:
|
||||
|
||||
```text
|
||||
pickup presentation
|
||||
└─ item definition
|
||||
├─ inline fragments (display, stats, equippable, reticle)
|
||||
└─ equipment definition
|
||||
├─ runtime instance class
|
||||
├─ capability packages
|
||||
└─ actors and components to spawn
|
||||
```
|
||||
|
||||
### The layer test
|
||||
|
||||
For every field ask:
|
||||
|
||||
1. Does it describe **launch and catalog UI**, or **runtime gameplay**?
|
||||
2. Is it common to several modes, or a mode-specific delta?
|
||||
3. Does it define an archetype, a capability, or presentation?
|
||||
4. Who consumes it, and at which lifecycle phase?
|
||||
|
||||
If one asset answers all four, split it.
|
||||
|
||||
---
|
||||
|
||||
## 2. Pick the container by lifecycle
|
||||
|
||||
The storage type is an architectural decision, not a preference.
|
||||
|
||||
| Need | Use | Why |
|
||||
|---|---|---|
|
||||
| Standalone immutable record referenced as an object | data asset | object identity, editable in Details |
|
||||
| Asset-manager discovery, bundles, async loading | primary data asset | stable identifier, bundle rules |
|
||||
| A class default, a class reference, inheritance, or a runtime instance of that class | zero-node Blueprint subclass | class identity plus editable defaults |
|
||||
| Polymorphic optional facet embedded in a definition | instanced object, edit-inline | composition without asset explosion |
|
||||
| Polymorphic load/activate/deactivate operation | instanced action object | behaviour in code, targets in data |
|
||||
| Dense homogeneous tabular records | data table | rows, import and export, bulk editing |
|
||||
|
||||
### A Blueprint with zero nodes is a data container
|
||||
|
||||
This is the most consequential and least visible fact in this skill. An entire
|
||||
definition layer can live in Blueprint classes that contain no graph at all,
|
||||
with values on their class defaults.
|
||||
|
||||
**An inventory that searches only for data-asset instances will not see any of
|
||||
it.** Count both, or your census of the data layer is wrong by however much that
|
||||
layer holds. Recipe: DD-01.
|
||||
|
||||
### When not to use a Blueprint class default
|
||||
|
||||
Do not use one merely because the editor offers it. Prefer a data asset when
|
||||
consumers need an object rather than a class, when no runtime instance of the
|
||||
subclass is created, when inherited defaults add nothing, and when asset-manager
|
||||
identity matters more than a class reference.
|
||||
|
||||
---
|
||||
|
||||
## 3. Pick reference semantics explicitly
|
||||
|
||||
Every reference chooses loading and ownership behaviour.
|
||||
|
||||
| Reference | Use when | Failure mode |
|
||||
|---|---|---|
|
||||
| primary asset identifier | the target must be addressed before it is loaded | type or name not registered; resolves to nothing |
|
||||
| soft object or class pointer | async load, optional content, role-specific content | not loaded when dereferenced; omitted from the build |
|
||||
| hard pointer | the target is required whenever the owner is loaded | dependency graph grows transitively |
|
||||
| class reference | data selects an implementation class or its defaults | construction and lifecycle become class-coupled |
|
||||
| instanced object | the owner exclusively owns a small polymorphic facet | duplication; cannot be shared independently |
|
||||
|
||||
Write the decision into the property metadata and the validation, not only into
|
||||
a design document.
|
||||
|
||||
### Asset bundles belong beside the soft reference
|
||||
|
||||
```cpp
|
||||
UPROPERTY(EditAnywhere, meta=(AssetBundles="Client"))
|
||||
TSoftClassPtr<UUserWidget> WidgetClass;
|
||||
|
||||
UPROPERTY(EditAnywhere, meta=(AssetBundles="Client,Server"))
|
||||
TSoftClassPtr<UGameplayAbility> AbilityType;
|
||||
```
|
||||
|
||||
The definition then declares both **composition** and the **load and cook
|
||||
boundary** in one place. Do not maintain a separate, undocumented table of which
|
||||
assets belong to which role.
|
||||
|
||||
Two failure modes follow directly, and both are silent:
|
||||
|
||||
- a soft reference with **no** bundle annotation is not pulled into the build by
|
||||
its owner, so it resolves to null in a packaged game and works in the editor;
|
||||
- a bare accessor on a soft reference assumes the bundle worked, and returns null
|
||||
without complaint when it did not.
|
||||
|
||||
Recipes: DD-03, DD-04.
|
||||
|
||||
### Cross-plugin dependency gate
|
||||
|
||||
For every reference crossing a feature-plugin boundary: is the target plugin
|
||||
active before this asset resolves? Does the owning bundle pull it into the
|
||||
build? Can the source deactivate while the target is still referenced? Is the
|
||||
dependency one-way, or did you just create a cycle?
|
||||
|
||||
---
|
||||
|
||||
## 4. Code consumer, data selection
|
||||
|
||||
A healthy definition has a narrow code consumer. The consumer owns when loading
|
||||
happens, authority and client rules, replication, initialization order, apply and
|
||||
rollback, validation and error handling.
|
||||
|
||||
Data owns which class or asset or tag is selected, numeric and presentation
|
||||
parameters, variant composition, and role flags where the consumer supports them.
|
||||
|
||||
### The switch test
|
||||
|
||||
If adding a new mode, weapon or pawn requires another branch in a manager, a
|
||||
definition, action or fragment boundary is probably missing.
|
||||
|
||||
### But do not move an invariant into data
|
||||
|
||||
These stay in code, always:
|
||||
|
||||
- "only the authority grants capabilities";
|
||||
- "the archetype must exist before this readiness state";
|
||||
- replication rules;
|
||||
- teardown order;
|
||||
- "one grant revokes only its own handles".
|
||||
|
||||
**The variant is data; the safety rule is code.** A rule that can be edited by
|
||||
opening an asset is a rule that will be edited by someone who does not know it is
|
||||
a rule.
|
||||
|
||||
---
|
||||
|
||||
## 5. Compose reuse; do not hide it in inheritance
|
||||
|
||||
Prefer flat composition:
|
||||
|
||||
```text
|
||||
experience = shared input + standard components + standard HUD + mode delta
|
||||
```
|
||||
|
||||
over an inheritance chain:
|
||||
|
||||
```text
|
||||
base → shooter → team → deathmatch
|
||||
```
|
||||
|
||||
Composition makes dependencies visible, reusable and diffable. Inheritance
|
||||
accumulates defaults silently and creates ordering questions nobody asked.
|
||||
|
||||
If your definition type supports Blueprint subclassing, consider **rejecting it
|
||||
in validation** and pointing the author at composition instead. A reference
|
||||
product does exactly this, with an error message naming the alternative
|
||||
— a validator that teaches rather than only refuses. Recipe: DD-05.
|
||||
|
||||
### The symmetry gate for actions
|
||||
|
||||
Before approving any action that adds something at runtime, fill this table:
|
||||
|
||||
| Resource added | Handle retained | Exact rollback |
|
||||
|---|---|---|
|
||||
| extension handler | request handle | release handle |
|
||||
| component | component request or instance | remove only the owned component |
|
||||
| capability grant | granted handles | revoke only those handles |
|
||||
| input binding | bind handles | remove each binding |
|
||||
| mapping context | player plus context | remove from the same player |
|
||||
| widget or layout | extension or layout handle | unregister or remove |
|
||||
|
||||
**A blank middle column means the action cannot be rolled back.** Do not
|
||||
advertise runtime-deactivatable modularity until the table is complete and
|
||||
tested. See `ue-modular-gameplay` for the failure modes this prevents.
|
||||
|
||||
---
|
||||
|
||||
## 6. Use fragments for optional facets
|
||||
|
||||
Avoid one definition carrying every possible field for every possible item.
|
||||
Use polymorphic inline fragments and let consumers query by class:
|
||||
|
||||
```cpp
|
||||
UCLASS(DefaultToInstanced, EditInlineNew, Abstract)
|
||||
class UItemFragment : public UObject {};
|
||||
|
||||
UPROPERTY(EditDefaultsOnly, Instanced)
|
||||
TArray<TObjectPtr<UItemFragment>> Fragments;
|
||||
```
|
||||
|
||||
The definition stays open to new facets without changing its base class,
|
||||
irrelevant fields do not appear on unrelated items, and designers compose facets
|
||||
in the editor.
|
||||
|
||||
Fragment rules: one responsibility each; state whether duplicates are legal;
|
||||
validate required combinations; never depend on array order unless the order is
|
||||
named and documented; fragments select and configure behaviour, they do not own
|
||||
lifecycle.
|
||||
|
||||
---
|
||||
|
||||
## 7. Put cross-cutting policy in a policy definition
|
||||
|
||||
When many capabilities repeat the same relationships, do not copy the same
|
||||
containers into all of them. Use a matrix selected by the archetype, so the same
|
||||
capabilities can run under different policies.
|
||||
|
||||
Good candidates: capability relationships, presentation parameter bags,
|
||||
sensitivity curves, class-to-widget mappings, quality thresholds.
|
||||
|
||||
Bad candidates: authority decisions, anti-cheat rules, irreversible lifecycle
|
||||
transitions, and the validation logic itself.
|
||||
|
||||
---
|
||||
|
||||
## 8. Validation is part of the architecture
|
||||
|
||||
A data graph creates failure classes that code-only systems do not have. Every
|
||||
definition family needs validation, and it needs to run somewhere other than the
|
||||
editor.
|
||||
|
||||
### Universal checks
|
||||
|
||||
- required references non-null;
|
||||
- arrays with no effective entries rejected;
|
||||
- class derives from the required base or implements the required interface;
|
||||
- tags valid and constrained to the expected category;
|
||||
- identifiers resolve;
|
||||
- soft targets belong to available, cooked plugins;
|
||||
- duplicate semantic keys detected;
|
||||
- client-only assets not bundled for the server, and the reverse;
|
||||
- every action has a complete rollback path;
|
||||
- production definitions reference no prototype or test content.
|
||||
|
||||
### Cross-family checks
|
||||
|
||||
These are the ones that catch real defects, because they cross an ownership
|
||||
boundary that no single reviewer owns:
|
||||
|
||||
```text
|
||||
playlist identifier → resolves to a registered gameplay definition
|
||||
gameplay definition → archetype has a class, or is explicitly empty
|
||||
archetype capabilities → input tags match a reachable input config
|
||||
item equippable fragment → resolves to an equipment definition
|
||||
equipment definition → capability packages valid for its instance type
|
||||
pickup presentation → matches the item definition it claims to present
|
||||
```
|
||||
|
||||
The last row is not theoretical. In the audited reference, two health pickups
|
||||
reference the pistol item definition. It compiles, cooks, and ships. Recipe:
|
||||
DD-07.
|
||||
|
||||
### Editor-only validation is feedback, not a boundary
|
||||
|
||||
If invalid data can reach a packaged build through an automated cook, the
|
||||
validation must run there too. In the audited reference **eleven of twelve
|
||||
validation entry points are compiled out of non-editor builds** — which is
|
||||
normal, correct, and exactly why a commandlet or automation test has to exist
|
||||
beside them. Recipe: DD-09.
|
||||
|
||||
---
|
||||
|
||||
## 9. Auditing an existing project
|
||||
|
||||
### A. Inventory both containers
|
||||
|
||||
1. Query the asset registry for instances deriving from the data-asset base.
|
||||
2. Query the class registry for zero-node classes whose parent is a definition
|
||||
type.
|
||||
3. Include every mount root — the primary content root alone is incomplete in a
|
||||
plugin-based project.
|
||||
4. Exclude world-partition external actor packages from asset counts.
|
||||
|
||||
**Do not read binary asset files from disk.** Use the asset registry and live
|
||||
editor APIs. This is a hard constraint, not a preference: a text search over
|
||||
binary assets produces confident nonsense.
|
||||
|
||||
### B. Build the reference graph
|
||||
|
||||
For each definition record the owner layer, the consumer, the identifier, the
|
||||
reference kinds, the bundle metadata, and the incoming and outgoing plugin
|
||||
boundaries.
|
||||
|
||||
### C. Measure behaviour density
|
||||
|
||||
A definition class should have zero or very few graph nodes. A definition
|
||||
carrying a large graph is secretly both data and manager.
|
||||
|
||||
### D. Trace one complete variant end to end
|
||||
|
||||
Do not stop at class declarations. Pick one real variant and follow it:
|
||||
|
||||
```text
|
||||
playlist → definition → action sets → archetype
|
||||
→ capability and input policy → item → equipment → grant
|
||||
```
|
||||
|
||||
Record the actual configured values. Generic reflection frequently hides
|
||||
inherited class-default fields and reports zero properties on a populated
|
||||
object; read known fields by name from the schema instead of concluding the
|
||||
object is empty. Recipe: DD-02.
|
||||
|
||||
---
|
||||
|
||||
## 10. Review checklist
|
||||
|
||||
- [ ] The definition's architectural layer has a name.
|
||||
- [ ] It has exactly one consumer owner.
|
||||
- [ ] Container type matches lifecycle.
|
||||
- [ ] Reference kinds are intentional and written in metadata.
|
||||
- [ ] Bundle metadata matches client and server use.
|
||||
- [ ] Reuse is composition, not a chain of inherited defaults.
|
||||
- [ ] Actions and fragments have focused responsibilities.
|
||||
- [ ] Apply and rollback symmetry is proven, not assumed.
|
||||
- [ ] Required links and semantic joins are validated.
|
||||
- [ ] Cross-plugin references are acyclic and cooked.
|
||||
- [ ] A real instance was traced end to end.
|
||||
- [ ] Production data contains no prototype or copy-pasted references.
|
||||
|
||||
If more than two boxes are unknown, the architecture is not data-driven yet — it
|
||||
is data-shaped.
|
||||
|
||||
---
|
||||
|
||||
## 11. When to stop
|
||||
|
||||
Introduce a definition layer when at least one is true: several modes reuse the
|
||||
same composition; content ships after launch; variants are authored by people who
|
||||
do not write code; a manager is accumulating per-variant branches.
|
||||
|
||||
Do not introduce one because the reference project has it. Every layer costs an
|
||||
asset type, a validation surface, a loading path and a debugging hop. The
|
||||
measured benefit of this architecture is that a mode becomes a list rather than a
|
||||
subclass — if you have one mode, you are paying for an answer to a question you
|
||||
do not have.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The measured material comes from an inspection of Epic's Lyra Starter Game on
|
||||
Unreal Engine 5.6, performed through a live editor bridge rather than by reading
|
||||
files: 84 native data assets across 21 runtime classes, 23 zero-node definition
|
||||
classes, and 37 inline actions across ten definitions and five action sets. That
|
||||
method matters — see DD-01, which exists because a purely text-based census of
|
||||
this layer is structurally incapable of seeing most of it.
|
||||
|
||||
Source addresses and asset paths stay in the research archive that produced this
|
||||
skill. Each entry carries a stable identifier (`DD-01`, `DD-03`, …) resolving
|
||||
back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Values obtained through the editor describe one project at one moment. Graph
|
||||
assets whose internals no safe reader exposed — environment queries, state trees,
|
||||
blackboards — were listed but not read, and are marked as such rather than
|
||||
summarised. Re-run the recipes against your own project; the counts here show the
|
||||
shape of a contrast, not a target.
|
||||
+401
@@ -0,0 +1,401 @@
|
||||
# Failure modes: data-driven architecture
|
||||
|
||||
Ten ways a data graph is wrong while remaining perfectly valid.
|
||||
|
||||
The shared property: **data does not compile, so nothing rejects it.** A null
|
||||
reference, an unregistered identifier, a missing bundle annotation, a
|
||||
copy-pasted prototype and a field nobody reads are all well-formed assets. They
|
||||
load. They cook. The only thing that can reject them is validation somebody
|
||||
chose to write, and that validation is usually compiled out of the build where
|
||||
it would matter.
|
||||
|
||||
A second property specific to this area, and the reason the first recipe is
|
||||
first: **most of this layer is invisible to text search.** Definitions live in
|
||||
binary assets; the values live on class defaults. A recipe that greps source
|
||||
answers a question about the *schema*, not about the *data*. Several recipes
|
||||
below therefore require the editor, and say so rather than pretending a grep
|
||||
will do.
|
||||
|
||||
---
|
||||
|
||||
## Census
|
||||
|
||||
### DD-01 - Half the definition layer is invisible to the census
|
||||
|
||||
**Mechanism.** A definition layer can be built from two containers at once: data
|
||||
asset instances, and Blueprint classes with no graph nodes whose values live on
|
||||
the class default object. An inventory that queries for one container reports a
|
||||
complete-looking number.
|
||||
|
||||
**Why it is silent.** The number is real, plausible, and internally consistent.
|
||||
Nothing indicates it is partial — an inventory has no way to report the
|
||||
containers it did not think to look in.
|
||||
|
||||
**Why the obvious check misses it.** "Find the data assets" is the correct
|
||||
phrasing of the wrong question. The second container is not a data asset by class,
|
||||
is not named like one, and appears in the content browser as a Blueprint. In the
|
||||
audited reference this was not a corner case: **every** gameplay definition, item
|
||||
definition and equipment definition lived in the second container.
|
||||
|
||||
**Symptom.** Audits, rename sweeps, memory budgets and dependency graphs that are
|
||||
each individually correct and collectively wrong. Decisions get made on a census
|
||||
that missed the most important layer.
|
||||
|
||||
**Detect.** Count both containers, from the editor — this cannot be done from
|
||||
source:
|
||||
|
||||
```python
|
||||
# editor-side, via the asset registry
|
||||
ar = unreal.AssetRegistryHelpers.get_asset_registry()
|
||||
# 1. instances deriving from the data asset base
|
||||
# 2. blueprint classes whose native parent is a definition type,
|
||||
# then check node count on the generated class
|
||||
```
|
||||
|
||||
From source you can only establish the *schema* — which types exist and which
|
||||
are meant to be subclassed:
|
||||
|
||||
```bash
|
||||
rg -n "class \w+ : public UPrimaryDataAsset|class \w+ : public UDataAsset" \
|
||||
--glob "*.h" .
|
||||
rg -n "UCLASS\([^)]*Blueprintable" --glob "*.h" . | rg -i "definition|data"
|
||||
```
|
||||
|
||||
**Guardrail.** Define "the data layer" as a set of *types*, not a set of asset
|
||||
classes, and enumerate instances per type. Record the count per container so a
|
||||
future reader can see which containers were searched.
|
||||
|
||||
---
|
||||
|
||||
### DD-02 - Reflection reports an empty object that is full
|
||||
|
||||
**Mechanism.** A generic property dump over a class default object returns
|
||||
nothing, or almost nothing, for an object whose inherited fields are populated.
|
||||
|
||||
**Why it is silent.** An empty result is a legitimate answer — plenty of objects
|
||||
genuinely have no properties set. The tool did not fail; it succeeded and
|
||||
returned nothing.
|
||||
|
||||
**Why the obvious check misses it.** The check *is* the dump. Its output is
|
||||
authoritative-looking structured data, and the natural conclusion — "this
|
||||
definition is empty, so the values must live elsewhere" — sends the investigation
|
||||
in a direction with no bottom.
|
||||
|
||||
**Symptom.** Hours spent looking for where the configuration "really" lives,
|
||||
followed by a confident and wrong claim that a definition carries no data.
|
||||
|
||||
**Detect.** Never conclude from a generic dump. Read known fields by name,
|
||||
driven by the schema you extracted from source:
|
||||
|
||||
```bash
|
||||
# 1. get the field names from the C++ declaration
|
||||
rg -n -A20 "class \w*ExperienceDefinition" --glob "*.h" . | rg "UPROPERTY" -A1
|
||||
```
|
||||
|
||||
```python
|
||||
# 2. editor-side, read each named field explicitly rather than enumerating
|
||||
for field in ("GameFeaturesToEnable", "DefaultPawnData", "Actions", "ActionSets"):
|
||||
value = cdo.get_editor_property(field)
|
||||
```
|
||||
|
||||
A field that a generic dump omitted and a named read returns is the finding, and
|
||||
it invalidates every conclusion drawn from the dump.
|
||||
|
||||
**Guardrail.** Treat generic reflection as a discovery aid and named reads as
|
||||
evidence. When reporting "this object has no data", state which method produced
|
||||
that answer.
|
||||
|
||||
---
|
||||
|
||||
## Loading and boundaries
|
||||
|
||||
### DD-03 - Soft reference with no bundle annotation
|
||||
|
||||
**Mechanism.** A definition holds a soft reference to content needed only by one
|
||||
role. Without a bundle annotation, nothing tells the asset manager to pull that
|
||||
content in when the owner loads.
|
||||
|
||||
**Why it is silent.** In the editor everything is loaded anyway — commonly both
|
||||
role bundles at once — so the reference resolves and the feature works. The
|
||||
failure requires a build where only what was asked for is present.
|
||||
|
||||
**Why the obvious check misses it.** The property is correct C++ and the
|
||||
reference is correct data. The annotation is optional metadata whose absence
|
||||
looks like every other property that legitimately does not need one. Reviewing
|
||||
the definition shows nothing missing.
|
||||
|
||||
**Symptom.** A feature that works in the editor and silently does nothing in a
|
||||
packaged build, because a class pointer resolved to null and the code path that
|
||||
uses it treats null as "nothing to do".
|
||||
|
||||
**Detect.** List soft references, list annotations, and compare:
|
||||
|
||||
```bash
|
||||
rg -n "TSoftObjectPtr<|TSoftClassPtr<" --glob "*.h" . | wc -l
|
||||
rg -no 'AssetBundles="[^"]*"' --glob "*.h" . | sed 's/.*AssetBundles=//' \
|
||||
| sort | uniq -c
|
||||
```
|
||||
|
||||
The gap between the two counts is your exposure. In the audited reference the
|
||||
whole project carried **ten** bundle annotations — two client-only and eight for
|
||||
both roles — against a much larger population of soft references. Most of those
|
||||
are loaded by other means; the point of the recipe is that it produces the short
|
||||
list worth checking rather than a verdict.
|
||||
|
||||
**Guardrail.** Annotate at the declaration, beside the reference, and validate
|
||||
that every soft reference in a definition family either carries an annotation or
|
||||
is documented as loaded by its consumer.
|
||||
|
||||
---
|
||||
|
||||
### DD-04 - Bare accessor on a soft reference
|
||||
|
||||
**Mechanism.** The consumer resolves a soft reference with an accessor that
|
||||
returns whatever is already loaded, and uses the result without checking.
|
||||
|
||||
**Why it is silent.** Inside the pipeline the annotation was designed for, the
|
||||
asset is always loaded. The null path is unreachable in every configuration
|
||||
anyone tests.
|
||||
|
||||
**Why the obvious check misses it.** The bundle annotation is visible on the
|
||||
property, so the loading question appears answered. The accessor call is one
|
||||
word and reads as a dereference, not as a decision.
|
||||
|
||||
**Symptom.** One feature's content silently missing, in one configuration, with
|
||||
no log line naming the asset.
|
||||
|
||||
**Detect.** Find bare accessors on soft references in consumer code:
|
||||
|
||||
```bash
|
||||
rg -n "\.Get\(\)" --glob "*.cpp" . | rg -i "class|widget|ability|component"
|
||||
rg -n -A2 "\.Get\(\)" --glob "*.cpp" . | rg -c "if \(|ensure|check"
|
||||
```
|
||||
|
||||
The finding is a resolve with no null branch. Compare against the same file's
|
||||
other resolves — in the audited reference one loop tested for null and the
|
||||
adjacent loop did not, which is the clearest possible evidence that the check was
|
||||
intended.
|
||||
|
||||
**Guardrail.** Either load explicitly with a failure path, or check and log at
|
||||
warning level with the asset name. "The bundle guarantees it" is a claim about
|
||||
one pipeline, not about the code.
|
||||
|
||||
---
|
||||
|
||||
## Composition
|
||||
|
||||
### DD-05 - Inheritance used where composition was intended
|
||||
|
||||
**Mechanism.** A definition type is Blueprint-subclassable, so authors subclass
|
||||
one definition from another to reuse its values, producing a chain of inherited
|
||||
defaults.
|
||||
|
||||
**Why it is silent.** It works. Inherited defaults are a supported feature, and
|
||||
the child behaves as the sum of the chain.
|
||||
|
||||
**Why the obvious check misses it.** Each subclass is individually reasonable and
|
||||
locally minimal — it changes two fields. The cost is structural: values now come
|
||||
from several assets that must be opened in sequence, and a change to a parent
|
||||
silently alters every descendant.
|
||||
|
||||
**Symptom.** A definition whose effective configuration cannot be read from the
|
||||
definition. Diffs that show nothing while behaviour changes.
|
||||
|
||||
**Detect.** Find whether the type permits subclassing, and whether validation
|
||||
objects:
|
||||
|
||||
```bash
|
||||
rg -n "UCLASS\([^)]*Blueprintable" --glob "*.h" . | rg -i "definition"
|
||||
rg -n -B2 -A6 "IsDataValid" --glob "*.cpp" . | rg -i "parent|subclass|inherit"
|
||||
```
|
||||
|
||||
The audited reference does this well and is worth copying exactly: its
|
||||
definition validation **rejects Blueprint subclasses of Blueprint definitions**
|
||||
and the error text names the alternative — use composition via action sets. A
|
||||
validator that says what to do instead is worth several that only refuse.
|
||||
|
||||
**Guardrail.** Decide per definition type whether subclassing is legal. If it is
|
||||
not, reject it in validation with a message naming the composition mechanism.
|
||||
|
||||
---
|
||||
|
||||
### DD-06 - An action set that is really a parent class
|
||||
|
||||
**Mechanism.** Reusable slices are introduced, and then one slice starts
|
||||
depending on being applied after another, or on values from a third.
|
||||
|
||||
**Why it is silent.** Order is stable in practice, because it comes from a list
|
||||
that rarely changes. The coupling produces correct behaviour for as long as
|
||||
nobody reorders anything.
|
||||
|
||||
**Why the obvious check misses it.** Composition was adopted specifically to
|
||||
avoid ordering questions, so nobody looks for one. The dependency is not
|
||||
expressed anywhere — it exists only in the fact that the current order works.
|
||||
|
||||
**Symptom.** Reordering a list, or reusing one slice without its neighbour,
|
||||
breaks a feature that has no visible relationship to the change.
|
||||
|
||||
**Detect.** Look for actions that read state another action writes:
|
||||
|
||||
```bash
|
||||
rg -n "class UGameFeatureAction_\w+" --glob "*.h" . -A30 \
|
||||
| rg "FindComponentByClass|GetSubsystem|Find\w+ForActor"
|
||||
```
|
||||
|
||||
Every hit is an action that consumes something it did not create — a candidate
|
||||
ordering dependency. For each, ask what happens if it runs first.
|
||||
|
||||
**Guardrail.** An action set is a set, not a sequence. If order matters, make it
|
||||
explicit — a phase, a prerequisite declaration, or a single combined action — and
|
||||
never leave it implied by list position.
|
||||
|
||||
---
|
||||
|
||||
## Content correctness
|
||||
|
||||
### DD-07 - Production data references the wrong definition
|
||||
|
||||
**Mechanism.** A presentation asset is duplicated to make a new one, and the
|
||||
reference inside it is not updated. The copy points at the original's definition.
|
||||
|
||||
**Why it is silent.** Both assets are valid, both load, and the reference is a
|
||||
real object of the right type. Nothing in the graph is broken; it simply says
|
||||
something untrue.
|
||||
|
||||
**Why the obvious check misses it.** Validation checks that references are
|
||||
non-null and correctly typed, which this passes. Confirming it means comparing
|
||||
the *meaning* of two assets — that a health pickup should reference a health
|
||||
item — and no type system encodes that.
|
||||
|
||||
**Symptom.** Picking up one thing and receiving another. In the audited
|
||||
reference, two health pickups reference the pistol item definition. It ships.
|
||||
|
||||
**Detect.** This one is a cross-family check, and it needs the editor:
|
||||
|
||||
```python
|
||||
# editor-side: for each pickup definition, compare its presentation identity
|
||||
# against the item definition it references
|
||||
for pickup in pickup_definitions:
|
||||
item = pickup.get_editor_property("InventoryItemDefinition")
|
||||
# heuristic that actually works: does the pickup asset's name share a token
|
||||
# with the item definition's name?
|
||||
```
|
||||
|
||||
The heuristic is crude and that is the point — a crude cross-family check finds
|
||||
copy-paste errors that no type check can. Add it to validation and let it produce
|
||||
warnings a human triages.
|
||||
|
||||
**Guardrail.** Every cross-family reference gets a semantic check, even a weak
|
||||
one. Also: keep prototype content out of production roots so that a naming-based
|
||||
check has a chance.
|
||||
|
||||
---
|
||||
|
||||
### DD-08 - Fields that exist, are authored, and are never read
|
||||
|
||||
**Mechanism.** A definition exposes extension fields for future use. They are
|
||||
never wired to a consumer.
|
||||
|
||||
**Why it is silent.** An unread field costs nothing and looks like configuration.
|
||||
Its default is usually the correct behaviour, so leaving it empty is right.
|
||||
|
||||
**Why the obvious check misses it.** The field is declared, is editable, and
|
||||
appears next to fields that do work. "Is this used?" answers yes on the
|
||||
declaration. Only searching for a *reader* answers correctly — and unlike the
|
||||
same failure in C++, the reader may be in a Blueprint graph that no text search
|
||||
can inspect, so a negative result is weaker here.
|
||||
|
||||
**Symptom.** A designer configures a field and observes no effect. Because the
|
||||
consumer may be in a binary asset, the investigation cannot conclude from source
|
||||
alone, and often stops with "it must do something".
|
||||
|
||||
**Detect.** From source, find declaration-without-reader candidates; then confirm
|
||||
in the editor:
|
||||
|
||||
```bash
|
||||
F='ExtraArgs'
|
||||
rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
A single hit — the declaration — makes the field a candidate. **Do not conclude
|
||||
it is dead**: check the editor's reference viewer for Blueprint consumers before
|
||||
deleting. In the audited reference several such fields were empty in every
|
||||
instance, which is evidence about usage but not about wiring.
|
||||
|
||||
**Guardrail.** Do not ship an editable field with no consumer. If it exists for
|
||||
forward compatibility, mark it as such where the designer sees it, not only in a
|
||||
commit message.
|
||||
|
||||
---
|
||||
|
||||
## Validation
|
||||
|
||||
### DD-09 - Validation compiled out of the build that needs it
|
||||
|
||||
**Mechanism.** Data validation is written as editor-only, so it runs when an
|
||||
author opens the asset and never during an automated cook.
|
||||
|
||||
**Why it is silent.** It is the correct default and it works — authors do get
|
||||
feedback. Nothing degrades; the check simply is not present in the pipeline where
|
||||
bad data would otherwise be caught.
|
||||
|
||||
**Why the obvious check misses it.** Searching for validation finds it, in
|
||||
quantity, well written. The question that matters is narrower: *does any of it
|
||||
run outside the editor?* — and answering that means reading preprocessor guards
|
||||
rather than validation logic.
|
||||
|
||||
**Symptom.** Invalid data reaches a packaged build. The team's belief that "we
|
||||
validate our data" is true and irrelevant.
|
||||
|
||||
**Detect.** Count validation entry points, then count the ones that survive a
|
||||
non-editor build, then look for a runner:
|
||||
|
||||
```bash
|
||||
rg -n "IsDataValid" --glob "*.h" . | wc -l
|
||||
rg -n -B3 "IsDataValid" --glob "*.h" . | rg -c "WITH_EDITOR"
|
||||
rg -n "Commandlet|AutomationTest|ValidateAssets" --glob "*.h" --glob "*.cs" .
|
||||
```
|
||||
|
||||
In the audited reference this gave twelve entry points, **eleven** of them
|
||||
editor-only, plus a content-validation commandlet and a validator base class —
|
||||
so the mechanism to run them in automation exists and the question becomes
|
||||
whether the pipeline invokes it.
|
||||
|
||||
**Guardrail.** Editor validation is feedback; a commandlet or automation test in
|
||||
the build pipeline is the boundary. Ensure the runner **fails the build** on
|
||||
error — a validator whose result is only printed is a log line, not a gate.
|
||||
|
||||
---
|
||||
|
||||
### DD-10 - The validator reports and the pipeline does not care
|
||||
|
||||
**Mechanism.** Validation runs in automation and prints errors. The process exit
|
||||
status does not reflect them, so the pipeline continues.
|
||||
|
||||
**Why it is silent.** Everything works: the check runs, the errors are found, the
|
||||
log contains them. The only missing link is the one nobody looks at until a bad
|
||||
asset ships.
|
||||
|
||||
**Why the obvious check misses it.** Both halves are present and correct — there
|
||||
is validation, and there is a pipeline step running it. The defect is in the
|
||||
contract between them, which is a single integer nobody reads.
|
||||
|
||||
**Symptom.** A green build over a log full of validation errors. Confidence
|
||||
inversely proportional to actual coverage.
|
||||
|
||||
**Detect.** Check what the runner returns and what the pipeline does with it:
|
||||
|
||||
```bash
|
||||
rg -n -A20 "class UContentValidationCommandlet|int32 Main\(" --glob "*.h" --glob "*.cpp" . \
|
||||
| rg "return |ExitCode|GIsCriticalError"
|
||||
rg -n "Commandlet" .github/ *.yml *.yaml 2>/dev/null
|
||||
```
|
||||
|
||||
A commandlet whose `Main` returns zero unconditionally is the finding. Confirm by
|
||||
deliberately breaking one asset and running the pipeline step: if it stays green,
|
||||
the gate does not exist.
|
||||
|
||||
**Guardrail.** Prove the gate by breaking something on purpose. **A check that
|
||||
has never failed is indistinguishable from a check that cannot fail** — this is
|
||||
the same principle as the poisoned-fixture rule for any linter, applied to your
|
||||
content pipeline.
|
||||
@@ -0,0 +1,268 @@
|
||||
# Patterns: a measured data-driven layer
|
||||
|
||||
A worked example of the layering rules from `SKILL.md`, measured in one reference
|
||||
product. Unlike most files in this bundle it was produced **through a live editor
|
||||
session**, not by reading source — because the thing being measured is data, and
|
||||
data lives in binary assets.
|
||||
|
||||
That method is the first lesson and it is not a technicality. Everything below
|
||||
would have been invisible, or wrong, from a text search.
|
||||
|
||||
Markers: **[measured]** — read from live objects through the editor;
|
||||
**[derived]** — conclusion from measured facts; **[open]** — not readable with
|
||||
the tools available.
|
||||
|
||||
---
|
||||
|
||||
## 1. The census, and why the obvious census is wrong
|
||||
|
||||
| Container | Count **[measured]** | Visible to a data-asset query |
|
||||
|---|---:|---|
|
||||
| Native data-asset instances, 21 runtime classes | 84 | yes |
|
||||
| Zero-node Blueprint classes used as definitions | 23 | **no** |
|
||||
| **Total inspected definition objects** | **107** | |
|
||||
|
||||
The 23 break down as ten gameplay definitions, seven item definitions and six
|
||||
equipment definitions **[measured]**. In other words: **every gameplay,
|
||||
item and equipment definition in the project is in the container that a
|
||||
data-asset query does not return.**
|
||||
|
||||
**[derived]** A census that asks "which data assets exist?" returns 84, is
|
||||
internally consistent, and misses the layer that actually composes the game. This
|
||||
is the single most portable warning in this file. Recipe: DD-01.
|
||||
|
||||
The same structure produced a second trap. A generic property dump over one of
|
||||
those class defaults reported almost nothing, because inherited fields were not
|
||||
enumerated **[measured]**. Reading the *named* fields from the schema returned
|
||||
full configuration. **[derived]** An empty reflection result is evidence about the
|
||||
reflection call, not about the object. Recipe: DD-02.
|
||||
|
||||
---
|
||||
|
||||
## 2. The layer split, as actually built
|
||||
|
||||
```text
|
||||
playlist (8 instances) map + definition + session and menu metadata
|
||||
└─ gameplay definition (10) plugins + archetype + action sets + delta
|
||||
├─ action set (5) reusable slice
|
||||
└─ archetype (6) pawn class, capabilities, input, policy, camera
|
||||
├─ capability package (12)
|
||||
├─ input config (5)
|
||||
└─ policy matrix (1, nine rules)
|
||||
```
|
||||
|
||||
Two properties of this split are worth copying.
|
||||
|
||||
**The playlist knows about maps and menus; the definition does not.** One mode
|
||||
can run on several maps, one map can host several modes, and menu metadata never
|
||||
enters the runtime composition **[measured]** — all eight playlists were verified
|
||||
to carry map, definition, visibility, default flag and player count, and nothing
|
||||
gameplay-facing.
|
||||
|
||||
**Reuse is a list, not a parent.** Two shooter modes differ only by a
|
||||
mode-specific delta — a different capability package, different scoring and music
|
||||
components, different score widgets — while sharing input, components and HUD as
|
||||
action sets **[measured]**. Adding a mode is editing a list.
|
||||
|
||||
That is enforced rather than merely encouraged: validation **rejects a Blueprint
|
||||
subclass of a Blueprint definition** and its error text names composition as the
|
||||
alternative **[measured]**. A validator that says what to do instead is worth
|
||||
several that only refuse. Recipe: DD-05.
|
||||
|
||||
---
|
||||
|
||||
## 3. Composition is dependency injection, spelled as data
|
||||
|
||||
37 inline actions across ten definitions and five action sets **[measured]**. The
|
||||
shape of the most common one:
|
||||
|
||||
```text
|
||||
(target actor class, component class, apply on client?, apply on server?)
|
||||
```
|
||||
|
||||
One mode adds scoring on both roles, music on the client only, and bot spawning,
|
||||
team setup and spawn rules on the server only — **without modifying the game state
|
||||
class** **[measured]**. A second mode swaps three of those and keeps the rest.
|
||||
|
||||
**[derived]** This is the payoff of the whole architecture, and it is worth
|
||||
stating plainly: the base classes contain no branch per mode, because the modes
|
||||
are not known to them. The switch test from the skill is passed by construction.
|
||||
|
||||
The same mechanism carries UI: one action pushes a layout into a layer, another
|
||||
registers eleven widgets into named slots **[measured]**. Mode-specific UI is two
|
||||
more rows in a list.
|
||||
|
||||
---
|
||||
|
||||
## 4. Where the role boundary is declared
|
||||
|
||||
Soft references in the action data carry bundle annotations, so the definition
|
||||
declares composition **and** the cook boundary in one place **[measured]**.
|
||||
|
||||
The measured population is small and the number is instructive: **ten annotations
|
||||
in the entire project — two for the client, eight for both roles** **[measured]**.
|
||||
|
||||
**[derived]** That is not a criticism; most content is loaded by other means. It
|
||||
is a statement about how narrow the annotated surface is, which matters because
|
||||
the consumers of those references use a bare accessor that returns null when the
|
||||
bundle did not deliver, and one of the two loops checks the result while its
|
||||
neighbour does not **[measured]**. The annotation and the null check are the two
|
||||
halves of one contract, and only one half is consistently present. Recipes:
|
||||
DD-03, DD-04.
|
||||
|
||||
---
|
||||
|
||||
## 5. The capability layer, and why input is not one-to-one with it
|
||||
|
||||
12 capability packages **[measured]**, ranging from an eleven-capability hero
|
||||
package down to packages that grant **only effects** and no capabilities at all.
|
||||
|
||||
The instructive detail: **not every granted capability has an input tag**
|
||||
**[measured]**. Death, spawn effects, auto-reload and auto-respawn are activated
|
||||
by events and policies, not by a button.
|
||||
|
||||
**[derived]** This is the concrete reason to keep the semantic input layer
|
||||
separate from the capability layer rather than folding one into the other. A
|
||||
design that assumes "capability implies binding" has no place to put the four
|
||||
above, and will grow a special case for each.
|
||||
|
||||
Six archetypes reuse the same capability packages across different pawn classes
|
||||
and input configurations **[measured]** — one archetype reuses the shooter
|
||||
capability package while changing pawn class and input config, which demonstrates
|
||||
that implementation and capability are genuinely independent axes.
|
||||
|
||||
One archetype has a **null** pawn class **[measured]**. **[derived]** It is a
|
||||
sentinel used as a configuration fallback, not an archetype — worth recognising
|
||||
because a validator that requires a non-null class must special-case it, and a
|
||||
validator that does not will be silently satisfied by every broken archetype.
|
||||
|
||||
---
|
||||
|
||||
## 6. Fragments: composition inside a definition
|
||||
|
||||
Item definitions carry a display name and a list of instanced polymorphic
|
||||
fragments; consumers query by fragment class **[measured]**. The measured pipeline:
|
||||
|
||||
```text
|
||||
pickup presentation (data asset)
|
||||
└─ item definition (zero-node class)
|
||||
├─ display, quickbar, pickup, reticle fragments
|
||||
├─ ammo statistics
|
||||
└─ equippable fragment → equipment definition (zero-node class)
|
||||
├─ runtime instance class
|
||||
├─ capability package
|
||||
└─ actor to spawn, at a named socket
|
||||
```
|
||||
|
||||
Three weapons differ only in these values **[measured]** — different instance
|
||||
class, capability package, spawned visual, reticle fragment and ammo numbers —
|
||||
with no change to the inventory or equipment managers.
|
||||
|
||||
**[derived]** The fragment list is what keeps the item definition from growing a
|
||||
boolean and five fields for every item type that will ever exist. It is the same
|
||||
move as the action list one layer up, applied inside a single asset.
|
||||
|
||||
---
|
||||
|
||||
## 7. The two content defects, and what they teach
|
||||
|
||||
**A health pickup that references the pistol item definition.** Two of them
|
||||
**[measured]**. Both assets are valid, correctly typed and non-null; the graph is
|
||||
intact and says something untrue.
|
||||
|
||||
**[derived]** No type check can catch this, and no validation in the project does.
|
||||
The only mechanism that would is a weak semantic check — does a pickup's identity
|
||||
share anything with the item it presents? — which feels too crude to write until
|
||||
you see what it catches. Recipe: DD-07.
|
||||
|
||||
**Prototype content in the production tree.** One pickup carries a display name
|
||||
and visual assets belonging to a different weapon entirely **[measured]**.
|
||||
|
||||
**[derived]** Prototype content in a production root is what makes the crude check
|
||||
above hard to write, because naming stops being a signal. The two defects
|
||||
compound: keep prototypes out of production roots and a naming heuristic becomes
|
||||
viable.
|
||||
|
||||
Both are recorded here for the reason the whole bundle exists: **these shipped in
|
||||
a reference project that many teams copy from.** Copying the architecture is
|
||||
correct; copying the content is not.
|
||||
|
||||
---
|
||||
|
||||
## 8. Fields that exist and are never set
|
||||
|
||||
All eight playlists carry extra launch arguments and a replay flag. In every one
|
||||
of the eight, the arguments are empty and the flag is false **[measured]**.
|
||||
|
||||
**[open]** Whether anything reads them cannot be settled from source, because a
|
||||
consumer could be in a Blueprint graph. Recorded as open rather than concluded —
|
||||
which is the honest form of this answer and the same discipline as any empty
|
||||
search result.
|
||||
|
||||
**[derived]** Regardless of wiring, a field that is empty in every instance is an
|
||||
extension surface rather than a feature, and a reader who assumes otherwise will
|
||||
look for behaviour that is not there. Recipe: DD-08.
|
||||
|
||||
---
|
||||
|
||||
## 9. Validation: present, well built, and mostly not in the pipeline
|
||||
|
||||
| Fact **[measured]** | Reading |
|
||||
|---|---|
|
||||
| 12 validation entry points across the definition families | The discipline exists and is applied broadly |
|
||||
| **11 of them compiled out of non-editor builds** | Feedback for authors, not a gate on the build |
|
||||
| A content-validation commandlet and a validator base class exist | The mechanism to run them in automation is present |
|
||||
|
||||
**[derived]** This is the correct default arrangement, and it means the whole
|
||||
question reduces to one that source cannot answer: does the build pipeline invoke
|
||||
the commandlet, and does it fail on error? A validator whose result is printed
|
||||
rather than returned is a log line.
|
||||
|
||||
Worth being precise about, because an earlier form of this material overstated
|
||||
it: **structural validation exists in this project.** What does not exist is
|
||||
semantic cross-family validation — nothing checks that a pickup presents the item
|
||||
it claims to. The correct criticism is narrow, and the broad version ("the data
|
||||
graph has no compiler") is wrong. Recipes: DD-09, DD-10.
|
||||
|
||||
---
|
||||
|
||||
## 10. What to copy, and in what order
|
||||
|
||||
1. **Separate the launch catalog from the runtime definition.** One asset type,
|
||||
large payoff, no dependencies on anything else here.
|
||||
2. **Reusable slices as flat lists, and reject definition inheritance in
|
||||
validation** with a message naming the alternative.
|
||||
3. **Fragments for optional facets**, so item definitions do not grow a field per
|
||||
item type.
|
||||
4. **Bundle annotations beside soft references**, plus a null check at every
|
||||
consumer — the two halves of one contract.
|
||||
5. **A cross-family semantic check**, however crude, run in automation with a
|
||||
non-zero exit status.
|
||||
|
||||
Items 1–3 are architecture and transfer directly. Items 4–5 are the parts the
|
||||
audited reference left incomplete, and are therefore the parts most likely to be
|
||||
inherited by anyone copying it.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6 through a live
|
||||
editor bridge on a single day: 84 native data assets across 21 runtime classes,
|
||||
23 zero-node definition classes, 37 inline actions, all eight playlists, six
|
||||
archetypes, twelve capability packages and five input configurations, read as
|
||||
live objects and class defaults. No binary asset was parsed from disk; no
|
||||
gameplay session was run.
|
||||
|
||||
Asset paths and configured values stay in the research archive that produced this
|
||||
skill. Each `DD-` identifier resolves back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Graph assets — environment queries, state trees, blackboards — were enumerated
|
||||
but their internals were **not** read, because no safe reader was available.
|
||||
They are listed as present and nothing is claimed about their contents.
|
||||
|
||||
Everything here describes one project at one moment through one tool. Counts are
|
||||
quoted to show the shape of a contrast, not as targets. Re-run the census against
|
||||
your own project, and count both containers.
|
||||
@@ -0,0 +1,242 @@
|
||||
---
|
||||
name: ue-evidence-discipline
|
||||
description: >-
|
||||
How to produce and check claims about an Unreal Engine codebase so they can be
|
||||
trusted later: separating measured from derived from open, why an empty search
|
||||
result proves nothing, verifying a headline claim yourself before repeating
|
||||
it, treating a delegated report as raw material, and writing detection recipes
|
||||
that have been run. Use when auditing a codebase, reviewing someone else's
|
||||
audit, adopting a reference project, or writing any document that will be
|
||||
quoted back at you.
|
||||
---
|
||||
|
||||
# Evidence discipline for codebase work
|
||||
|
||||
The invariant this whole bundle rests on:
|
||||
|
||||
> A claim without a source you can point at is not a result. A skill that is
|
||||
> confident and wrong is worse than a missing one.
|
||||
|
||||
Every other skill here is normative — it tells you how to design something. This
|
||||
one is methodological: it tells you how to know whether what you just wrote is
|
||||
true. Read it once before using the others, and again before writing an audit
|
||||
someone else will act on.
|
||||
|
||||
Ten ways this discipline fails in practice, each with a detection recipe:
|
||||
[failure modes](references/failure-modes.md). They are all mistakes made during
|
||||
the work that produced this bundle, which is the only reason they are specific
|
||||
enough to be useful.
|
||||
|
||||
---
|
||||
|
||||
## 1. Three kinds of statement, never mixed in one sentence
|
||||
|
||||
| Marker | Meaning | Test |
|
||||
|---|---|---|
|
||||
| **measured** | read in the source, config or a live object | someone else can open the same place and see it |
|
||||
| **derived** | a conclusion from measured facts plus engine semantics | the facts are cited and the inference is stated |
|
||||
| **open** | not answerable with the tools used | says what would answer it |
|
||||
|
||||
**Do not mix two of these in one sentence.** The overwhelming majority of false
|
||||
conclusions in codebase work are born exactly there: a measured fact and an
|
||||
inference joined by a comma, where the reader inherits the confidence of the
|
||||
first half for the second.
|
||||
|
||||
Concretely, this is wrong:
|
||||
|
||||
> The pool is only configured for one tier, so the other tiers are broken.
|
||||
|
||||
and this is right:
|
||||
|
||||
> The pool is configured for one tier **[measured]**. The other tiers therefore
|
||||
> inherit engine values **[derived]** — whether that is wrong for this product is
|
||||
> **[open]** until someone measures the target hardware.
|
||||
|
||||
The second version is longer and it is the only one you can act on.
|
||||
|
||||
### Corollary: an absence has a marker too
|
||||
|
||||
"There is no lag compensation in this project" is a measured claim **only** if
|
||||
you say which searches you ran. Otherwise it is derived from your search
|
||||
strategy, which brings us to the next rule.
|
||||
|
||||
---
|
||||
|
||||
## 2. An empty search result proves the pattern, not the codebase
|
||||
|
||||
This is the single most expensive mistake available in this work, because it
|
||||
produces confident negative claims that survive review.
|
||||
|
||||
A search returns nothing when the code is absent — and also when the pattern was
|
||||
wrong, when the path was wrong, when the file type was excluded, when ignore
|
||||
rules filtered the directory, or when the identifier lives in a binary asset.
|
||||
The tool reports the same empty result for all six.
|
||||
|
||||
**Before writing "this does not exist":**
|
||||
|
||||
1. widen the pattern and confirm it can match something you know is there;
|
||||
2. check the path and include plugins, not only the primary source directory;
|
||||
3. check file-type filters in both directions — too narrow hides code, too wide
|
||||
counts compiled debug symbols as source;
|
||||
4. for anything that could live in a binary asset, state that the search cannot
|
||||
see it.
|
||||
|
||||
Two real corrections from the audit behind this bundle, both caught by re-running:
|
||||
|
||||
- a search for one declaration form found 99 entries; the correct pattern found
|
||||
**212 across five files**, because plugin configuration uses a syntax that
|
||||
differs by one character;
|
||||
- a search for a suspected dead field returned four hits; three were compiled
|
||||
debug symbol files and the source truth was **one**.
|
||||
|
||||
Both would have shipped as facts.
|
||||
|
||||
---
|
||||
|
||||
## 3. Verify a headline claim yourself before repeating it
|
||||
|
||||
Any statement that will lead a section, appear in a summary, or be quoted gets
|
||||
re-read at its source **by the person publishing it** — not accepted from a
|
||||
teammate, a tool, or an earlier version of the same document.
|
||||
|
||||
This is not distrust. It is that a claim changes shape as it travels: a specific
|
||||
finding about one function becomes a general statement about a subsystem in one
|
||||
retelling, and nobody notices because each step was small.
|
||||
|
||||
Three corrections that came from exactly this re-reading, during the work behind
|
||||
this bundle:
|
||||
|
||||
- "a field that is never read" was actually **written once in a constructor and
|
||||
never read** — the precise version names a real defect class; the loose version
|
||||
hides it;
|
||||
- "the data graph has no compiler" was wrong: structural validation existed, in
|
||||
quantity. The true statement was narrower — **semantic cross-family validation
|
||||
is absent** — and the narrow version is the one that tells you what to build;
|
||||
- a count of definitions "spread over 40+ files" was **31 files**, and the
|
||||
original number had never been measured.
|
||||
|
||||
Note that in all three the correction made the claim *more* useful, not less. A
|
||||
weaker true statement beats a stronger false one, and usually points at the fix.
|
||||
|
||||
---
|
||||
|
||||
## 4. A delegated report is raw material
|
||||
|
||||
Output from a subagent, a script, or a search tool is an input to your judgement,
|
||||
not a result. Re-verify every headline claim by reading the source directly
|
||||
before it reaches a document or a conclusion.
|
||||
|
||||
The same rule with a sharper edge: **a report that agrees with what you expected
|
||||
deserves more scrutiny, not less.** During the work behind this bundle a delegated
|
||||
search reported three instances of a defect; direct measurement found nineteen.
|
||||
The report was not wrong in kind, only in scale, which is the hardest error to
|
||||
notice because nothing about it looks incorrect.
|
||||
|
||||
---
|
||||
|
||||
## 5. A recipe you have not run is a hypothesis
|
||||
|
||||
Every detection recipe in this bundle was executed against a real codebase before
|
||||
being written down. That is not thoroughness for its own sake — it is the only
|
||||
way to find the recipes that are confidently wrong.
|
||||
|
||||
Two examples from this bundle, both caught by running them:
|
||||
|
||||
- a search for "is this field ever assigned" matched the **declaration**
|
||||
`double LastFireTime = 0.0;`, so it reported a writer where none exists — a
|
||||
false negative, in the exact case the recipe was written to catch;
|
||||
- a search for "is this ordering value ever compared" matched the arrow operator
|
||||
in `Entry->Priority`, so two assignments were counted as comparisons.
|
||||
|
||||
Both recipes looked correct. Both would have told the reader "no problem here" at
|
||||
precisely the moment there was one.
|
||||
|
||||
**A detector that fails silently in the case it exists for is worse than no
|
||||
detector**, because it converts an open question into a wrong answer. Run the
|
||||
recipe against a codebase where you already know the answer, both when the answer
|
||||
is yes and when it is no.
|
||||
|
||||
---
|
||||
|
||||
## 6. A check that has never failed may not be a check
|
||||
|
||||
Write the failing case for every gate you add, and confirm it fails.
|
||||
|
||||
The worked example is from this bundle's own tooling. A validator ran green for
|
||||
months. It checked exactly one condition: that skill documents did not contain an
|
||||
absolute path from **a different workspace than the one it ran in** — a literal
|
||||
that could never appear. The check existed, the consumer existed, the report said
|
||||
zero errors, and nineteen real leaks were present.
|
||||
|
||||
That is category C2 from `ue-reference-project-adoption` — a mechanism whose
|
||||
consumer exists and whose condition can never become true — found in our own
|
||||
code rather than in someone else's. It is the reason the delivery gate for this
|
||||
bundle ships with one poisoned fixture per rule and refuses to pass if any rule
|
||||
has never reddened.
|
||||
|
||||
Apply the same rule to content-validation commandlets, asserts, editor
|
||||
validation and CI steps: **break something on purpose and confirm the pipeline
|
||||
notices.** A gate whose failure path has never executed is a log line with
|
||||
ambitions.
|
||||
|
||||
---
|
||||
|
||||
## 7. Write findings down where they survive
|
||||
|
||||
A finding that exists only in a conversation is lost. Context ends; files do not.
|
||||
|
||||
- write the observation at the moment it is made, not at the end of the session;
|
||||
- put the citation next to the claim, not in a separate index;
|
||||
- record open questions **as open**, rather than closing them with a plausible
|
||||
guess;
|
||||
- when a later measurement contradicts an earlier note, add the correction and
|
||||
keep both — the fact that the number changed is itself information about how
|
||||
it was obtained.
|
||||
|
||||
That last point matters more than it looks. Two of the corrections in §3 were
|
||||
only findable because the original number had been written down precisely enough
|
||||
to be checked.
|
||||
|
||||
---
|
||||
|
||||
## 8. Publication checklist
|
||||
|
||||
Before an audit, a review or a skill leaves your hands:
|
||||
|
||||
- [ ] every headline claim re-read at its source by you;
|
||||
- [ ] measured, derived and open marked, and never mixed in one sentence;
|
||||
- [ ] every negative claim states the searches that produced it;
|
||||
- [ ] every recipe executed against a real codebase, in both directions;
|
||||
- [ ] every gate has a failing case that has been observed to fail;
|
||||
- [ ] open questions listed as open;
|
||||
- [ ] counts re-measured rather than carried over from an earlier draft;
|
||||
- [ ] corrections recorded rather than silently applied.
|
||||
|
||||
---
|
||||
|
||||
## 9. What this does not give you
|
||||
|
||||
None of the above makes a claim true. It makes a false claim **findable** — by a
|
||||
reader, by a later measurement, by the person who inherits the document. That is
|
||||
the whole ambition, and it is worth stating plainly because the alternative
|
||||
belief is dangerous: a document full of markers and citations can still be
|
||||
wrong, and a green gate means the form is right, not the content.
|
||||
|
||||
The one thing that reliably catches the rest is the habit in §3: read it yourself
|
||||
before you say it.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Every rule here was written after violating it. The examples are real, they come
|
||||
from the audit that produced this bundle, and they are included specifically
|
||||
because rules stated without the failure that motivated them do not survive
|
||||
contact with a deadline.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
This skill contains no claims about any codebase. Its examples describe mistakes
|
||||
made during one research project and the corrections that followed; they are
|
||||
offered as illustrations of a class, not as findings about the project they came
|
||||
from.
|
||||
@@ -0,0 +1,327 @@
|
||||
# Failure modes: evidence discipline
|
||||
|
||||
Ten ways an audit produces a confident wrong answer.
|
||||
|
||||
Every entry below happened during the work that produced this skill bundle. Not
|
||||
"could happen" — happened, was caught, and cost real time. That is the only
|
||||
provenance this document has and the only one it needs: a method document written
|
||||
from imagination describes a method nobody has used.
|
||||
|
||||
The shared property is uncomfortable: **each of these produces output that looks
|
||||
like evidence.** A count, a file list, a quoted line, a green check. The failure
|
||||
is never a missing answer; it is a well-formed answer to a question you did not
|
||||
ask.
|
||||
|
||||
Identifiers (`ED-01` and up) are stable.
|
||||
|
||||
---
|
||||
|
||||
## Searching
|
||||
|
||||
### ED-01 - An empty result read as absence
|
||||
|
||||
**Mechanism.** A search returns nothing, and "nothing" is recorded as "this does
|
||||
not exist in the codebase".
|
||||
|
||||
**Why it is silent.** Zero hits is a definite-looking answer. It has no error
|
||||
state, no partial result, and nothing that suggests the pattern rather than the
|
||||
tree was at fault.
|
||||
|
||||
**Why the obvious check misses it.** The obvious check *is* the search. Verifying
|
||||
it requires a second, differently-shaped search — which feels redundant precisely
|
||||
when it is most needed, because the first one was so clear.
|
||||
|
||||
**Symptom.** A claim of the form "there is no X here" that turns out to be "my
|
||||
pattern did not match X". In this project it happened three times: a structure
|
||||
assertion missed because the namespace prefix was omitted; a subsystem file
|
||||
missed because the path was guessed rather than found; and — the one that would
|
||||
have shipped — a leaked-path audit whose pattern used doubled backslashes and
|
||||
returned exactly one hit, which was nearly reported as "only one leak" when the
|
||||
real count was nineteen across two donors.
|
||||
|
||||
**Detect.** Before recording an absence, widen deliberately:
|
||||
|
||||
```bash
|
||||
rg -n "Scalability::FQualityLevels" . # first attempt: 0 hits
|
||||
rg -n "FQualityLevels" . # widened: 5 hits
|
||||
rg -c "PatternPart" . ; rg -c "OtherPart" . # both halves separately
|
||||
```
|
||||
|
||||
Then check what your tool excludes by default — ignore files, binary files,
|
||||
hidden directories — and whether the path you searched is the path that exists.
|
||||
|
||||
**Guardrail.** **An empty result proves the pattern, not the absence.** Record
|
||||
absences only after a widened search, a path check, and a statement of what the
|
||||
search could not see.
|
||||
|
||||
---
|
||||
|
||||
### ED-02 - Build output counted as source
|
||||
|
||||
**Mechanism.** An unrestricted search matches compiled artefacts — debug symbols,
|
||||
binaries, intermediate files — and their hits are counted alongside source.
|
||||
|
||||
**Why it is silent.** The matches are real. The count is arithmetically correct.
|
||||
Nothing marks a hit as coming from a file that is generated rather than written.
|
||||
|
||||
**Why the obvious check misses it.** The result list is usually long enough that
|
||||
nobody reads every line, and the summary count is what gets quoted. The tool even
|
||||
labels binary matches — in a line most people skip.
|
||||
|
||||
**Symptom.** A field described as "appearing four times" when it appears once. In
|
||||
this project exactly that: three of four hits were debug symbol files, and the
|
||||
source truth was a single declaration — which is a much stronger finding, since a
|
||||
field declared once and never read is a cleaner defect than one used four times.
|
||||
|
||||
**Detect.** Restrict to source globs, always, and check the difference:
|
||||
|
||||
```bash
|
||||
rg -c "Symbol" . # everything
|
||||
rg -c "Symbol" --glob "*.cpp" --glob "*.h" . # source only
|
||||
```
|
||||
|
||||
If the two numbers differ, the first one was never a fact about your code.
|
||||
|
||||
**Guardrail.** Default to source globs. When a count matters, state which file
|
||||
types it covers.
|
||||
|
||||
---
|
||||
|
||||
### ED-03 - A count taken over the wrong root
|
||||
|
||||
**Mechanism.** A census runs over the primary content directory in a project whose
|
||||
plugins mount their own roots, and reports a total.
|
||||
|
||||
**Why it is silent.** The number is real and internally consistent. Nothing in
|
||||
the result indicates which roots were not visited.
|
||||
|
||||
**Why the obvious check misses it.** The primary root is the obvious root, and it
|
||||
is where nearly all hand-authored content lives in a small project. The error only
|
||||
appears at a scale where checking is expensive.
|
||||
|
||||
**Symptom.** Budgets, dependency graphs and audits that are individually correct
|
||||
and collectively wrong. In this project the primary root held 2 838 assets while
|
||||
the full set of 105 mount roots held 17 256 — a factor of six. Separately, a
|
||||
native-symbol count restricted to the main source directory undercounted by
|
||||
eighteen for the same reason.
|
||||
|
||||
**Detect.** Enumerate roots before counting within them:
|
||||
|
||||
```python
|
||||
roots = ar.get_sub_paths("/", False) # returns all mount roots
|
||||
```
|
||||
|
||||
```bash
|
||||
rg -c "PATTERN" Source/ # one root
|
||||
rg -c "PATTERN" Source/ Plugins/ # all of them
|
||||
```
|
||||
|
||||
**Guardrail.** A count states its scope in the same sentence as its value. "3 876
|
||||
project-owned assets across 105 mount roots" is a fact; "3 876 assets" is a number
|
||||
waiting to be misused.
|
||||
|
||||
---
|
||||
|
||||
## Reasoning
|
||||
|
||||
### ED-04 - Measured and inferred mixed in one sentence
|
||||
|
||||
**Mechanism.** A sentence contains something read in source and something
|
||||
concluded from it, with no marker separating them.
|
||||
|
||||
**Why it is silent.** The sentence is true. Both halves are defensible. The reader
|
||||
inherits the conclusion with the same confidence as the observation, which is one
|
||||
level of confidence too many.
|
||||
|
||||
**Why the obvious check misses it.** Review checks whether claims are correct, not
|
||||
whether their epistemic status is labelled. A correct inference passes.
|
||||
|
||||
**Symptom.** An inference propagating into other documents as a measurement,
|
||||
where it can no longer be traced back and questioned. Most false conclusions in
|
||||
this project's audit originated at exactly this seam.
|
||||
|
||||
**Detect.** This one is answered by format, not by search. Require a marker per
|
||||
claim — measured, derived, open — and treat an unmarked claim as unreviewed.
|
||||
Then check the derived ones for the strongest available failure: what would have
|
||||
to be true for this inference to be wrong, and did anyone check?
|
||||
|
||||
**Guardrail.** Three categories, never mixed within a sentence. An open question
|
||||
stays written as open rather than closed with a plausible guess — a plausible
|
||||
guess is indistinguishable from a finding six months later.
|
||||
|
||||
---
|
||||
|
||||
### ED-05 - A subagent's headline accepted without re-reading
|
||||
|
||||
**Mechanism.** Delegated investigation returns a confident summary. The summary is
|
||||
used directly.
|
||||
|
||||
**Why it is silent.** Reports are fluent, structured and usually mostly right.
|
||||
Errors arrive in the same register as correct findings.
|
||||
|
||||
**Why the obvious check misses it.** Reading the report *is* the check, and the
|
||||
report is internally coherent. Detecting the error requires re-reading the source
|
||||
the report was derived from — that is, redoing the delegated work at the points
|
||||
that matter.
|
||||
|
||||
**Symptom.** In this project: an exploration agent reported three sites carrying
|
||||
absolute workstation paths; the real count was nineteen. Separately, a claim that
|
||||
the game-phase system lived in a feature plugin, when it lives in the main module.
|
||||
|
||||
**Detect.** For each headline claim in a report, open the cited location and read
|
||||
it. If the report cites no location, the claim is unverifiable and is dropped
|
||||
rather than softened.
|
||||
|
||||
**Guardrail.** Delegated output is **raw material**, not a result. Every headline
|
||||
claim is re-read at the source before it reaches a document or a conclusion.
|
||||
|
||||
---
|
||||
|
||||
### ED-06 - A number carried forward without re-measurement
|
||||
|
||||
**Mechanism.** A count measured once is quoted in later documents. The tree
|
||||
changes, or the original measurement was scoped differently, and the number
|
||||
persists.
|
||||
|
||||
**Why it is silent.** Numbers do not expire visibly. A stale count looks exactly
|
||||
like a fresh one, and having a number at all suppresses the impulse to take one.
|
||||
|
||||
**Why the obvious check misses it.** Review checks whether the document is
|
||||
coherent, and it is. The original measurement was correct when taken.
|
||||
|
||||
**Symptom.** In this project, an audit stated "59 native definitions across 40+
|
||||
files". Re-measured during depersonalization: **94 definitions across 31 files**,
|
||||
with the discrepancy caused by ED-03 in the original. The correction was recorded
|
||||
alongside the original rather than replacing it, because which of the two readings
|
||||
was direct is itself information.
|
||||
|
||||
**Detect.** Re-run the measurement when quoting it in a new context, and diff:
|
||||
|
||||
```bash
|
||||
rg -c "PATTERN" Source/ Plugins/ | awk -F: '{s+=$2} END {print s}'
|
||||
```
|
||||
|
||||
**Guardrail.** A quoted number carries the command that produced it, so the next
|
||||
reader can re-run it in one paste. Corrections are recorded as corrections — an
|
||||
archive that silently self-heals loses the trail of which reading was direct.
|
||||
|
||||
---
|
||||
|
||||
## Recording
|
||||
|
||||
### ED-07 - A finding that never left the conversation
|
||||
|
||||
**Mechanism.** Something real is discovered, discussed, and not written down.
|
||||
|
||||
**Why it is silent.** At the moment of discovery it feels known. The cost arrives
|
||||
later, when the context holding it is gone.
|
||||
|
||||
**Why the obvious check misses it.** There is nothing to check. The absence of a
|
||||
note is not visible from anywhere except the future.
|
||||
|
||||
**Symptom.** The same investigation performed twice. A conclusion remembered
|
||||
without its evidence, which then cannot be defended or corrected.
|
||||
|
||||
**Detect.** At the end of a session, diff what was concluded against what was
|
||||
written. Anything in the first list and not the second is already lost — writing
|
||||
it down later is a reconstruction, not a record.
|
||||
|
||||
**Guardrail.** **Found and not written down is lost.** Write the observation
|
||||
immediately, with its address, before continuing. A conversation is not a store.
|
||||
|
||||
---
|
||||
|
||||
### ED-08 - A check whose condition cannot become true
|
||||
|
||||
**Mechanism.** A validator tests for a literal that could never appear — a path
|
||||
from a different workspace, a name from a previous project, a case the current
|
||||
tree cannot produce.
|
||||
|
||||
**Why it is silent.** The check runs, passes, and reports success. Green is the
|
||||
outcome it was built to produce, and it produces it forever.
|
||||
|
||||
**Why the obvious check misses it.** The validator exists, is invoked, and is
|
||||
listed in the process documentation. Everything about it says "this is checked".
|
||||
Nobody re-reads a passing check.
|
||||
|
||||
**Symptom.** In this project: a skill validator searched for one hardcoded
|
||||
absolute path from a *different* project, and only inside one file type. It ran
|
||||
green for months while nineteen leaked paths from two donors sat in the tree.
|
||||
That validator became this bundle's worked example — our own code, not the
|
||||
reference's.
|
||||
|
||||
**Detect.** For every check, produce the input that makes it fail. If you cannot
|
||||
construct one, the check does not exist:
|
||||
|
||||
```bash
|
||||
python validate.py # green
|
||||
# now poison a fixture with the exact thing the rule forbids
|
||||
python validate.py # must be red
|
||||
```
|
||||
|
||||
**Guardrail.** **A rule with no fixture that reddens it is not a rule.** Keep one
|
||||
poisoned fixture per rule, plus a clean baseline, and assert the rule count so a
|
||||
rule cannot be dropped silently.
|
||||
|
||||
---
|
||||
|
||||
### ED-09 - A namespace collision that collapses two sets
|
||||
|
||||
**Mechanism.** Two independent collections use the same identifier prefix. A tool
|
||||
that merges them by key silently overwrites, and reports success on the survivors.
|
||||
|
||||
**Why it is silent.** Every remaining entry resolves. The tool's output is a
|
||||
consistent, complete-looking mapping — of a smaller set than it was given.
|
||||
|
||||
**Why the obvious check misses it.** The check verifies that everything present
|
||||
resolves, which is true. Nothing verifies that everything given is still present.
|
||||
The count is the only tell, and only if someone compares it against the source.
|
||||
|
||||
**Symptom.** In this project, two skills adopted the same two-letter prefix. Nine
|
||||
entries vanished from the resolver, which reported "173 shipped, 173 mapped, 0
|
||||
unresolved" — a perfectly green run against a set that had lost nine members. It
|
||||
was caught only because the total failed to grow after adding nine entries.
|
||||
|
||||
**Detect.** Count both sides independently and compare:
|
||||
|
||||
```bash
|
||||
rg -c "^### [A-Z]{2,4}-" skills/*/references/*.md | awk -F: '{s+=$2} END {print s}'
|
||||
# compare against what the merging tool reports
|
||||
```
|
||||
|
||||
Then check prefix uniqueness directly, and make it a rule rather than a habit.
|
||||
|
||||
**Guardrail.** Any tool that merges by key asserts that its output cardinality
|
||||
equals its input cardinality. A resolver that cannot lose entries is worth more
|
||||
than one that reports zero failures.
|
||||
|
||||
---
|
||||
|
||||
### ED-10 - A recipe published without being run
|
||||
|
||||
**Mechanism.** A detection recipe is written from understanding of the defect
|
||||
rather than from executing it against a real tree.
|
||||
|
||||
**Why it is silent.** The recipe is plausible, well-formed, and would work if the
|
||||
world were slightly simpler. It is published in a document whose whole purpose is
|
||||
to be trusted.
|
||||
|
||||
**Why the obvious check misses it.** Reading the recipe confirms it expresses the
|
||||
right idea. Only running it reveals that the pattern matches something it should
|
||||
not, or misses something it should catch.
|
||||
|
||||
**Symptom.** In this project, three recipes were wrong on first draft and each was
|
||||
wrong in a way that produced a **false negative** — the worst direction. A search
|
||||
for a field's writer matched the declaration's own initializer and reported "there
|
||||
is a writer" in exactly the case the recipe exists to catch. A character class
|
||||
intended to find comparisons matched the arrow operator and reported assignments
|
||||
as comparisons. A search for a namespace matched nothing because real names carry
|
||||
a prefix — and zero hits would have read as "this problem is absent here".
|
||||
|
||||
**Detect.** Run every recipe against a tree where you already know the answer, and
|
||||
check both directions: does it find the known instance, and does it stay quiet
|
||||
where there is none?
|
||||
|
||||
**Guardrail.** **A recipe that has not been run is a hypothesis.** Publish the
|
||||
corrected form, and where the first draft failed in an instructive way, publish
|
||||
that too — the trap is often more useful than the recipe.
|
||||
@@ -0,0 +1,321 @@
|
||||
---
|
||||
name: ue-game-settings-architecture
|
||||
description: >-
|
||||
Design or review game settings and performance architecture in Unreal Engine:
|
||||
setting model trees separated from widgets, reflection-backed property paths,
|
||||
declarative edit conditions, apply and cancel as a transaction, machine-local
|
||||
versus player-shared persistence, scalability and console variables and device
|
||||
profiles, hardware benchmark wiring, and sampled performance statistics. Use
|
||||
when adding video, audio, input or accessibility settings, cloud saves,
|
||||
quality presets, frame-rate modes, device tuning or performance overlays.
|
||||
---
|
||||
|
||||
# UE game settings architecture
|
||||
|
||||
The invariant:
|
||||
|
||||
> A setting is a **model** with persistence, validation and application
|
||||
> semantics. A widget is one view of that model.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes: [failure modes](references/failure-modes.md).
|
||||
|
||||
Settings look like the least architectural part of a game and are one of the
|
||||
most: they span a UI framework, a reflection layer, two persistence stores, the
|
||||
console-variable system, device profiles and the engine's scalability tables.
|
||||
Every one of those boundaries fails silently.
|
||||
|
||||
Related skills: `ue-data-driven-architecture`, `ue-ui-architecture`,
|
||||
`ue-input-architecture`, `ue-modular-gameplay`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Separate model, registry and presentation
|
||||
|
||||
```text
|
||||
registry (per local player)
|
||||
└─ collections and pages
|
||||
└─ setting models
|
||||
├─ data source: getter and setter
|
||||
├─ display text and description
|
||||
├─ edit conditions
|
||||
├─ options and validation
|
||||
└─ apply / store-initial / restore lifecycle
|
||||
|
||||
visual data asset
|
||||
└─ setting model class → entry widget class
|
||||
```
|
||||
|
||||
The screen renders a tree of models through a class-to-widget map; it does not
|
||||
hand-build a widget per option. The payoff: one model renders in several screens,
|
||||
search and grouping operate on models, edit conditions are testable without any
|
||||
UI, and settings code depends on no concrete widget class.
|
||||
|
||||
Resolve the widget by walking the model's **superclass chain**, so mapping one
|
||||
base class covers every derived setting and a specific override remains possible.
|
||||
|
||||
---
|
||||
|
||||
## 2. Reflection-backed data sources, and their one cost
|
||||
|
||||
A setting stores a **path** to a getter and setter rooted at the local player,
|
||||
and resolves it through reflection:
|
||||
|
||||
```text
|
||||
local player → GetLocalSettings() → GetShadowQuality / SetShadowQuality
|
||||
```
|
||||
|
||||
This removes a class per setting. It also removes the compiler from the
|
||||
relationship.
|
||||
|
||||
### The cost, stated precisely
|
||||
|
||||
A macro that stringifies a function name checks that the **name** exists. It does
|
||||
not check the return type, the parameter count, or constness. Those failures
|
||||
surface as a runtime assertion during registry initialization — **and assertions
|
||||
are compiled out of shipping**, so in a shipping build a broken path is a setting
|
||||
that silently does nothing.
|
||||
|
||||
Three consequences worth designing around:
|
||||
|
||||
1. validate every path during registry initialization and **fail loudly with the
|
||||
setting's developer name**, not with a reflection error;
|
||||
2. run that validation in an automated test, not only when a developer opens the
|
||||
screen;
|
||||
3. prefer explicit binding where conversion is complex, where several values
|
||||
change atomically, or where the setter has side effects worth reading.
|
||||
|
||||
Recipe: GS-01.
|
||||
|
||||
---
|
||||
|
||||
## 3. Edit conditions are declarative dependencies
|
||||
|
||||
Availability depends on platform traits, primary-player status, other settings,
|
||||
capabilities, privileges, input device and build configuration. Express it in the
|
||||
model, never in widget visibility.
|
||||
|
||||
Return more than a boolean:
|
||||
|
||||
```text
|
||||
editable
|
||||
disabled (reason shown to the player)
|
||||
hidden (reason recorded for the developer)
|
||||
option disabled (one entry removed from a list)
|
||||
killed (hidden, not resettable, not reported)
|
||||
```
|
||||
|
||||
A condition also declares which settings it depends on, so that changing one
|
||||
setting re-evaluates the others.
|
||||
|
||||
**The two-reason split is the part worth copying.** A disabled setting shows the
|
||||
player why; a hidden one records why for whoever later asks where it went. Most
|
||||
hand-built settings screens have neither.
|
||||
|
||||
---
|
||||
|
||||
## 4. Persistence: machine-local versus player-shared
|
||||
|
||||
| | Machine-local | Player-shared |
|
||||
|---|---|---|
|
||||
| Holds | resolution, quality, frame rate, device profile, benchmark result | accessibility, sensitivities, subtitles, language, bindings |
|
||||
| Backed by | engine user settings, local ini | per-player save game, cloud-syncable |
|
||||
| Scope | one per machine | one per player |
|
||||
|
||||
### The ownership test
|
||||
|
||||
1. Should this follow the user to another machine?
|
||||
2. Can two local players need different values at once?
|
||||
3. Does it change global renderer or audio device state?
|
||||
4. Can a secondary local player safely apply it?
|
||||
|
||||
### The leak to watch for
|
||||
|
||||
A per-player accessor that returns a **process-global singleton** is the standard
|
||||
way this abstraction breaks. It compiles, it reads naturally at every call site,
|
||||
and it silently makes split-screen impossible for every value behind it. The
|
||||
symptom in a codebase that has it: a proliferation of "only the primary player
|
||||
may edit this" conditions, which is the architecture apologising for itself.
|
||||
|
||||
Recipe: GS-03. If you find it, the fix is not to remove the conditions.
|
||||
|
||||
---
|
||||
|
||||
## 5. Apply and cancel is a transaction
|
||||
|
||||
1. store initial values on open;
|
||||
2. edits update the model, and optionally preview;
|
||||
3. apply commits and invokes application;
|
||||
4. cancel restores initial values **and reapplies the prior runtime state**;
|
||||
5. reset-to-defaults is another staged mutation, not an immediate save.
|
||||
|
||||
### Classify application timing deliberately
|
||||
|
||||
| Setting | Timing |
|
||||
|---|---|
|
||||
| volume, UI preview | immediate, restored on cancel |
|
||||
| scalability channel | preview or apply, per policy |
|
||||
| resolution and window mode | apply with a confirmation countdown |
|
||||
| key binding | staged profile plus collision check |
|
||||
| restart-required | save now, mark pending |
|
||||
|
||||
Applying a heavyweight global change from inside a change notification is a
|
||||
decision, not a default. If it is done because staging was not written yet, mark
|
||||
it as such — and treat a comment saying "for now" on a global apply path as a
|
||||
finding rather than a note. Recipe: GS-05.
|
||||
|
||||
---
|
||||
|
||||
## 6. From setting to console variable
|
||||
|
||||
The preferred path keeps a typed object in the middle:
|
||||
|
||||
```text
|
||||
model → typed settings setter → staged quality levels
|
||||
→ apply → engine scalability API → expanded console variables
|
||||
→ clamped by device profile
|
||||
```
|
||||
|
||||
A typed object gives range validation, persistence, transactional restore,
|
||||
platform clamping, change notification and a test seam. A widget writing a
|
||||
console variable directly gives none of them.
|
||||
|
||||
When a setting genuinely maps to a custom variable: resolve and cache safely,
|
||||
set a defined priority source, clamp, restore on cancel, and document shipping
|
||||
availability. Never write to a read-only variable and assume it took.
|
||||
|
||||
---
|
||||
|
||||
## 7. Device profiles are constraints, not presets
|
||||
|
||||
```text
|
||||
platform default → base profile → mode suffix → user request
|
||||
→ platform clamp → effective value
|
||||
```
|
||||
|
||||
Keep **requested** and **effective** separate, and be able to explain to the
|
||||
player why a request was clamped.
|
||||
|
||||
Profile selection commonly composes a name from parts — base, mode, user choice —
|
||||
and falls back through progressively shorter combinations until one exists. That
|
||||
is a good pattern with one specific hazard: a composition part that is declared,
|
||||
read, and never assigned. The fallback then silently always takes the shorter
|
||||
branch, and the feature the part represents does not exist while appearing to.
|
||||
Recipe: GS-07.
|
||||
|
||||
---
|
||||
|
||||
## 8. Wire the benchmark, do not merely implement it
|
||||
|
||||
```text
|
||||
can run? → run benchmark → map indices through thresholds
|
||||
→ set quality and resolution scale → save and apply
|
||||
```
|
||||
|
||||
A function named "should this run at startup" with **no caller** means startup
|
||||
autodetection does not exist, however complete the rest of the pipeline is.
|
||||
|
||||
Decide explicitly: automatic on first launch or manual only; how hardware changes
|
||||
are detected; whether user overrides survive; what happens on failure. Then add a
|
||||
**call-site test**, not only a unit test of the threshold mapping — the mapping
|
||||
was never the part that was missing. Recipe: GS-08.
|
||||
|
||||
---
|
||||
|
||||
## 9. Performance statistics: collect once, present many
|
||||
|
||||
```text
|
||||
engine frame data → one consumer → latest values + sampled ring buffers → widgets
|
||||
```
|
||||
|
||||
Rules that keep this honest:
|
||||
|
||||
- one collector, never one per widget;
|
||||
- a fixed-capacity ring buffer with a **documented sample cadence** — a buffer
|
||||
size without a cadence describes no window at all;
|
||||
- units written down: milliseconds, frames per second, bytes per second, percent;
|
||||
- an unavailable statistic returns *unavailable*, never zero;
|
||||
- availability in shipping is a deliberate setting.
|
||||
|
||||
A frame-rate counter is not evidence about rendering cost. Use it to notice, not
|
||||
to conclude.
|
||||
|
||||
---
|
||||
|
||||
## 10. Reapply after composition changes
|
||||
|
||||
Global scalability and frame pacing may need reapplying after a game mode or
|
||||
feature set finishes loading, because features can change device profile,
|
||||
console variables and frontend policy.
|
||||
|
||||
Do it **once, at one named lifecycle point, after the composition is fully
|
||||
known** — not from each feature. The same entry point should serve the other
|
||||
trigger that changes the same state, such as a hotfixed device profile. Two
|
||||
causes, one reaction, one idempotent function.
|
||||
|
||||
---
|
||||
|
||||
## 11. Guard engine-struct assumptions
|
||||
|
||||
Code that enumerates every member of an engine struct — copying, clamping or
|
||||
maxing quality levels field by field — should carry a compile-time size guard
|
||||
with a message telling the upgrader what to review.
|
||||
|
||||
A failing size assertion after an engine upgrade is a **feature**. It is the only
|
||||
thing standing between a newly added quality channel and being silently ignored
|
||||
by five functions. Update the guard only after reviewing every one of them, and
|
||||
know how many there are before you start. Recipe: GS-09.
|
||||
|
||||
---
|
||||
|
||||
## 12. Review checklist
|
||||
|
||||
- [ ] Setting is a model; the widget is a view.
|
||||
- [ ] Stable developer name is separate from localized display text.
|
||||
- [ ] Every property path is validated at initialization and in a test.
|
||||
- [ ] Machine-local versus player-shared ownership is explicit per value.
|
||||
- [ ] No per-player accessor returns a global singleton.
|
||||
- [ ] Apply and cancel restore runtime state, not only stored values.
|
||||
- [ ] Expensive apply operations have deliberate, documented timing.
|
||||
- [ ] Requested and effective values are distinguishable.
|
||||
- [ ] Direct console-variable writes are justified and restorable.
|
||||
- [ ] The benchmark has a real caller and a stated policy.
|
||||
- [ ] Statistics are collected once, sampled at a documented cadence.
|
||||
- [ ] Unavailable is distinct from zero.
|
||||
- [ ] Reapplication after composition change is centralized and idempotent.
|
||||
- [ ] Engine-struct assumptions carry compile-time guards.
|
||||
|
||||
---
|
||||
|
||||
## 13. When to simplify
|
||||
|
||||
A prototype can bind widgets straight to engine user settings. Adopt a model
|
||||
registry when at least one holds: platforms hide or clamp settings differently;
|
||||
apply and cancel are required; several local users need separate preferences;
|
||||
search, categories or dynamic conditions matter; gamepad navigation needs
|
||||
reusable rows; settings span ini, save games, console variables and subsystems.
|
||||
|
||||
Do not adopt a five-thousand-line framework for three toggles. Adopt it when the
|
||||
settings screen has become an application subsystem rather than a panel.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings come from a source audit of Epic's Lyra Starter Game on Unreal
|
||||
Engine 5.6 and the settings plugin it ships with — roughly 5700 lines of
|
||||
framework plus 2100 of project-specific settings backends, read rather than run.
|
||||
The framework itself contains no reference to the game and is the most portable
|
||||
component examined in the whole research project; the findings below are about
|
||||
its **wiring**, not its design.
|
||||
|
||||
Source addresses stay in the research archive. Each entry carries a stable
|
||||
identifier (`GS-01`, `GS-03`, …) resolving back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version. Concrete widget assets are binary and were not
|
||||
read, so claims here concern the model layer and its bindings. Values quoted from
|
||||
configuration describe that project on that day.
|
||||
+348
@@ -0,0 +1,348 @@
|
||||
# Failure modes: game settings and performance
|
||||
|
||||
Nine ways a settings layer stores a value the player chose and does nothing with
|
||||
it.
|
||||
|
||||
The shared property is unusual and worth stating carefully: **in this area the
|
||||
framework is usually right and the wiring is usually missing.** Every finding
|
||||
below sits at a seam — between a model and a property path, between a stored
|
||||
value and the code that applies it, between a complete pipeline and the caller it
|
||||
never got. None of them is a design flaw in the mechanism they belong to.
|
||||
|
||||
That matters for adoption, which is what this skill is for: the correct response
|
||||
to this list is *copy the framework, audit the wiring*, not *write your own*.
|
||||
|
||||
Recipes use `rg` from a project source root and were executed against the audited
|
||||
project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## Binding
|
||||
|
||||
### GS-01 - A property path the compiler does not check
|
||||
|
||||
**Mechanism.** Settings bind to storage through a stringified path to a getter
|
||||
and setter, resolved by reflection at runtime. A macro checks that the function
|
||||
*name* exists; nothing checks the return type, the parameter count or constness.
|
||||
|
||||
**Why it is silent.** A mismatched path fails at initialization with an
|
||||
assertion — and assertions are compiled out of shipping builds. In shipping, a
|
||||
broken binding is a setting that reads a default and writes nowhere.
|
||||
|
||||
**Why the obvious check misses it.** The macro's name contains "checked", and it
|
||||
genuinely does check something. Review sees a compile-time-looking construct and
|
||||
stops. The remaining gap is invisible until the types diverge, which happens
|
||||
during a refactor rather than at authoring time.
|
||||
|
||||
**Symptom.** In development, an assertion when the settings screen opens — often
|
||||
dismissed as unrelated. In shipping, a setting that silently does nothing, which
|
||||
is indistinguishable from a setting that is not implemented.
|
||||
|
||||
**Detect.** Find the path-building macros and confirm the failure mode of the
|
||||
resolver, then check whether anything validates paths outside the editor:
|
||||
|
||||
```bash
|
||||
rg -n "GET_FUNCTION_NAME_STRING_CHECKED|FCachedPropertyPath|PropertyPathHelpers" \
|
||||
--glob "*.h" --glob "*.cpp" .
|
||||
rg -n -B4 "ensure|checkf" --glob "*.cpp" . | rg -i "getter|setter|propertypath"
|
||||
rg -n "UE_BUILD_SHIPPING" --glob "*.cpp" . | rg -i "ensure"
|
||||
```
|
||||
|
||||
**Guardrail.** Validate every path during registry initialization, report the
|
||||
**setting's developer name** rather than a reflection error, and run that
|
||||
validation in an automated test so it fails a build rather than a play session.
|
||||
|
||||
---
|
||||
|
||||
### GS-02 - Uniqueness enforced on the identifier, nothing enforced on the pair
|
||||
|
||||
**Mechanism.** Settings carry a stable developer name, and the registry checks it
|
||||
is unique. Nothing checks that the name still matches the property path it binds.
|
||||
|
||||
**Why it is silent.** Both halves are individually valid: a unique name and a
|
||||
resolvable path. A setting named for one property and bound to another works
|
||||
perfectly and is wrong.
|
||||
|
||||
**Why the obvious check misses it.** The uniqueness check exists and passes,
|
||||
which answers "are the identifiers sane?". Copy-paste between two settings
|
||||
changes the display text and leaves the path — and the resulting setting looks
|
||||
correct in every view except behaviour.
|
||||
|
||||
**Symptom.** Two settings that move together, or one that changes something the
|
||||
player did not select. Reported as "the wrong option is wired up", investigated
|
||||
in the widget layer.
|
||||
|
||||
**Detect.** Extract name/path pairs and eyeball the mismatches:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A6 "SetDevName\(" --glob "*.cpp" . \
|
||||
| rg "SetDevName|SetDynamicGetter|SetDynamicSetter"
|
||||
```
|
||||
|
||||
Read the triples. A developer name and a getter that share no token is the
|
||||
finding — crude, and it catches the copy-paste error that produces this.
|
||||
|
||||
**Guardrail.** Assert the correspondence in a test that walks the registry, or
|
||||
generate the name from the path. A convention nothing enforces is a convention
|
||||
that has already been broken somewhere.
|
||||
|
||||
---
|
||||
|
||||
## Ownership
|
||||
|
||||
### GS-03 - A per-player accessor that returns a global singleton
|
||||
|
||||
**Mechanism.** The local player exposes an accessor for its settings; the
|
||||
implementation returns a process-wide singleton.
|
||||
|
||||
**Why it is silent.** With one local player it is correct in every observable
|
||||
way. The abstraction is only wrong when a second player exists, which is a
|
||||
configuration most testing never enters.
|
||||
|
||||
**Why the obvious check misses it.** Every call site reads correctly — a player
|
||||
asking for its settings. The defect is one line inside an accessor that nobody
|
||||
reads twice, because accessors are the least interesting code in a file.
|
||||
|
||||
**Symptom.** In split-screen, the second player edits the first player's
|
||||
settings. Usually discovered not as a bug but as a growing population of "only
|
||||
the primary player may change this" conditions, added one at a time by people
|
||||
working around a cause nobody named.
|
||||
|
||||
**Detect.** Read the body of every per-player settings accessor:
|
||||
|
||||
```bash
|
||||
rg -n -A4 "::Get\w*Settings\(\) const|::Get\w*Settings\(\)" --glob "*.cpp" . \
|
||||
| rg -i "::Get\(\)|GEngine->|StaticClass"
|
||||
rg -n "PlayingAsPrimaryPlayer|IsPrimaryPlayer" --glob "*.cpp" . | wc -l
|
||||
```
|
||||
|
||||
The second count is the corroborating signal: a large number of primary-player
|
||||
conditions in a project that claims per-player settings means the claim is not
|
||||
true. In the audited project the accessor returns a global object, and the
|
||||
conditions are numerous.
|
||||
|
||||
**Guardrail.** Decide per value whether it is machine-scoped or player-scoped,
|
||||
and let the accessor's return type say which. If a value must be global, name the
|
||||
accessor globally — the honest name prevents the workarounds.
|
||||
|
||||
---
|
||||
|
||||
### GS-04 - Player preferences stored in machine-scoped storage
|
||||
|
||||
**Mechanism.** A value that is clearly a personal preference — volume, for
|
||||
instance — lives in the machine-local store because that is where the subsystem
|
||||
applying it happens to read from.
|
||||
|
||||
**Why it is silent.** It works for one player on one machine, which is the
|
||||
overwhelmingly common case. The value persists, applies, and survives restart.
|
||||
|
||||
**Why the obvious check misses it.** The placement follows the *implementation*
|
||||
rather than the *semantics*, and the implementation reason is real: the applying
|
||||
code is machine-global. A reviewer asking "does this work?" gets yes.
|
||||
|
||||
**Symptom.** Preferences that do not follow the player to another machine, and
|
||||
two local players who cannot have different values. Discovered when cloud saves
|
||||
or split-screen are added, long after the placement was set.
|
||||
|
||||
**Detect.** Classify the fields of both stores against the ownership test:
|
||||
|
||||
```bash
|
||||
rg -n "UPROPERTY\(Config\)" -A2 --glob "*.h" <local-settings>.h | rg "\w+ \w+;"
|
||||
rg -n "UPROPERTY\(\)" -A2 --glob "*.h" <shared-settings>.h | rg "\w+ \w+;"
|
||||
```
|
||||
|
||||
For each field ask the four ownership questions from the skill. In the audited
|
||||
project the audio volumes sit in machine-local storage, and the registry that
|
||||
builds the audio screen reaches into both stores from adjacent lines.
|
||||
|
||||
**Guardrail.** Place by semantics, then solve the application problem. If the
|
||||
applying subsystem is global, the setting can still be player-scoped with the
|
||||
primary player's value applied — that is a deliberate policy rather than an
|
||||
accident of storage.
|
||||
|
||||
---
|
||||
|
||||
## Application
|
||||
|
||||
### GS-05 - A global apply invoked from a change notification
|
||||
|
||||
**Mechanism.** Changing a setting immediately calls the heavyweight global apply
|
||||
path from inside the change handler, rather than staging the value until the
|
||||
player presses apply.
|
||||
|
||||
**Why it is silent.** The result is correct and even feels responsive — the
|
||||
player sees the change instantly. Cancel still restores the value, so the
|
||||
transaction is not broken, only leaky.
|
||||
|
||||
**Why the obvious check misses it.** It is three lines in an edit condition,
|
||||
placed where a reviewer expects a condition rather than an action. The cost only
|
||||
appears with a setting a player can scrub — a slider makes it one global reapply
|
||||
per tick.
|
||||
|
||||
**Symptom.** Stutter while adjusting a quality setting; a visible flash on values
|
||||
the player then cancels; and, for expensive channels, a hitch per change.
|
||||
|
||||
**Detect.** Find application calls inside notification paths, and check for the
|
||||
tell:
|
||||
|
||||
```bash
|
||||
rg -n -B6 "ApplyScalabilitySettings|ApplySettings|SetQualityLevels" --glob "*.cpp" . \
|
||||
| rg "SettingChanged|OnSettingChanged|NotifySettingChanged"
|
||||
rg -n -i "//\s*TODO.*(for now|immediately)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
The second command is the high-yield one: in the audited project the immediate
|
||||
apply carries a comment saying "for now", which converts a judgement call into a
|
||||
confirmed finding.
|
||||
|
||||
**Guardrail.** Stage by default; preview deliberately and cheaply; apply on
|
||||
apply. Treat "for now" on a global apply path as a scheduled item, not a note.
|
||||
|
||||
---
|
||||
|
||||
### GS-06 - A stored preference with no application code
|
||||
|
||||
**Mechanism.** A value is exposed, edited, persisted and loaded. The function
|
||||
that would apply it has an empty body.
|
||||
|
||||
**Why it is silent.** Every visible part of the round trip works: the slider
|
||||
moves, the value saves, it comes back after restart. Only the effect is missing,
|
||||
and the effect is subjective — sensitivity, in the audited case.
|
||||
|
||||
**Why the obvious check misses it.** The apply function exists and is called. A
|
||||
review of the settings layer finds a complete implementation; the emptiness is in
|
||||
a body that the caller has no reason to open. And because the value *is* readable,
|
||||
game code may consume it directly elsewhere, making the empty function neither
|
||||
wrong nor sufficient.
|
||||
|
||||
**Symptom.** A player changes a preference, the UI confirms it, and nothing feels
|
||||
different. Support cannot reproduce it because the value really is stored.
|
||||
|
||||
**Detect.** Find apply functions with empty bodies:
|
||||
|
||||
```bash
|
||||
rg -n -A3 "void \w+::Apply\w+\(\)" --glob "*.cpp" . | rg -B2 "^\s*\}"
|
||||
```
|
||||
|
||||
Then, for each stored preference, find its consumer:
|
||||
|
||||
```bash
|
||||
V='MouseSensitivityX'
|
||||
rg -n "\b$V\b" --glob "*.cpp" --glob "*.h" . | rg -v "UPROPERTY|Set$V|Get$V"
|
||||
```
|
||||
|
||||
No consumer outside the accessors means the value goes nowhere.
|
||||
|
||||
**Guardrail.** A setting ships with its consumer, in the same change. If
|
||||
application is intentionally the game's responsibility rather than the
|
||||
framework's, delete the empty function — an empty function is a claim.
|
||||
|
||||
---
|
||||
|
||||
## Wiring
|
||||
|
||||
### GS-07 - A composition part that is read six times and never assigned
|
||||
|
||||
**Mechanism.** A device-profile name is composed from parts — base platform, a
|
||||
mode suffix, a user choice — with fallback through progressively shorter
|
||||
combinations. One part is declared as a local, read repeatedly, and never
|
||||
written.
|
||||
|
||||
**Why it is silent.** The fallback is designed to handle a missing part, and it
|
||||
does. Every path produces a valid profile name; the composition simply never
|
||||
takes the branches that use the absent part.
|
||||
|
||||
**Why the obvious check misses it.** The variable appears in six places,
|
||||
including a log line that prints it. Any "is this used?" search answers yes
|
||||
loudly. The absent half is the **writer**, and no default tool asks that
|
||||
question — this is the canonical shape from `ue-reference-project-adoption`,
|
||||
found here in a subsystem where the consequence is a whole feature that does not
|
||||
exist.
|
||||
|
||||
**Symptom.** Documentation and code structure both indicate that a game mode can
|
||||
select its own device profile. It cannot. A team planning per-mode performance
|
||||
profiles discovers this after designing around it.
|
||||
|
||||
**Detect.** For any composed name, count reads against writes:
|
||||
|
||||
```bash
|
||||
V='ExperienceSuffix'
|
||||
rg -n "\b$V\b" --glob "*.cpp" .
|
||||
rg -n "\b$V\b\s*=[^=]" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Reads with no assignment is the finding. In the audited project the first command
|
||||
returns six lines — a declaration, four reads and a log — and the second returns
|
||||
nothing but the declaration itself, which the authors annotated with a comment
|
||||
saying nothing sets it.
|
||||
|
||||
**Guardrail.** Do not ship a composition part with no producer. If it is
|
||||
scaffolding for a planned feature, say so where a reader of the *composition*
|
||||
will see it, not only at the declaration.
|
||||
|
||||
---
|
||||
|
||||
### GS-08 - A complete pipeline with no caller
|
||||
|
||||
**Mechanism.** Automatic quality detection is fully implemented: a capability
|
||||
check, a benchmark, threshold mapping, application and save. The function that
|
||||
decides whether to run it at startup has no caller anywhere.
|
||||
|
||||
**Why it is silent.** The manual path works. A player who presses the button gets
|
||||
correct auto-detection, so the feature demonstrably functions — just never on its
|
||||
own.
|
||||
|
||||
**Why the obvious check misses it.** Every component is present and testable, and
|
||||
unit tests of the threshold mapping pass. "Do we support automatic quality?"
|
||||
answers yes from any angle except the one that matters: who starts it.
|
||||
|
||||
**Symptom.** First-run experience uses default quality on every machine. Nobody
|
||||
notices, because developers' machines default acceptably and the button exists.
|
||||
|
||||
**Detect.** For the decision function, distinguish declaration from call:
|
||||
|
||||
```bash
|
||||
F='ShouldRunAutoBenchmarkAtStartup'
|
||||
rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Exactly two hits — the declaration and the definition — is the finding. Extend to
|
||||
content if the project can call functions from assets, and record the result as
|
||||
open if you cannot inspect them.
|
||||
|
||||
**Guardrail.** Every entry-point function gets a **call-site test**, not only a
|
||||
unit test of what it computes. The mapping was never the part at risk.
|
||||
|
||||
---
|
||||
|
||||
### GS-09 - A size guard whose count nobody knows
|
||||
|
||||
**Mechanism.** Code enumerates every member of an engine struct — copying,
|
||||
clamping, maximising quality levels field by field — and protects itself with a
|
||||
compile-time assertion on the struct's size.
|
||||
|
||||
**Why it is silent.** It is correct, and it is good practice: after an engine
|
||||
upgrade adds a channel, the build breaks instead of silently ignoring it.
|
||||
|
||||
**Why the obvious check misses it.** Nothing is wrong here. The hazard is
|
||||
procedural: the guard fires during an upgrade, under time pressure, and the
|
||||
person who bumps the number needs to know **how many functions** must be reviewed.
|
||||
That number is not written anywhere near the assertion.
|
||||
|
||||
**Symptom.** An engine upgrade where the size constant is updated and one of the
|
||||
five enumerating functions is not, silently dropping a quality channel — exactly
|
||||
the failure the guard was written to prevent.
|
||||
|
||||
**Detect.** Count the guards and the functions they protect before upgrading:
|
||||
|
||||
```bash
|
||||
rg -n "static_assert\(sizeof\(" --glob "*.cpp" --glob "*.h" .
|
||||
rg -c "static_assert\(sizeof\(Scalability::FQualityLevels\)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
In the audited project this returns **five** sites for one struct. Five is the
|
||||
number an upgrader needs and would otherwise have to discover by grepping mid-fix.
|
||||
|
||||
**Guardrail.** Write the count and the list of functions into the assertion
|
||||
message. The guard is only as good as the instruction it gives the person it
|
||||
stops.
|
||||
@@ -0,0 +1,229 @@
|
||||
# Patterns: a measured settings and performance layer
|
||||
|
||||
A worked example of `SKILL.md` against one reference product. The conclusion is
|
||||
unusual enough to state at the top, because it changes what you should do with
|
||||
the rest of the file:
|
||||
|
||||
> **The framework is the best thing in the audited project. The wiring around it
|
||||
> is where every defect is.**
|
||||
|
||||
Roughly 5700 lines of settings framework contain no reference to the game that
|
||||
ships it, and no finding in this file is about its design. The nine failure modes
|
||||
are all in the 2100 lines of project-specific backend and registry that bind it —
|
||||
paths that the compiler cannot check, a pipeline with no caller, a composition
|
||||
part with no writer. That distinction is the adoption advice: **take the
|
||||
framework, audit the wiring**.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from
|
||||
measured facts; **[open]** — not answerable from source.
|
||||
|
||||
---
|
||||
|
||||
## 1. What the framework gives you, concretely
|
||||
|
||||
| Capability **[measured]** | Why it is hard to rebuild |
|
||||
|---|---|
|
||||
| Setting model fully separated from widgets | The separation is what makes search, grouping and testing possible at all |
|
||||
| Widget resolved by walking the model's **superclass chain** | Map one base class, cover every derived setting, keep specific overrides |
|
||||
| Apply / cancel / reset as a real transaction over a dirty set | Correct restore of *runtime* state, not only stored values |
|
||||
| Edit conditions with five outcomes, not a boolean | disabled-with-player-reason, hidden-with-developer-reason, single option disabled, killed |
|
||||
| Cascading dependencies between settings | Changing one setting re-evaluates the others without hand-written wiring |
|
||||
| Full-text search over rich text reduced to plain text | Falls out of the model separation |
|
||||
| Asynchronous readiness before the screen renders | Audio device lists and cloud values arrive late; the screen waits |
|
||||
| A debug mode that shows developer names and classes in the UI | Makes every other finding in this file findable |
|
||||
|
||||
**[derived]** The five-outcome edit condition is the piece most worth
|
||||
internalising even if you write your own framework. Hand-built settings screens
|
||||
almost universally have a boolean, and the two missing things — *tell the player
|
||||
why it is disabled* and *tell the developer why it is hidden* — are exactly the
|
||||
two questions asked later at cost.
|
||||
|
||||
---
|
||||
|
||||
## 2. The trade the framework makes, and where it lands
|
||||
|
||||
Bindings are **string paths resolved by reflection** — a getter and setter path
|
||||
rooted at the local player **[measured]**. That removes a class per setting and
|
||||
removes the compiler from the relationship.
|
||||
|
||||
The failure mode is precise: the stringification macro checks that a function
|
||||
*name* exists, not its type, arity or constness; mismatches surface as an
|
||||
assertion at registry initialization; **assertions are compiled out of shipping**
|
||||
**[measured]**.
|
||||
|
||||
**[derived]** So the worst case is a shipping build where a setting reads a
|
||||
default and writes nowhere, with no diagnostic anywhere. That is not an argument
|
||||
against the design — it is an argument for one automated test that walks the
|
||||
registry and resolves every path. Recipes: GS-01, GS-02.
|
||||
|
||||
---
|
||||
|
||||
## 3. The ownership split, and the leak inside it
|
||||
|
||||
The design is a clean two-store split **[measured]**:
|
||||
|
||||
| | Machine-local | Player-shared |
|
||||
|---|---|---|
|
||||
| Base | engine user settings, ini | per-player save game, cloud-syncable |
|
||||
| Holds | resolution, quality, frame limits, device profile, benchmark result | colour-blind mode, sensitivities, subtitles, language, haptics |
|
||||
|
||||
The authors state the rationale in a comment: shared settings are safe to put in
|
||||
the cloud and are stored per player, so controller preferences belong there
|
||||
rather than in local settings where every user would receive them **[measured]**.
|
||||
|
||||
Then the accessor breaks it. The per-player getter for local settings returns a
|
||||
**process-global singleton** **[measured]**.
|
||||
|
||||
**[derived]** "This player's local settings" is a fiction; there is one object per
|
||||
process. The corroborating evidence is structural rather than a second defect:
|
||||
the registry is dense with "only while playing as the primary player" conditions,
|
||||
which is the architecture apologising for the accessor. Recipe: GS-03.
|
||||
|
||||
A second, milder leak in the same area: audio volumes live in machine-local
|
||||
storage **[measured]** — a pure player preference placed by where the applying
|
||||
subsystem reads, not by what the value means. The audio registry reaches into
|
||||
both stores from adjacent lines. Recipe: GS-04.
|
||||
|
||||
**[derived]** Both are worth seeing together, because they are the same mistake at
|
||||
two scales: **storage placed by implementation convenience rather than by
|
||||
semantics**, then papered over at the call sites.
|
||||
|
||||
---
|
||||
|
||||
## 4. Three findings that are the same shape
|
||||
|
||||
This is the most transferable observation in the file. Three unrelated features,
|
||||
one structure:
|
||||
|
||||
| Feature | Present **[measured]** | Missing **[measured]** |
|
||||
|---|---|---|
|
||||
| Automatic quality detection at first launch | capability check, benchmark, threshold mapping, apply, save | **any caller** of the decision function |
|
||||
| Per-game-mode device profiles | name composition, fallback chain, four candidate forms, logging | **any writer** of the mode suffix |
|
||||
| Input sensitivity application | storage, dirty tracking, save, load, an apply function | **a body** in the apply function |
|
||||
|
||||
**[derived]** Each is complete except for one link, and in each case the missing
|
||||
link is invisible to every check that the rest of the feature passes. The
|
||||
decision function has two occurrences — its declaration and its definition. The
|
||||
suffix has six occurrences, of which one is a log line printing it and none is an
|
||||
assignment. The apply function is called correctly from the right place and
|
||||
contains nothing.
|
||||
|
||||
The suffix case deserves emphasis because it is the canonical shape from
|
||||
`ue-reference-project-adoption` appearing in a subsystem where the consequence is
|
||||
architectural: **a variable read four times and written zero times**, with the
|
||||
authors' own comment noting that nothing sets it. Any search asking "is this
|
||||
used?" answers yes, loudly. Recipes: GS-06, GS-07, GS-08.
|
||||
|
||||
---
|
||||
|
||||
## 5. Where the transaction leaks
|
||||
|
||||
Video settings apply **immediately from inside a change notification**, not on
|
||||
apply, and the code says so with a comment reading "for now" **[measured]**.
|
||||
|
||||
**[derived]** Cancel still restores the value, so the transaction is not broken —
|
||||
it is leaky in a way the player sees: a preview they did not ask for, on a global
|
||||
apply path, potentially once per slider tick.
|
||||
|
||||
The comment is what turns this from a judgement call into a finding. **A "for
|
||||
now" on a global apply path is a scheduled item.** Recipe: GS-05.
|
||||
|
||||
---
|
||||
|
||||
## 6. The upgrade guard, and the number it does not tell you
|
||||
|
||||
Manual enumeration of an engine struct's members — copying, clamping and
|
||||
maximising quality levels field by field — is protected by a compile-time
|
||||
assertion on the struct's size, at **five separate sites** for one struct
|
||||
**[measured]**.
|
||||
|
||||
**[derived]** This is good practice and it works: after an engine upgrade adds a
|
||||
quality channel, the build breaks rather than silently ignoring it. The gap is
|
||||
procedural. The person who hits the assertion during an upgrade needs to know
|
||||
*how many functions* must be reviewed, and that number is nowhere near the
|
||||
message. Recipe: GS-09.
|
||||
|
||||
**Put the count in the assertion text.** A guard is only as useful as the
|
||||
instruction it gives the person it stops.
|
||||
|
||||
---
|
||||
|
||||
## 7. The reapplication point, and why it is where it is
|
||||
|
||||
Global scalability is reapplied at the **last line** of the mode-load completion,
|
||||
after the load state is set and after all three waves of load notification
|
||||
**[measured]**. The same function serves the other trigger that changes the same
|
||||
state — a hotfixed device profile **[measured]**.
|
||||
|
||||
Four reasons, all measured, and together they are a small lesson in ordering:
|
||||
|
||||
1. feature plugins activated by the mode can change console variables and
|
||||
frontend policy, so effective values are unknown until they have all run;
|
||||
2. the device profile itself may change, and the reapplication reselects it;
|
||||
3. scalability is process-global, not per world — the authors note this in a
|
||||
comment explaining why they do not track it per world in multi-instance
|
||||
editor play;
|
||||
4. it is guarded off for servers, where render settings are meaningless.
|
||||
|
||||
**[derived]** Two different causes — mode loaded, profile hotfixed — reduced to
|
||||
one idempotent reaction at one named point. That arrangement is worth copying
|
||||
directly, and it is the opposite of letting each feature reapply what it thinks
|
||||
it changed.
|
||||
|
||||
---
|
||||
|
||||
## 8. Performance statistics: one collector, and the honest limits
|
||||
|
||||
Statistics come from the engine's own frame-data consumer interface rather than
|
||||
from hand-rolled timers **[measured]**, feeding a ring buffer for graphing.
|
||||
Eighteen statistics, with a compile-time count assertion in three places
|
||||
**[measured]**.
|
||||
|
||||
**[derived]** The design point worth copying is that presentation and collection
|
||||
are separate: widgets read a cache, and no widget collects. The design point worth
|
||||
questioning before copying is that a buffer size is meaningless without a
|
||||
documented sample cadence — a hundred-odd samples describes a window only once you
|
||||
know the interval.
|
||||
|
||||
Two hard couplings to be aware of before lifting the subsystem: one statistic
|
||||
reads a game-specific state class, and latency statistics are gated on a specific
|
||||
GPU vendor because the project integrates one vendor's plugin **[measured]**.
|
||||
Both are small and both are visible; neither is a reason to rewrite the rest.
|
||||
|
||||
---
|
||||
|
||||
## 9. What to copy, in order
|
||||
|
||||
1. **The framework itself**, if the UI dependency is acceptable. Five-outcome edit
|
||||
conditions, superclass-chain widget resolution, the change tracker, async
|
||||
readiness.
|
||||
2. **The reapplication point** (§7) — one idempotent function, two causes.
|
||||
3. **The two-store split** (§3) — but decide placement by semantics and check
|
||||
your accessor's body before trusting the word "local".
|
||||
4. **The size-guard discipline** (§6), improved by writing the count into the
|
||||
message.
|
||||
5. **Collector/presenter separation** for statistics (§8).
|
||||
|
||||
Audit before shipping: every property path resolves (GS-01, GS-02); the accessor
|
||||
is not a singleton (GS-03); every apply function has a body (GS-06); every
|
||||
composition part has a writer (GS-07); every pipeline has a caller (GS-08).
|
||||
|
||||
**[derived]** That list is short, mechanical, and would have caught every defect
|
||||
in this file. None of it requires understanding the framework — which is the
|
||||
point of shipping recipes rather than conclusions.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6 and the settings
|
||||
plugin shipped with it: roughly 5700 lines of framework, 2100 of project settings
|
||||
backends, 2000 of registry, and 800 of performance code, read as source in a
|
||||
single workspace. Source addresses stay in the research archive that produced
|
||||
this skill; each `GS-` identifier resolves back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Widget assets are binary and were not read, so nothing here describes the
|
||||
presentation layer beyond its base classes. Configuration values quoted describe
|
||||
that project on that day. Re-run the recipes against your own tree.
|
||||
@@ -0,0 +1,376 @@
|
||||
---
|
||||
name: ue-gameplay-cues
|
||||
description: >-
|
||||
Design, implement or review the GameplayCue presentation layer in Unreal
|
||||
Engine GAS: what a cue is and when it is the right mechanism, Executed versus
|
||||
OnActive/WhileActive/OnRemove forms, the three addressing paths, Static versus
|
||||
Actor notifies, the dedicated-server rule, cue asset discovery and loading
|
||||
cost, registering cue paths from feature plugins, teardown and leaked looping
|
||||
cues, and the tag-to-asset validation gap. Use when adding VFX or audio
|
||||
feedback to abilities or effects, decoupling presentation from gameplay
|
||||
assets, debugging a cue that never fires or never stops, or auditing cue
|
||||
loading cost.
|
||||
---
|
||||
|
||||
# UE gameplay cues
|
||||
|
||||
The invariant:
|
||||
|
||||
> A cue is a one-way, cosmetic, tag-addressed broadcast. Gameplay must produce
|
||||
> exactly the same result if every cue in the project fails to fire.
|
||||
|
||||
Worked analysis of a real cue subsystem, including the parts that were built and
|
||||
never reached: [patterns](references/patterns.md).
|
||||
Twenty detection recipes for silent cue failures: [failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting or reviewing a cue-based design. Each entry
|
||||
names the mechanism, why it stays silent, why the obvious check misses it, and a
|
||||
command you can run against your own tree.
|
||||
|
||||
Related skills: `ue-gas-architecture`, `ue-gameplay-tag-governance`,
|
||||
`ue-modular-gameplay`, `ue-asset-loading-and-memory`, `ue-gameplay-messaging`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. What a GameplayCue is, and the problem it solves
|
||||
|
||||
A GameplayCue is a **presentation event addressed by a GameplayTag** rather than
|
||||
by an asset reference.
|
||||
|
||||
Without it, the chain looks like this:
|
||||
|
||||
```text
|
||||
GE_Fire_Damage --hard reference--> NS_Fire_Impact (Niagara)
|
||||
SW_Fire_Impact (audio)
|
||||
```
|
||||
|
||||
The gameplay asset now owns the VFX and the audio. Loading the effect loads the
|
||||
particles. A dedicated server, which needs the damage numbers and nothing else,
|
||||
loads them too. An artist changing the impact effect touches a gameplay asset.
|
||||
|
||||
With a cue, the chain is cut:
|
||||
|
||||
```text
|
||||
GE_Fire_Damage --tag--> "GameplayCue.Fire.Impact" <--tag-- GCN_Fire_Impact
|
||||
```
|
||||
|
||||
The gameplay asset references a **string-like identifier**. A separate registry
|
||||
maps that tag to a notify asset, discovered by scanning content directories. The
|
||||
gameplay asset has no idea the presentation exists.
|
||||
|
||||
What this buys:
|
||||
|
||||
- gameplay assets carry no presentation payload, so they are small and
|
||||
server-safe;
|
||||
- presentation can be added, replaced or shipped in a separate plugin without
|
||||
touching gameplay;
|
||||
- the whole presentation layer can be absent on a dedicated server.
|
||||
|
||||
What it costs — and this is the part that gets skipped:
|
||||
|
||||
- **the join is no longer checked by anything.** A tag with no notify is silence.
|
||||
A notify with no matching tag is a never-loaded asset. Neither is an error.
|
||||
- discovery is directory-scan based, so cue assets must be found, indexed and
|
||||
loaded by a manager that has its own lifecycle and its own failure modes;
|
||||
- the mechanism is one-way and unreliable by design, which makes it wrong for
|
||||
anything that must be correct.
|
||||
|
||||
Treat cues as an application of `ue-gameplay-tag-governance` with a dedicated
|
||||
loading subsystem attached. Every rule about tag joins applies here first.
|
||||
|
||||
---
|
||||
|
||||
## 2. The three forms
|
||||
|
||||
| Form | Fires when | Use for |
|
||||
|---|---|---|
|
||||
| **Executed** | a one-shot event | impacts, hit reactions, one-off flashes |
|
||||
| **OnActive** | a duration effect starts, on clients present at that moment | start of a looping effect |
|
||||
| **WhileActive** | an actor becomes relevant while the effect is already running | joining/relevancy catch-up for the same loop |
|
||||
| **OnRemove** | the duration effect ends | stopping the loop, cleanup |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **`OnActive` and `WhileActive` must produce the same visual state.** A client
|
||||
who was present and a client who arrived late must see the same thing. Putting
|
||||
the "start" logic only in `OnActive` means late joiners see nothing.
|
||||
2. **Everything `OnActive`/`WhileActive` spawns, `OnRemove` destroys.** Looping
|
||||
cues are the main source of leaked VFX in GAS projects.
|
||||
3. **Never use `Executed` for anything with a duration**, and never use the
|
||||
duration forms for a one-shot — the recovery paths differ.
|
||||
4. **Assume `OnRemove` may run without `OnActive` having run** on that client, and
|
||||
the reverse. Write both defensively.
|
||||
|
||||
---
|
||||
|
||||
## 3. Static versus Actor notifies
|
||||
|
||||
| Notify type | Instanced | State | Use for |
|
||||
|---|---|---|---|
|
||||
| **Static** | no — a CDO handles the call | none | one-shot effects, sounds, decals |
|
||||
| **Actor** | yes — an actor is spawned per instance | yes | looping effects, anything requiring cleanup |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Default to Static.** It allocates nothing and cannot leak.
|
||||
2. **A Static notify must not store per-instance state.** It runs on the class
|
||||
default object; a member written during one invocation is shared by every
|
||||
subsequent one, including other players' invocations.
|
||||
3. **Use an Actor notify only when there is something to tear down** — a looping
|
||||
particle system, an attached component, a timeline.
|
||||
4. **Actor notifies are actors**: they have relevancy, they cost replication
|
||||
consideration if not marked otherwise, and they can outlive their owner if
|
||||
removal is missed.
|
||||
|
||||
---
|
||||
|
||||
## 4. The three ways a cue gets addressed
|
||||
|
||||
```text
|
||||
1. GameplayEffect GE lists cue tags; ASC fires them with the effect
|
||||
2. Direct ASC call ExecuteGameplayCue / AddGameplayCue / RemoveGameplayCue
|
||||
3. IGameplayCueInterface the actor handles cue tags itself
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Prefer route 1 for anything tied to an effect's lifetime.** The effect and
|
||||
its presentation then start and stop together by construction.
|
||||
2. **Route 2 is for presentation with no effect behind it** — a UI-driven flash, a
|
||||
client-local confirmation. It is manual, so its removal is manual too.
|
||||
3. **Route 3 is for actor-specific handling** that no generic notify can express.
|
||||
It couples the actor to the cue vocabulary; use sparingly.
|
||||
4. **Do not mix routes for one cue tag.** Two owners of the same presentation
|
||||
produces double effects that appear only when both paths trigger.
|
||||
|
||||
---
|
||||
|
||||
## 5. The network rule
|
||||
|
||||
**Cues never run on a dedicated server.** The presentation layer does not exist
|
||||
there.
|
||||
|
||||
Consequences that must be designed for, not discovered:
|
||||
|
||||
1. **A cue may not carry gameplay side effects.** Damage, score, state changes,
|
||||
spawning — none of it may live in a notify, because on a dedicated server that
|
||||
code never executes. This is the single most important rule in this document.
|
||||
2. **Cue parameters are a lossy channel.** They are packed for replication;
|
||||
whatever your conversion helper drops is dropped silently on every client.
|
||||
Verify the round trip for every field you rely on, especially tag containers.
|
||||
3. **Cue delivery is unreliable.** Packet loss, relevancy changes and joining
|
||||
mid-effect all produce missed cues. Design for a missed cue being invisible,
|
||||
not desynchronising.
|
||||
4. **Cue assets belong in the client asset bundle only.** Putting them in the
|
||||
server bundle loads presentation content on a machine that will never use it.
|
||||
|
||||
### Test that the rule holds
|
||||
|
||||
Run the feature on a dedicated server with the presentation content deliberately
|
||||
absent. Gameplay results must be byte-identical.
|
||||
|
||||
---
|
||||
|
||||
## 6. Registering cue paths from feature plugins
|
||||
|
||||
A feature plugin that ships cue assets must register its content directory with
|
||||
the cue manager, or its cues are invisible.
|
||||
|
||||
The important structural point: **cue path registration must happen earlier than
|
||||
feature activation.** The manager builds its tag-to-asset index during its object
|
||||
library scan; a path added after that scan is not indexed until the next rescan.
|
||||
This is earlier than any of the standard activation hooks.
|
||||
|
||||
The pattern that solves it:
|
||||
|
||||
> **The feature action is a pure data declaration; a lifecycle observer is the
|
||||
> executor.**
|
||||
|
||||
The action object declares only the directory list and validates it in the
|
||||
editor. A separate observer, registered with the feature policy, listens for the
|
||||
earliest lifecycle phase and performs the registration for every action of that
|
||||
type it finds.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Register on the earliest available phase**, not on activation.
|
||||
2. **Batch the rescan.** Add all paths, then rebuild the index once — one rescan
|
||||
per feature, not one per directory.
|
||||
3. **Refresh the primary-asset registration when the path set changes**, on the
|
||||
way in *and* on the way out. Skipping it on removal leaves the asset manager
|
||||
describing directories that no longer exist.
|
||||
4. **Implement removal and assert it.** Count the paths removed and compare
|
||||
against the paths added; a mismatch means the registry is drifting across
|
||||
activation cycles.
|
||||
5. **Use the same manager instance on both sides of the lifecycle.** Registering
|
||||
through your subclassed manager and unregistering through the engine base
|
||||
class is a silent asymmetry — it compiles, runs, and leaves entries behind.
|
||||
|
||||
---
|
||||
|
||||
## 7. Loading and cost
|
||||
|
||||
Cue assets are discovered by scanning directories and are, by default, loaded on
|
||||
demand — which means a hitch the first time each cue fires.
|
||||
|
||||
The mitigation is preloading, and it is where most complexity in a cue manager
|
||||
lives: async library loading, a preload set, always-loaded tags, reference
|
||||
tracking, and hooks for GC and map transitions.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Decide the load policy explicitly and verify it is reachable.** A load-mode
|
||||
switch that short-circuits every branch renders the entire preload subsystem
|
||||
dead code — the machinery exists, is maintained, is never executed, and any
|
||||
diagnostic command reports zero because there is genuinely nothing there.
|
||||
2. **A constant is not a knob.** A local `const bool` inside a namespace whose
|
||||
name ends in `Cvars` is not a console variable and cannot be changed at
|
||||
runtime. If a policy must be tunable, register it as an actual console
|
||||
variable; otherwise do not imply that it is one.
|
||||
3. **Preload the small set that must never hitch** — the cues on the critical
|
||||
path, hit feedback above all — and let the rest load on demand.
|
||||
4. **Put cue paths in the client bundle only** (see §5).
|
||||
5. **Measure the cue set size before optimising it.** Discovery cost scales with
|
||||
the number of scanned directories; load cost scales with the assets actually
|
||||
referenced.
|
||||
|
||||
---
|
||||
|
||||
## 8. Teardown
|
||||
|
||||
Cues attach to an avatar. When the avatar changes — death, possession change,
|
||||
respawn — cues that are still active belong to nothing.
|
||||
|
||||
The correct teardown sequence when unbinding an ability system from an avatar:
|
||||
|
||||
```text
|
||||
1. verify the ASC's avatar is actually this actor
|
||||
2. cancel abilities
|
||||
3. clear ability input bindings
|
||||
4. remove all gameplay cues
|
||||
5. clear the avatar
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Check avatar identity before tearing down.** A component that unbinds
|
||||
unconditionally can rip the ASC away from an avatar it no longer owns.
|
||||
2. **Remove all cues explicitly.** Do not rely on the removal of the effects that
|
||||
added them; direct-call cues (route 2) have no effect behind them.
|
||||
3. **Test the cycle twice.** Leaks and duplicates appear on the second respawn,
|
||||
not the first.
|
||||
|
||||
---
|
||||
|
||||
## 9. The validation gap, and what to do about it
|
||||
|
||||
Nothing in the engine checks that:
|
||||
|
||||
- a cue tag referenced by an effect has a notify asset;
|
||||
- a notify asset's tag is referenced by anything;
|
||||
- a cue tag is spelled the way the notify expects.
|
||||
|
||||
All three failures present identically: no visual, no log, no error.
|
||||
|
||||
### Minimum guardrails
|
||||
|
||||
1. **Constrain cue tag properties** with `meta=(Categories="GameplayCue")` so
|
||||
authoring uses a dropdown.
|
||||
2. **Write a content validator** that resolves every cue tag in every effect
|
||||
against the manager's index, and reports both directions of orphan.
|
||||
3. **Log once, in development builds, when a cue fires with no handler.** This
|
||||
converts the silent case into a searchable one.
|
||||
4. **Keep cue tags in one namespace with one owner.** A cue tag placed in the
|
||||
wrong namespace because "the notify happens to live there" is a rename waiting
|
||||
to happen — and renames need redirects (see `ue-gameplay-tag-governance`).
|
||||
|
||||
---
|
||||
|
||||
## 10. Review checklist
|
||||
|
||||
- [ ] No gameplay logic inside any notify.
|
||||
- [ ] Feature works identically with all cue assets absent.
|
||||
- [ ] `OnActive` and `WhileActive` converge on the same visual state.
|
||||
- [ ] Everything spawned in the duration forms is destroyed in `OnRemove`.
|
||||
- [ ] Static notifies hold no per-instance state.
|
||||
- [ ] Actor notifies exist only where there is something to tear down.
|
||||
- [ ] One addressing route per cue tag.
|
||||
- [ ] Cue assets are in the client bundle only.
|
||||
- [ ] Cue parameter round-trip verified for every field used.
|
||||
- [ ] Feature-plugin cue paths registered at the earliest lifecycle phase.
|
||||
- [ ] Path removal implemented, asserted, and using the same manager instance.
|
||||
- [ ] Load policy is reachable, and any "knob" is a real console variable.
|
||||
- [ ] Teardown checks avatar identity and removes all cues.
|
||||
- [ ] A validator reports orphan tags and orphan notifies.
|
||||
|
||||
---
|
||||
|
||||
## 11. Tests
|
||||
|
||||
### Behaviour
|
||||
|
||||
- one-shot cue fires once per execution, not per client-side prediction retry;
|
||||
- looping cue starts on `OnActive`, is matched by `WhileActive` for a late
|
||||
joiner, and stops on `OnRemove`;
|
||||
- `OnRemove` without a preceding `OnActive` does not crash or leak;
|
||||
- cue with no registered notify produces no error and no side effect.
|
||||
|
||||
### Network
|
||||
|
||||
- dedicated server + remote client: gameplay results identical with and without
|
||||
cue content;
|
||||
- client joining mid-effect sees the same state as a client present at the start;
|
||||
- relevancy lost and regained during a looping cue converges to the correct state.
|
||||
|
||||
### Lifecycle
|
||||
|
||||
- death and respawn twice: no duplicated or orphaned cues;
|
||||
- feature plugin activate, deactivate, activate: cue path count returns to its
|
||||
original value;
|
||||
- possession change mid-cue.
|
||||
|
||||
### Cost
|
||||
|
||||
- first-fire hitch measured for a non-preloaded cue;
|
||||
- preloaded set size and residency reported by a diagnostic command that has been
|
||||
verified to read a live code path.
|
||||
|
||||
---
|
||||
|
||||
## 12. When not to use a cue
|
||||
|
||||
Use something else when:
|
||||
|
||||
- the effect must be guaranteed — it is gameplay, not presentation;
|
||||
- the receiver needs a return value or acknowledgement — direct call;
|
||||
- the event drives UI state that must survive a missed message — replicated
|
||||
state plus a local message (`ue-gameplay-messaging`);
|
||||
- exactly one known actor reacts, in the same module — direct call or delegate.
|
||||
|
||||
A cue earns its cost when presentation is genuinely optional, genuinely
|
||||
many-to-many, and genuinely absent on the server. Used for anything else, it is a
|
||||
tag join with no validation and unreliable delivery — which is a description of a
|
||||
bug that has not happened yet.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The failure modes and the worked analysis in `references/` come from a
|
||||
line-by-line audit of Epic's Lyra Starter Game on Unreal Engine 5.6, read as
|
||||
source rather than run. The cue subsystem there is unusually instructive because
|
||||
much of its preload machinery is present, maintained, and unreachable in the
|
||||
shipped configuration — a shape that reading the class list will never reveal.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; what ships
|
||||
is the detection recipe. Each entry carries a stable identifier (`GC-01` and up)
|
||||
that resolves back to the audited location, so any specific claim can be produced
|
||||
on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one project on one engine version. It is evidence
|
||||
of failure modes, not a guarantee about other versions. Re-run the detection
|
||||
recipes against your own tree before acting on any specific claim.
|
||||
@@ -0,0 +1,665 @@
|
||||
# Failure modes: GameplayCues
|
||||
|
||||
Twenty ways a cue-based presentation layer fails without saying so.
|
||||
|
||||
The shared property of this list, stated once so it is not repeated twenty times:
|
||||
**a cue defect is silent by construction.** The mechanism is a tag join with no
|
||||
validation, delivering an event that is *allowed* to be missed, to a layer that
|
||||
does not exist on the server. Every diagnostic has to be built deliberately,
|
||||
because the engine will not volunteer one.
|
||||
|
||||
Each entry gives the mechanism, why it stays silent, why the obvious check misses
|
||||
it, the symptom a human reports, a `Detect` recipe you can run against your own
|
||||
tree, and the guardrail. Identifiers (`GC-01` and up) are stable and resolve back
|
||||
to the audited source in the research archive.
|
||||
|
||||
Detection recipes use `rg` (ripgrep). They were run against the audited project
|
||||
while this file was written; where a first-draft recipe gave a wrong answer, the
|
||||
corrected form is the one printed here and the trap is called out.
|
||||
|
||||
---
|
||||
|
||||
### GC-01 - Gameplay logic inside a notify
|
||||
|
||||
**Mechanism.** Damage, scoring, state changes or spawning placed in a cue notify,
|
||||
because the notify is where the designer was already working.
|
||||
|
||||
**Why it is silent.** In the editor and on a listen server the notify runs on the
|
||||
same process that owns gameplay, so the write lands and the feature appears
|
||||
correct. Nothing anywhere declares that a notify is presentation-only.
|
||||
|
||||
**Why the obvious check misses it.** Code review of the notify reads as reasonable
|
||||
gameplay code, and it *is* reasonable — in the wrong file. The defect is not in the
|
||||
statement, it is in the file the statement is in. No compiler, linter or test that
|
||||
runs in-editor can see it.
|
||||
|
||||
**Symptom.** The feature works for everyone during development and simply does not
|
||||
happen on a dedicated server. Reported as "the server build is broken".
|
||||
|
||||
**Detect.** Look for writes inside notify classes:
|
||||
|
||||
```bash
|
||||
rg -n --type cpp -g '*Cue*' \
|
||||
"(ApplyGameplayEffect|AddScore|->Server|SpawnActor|SetHealth|=\s*true;)" \
|
||||
Source/
|
||||
```
|
||||
|
||||
Stronger, and the one that actually settles it: run the feature on a dedicated
|
||||
server with every cue asset removed. Gameplay results must be identical.
|
||||
|
||||
**Guardrail.** A notify may read state and produce audiovisual output, nothing
|
||||
else. Review every notify for writes, and keep the dedicated-server-without-cues
|
||||
run in the acceptance list for any feature that ships a cue.
|
||||
|
||||
---
|
||||
|
||||
### GC-02 - Cue tag with no notify asset
|
||||
|
||||
**Mechanism.** An effect references a cue tag; no notify is registered for it —
|
||||
a typo, an unscanned directory, or a plugin that failed to register its path.
|
||||
|
||||
**Why it is silent.** An unhandled cue is a legal runtime state: the engine
|
||||
delivers the event to a registry that has no entry, and returns. Missing
|
||||
presentation is exactly what "presentation is optional" means, so there is nothing
|
||||
for the engine to complain about.
|
||||
|
||||
**Why the obvious check misses it.** The tag exists, the effect compiles, the asset
|
||||
cooks. Both halves of the join are individually valid; only their *relation* is
|
||||
broken, and nothing in the toolchain owns that relation.
|
||||
|
||||
**Symptom.** Silence. No visual, no log, no error. The artist is told the effect
|
||||
"doesn't work" and starts editing a notify that was never being called.
|
||||
|
||||
**Detect.** Compare the two sides of the join. Tags referenced by content versus
|
||||
tags claimed by notifies:
|
||||
|
||||
```bash
|
||||
rg -no "GameplayCue\.[A-Za-z0-9_.]+" Content/ Source/ | sort -u
|
||||
```
|
||||
|
||||
Feed both sets into a diff. Then prove your tooling works by deliberately
|
||||
misspelling one cue tag and confirming the check reports it — a validator nobody
|
||||
has seen fail is not known to work.
|
||||
|
||||
**Guardrail.** Constrain cue tag properties with `meta=(Categories="GameplayCue")`
|
||||
so authoring uses a dropdown; add a content validator that resolves every
|
||||
referenced tag against the runtime index; log once per unhandled cue in
|
||||
development builds.
|
||||
|
||||
---
|
||||
|
||||
### GC-03 - Notify asset that no tag references
|
||||
|
||||
**Mechanism.** The reverse orphan: a notify exists for a tag nobody fires,
|
||||
usually left behind by a rename or a cancelled feature.
|
||||
|
||||
**Why it is silent.** An unreferenced notify is still scanned, indexed and
|
||||
sometimes preloaded. It behaves exactly like a notify waiting for its first
|
||||
trigger, which is a normal state.
|
||||
|
||||
**Why the obvious check misses it.** Reference-viewer style checks ask "what does
|
||||
this asset reference?", not "who addresses this asset by string?" Nothing
|
||||
references a notify by hard pointer — that is the entire point of the design.
|
||||
|
||||
**Symptom.** No symptom at all, which is why it accumulates. It surfaces as
|
||||
unexplained package count and memory in the cue set.
|
||||
|
||||
**Detect.** Same recipe as GC-02, read in the other direction: tags claimed by
|
||||
notify assets minus tags referenced by effects. Run the validator both ways in one
|
||||
pass, and report both orphan directions separately.
|
||||
|
||||
**Guardrail.** One validator, two directions, in CI. Orphan notifies are deleted
|
||||
in the same commit that finds them, or the list stops being read.
|
||||
|
||||
---
|
||||
|
||||
### GC-04 - Per-instance state on a Static notify
|
||||
|
||||
**Mechanism.** A Static notify runs on the class default object. A member written
|
||||
during one invocation persists into every later invocation, including other
|
||||
players'.
|
||||
|
||||
**Why it is silent.** Single-player and single-source testing never overlaps two
|
||||
invocations, so the shared member always holds the value the current caller just
|
||||
wrote. The bug requires concurrency to appear.
|
||||
|
||||
**Why the obvious check misses it.** The class looks like an ordinary object with
|
||||
ordinary members. Nothing at the declaration site says "this is a singleton" — the
|
||||
CDO behaviour comes from how the cue manager dispatches, which is in a different
|
||||
file entirely.
|
||||
|
||||
**Symptom.** One player's cue reads another player's data. Effects that are correct
|
||||
alone and wrong in multiplayer, intermittently.
|
||||
|
||||
**Detect.** Find mutable members on Static notify classes:
|
||||
|
||||
```bash
|
||||
rg -n -A40 "class .*GameplayCueNotify_Static" Source/ \
|
||||
| rg -n "^\s*(float|int32|bool|F[A-Z]\w+|T\w+<)[^()]*;"
|
||||
```
|
||||
|
||||
Any non-const, non-config member on a Static notify is the defect.
|
||||
|
||||
**Guardrail.** Static notifies hold no mutable state. If state is needed, it is an
|
||||
Actor notify — that is what the distinction is for.
|
||||
|
||||
---
|
||||
|
||||
### GC-05 - Looping cue that is never removed
|
||||
|
||||
**Mechanism.** `OnActive`/`WhileActive` spawns a persistent effect; `OnRemove`
|
||||
does not destroy it, or never runs because the owning effect was removed by a path
|
||||
that skipped cue removal.
|
||||
|
||||
**Why it is silent.** A spawned particle system is a legitimate long-lived object.
|
||||
Nothing distinguishes "still playing because the effect is active" from "still
|
||||
playing because nobody stopped it".
|
||||
|
||||
**Why the obvious check misses it.** The first death-respawn cycle usually looks
|
||||
clean, because the leak needs a second cycle to become visible as duplication.
|
||||
Most manual testing stops at one cycle.
|
||||
|
||||
**Symptom.** Accumulating VFX and audio across respawns; frame time degrading over
|
||||
a match; a bug report that says "it gets worse the longer you play".
|
||||
|
||||
**Detect.** For each duration-form notify, check that the removal path destroys
|
||||
what the activation path spawned:
|
||||
|
||||
```bash
|
||||
rg -n "OnActive|WhileActive|OnRemove" Source/ -g '*Cue*' -A15 \
|
||||
| rg -n "(SpawnSystem|SpawnEmitter|SpawnSound|Destroy|Deactivate)"
|
||||
```
|
||||
|
||||
Then count active cue actors before and after two full death-respawn cycles. The
|
||||
count must return to baseline, not merely stop growing.
|
||||
|
||||
**Guardrail.** Everything the duration forms spawn, `OnRemove` destroys. Write
|
||||
`OnRemove` to be safe when `OnActive` never ran on that client.
|
||||
|
||||
---
|
||||
|
||||
### GC-06 - `OnActive` without `WhileActive`
|
||||
|
||||
**Mechanism.** Start-of-loop logic placed only in `OnActive`, which fires only for
|
||||
clients present at the moment the effect began.
|
||||
|
||||
**Why it is silent.** For every client who was present, the behaviour is perfect.
|
||||
The failure exists only in the experience of a client who was not, and that client
|
||||
sees an absence — the hardest thing to notice.
|
||||
|
||||
**Why the obvious check misses it.** Both callbacks exist on the class, so a
|
||||
structural review ticks the box. The difference between them is a delivery-timing
|
||||
property of the engine, not a property of the code you are reading.
|
||||
|
||||
**Symptom.** Players who join, or who lose and regain relevancy, see nothing while
|
||||
everyone else sees the effect. Reproduces only with a real client join, never in a
|
||||
single-process test.
|
||||
|
||||
**Detect.** Find notifies that implement one callback and not the other:
|
||||
|
||||
```bash
|
||||
rg -l "OnActive" Source/ -g '*Cue*' > /tmp/active.txt
|
||||
rg -l "WhileActive" Source/ -g '*Cue*' > /tmp/while.txt
|
||||
comm -23 <(sort /tmp/active.txt) <(sort /tmp/while.txt)
|
||||
```
|
||||
|
||||
Anything listed implements the start and not the catch-up.
|
||||
|
||||
**Guardrail.** `OnActive` and `WhileActive` must converge on the same visual state.
|
||||
Test by joining a match mid-effect, and by losing and regaining relevancy during a
|
||||
loop.
|
||||
|
||||
---
|
||||
|
||||
### GC-07 - Two addressing routes for one tag
|
||||
|
||||
**Mechanism.** The same cue tag is fired both by a gameplay effect and by a direct
|
||||
ability-system call, or additionally handled through the cue interface on the
|
||||
actor.
|
||||
|
||||
**Why it is silent.** Each route is individually correct and each was added by
|
||||
someone solving a real problem. The duplication is a property of the pair, which
|
||||
no single author sees.
|
||||
|
||||
**Why the obvious check misses it.** Searching for the tag finds both sites, but
|
||||
they look like they belong to different features — one in an effect asset, one in
|
||||
C++. Nothing marks a cue tag as having an owner.
|
||||
|
||||
**Symptom.** Doubled effects, appearing only when both paths trigger, often only in
|
||||
one game mode. Reported as "the explosion is twice as loud in Team Deathmatch".
|
||||
|
||||
**Detect.** For each cue tag, enumerate every firing route:
|
||||
|
||||
```bash
|
||||
TAG="GameplayCue.Fire.Impact"
|
||||
rg -n "$TAG" Source/ Content/ Config/
|
||||
rg -n "(ExecuteGameplayCue|AddGameplayCue|RemoveGameplayCue)" Source/ -A2 | rg -n "$TAG"
|
||||
```
|
||||
|
||||
More than one owner is the defect, even when the current behaviour looks right.
|
||||
|
||||
**Guardrail.** One owner per cue tag, documented at the tag declaration. Search
|
||||
every route before adding a firing site.
|
||||
|
||||
---
|
||||
|
||||
### GC-08 - Manual cue added without a manual removal
|
||||
|
||||
**Mechanism.** A direct add call has no effect behind it, so no effect removal will
|
||||
ever clean it up.
|
||||
|
||||
**Why it is silent.** The add succeeds and the presentation appears. Removal is a
|
||||
separate call that nobody is forced to write, and the object model does not record
|
||||
that a removal is owed.
|
||||
|
||||
**Why the obvious check misses it.** The code that adds is present and correct. The
|
||||
defect is the absence of a second call on every exit path, and absence is what
|
||||
review is worst at seeing.
|
||||
|
||||
**Symptom.** A looping cue that survives the state that created it — including
|
||||
across death, possession change and disconnect.
|
||||
|
||||
**Detect.** Count adds against removes:
|
||||
|
||||
```bash
|
||||
rg -c "AddGameplayCue" Source/
|
||||
rg -c "RemoveGameplayCue" Source/
|
||||
```
|
||||
|
||||
An imbalance is a lead, not a verdict — then check each add site for a paired
|
||||
removal on *every* exit path, not just the happy one.
|
||||
|
||||
**Guardrail.** Direct-call cues need an explicit paired removal on every exit path.
|
||||
Prefer attaching the cue to an effect's lifetime instead, so removal is structural
|
||||
rather than remembered.
|
||||
|
||||
---
|
||||
|
||||
### GC-09 - Teardown without an avatar identity check
|
||||
|
||||
**Mechanism.** A component unbinds an ability system component unconditionally,
|
||||
without checking that the ASC's current avatar is still this actor.
|
||||
|
||||
**Why it is silent.** The unbind succeeds. The damage is done to a *different*
|
||||
actor's state — the one the ASC has already moved to — so the error surfaces far
|
||||
from its cause.
|
||||
|
||||
**Why the obvious check misses it.** The teardown code is correct in isolation and
|
||||
correct in the common case. It is wrong only in the ordering where the avatar
|
||||
changed first, which single-process testing rarely produces.
|
||||
|
||||
**Symptom.** Abilities and cues on the *current* avatar are cancelled when an old
|
||||
one is cleaned up. Looks like a spontaneous ability cancellation with no cause.
|
||||
|
||||
**Detect.** Find unbind paths and check for an identity guard above them:
|
||||
|
||||
```bash
|
||||
rg -n -B8 "(ClearAbilityInput|CancelAbilities|RemoveAllGameplayCues|SetAvatarActor\(nullptr)" Source/ \
|
||||
| rg -n "GetAvatarActor\(\)\s*==|IsValid\(.*Avatar"
|
||||
```
|
||||
|
||||
An unbind path with no `GetAvatarActor()` comparison above it is the defect.
|
||||
|
||||
**Guardrail.** Verify the ASC's avatar is this actor before cancelling abilities,
|
||||
clearing input and removing cues. The audited project does this correctly; the
|
||||
guard is one comparison and it is worth copying verbatim.
|
||||
|
||||
---
|
||||
|
||||
### GC-10 - Cue assets in the server bundle
|
||||
|
||||
**Mechanism.** Cue content is assigned to an asset bundle that the dedicated
|
||||
server loads.
|
||||
|
||||
**Why it is silent.** Loading extra assets is never an error. The server runs
|
||||
correctly, just fatter, and nothing reports "you loaded a particle system you can
|
||||
never play".
|
||||
|
||||
**Why the obvious check misses it.** Bundle assignment lives in asset manager
|
||||
configuration, not next to the cue content. Reviewing the cue does not show you
|
||||
which bundle it is in.
|
||||
|
||||
**Symptom.** Server memory spent on particles and audio. Discovered, if ever,
|
||||
during a memory investigation months later.
|
||||
|
||||
**Detect.** Find where cue paths are registered and which bundle constant is used:
|
||||
|
||||
```bash
|
||||
rg -n "AddGameplayCueNotifyPath|GameplayCueRefs|LoadState" Source/ -B3 -A3
|
||||
```
|
||||
|
||||
The bundle name at that site should be the client bundle. Cross-check against
|
||||
the asset manager's bundle definitions.
|
||||
|
||||
**Guardrail.** Cue paths belong to the client bundle only. Make the bundle constant
|
||||
explicit at the registration site so review can see it without opening
|
||||
configuration.
|
||||
|
||||
---
|
||||
|
||||
### GC-11 - Load policy that short-circuits the whole subsystem
|
||||
|
||||
**Mechanism.** A load-mode constant causes an early return at every switch site,
|
||||
including the one that binds the subsystem's delegates.
|
||||
|
||||
**Why it is silent.** Every branch is a legal configuration. "Load everything
|
||||
upfront" is a sensible mode; the fact that choosing it disables an entire preload
|
||||
subsystem is an emergent property of three separate switch statements, none of
|
||||
which says so.
|
||||
|
||||
**Why the obvious check misses it.** The preload machinery exists, is maintained,
|
||||
and is full of correct code. Reading it tells you nothing about whether it runs.
|
||||
Worse, the diagnostic command built for it reports **zero**, which reads as "there
|
||||
is nothing to preload" rather than "this code never executes".
|
||||
|
||||
**Symptom.** A named feature that has no effect. Engineers investigate the preload
|
||||
implementation for a bug that is not in it.
|
||||
|
||||
**Detect.** Find the mode constant, then find every switch on it, then check which
|
||||
branch each falls into:
|
||||
|
||||
```bash
|
||||
rg -n "static .*LoadMode\s*=" Source/
|
||||
rg -n "switch\s*\(.*LoadMode" Source/ -A12
|
||||
```
|
||||
|
||||
If the configured value hits an early `return` at a site that also performs
|
||||
delegate binding, everything below that binding is unreachable. In the audited
|
||||
project three switch sites short-circuit, and the third returns before its
|
||||
delegate bindings.
|
||||
|
||||
**Guardrail.** Verify the load policy reaches its intended branch. Treat "the dump
|
||||
command shows nothing" as a question about reachability, not about content.
|
||||
|
||||
---
|
||||
|
||||
### GC-12 - A constant presented as a runtime knob
|
||||
|
||||
**Mechanism.** A local `const bool` or a `static` variable placed in a namespace
|
||||
named after console variables, with no console-variable registration.
|
||||
|
||||
**Why it is silent.** It compiles and behaves as the constant it is. The naming is
|
||||
the only thing that lies, and naming is not checked by anything.
|
||||
|
||||
**Why the obvious check misses it.** The namespace name contains the word an
|
||||
engineer greps for. Searching for the knob finds the declaration, which looks
|
||||
exactly like the seven real registrations elsewhere in the project.
|
||||
|
||||
**Symptom.** Engineers try to change behaviour from the console and cannot. Time is
|
||||
lost concluding that the console, the build, or the module is broken.
|
||||
|
||||
**Detect.** Compare pseudo-knobs against real registrations:
|
||||
|
||||
```bash
|
||||
rg -n "^namespace \w*Cvar\w*" Source/ # namespaces that promise knobs
|
||||
rg -c "FAutoConsoleVariableRef" Source/ # registrations that deliver them
|
||||
```
|
||||
|
||||
Note the trap: searching for the literal `namespace Cvars` finds nothing, because
|
||||
real names are prefixed (`FooManagerCvars`). The first draft of this recipe
|
||||
returned zero hits and would have been read as "no such problem here". In the
|
||||
audited project the corrected recipe finds one such namespace against eight real
|
||||
registrations, and inside that namespace only a console *command* is registered —
|
||||
the two variables beside it are plain constants.
|
||||
|
||||
**Guardrail.** If it must be tunable, register it as a real console variable.
|
||||
Otherwise do not name its scope after one.
|
||||
|
||||
---
|
||||
|
||||
### GC-13 - A collection that is iterated but never filled
|
||||
|
||||
**Mechanism.** A local container is declared, immediately looped over, and never
|
||||
populated — usually with a comment where the population belongs.
|
||||
|
||||
**Why it is silent.** A loop over an empty container is a no-op, which is
|
||||
indistinguishable from a loop whose input happens to be empty right now.
|
||||
|
||||
**Why the obvious check misses it.** The loop body is real code, often including
|
||||
error handling and logging, so the feature reads as implemented. There is even a
|
||||
warning path for invalid entries — which can never fire, because there are no
|
||||
entries.
|
||||
|
||||
**Symptom.** A named feature ("always loaded cues") that has no inputs and
|
||||
therefore no effect, while appearing complete in review.
|
||||
|
||||
**Detect.** Find containers declared and iterated within a few lines, with no
|
||||
insertion between:
|
||||
|
||||
```bash
|
||||
rg -n -A6 "T(Array|Set)<\w+>\s+\w+;" Source/ | rg -n "for\s*\("
|
||||
```
|
||||
|
||||
For each hit, search the same scope for `Add|Emplace|Append|Insert` on that name.
|
||||
None means the loop is decorative.
|
||||
|
||||
**Guardrail.** A loop over an empty local is a review blocker. If the population is
|
||||
future work, the loop is future work too.
|
||||
|
||||
---
|
||||
|
||||
### GC-14 - Feature-plugin cue path registered too late
|
||||
|
||||
**Mechanism.** Registration is performed on feature activation, after the manager
|
||||
has already built its tag-to-asset index.
|
||||
|
||||
**Why it is silent.** Registration succeeds. The path is in the list. It is simply
|
||||
not in the *index*, and no error distinguishes those two states.
|
||||
|
||||
**Why the obvious check misses it.** In the editor, incidental rescans — asset
|
||||
saves, hot reloads, content browser operations — rebuild the index often enough
|
||||
that the cues appear to work. The failure is specific to a cold packaged run.
|
||||
|
||||
**Symptom.** A plugin's cues are invisible in a packaged build and fine in the
|
||||
editor. The classic "works on my machine" that is actually "works in my editor".
|
||||
|
||||
**Detect.** Find where cue paths are added and which lifecycle phase the call sits
|
||||
in:
|
||||
|
||||
```bash
|
||||
rg -n "AddGameplayCueNotifyPath" Source/ -B12 \
|
||||
| rg -n "(OnGameFeatureRegistering|OnGameFeatureLoading|OnGameFeatureActivating|Activate)"
|
||||
```
|
||||
|
||||
`Registering` is early enough; `Activating` is not. Also check the rescan flag: one
|
||||
rescan per feature, not one per directory.
|
||||
|
||||
**Guardrail.** Register on the earliest lifecycle phase available, via a policy
|
||||
observer rather than an activation hook. Batch the rescan: add all paths, rebuild
|
||||
once. The audited project does exactly this, and its feature action deliberately
|
||||
has no activation body at all.
|
||||
|
||||
---
|
||||
|
||||
### GC-15 - Asymmetric manager access across the lifecycle
|
||||
|
||||
**Mechanism.** Registration goes through a subclassed manager; removal goes through
|
||||
the engine base class.
|
||||
|
||||
**Why it is silent.** Both calls compile and both succeed, because the subclass
|
||||
inherits the base method. As long as the subclass adds no bookkeeping, the two
|
||||
paths are equivalent — today.
|
||||
|
||||
**Why the obvious check misses it.** The two lines are in different functions,
|
||||
often hundreds of lines apart, and both read as "get the manager, call the
|
||||
method". The asymmetry is only visible when you put the two resolution expressions
|
||||
side by side.
|
||||
|
||||
**Symptom.** Nothing, until the subclass gains state. Then removal silently leaves
|
||||
entries behind and the registry drifts across activation cycles.
|
||||
|
||||
**Detect.** Put the add and remove sites next to each other:
|
||||
|
||||
```bash
|
||||
rg -n "(AddGameplayCueNotifyPath|RemoveGameplayCueNotifyPath)" Source/ -B4 \
|
||||
| rg -n "(GetGameplayCueManager|UAbilitySystemGlobals|Cast<)"
|
||||
```
|
||||
|
||||
Different resolution expressions on the two sides is the defect.
|
||||
|
||||
**Guardrail.** Resolve the manager the same way on both sides, and assert the
|
||||
removal count. The audited project asserts the count — that assertion is the
|
||||
mitigating factor and the part worth copying.
|
||||
|
||||
---
|
||||
|
||||
### GC-16 - Registry refresh on the way in but not on the way out
|
||||
|
||||
**Mechanism.** The primary-asset registration is refreshed when paths are added and
|
||||
not when they are removed.
|
||||
|
||||
**Why it is silent.** A stale registration describes directories that no longer
|
||||
exist. Asset lookups against them fail the same way a genuinely empty directory
|
||||
fails: nothing found, no error.
|
||||
|
||||
**Why the obvious check misses it.** The refresh call on the add side is visible
|
||||
and correct. Its absence on the remove side is one missing line in a function that
|
||||
otherwise does everything right, including asserting its own removal count.
|
||||
|
||||
**Symptom.** Divergence that accumulates across feature activation cycles. Path
|
||||
counts do not return to baseline after activate, deactivate, activate.
|
||||
|
||||
**Detect.** Count refresh calls against path-set mutations:
|
||||
|
||||
```bash
|
||||
rg -n "RefreshGameplayCuePrimaryAsset" Source/ -B15 \
|
||||
| rg -n "(Registering|Unregistering|Add|Remove)"
|
||||
```
|
||||
|
||||
A refresh on the registering path with no counterpart on unregistering is the
|
||||
defect. Confirm behaviourally: activate, deactivate, activate a feature and
|
||||
compare path counts.
|
||||
|
||||
**Guardrail.** Refresh on every path-set change, in both directions.
|
||||
|
||||
---
|
||||
|
||||
### GC-17 - Lossy payload conversion helper
|
||||
|
||||
**Mechanism.** A helper that converts between a message struct and cue parameters
|
||||
drops fields.
|
||||
|
||||
**Why it is silent.** The conversion produces a valid, well-formed result. The
|
||||
missing field arrives as its default value, which for a tag container is empty —
|
||||
and empty is a legal value that means "no tags".
|
||||
|
||||
**Why the obvious check misses it.** The function has a name that promises a
|
||||
conversion, a signature that mentions both types, and a body that assigns most
|
||||
fields. Review confirms "yes, this converts". Only a field-by-field round trip
|
||||
shows the hole.
|
||||
|
||||
**Symptom.** A field populated at the sender and empty at the receiver, on every
|
||||
client, with no warning. Investigated as a replication bug for as long as it takes
|
||||
to read the helper.
|
||||
|
||||
**Detect.** Read every conversion function field by field, and treat every
|
||||
unfinished-work comment inside one as a declaration that the field is not carried:
|
||||
|
||||
```bash
|
||||
rg -n -A20 "(ToCueParameters|FromCueParameters|VerbMessageTo|CueParametersTo)" Source/ \
|
||||
| rg -n "(TODO|FIXME|\?\?\?)"
|
||||
```
|
||||
|
||||
In the audited project this finds the same field dropped in both directions, each
|
||||
marked with its own comment.
|
||||
|
||||
**Guardrail.** Verify the round trip field by field before relying on any field.
|
||||
Treat any unfinished-work comment in a conversion function as "this field is not
|
||||
carried".
|
||||
|
||||
---
|
||||
|
||||
### GC-18 - Concluding "unused" from absent C++ callers
|
||||
|
||||
**Mechanism.** A Blueprint-callable function's call sites can live in binary asset
|
||||
graphs, which are not text-searchable.
|
||||
|
||||
**Why it is silent.** The search is clean. Zero hits is a definite-looking answer,
|
||||
and definite-looking answers are the ones people act on.
|
||||
|
||||
**Why the obvious check misses it.** The obvious check *is* the problem: `rg`
|
||||
searches text, and the callers are not text. The tool answers the question it was
|
||||
asked — "where does this string appear in source?" — which is not the question the
|
||||
engineer meant.
|
||||
|
||||
**Symptom.** Code deleted as dead breaks a Blueprint at runtime, discovered after
|
||||
the fact, usually by a designer.
|
||||
|
||||
**Detect.** For any Blueprint-exposed symbol, treat source search as inconclusive
|
||||
by construction:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\(.*Blueprint(Callable|ImplementableEvent|NativeEvent)" Source/ -A2
|
||||
```
|
||||
|
||||
Anything on that list needs the editor's reference viewer before removal. An empty
|
||||
`rg` result proves the pattern, not the absence.
|
||||
|
||||
**Guardrail.** "No C++ callers" is not evidence of disuse for a Blueprint-exposed
|
||||
symbol. Record it as an open question, not as dead code — which is exactly what the
|
||||
audit did with the helpers from GC-17.
|
||||
|
||||
---
|
||||
|
||||
### GC-19 - Cue tag parked in the wrong namespace
|
||||
|
||||
**Mechanism.** A tag is declared next to the notify asset that happens to implement
|
||||
it, rather than in the namespace that owns the concept.
|
||||
|
||||
**Why it is silent.** The tag works. Its location is a naming decision, and naming
|
||||
decisions have no runtime behaviour — until someone tries to fix one.
|
||||
|
||||
**Why the obvious check misses it.** Nothing validates that a tag's namespace
|
||||
matches its meaning. The only record that the placement is wrong is, at best, a
|
||||
comment written by the person who did it.
|
||||
|
||||
**Symptom.** Moving it later requires a redirect. Without redirects, every asset
|
||||
referencing it silently loses the reference — the rename is not blocked, it is
|
||||
merely destructive.
|
||||
|
||||
**Detect.** Read the tag declarations for confessions, then check whether the
|
||||
project has any migration mechanism at all:
|
||||
|
||||
```bash
|
||||
rg -n "DevComment=.*(needs to move|wrong|temporary|should be)" Config/
|
||||
rg -c "GameplayTagRedirects" Config/
|
||||
```
|
||||
|
||||
In the audited project the first search finds a tag whose own comment says it is
|
||||
misplaced, and the second returns zero — the tag cannot be moved without breaking
|
||||
references.
|
||||
|
||||
**Guardrail.** Cue tags follow the concept, not the asset location. Add the
|
||||
redirect in the same commit as any move, and make sure a redirect mechanism exists
|
||||
before you need it.
|
||||
|
||||
---
|
||||
|
||||
### GC-20 - Editor measurements treated as representative of cue loading cost
|
||||
|
||||
**Mechanism.** The editor loads both the client and the server asset bundles.
|
||||
|
||||
**Why it is silent.** The measurement completes and produces a number. Numbers
|
||||
from a profiler carry an authority that their provenance does not.
|
||||
|
||||
**Why the obvious check misses it.** The condition that causes it is one clause in
|
||||
one branch of the bundle-selection logic, in a different subsystem from the one
|
||||
being measured. Nobody profiling cue residency reads the experience loader.
|
||||
|
||||
**Symptom.** Memory and load-time figures that describe no shipping configuration.
|
||||
Cue residency in particular is overstated, and budgets built on it are wrong in the
|
||||
safe direction until the day they are not.
|
||||
|
||||
**Detect.** Read the bundle-selection branch before trusting any in-editor asset
|
||||
memory number:
|
||||
|
||||
```bash
|
||||
rg -n "GIsEditor\s*\|\|" Source/ -A4
|
||||
```
|
||||
|
||||
If the editor branch unions both bundle sets, editor measurements are usable for
|
||||
direction of change only.
|
||||
|
||||
**Guardrail.** Use editor measurements for direction of change. Absolute cue
|
||||
loading cost requires a packaged build, and any budget quoted without one carries
|
||||
that caveat in writing.
|
||||
@@ -0,0 +1,200 @@
|
||||
# Patterns: a cue subsystem read end to end
|
||||
|
||||
A worked reading of one real GameplayCue subsystem, audited as source. It is here
|
||||
because the cue layer has a property that makes reading a class list useless:
|
||||
**large parts of it can be present, correct, maintained, and unreachable**, and
|
||||
nothing in the structure says which parts.
|
||||
|
||||
Individual failures are in [failure-modes.md](failure-modes.md), one entry each,
|
||||
with a recipe. This file is about the shapes — what the subsystem is for, where
|
||||
its cost lives, and which of its parts were actually running.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not settled by source reading, and left open.
|
||||
|
||||
---
|
||||
|
||||
## 1. The trade the mechanism makes
|
||||
|
||||
A cue replaces an asset reference with a string-like tag, and a registry resolves
|
||||
the tag by scanning content directories.
|
||||
|
||||
What you buy is real: gameplay assets carry no presentation payload, presentation
|
||||
ships in a separate plugin, and the whole layer is absent on a dedicated server.
|
||||
|
||||
What you pay is one specific thing, and it is worth naming precisely:
|
||||
|
||||
> **You have exchanged a reference the toolchain checks for a join the toolchain
|
||||
> does not check.**
|
||||
|
||||
A hard pointer that dangles is a cook error. A tag with no notify is silence.
|
||||
Both halves of the join stay individually valid — the tag exists, the asset
|
||||
exists — and no tool owns the relation between them. Everything in this file
|
||||
follows from that one exchange.
|
||||
|
||||
This is why cues are an application of tag governance with a loading subsystem
|
||||
attached, rather than a feature in their own right. Every rule about string joins
|
||||
applies here first, and the loading subsystem is where the surprises live.
|
||||
|
||||
---
|
||||
|
||||
## 2. Where the complexity actually is
|
||||
|
||||
Reading the audited manager, the code divides into three very unequal parts:
|
||||
|
||||
| Part | Size | Status in the shipped configuration |
|
||||
|---|---|---|
|
||||
| Dispatch: tag arrives, notify is invoked | small | running |
|
||||
| Discovery: scan directories, build the tag-to-asset index | moderate | running |
|
||||
| Preload: async library load, preload set, always-loaded tags, reference tracking, GC and map-transition hooks | **most of the file** | **unreachable** |
|
||||
|
||||
The third row is the finding, and it is a shape rather than a bug.
|
||||
|
||||
A load-mode constant is set to "load everything upfront" **[measured]**. That
|
||||
value short-circuits three separate switch sites **[measured]**, and the third of
|
||||
them returns before the statements that bind the subsystem's delegates
|
||||
**[measured]**.
|
||||
|
||||
The consequence **[derived]**: tag-loaded, post-garbage-collect and post-load-map
|
||||
handlers are never bound; the preload set, the always-loaded set and the reference
|
||||
tracker are never populated; and the project's own diagnostic command reports
|
||||
zero, forever.
|
||||
|
||||
**That last detail is the trap.** A diagnostic reporting zero reads as "there is
|
||||
nothing to preload". It actually means "this counter is dead". Anyone
|
||||
investigating cue memory starts from a number that is not measuring anything, and
|
||||
the natural next step — reading the preload implementation for a bug — is a day
|
||||
spent in correct code. Recipe: GC-11.
|
||||
|
||||
The transferable rule is not about cues at all:
|
||||
|
||||
> Before debugging a subsystem, establish that its code runs. A diagnostic that
|
||||
> reports zero is a claim about reachability until proven otherwise.
|
||||
|
||||
---
|
||||
|
||||
## 3. Knobs that are not knobs
|
||||
|
||||
In the same file, a namespace named after console variables contains three
|
||||
declarations **[measured]**. Exactly one — a console *command* — is registered
|
||||
with the engine. The other two, including the load mode from §2, are plain
|
||||
statics **[measured]**.
|
||||
|
||||
Project-wide the ratio is one such namespace against eight real console-variable
|
||||
registrations **[measured]**.
|
||||
|
||||
Two things make this worth an entry rather than a footnote:
|
||||
|
||||
1. **The naming is the entire defect.** The code is correct; it is a constant and
|
||||
behaves as one. What lies is the scope name, and nothing checks scope names.
|
||||
2. **The obvious search fails.** Grepping for the literal `namespace Cvars` finds
|
||||
nothing, because real names are prefixed with the owning type. The first draft
|
||||
of the detection recipe for this returned zero hits and would have been read as
|
||||
"we do not have this problem". The corrected recipe searches for any namespace
|
||||
whose name *contains* the token. This is recorded in GC-12 with the trap
|
||||
spelled out, because the trap generalises further than the finding does.
|
||||
|
||||
---
|
||||
|
||||
## 4. Feature-plugin registration: the one thing done right
|
||||
|
||||
Cue path registration from a feature plugin has a timing constraint that is easy
|
||||
to get wrong and invisible when you do: the manager builds its index during its
|
||||
object library scan, so a path added after that scan is not indexed until
|
||||
something triggers a rescan. In the editor, incidental rescans hide it. In a cold
|
||||
packaged run, nothing does. Recipe: GC-14.
|
||||
|
||||
The audited project solves it with a pattern worth copying verbatim
|
||||
**[measured]**:
|
||||
|
||||
> **The feature action is a pure data declaration; a lifecycle observer is the
|
||||
> executor.**
|
||||
|
||||
The action object declares the directory list and validates it in the editor — and
|
||||
has *no activation body at all* **[measured]**. A separate observer registered
|
||||
with the feature policy listens for the earliest lifecycle phase and performs the
|
||||
registration for every action of that type it finds. Paths are added with the
|
||||
rescan flag off, and a single index rebuild follows **[measured]**.
|
||||
|
||||
The empty activation body is the part people delete when adopting this, because an
|
||||
action with no `Activate` looks unfinished. It is not: it is the point.
|
||||
|
||||
### And the asymmetry that survived inside it
|
||||
|
||||
The same function pair is also the audit's best example of a near-miss
|
||||
**[measured]**:
|
||||
|
||||
| Direction | Manager resolution | Refresh call |
|
||||
|---|---|---|
|
||||
| Registering | project's own manager subclass | present |
|
||||
| Unregistering | engine base class | **absent** |
|
||||
|
||||
Both calls compile and both work, because the subclass inherits the method. They
|
||||
are equivalent exactly as long as the subclass adds no bookkeeping **[derived]**.
|
||||
And the refresh performed on the way in has no counterpart on the way out, so the
|
||||
asset manager keeps describing directories that have been removed **[derived]**.
|
||||
|
||||
Recipes: GC-15, GC-16.
|
||||
|
||||
**What redeems it, and what to actually copy:** the removal path counts what it
|
||||
removed and asserts the count against what was added **[measured]**. In an audit
|
||||
whose dominant finding across the whole project was "add works, remove is
|
||||
incomplete", this was the one place where teardown was both implemented and
|
||||
asserted. Copy the assertion; fix the asymmetry it sits next to.
|
||||
|
||||
---
|
||||
|
||||
## 5. The network rule, and why it is first
|
||||
|
||||
Cues never run on a dedicated server, because the presentation layer does not
|
||||
exist there.
|
||||
|
||||
Every consequence of that is a design constraint rather than a caution:
|
||||
|
||||
- a notify may not carry gameplay side effects — on a dedicated server that code
|
||||
never executes (GC-01);
|
||||
- cue parameters are a lossy channel, and what a conversion helper drops is
|
||||
dropped silently on every client. The audited helpers drop a context tag
|
||||
container in **both** directions, each with its own unfinished-work comment
|
||||
**[measured]** (GC-17);
|
||||
- delivery is unreliable by design: packet loss, relevancy changes and mid-effect
|
||||
joins all produce missed cues;
|
||||
- cue assets belong in the client bundle only.
|
||||
|
||||
The acceptance test that covers all four at once is cheap and nobody runs it: **run
|
||||
the feature on a dedicated server with every cue asset removed, and require
|
||||
byte-identical gameplay results.**
|
||||
|
||||
---
|
||||
|
||||
## 6. What source reading could not settle
|
||||
|
||||
Two limits, stated because they bound the claims above:
|
||||
|
||||
1. **Blueprint call sites are invisible to text search.** The conversion helpers
|
||||
in §5 have zero C++ callers and are Blueprint-callable **[measured]**. That is
|
||||
recorded as **[open]** — an open question, not dead code. Deleting them on the
|
||||
strength of an empty `rg` result is GC-18, committed by the person who wrote
|
||||
the audit. (It was not.)
|
||||
2. **Editor memory figures do not describe shipping.** The experience loader's
|
||||
bundle selection unions client and server bundles when running in the editor
|
||||
**[measured]**, so in-editor cue residency is overstated by whatever the server
|
||||
set contains. Editor measurements are good for direction of change and nothing
|
||||
else (GC-20).
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source
|
||||
rather than run, in a single workspace. Source addresses stay in the research
|
||||
archive that produced this skill; each `GC-` identifier resolves back to the
|
||||
audited location there, so any specific claim above can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. These are examples and failure
|
||||
evidence, not guarantees about other engine versions or other samples. Several
|
||||
claims are explicitly marked open and stay open. Re-run the recipes in
|
||||
[failure-modes.md](failure-modes.md) against your own tree before acting on
|
||||
anything here.
|
||||
@@ -0,0 +1,435 @@
|
||||
---
|
||||
name: ue-gameplay-messaging
|
||||
description: >-
|
||||
Design, implement or review a typed tag-addressed pub/sub message bus in
|
||||
Unreal Engine: subsystem scope and lifetime, struct type checks, exact versus
|
||||
hierarchical channel matching, listener handles and ownership, parity between
|
||||
the C++ and Blueprint APIs, the local-versus-network boundary, and migration
|
||||
from direct dependencies. Use when decoupling gameplay from UI, feedback,
|
||||
telemetry or processors, or when debugging duplicate, stale, mismatched or
|
||||
wrongly replicated messages.
|
||||
---
|
||||
|
||||
# UE gameplay messaging
|
||||
|
||||
The invariant:
|
||||
|
||||
> Messages announce facts inside one process. They do not own state, ordering,
|
||||
> persistence, authority or replication.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Seventeen detection recipes for silent bus failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting a bus implementation. Each entry names the
|
||||
mechanism, why it stays silent, why the obvious check misses it, and a command you
|
||||
can run against your own tree.
|
||||
|
||||
Related skills: `ue-gameplay-tag-governance`, `ue-ui-architecture`,
|
||||
`ue-gas-architecture`, `ue-multiplayer-authority`, `ue-data-driven-architecture`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Choose the bus scope explicitly
|
||||
|
||||
A bus implemented as a GameInstance subsystem:
|
||||
|
||||
- is shared by every world belonging to that GameInstance;
|
||||
- survives map travel;
|
||||
- is destroyed at GameInstance teardown;
|
||||
- is local to one process;
|
||||
- does **not** replicate.
|
||||
|
||||
That scope suits:
|
||||
|
||||
- gameplay event → HUD and feedback;
|
||||
- local processors and telemetry;
|
||||
- decoupling feature plugins from the base module;
|
||||
- cross-system notifications where no receiver owns the sender.
|
||||
|
||||
It is unsuitable as the source of truth for:
|
||||
|
||||
- inventory, team or score state;
|
||||
- guaranteed ordered workflows;
|
||||
- cross-process communication;
|
||||
- persistence;
|
||||
- anything an RPC should carry.
|
||||
|
||||
### The state-versus-event test
|
||||
|
||||
Ask: **"if a listener subscribes one second late, must it know the current
|
||||
value?"**
|
||||
|
||||
- **yes** → replicated property or state component, plus a change notification;
|
||||
- **no** → a transient message may be appropriate.
|
||||
|
||||
"An elimination occurred" is an event. "The current team score" is state. A bus
|
||||
that is asked to answer the second question will appear to work for every
|
||||
listener that happened to be subscribed early, which is every listener during
|
||||
development.
|
||||
|
||||
---
|
||||
|
||||
## 2. Storage model and handle identity
|
||||
|
||||
```text
|
||||
map: channel tag -> channel listener list
|
||||
|
||||
channel listener list
|
||||
├─ monotonically increasing handle ID (local to THIS channel)
|
||||
└─ listeners[]
|
||||
├─ callback
|
||||
├─ weak expected struct type
|
||||
├─ had-valid-type flag
|
||||
├─ handle ID
|
||||
└─ exact / partial match
|
||||
```
|
||||
|
||||
Registration returns a handle carrying the subsystem, the channel, and the
|
||||
channel-local ID. **The handle is the ownership receipt.** The caller, or an RAII
|
||||
wrapper, retains it and unregisters when its lifecycle ends.
|
||||
|
||||
### The handle identity rule
|
||||
|
||||
If IDs are allocated per channel, then **every internal removal must use both the
|
||||
registration channel and the ID.** Never search another channel by ID alone.
|
||||
|
||||
This is not a stylistic point. A hierarchical broadcast walks several channels in
|
||||
one call; code written inside that walk has two plausible variables in scope — the
|
||||
channel the message was sent on, and the channel currently being visited — and
|
||||
they are the same value only on the first iteration. Recipe: MB-01.
|
||||
|
||||
---
|
||||
|
||||
## 3. Typed broadcast
|
||||
|
||||
Template the C++ API and resolve the struct type at compile time on both sides.
|
||||
Store the expected type at registration.
|
||||
|
||||
On broadcast, the payload is deliverable when:
|
||||
|
||||
```text
|
||||
broadcast struct IsChildOf listener expected struct
|
||||
```
|
||||
|
||||
which permits a listener registered for a base struct to receive derived
|
||||
payloads, and makes exact equality a special case.
|
||||
|
||||
### Type mismatch policy
|
||||
|
||||
- log the channel, the sent struct and the expected struct;
|
||||
- skip that listener;
|
||||
- never reinterpret incompatible bytes;
|
||||
- never abort the rest of the broadcast;
|
||||
- add development-only assertions and tests.
|
||||
|
||||
A path that disables the type check entirely — registering with a null expected
|
||||
type — must be private and documented. If it is reachable from Blueprint through
|
||||
an empty type pin, it is not private. Recipe: MB-03.
|
||||
|
||||
### Payload lifetime
|
||||
|
||||
The bus should pass a pointer to the caller's stack object rather than copying.
|
||||
That is the right performance decision and it creates one hard rule:
|
||||
**a listener may not retain the payload pointer after its callback returns.**
|
||||
Copy what you need. Recipe: MB-04.
|
||||
|
||||
---
|
||||
|
||||
## 4. Exact versus hierarchical matching
|
||||
|
||||
Channels are tags, so a broadcast can walk from the sent channel up through its
|
||||
parents:
|
||||
|
||||
```text
|
||||
Event.Combat.Elimination
|
||||
→ Event.Combat
|
||||
→ Event
|
||||
```
|
||||
|
||||
On the first tag every listener is called regardless of match type; on ancestors,
|
||||
only listeners that registered for partial matching.
|
||||
|
||||
**Default to exact.** It minimises accidental traffic and type ambiguity. Use
|
||||
partial matching for analytics, debug logging, and aggregators that genuinely
|
||||
want a subtree.
|
||||
|
||||
### The hierarchy type gate
|
||||
|
||||
Every descendant channel observed through one parent must share a compatible base
|
||||
payload struct. Otherwise partial matching converts a semantic hierarchy into a
|
||||
stream of type errors. Document the payload contract at each parent tag — this is
|
||||
the same rule as parent-tag contracts in `ue-gameplay-tag-governance`, applied to
|
||||
payload types instead of semantics.
|
||||
|
||||
### Cost
|
||||
|
||||
Walking parents costs one map lookup per tag depth on **every** broadcast, even
|
||||
when no partial listeners exist anywhere. If the listener array is copied per
|
||||
non-empty bucket for mutation safety, that copy includes a callable object per
|
||||
listener, which may allocate. On a high-frequency channel this is measurable.
|
||||
Recipe: MB-16.
|
||||
|
||||
---
|
||||
|
||||
## 5. Broadcast safely under mutation
|
||||
|
||||
Callbacks may unregister themselves or others. Copy the listener array before
|
||||
iterating, then invoke the copy — and then **define the resulting semantics
|
||||
explicitly**, because they are observable:
|
||||
|
||||
- a listener removed during this broadcast may still receive this one message;
|
||||
- a listener added during this broadcast receives from the next one;
|
||||
- ordering is not guaranteed and must not be assumed.
|
||||
|
||||
Do not let registration order become an accidental priority contract. If removal
|
||||
uses a swap-with-last strategy, order changes as a side effect of unrelated
|
||||
unsubscribes. Recipe: MB-11.
|
||||
|
||||
---
|
||||
|
||||
## 6. Registration forms and ownership
|
||||
|
||||
Provide three forms: a raw callable, a UObject member bound through a weak
|
||||
reference, and a parameter struct for match type and future options.
|
||||
|
||||
The trap is that these three are not equally expressive. If the most convenient
|
||||
form — the UObject member overload — cannot express the match type, then the
|
||||
match type is effectively unavailable to the code that would most benefit from
|
||||
it, and a project can ship a fully implemented hierarchical matching feature that
|
||||
nothing uses. Recipe: MB-05.
|
||||
|
||||
### Ownership patterns
|
||||
|
||||
| Owner | Retain | Release |
|
||||
|---|---|---|
|
||||
| C++ component | handle as a member | `EndPlay` / deinit |
|
||||
| subsystem | handles in a container | `Deinitialize` |
|
||||
| feature action | per activation context | on deactivation |
|
||||
| async operation | inside the action object | completion, cancel or destruction |
|
||||
|
||||
**A weak callback is not an unsubscription.** It prevents a call into a dead
|
||||
object; the listener record and its callable stay in the map, and accumulate.
|
||||
Recipe: MB-08.
|
||||
|
||||
### Teardown must not assert
|
||||
|
||||
An accessor that asserts when there is no world is the wrong tool inside
|
||||
`EndPlay`, `NativeDestruct` or any other teardown path, where the world may
|
||||
already be gone. Prefer unregistering through the handle, which resolves the
|
||||
subsystem weakly and no-ops when it is dead. Recipe: MB-09.
|
||||
|
||||
---
|
||||
|
||||
## 7. Blueprint parity
|
||||
|
||||
A wildcard struct pin lets Blueprint send arbitrary payloads while C++ receives a
|
||||
type plus bytes. An async listen node stores the channel, payload type and match
|
||||
type, registers on activation, copies the payload into typed output storage, and
|
||||
unregisters on cancel or destruction.
|
||||
|
||||
### Keep the two APIs semantically identical
|
||||
|
||||
This is the rule most often broken, because the two paths are written at
|
||||
different times by different people:
|
||||
|
||||
- if C++ accepts derived payloads, Blueprint must too;
|
||||
- if C++ logs a mismatch, Blueprint must log it as well.
|
||||
|
||||
A silent Blueprint skip produces "the node never fires", which every developer
|
||||
first investigates as a wrong tag. Pick one contract, document it, and test both
|
||||
APIs against the same matrix. Recipe: MB-02.
|
||||
|
||||
---
|
||||
|
||||
## 8. Messaging is not replication
|
||||
|
||||
No broadcast crosses the network. For networked facts the shape is always:
|
||||
|
||||
```text
|
||||
server-authoritative state change
|
||||
→ replicated property / RPC / replicated fast array
|
||||
→ receiving side broadcasts a LOCAL message
|
||||
→ local UI, feedback and processors react
|
||||
```
|
||||
|
||||
This keeps transport and local fan-out separate, and it is the only arrangement
|
||||
in which a late joiner is correct: the state replicates, and the message is
|
||||
merely a notification that it changed.
|
||||
|
||||
**Never treat a client-local message as proof of a server-side fact.** A message
|
||||
may request UI behaviour. It is not evidence of damage, score, ownership or team
|
||||
membership. Recipe: MB-15.
|
||||
|
||||
---
|
||||
|
||||
## 9. Payload design
|
||||
|
||||
A message struct is a compact immutable fact:
|
||||
|
||||
```text
|
||||
verb / channel context
|
||||
instigator
|
||||
target
|
||||
magnitude
|
||||
context tags
|
||||
optional IDs or weak references
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- prefer IDs and weak references over large object graphs;
|
||||
- carry enough context that a listener needs no callback into the sender;
|
||||
- do not embed mutable ownership;
|
||||
- initialise every field — a payload never passes through serialisation, so
|
||||
there are no free defaults;
|
||||
- version deliberately if the same struct is persisted or replicated elsewhere;
|
||||
- one parent channel hierarchy shares one base payload.
|
||||
|
||||
### Do not send commands disguised as events
|
||||
|
||||
```text
|
||||
bad: message "please give the player a weapon"
|
||||
good: service call performs the grant; message "weapon equipped"
|
||||
```
|
||||
|
||||
A command has exactly one accountable handler and a failure result. An event has
|
||||
zero or many listeners and no failure path at all. Sending a command as an event
|
||||
means nobody is responsible for it happening, and nobody notices when it does
|
||||
not. Recipe: MB-14.
|
||||
|
||||
---
|
||||
|
||||
## 10. Channel naming and the string contract
|
||||
|
||||
Name channels so bus traffic is distinguishable from other tags:
|
||||
|
||||
```text
|
||||
<System>.<Domain>.Message[.<Detail>]
|
||||
```
|
||||
|
||||
The word in the middle separates channels from the rest of the vocabulary and
|
||||
gives a natural place for a partial-match subscription.
|
||||
|
||||
**The channel string is the public contract, and the compiler does not check it.**
|
||||
If publisher and subscribers each declare their own file-local tag constant for
|
||||
the same string, the engine deduplicates by string and everything works — until
|
||||
one of them is edited. Put shared channels in one header owned by the publisher.
|
||||
Recipe: MB-10.
|
||||
|
||||
---
|
||||
|
||||
## 11. Migration recipe
|
||||
|
||||
When decoupling a direct dependency:
|
||||
|
||||
1. Name the fact in the past tense.
|
||||
2. Define its payload struct and channel tag.
|
||||
3. Keep authoritative mutation in the existing owner.
|
||||
4. Broadcast only **after** the mutation succeeds.
|
||||
5. Move independent side effects to listeners, one at a time.
|
||||
6. Retain and release every handle.
|
||||
7. Test zero, one and multiple listeners.
|
||||
8. Verify no late subscriber needs the state it missed.
|
||||
|
||||
Good first migrations: combat event to HUD feed and audio; inventory delta to a
|
||||
notification widget; phase transition to local presentation.
|
||||
|
||||
Do not migrate tight request/response code merely to avoid a function reference.
|
||||
|
||||
---
|
||||
|
||||
## 12. Review checklist
|
||||
|
||||
- [ ] Message is an event, not durable state and not a command.
|
||||
- [ ] Subsystem scope matches the required lifetime.
|
||||
- [ ] Payload is a struct with a documented contract and initialised fields.
|
||||
- [ ] Parent channels have a compatible payload hierarchy.
|
||||
- [ ] Exact matching is the default; partial matching is justified per site.
|
||||
- [ ] Registration returns a handle and the caller retains it.
|
||||
- [ ] UObject callbacks bind weakly.
|
||||
- [ ] Unregistration happens on every lifecycle exit, including teardown paths.
|
||||
- [ ] Teardown does not call an asserting accessor.
|
||||
- [ ] Broadcast tolerates mutation from inside a callback.
|
||||
- [ ] C++ and Blueprint type semantics are identical and both log mismatches.
|
||||
- [ ] No listener ordering is assumed without an explicit contract.
|
||||
- [ ] Replication and RPC transport are separate from local fan-out.
|
||||
- [ ] Shared channel tags live in one owned header.
|
||||
|
||||
---
|
||||
|
||||
## 13. Tests
|
||||
|
||||
### Type matrix
|
||||
|
||||
- exact same struct;
|
||||
- derived broadcast to base listener;
|
||||
- base broadcast to derived listener — rejected;
|
||||
- unrelated struct — logged and skipped;
|
||||
- expired expected struct type;
|
||||
- C++ and Blueprint produce the same result for every row above.
|
||||
|
||||
### Channel matrix
|
||||
|
||||
- exact listener, exact broadcast;
|
||||
- parent with partial matching receives a child broadcast;
|
||||
- parent with exact matching does not;
|
||||
- sibling does not;
|
||||
- a nested parent chain invokes each intended listener exactly once.
|
||||
|
||||
### Lifecycle
|
||||
|
||||
- unregister before broadcast;
|
||||
- unregister from inside a callback;
|
||||
- owner destroyed without an explicit unregister;
|
||||
- async node cancelled;
|
||||
- map travel with a GameInstance-scoped bus — listener count returns to baseline;
|
||||
- feature activation and deactivation cycle, twice;
|
||||
- duplicate registration is either documented or detected.
|
||||
|
||||
### Network boundary
|
||||
|
||||
- a server-side broadcast is invisible to clients;
|
||||
- replicated transport triggers exactly one client-local rebroadcast;
|
||||
- a listen-server process does not process both the server and client path.
|
||||
|
||||
---
|
||||
|
||||
## 14. When not to add a bus
|
||||
|
||||
Use a direct interface, service or delegate when:
|
||||
|
||||
- exactly one receiver must handle the call;
|
||||
- the caller needs a return value or a failure reason;
|
||||
- ordering is part of correctness;
|
||||
- sender and receiver share ownership and lifetime;
|
||||
- the state must be queryable after the event.
|
||||
|
||||
A bus reduces coupling only when the fact genuinely has independent, optional
|
||||
observers. Used anywhere else it hides control flow and weakens correctness —
|
||||
it converts a compile-time dependency into a runtime one that nothing verifies.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
the gameplay message router plugin shipped with Epic's Lyra Starter Game on
|
||||
Unreal Engine 5.6, together with its use in that project. The plugin's runtime
|
||||
module is roughly a thousand lines and was read in full, which is why several
|
||||
claims here are negative results ("there is no networking in this module") stated
|
||||
with confidence rather than hedged.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; what
|
||||
ships is the detection recipe. Each entry carries a stable identifier (`MB-01`,
|
||||
`MB-02`, …) that resolves back to the audited location in that archive.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one plugin at one engine version in one
|
||||
workspace. Blueprint graphs were not read — they are binary — so every statement
|
||||
about Blueprint *usage* in that project is bounded by what C++ could show, and is
|
||||
marked as open rather than concluded. Re-run the recipes against your own tree
|
||||
before trusting any specific claim.
|
||||
@@ -0,0 +1,646 @@
|
||||
# Failure modes: gameplay messaging
|
||||
|
||||
Seventeen ways a tag-addressed message bus stops delivering, over-delivers, or
|
||||
delivers into a void, without reporting anything useful.
|
||||
|
||||
The shared property of this list: **a bus converts a compile-time dependency into
|
||||
a runtime one, and then does not check it.** The publisher compiles whether or not
|
||||
anyone listens; the listener compiles whether or not anyone publishes. Every
|
||||
entry below is a consequence of that trade, and the observable is almost always
|
||||
the same — nothing happens, or something happens twice.
|
||||
|
||||
Recipes use `rg`. Two conventions worth stating once, because both cost real time
|
||||
during the audit that produced this file:
|
||||
|
||||
- **Restrict searches to source globs.** An unrestricted search of a built tree
|
||||
matches compiled debug symbols, which are not code. In the measured reference a
|
||||
search for a suspected dead field returned four hits; three were `.pdb` files
|
||||
and the source truth was one. A count that includes build output is not a count
|
||||
of your code.
|
||||
- **Most recipes are a pair of searches.** The defect class here is "the thing is
|
||||
present and the thing that would give it meaning is absent", and one search can
|
||||
only show the first half.
|
||||
|
||||
---
|
||||
|
||||
## Delivery correctness
|
||||
|
||||
### MB-01 - Listener removed using the wrong channel key
|
||||
|
||||
**Mechanism.** A hierarchical broadcast walks from the sent channel up through its
|
||||
parents. Inside that walk, code that removes a listener addresses the removal by
|
||||
the *original broadcast channel* rather than the ancestor tag currently being
|
||||
visited. Handle IDs are allocated per channel, so the pair (channel, ID) is the
|
||||
only unique identity.
|
||||
|
||||
**Why it is silent.** On the first iteration of the walk the two variables hold
|
||||
the same value, so exact-match listeners — which is nearly all of them in most
|
||||
projects — are removed correctly. The bug requires a listener registered on a
|
||||
parent with partial matching *and* an expired payload type, simultaneously. Both
|
||||
conditions are rare, and neither is an error on its own.
|
||||
|
||||
**Why the obvious check misses it.** The line reads correctly in isolation:
|
||||
removing by "channel" is obviously right, and the variable is named `Channel`.
|
||||
The defect is that the enclosing loop rebinds the meaning of "the channel we are
|
||||
talking about" on every iteration, and the correct variable is three lines up in
|
||||
the `for` header. Worse, the error-log line a few lines below uses the loop
|
||||
variable correctly — so the file contains both the right and the wrong idiom, and
|
||||
the wrong one is the one that mutates state.
|
||||
|
||||
**Symptom.** Two outcomes, both bad and neither loud. Either the stale listener is
|
||||
never removed and a warning is emitted on every subsequent broadcast forever, or —
|
||||
worse — an unrelated listener that happens to hold the same channel-local ID in
|
||||
the broadcast channel's bucket is removed instead, and its owner silently stops
|
||||
receiving messages.
|
||||
|
||||
**Detect.** Compare the loop variable against the key used for removal:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "for\s*\(.*RequestDirectParent" --glob "*.cpp" . \
|
||||
| rg "UnregisterListenerInternal|Remove\("
|
||||
```
|
||||
|
||||
Read the output next to the `for` header. If the removal uses the loop's starting
|
||||
value rather than the loop variable, that is this defect. In the measured
|
||||
reference the walk iterated over an ancestor tag while removal passed the
|
||||
broadcast channel.
|
||||
|
||||
**Guardrail.** When handle IDs are channel-local, make the removal API take the
|
||||
pair and nothing else, and assert inside it that the pair was found. A removal
|
||||
that silently finds nothing is how this defect stays invisible.
|
||||
|
||||
---
|
||||
|
||||
### MB-02 - Covariance in C++, equality in Blueprint
|
||||
|
||||
**Mechanism.** The C++ delivery path accepts a payload whose struct derives from
|
||||
the listener's expected struct. The Blueprint async listener compares types with
|
||||
strict equality.
|
||||
|
||||
**Why it is silent.** Both behaviours are individually reasonable and both are
|
||||
implemented correctly. A Blueprint listener that does not match simply does not
|
||||
fire, and not firing is the normal state of a listener most of the time.
|
||||
|
||||
**Why the obvious check misses it.** The two implementations live in different
|
||||
files, written for different consumers, and each is correct against its own local
|
||||
reasoning. There is no shared test matrix, because the two APIs are usually tested
|
||||
by different people — the C++ one by whoever wrote the bus, the Blueprint one by
|
||||
whoever first used the node. Nothing in either file mentions the other.
|
||||
|
||||
**Symptom.** A Blueprint listener subscribed to a base struct never fires for
|
||||
derived payloads, while a C++ listener on the same base receives them. The
|
||||
designer reports "the node never fires", which everyone investigates as a wrong
|
||||
tag — the one hypothesis that is cheap to test and wrong.
|
||||
|
||||
**Detect.** Find both type gates and compare the operators:
|
||||
|
||||
```bash
|
||||
rg -n "IsChildOf" --glob "*.cpp" . # covariant gate
|
||||
rg -n "StructType\s*==|== \w*StructType" --glob "*.cpp" . # equality gate
|
||||
```
|
||||
|
||||
Two different operators guarding the same conceptual check, in the same plugin,
|
||||
is the finding. In the measured reference the C++ path used covariance and the
|
||||
async Blueprint path used equality, on the same registration.
|
||||
|
||||
**Guardrail.** Pick one contract, write it down, and run one test matrix through
|
||||
both APIs. If they must differ, the difference belongs in the node's tooltip, not
|
||||
only in the source.
|
||||
|
||||
---
|
||||
|
||||
### MB-03 - Type check disabled by an empty type pin
|
||||
|
||||
**Mechanism.** Registering with a null expected struct type sets a flag that
|
||||
disables the type check entirely, so the listener receives any payload. The path
|
||||
is documented as internal — and is reachable from a Blueprint node whose payload
|
||||
type pin was left empty.
|
||||
|
||||
**Why it is silent.** Nothing fails. The listener receives messages, possibly
|
||||
several unrelated kinds, and the Blueprint graph downstream either uses the
|
||||
payload or does not. There is no log, because from the bus's point of view the
|
||||
listener asked for exactly this.
|
||||
|
||||
**Why the obvious check misses it.** The comment on the branch says "for internal
|
||||
use", and reviewers believe comments about intent. Confirming reachability means
|
||||
tracing a Blueprint pin's default value into a factory function into a private
|
||||
registration call — three hops across two modules, none of which look
|
||||
interesting.
|
||||
|
||||
**Symptom.** In Blueprint, usually nothing worse than a payload extraction that
|
||||
returns false. In C++, the same path combined with a templated callback would
|
||||
reinterpret unrelated bytes as the expected struct, which is undefined behaviour
|
||||
with no diagnostic.
|
||||
|
||||
**Detect.** Find the escape hatch and then find who can reach it:
|
||||
|
||||
```bash
|
||||
rg -n "bHadValidType|StructType\s*==\s*nullptr" --glob "*.cpp" --glob "*.h" .
|
||||
rg -n -B4 "RegisterListenerInternal\(" --glob "*.cpp" . | rg -i "Get\(\)|nullptr"
|
||||
```
|
||||
|
||||
Any caller that can pass a null type is a caller that can disable the check.
|
||||
|
||||
**Guardrail.** If a code path must exist for internal reasons, make it
|
||||
inaccessible rather than merely undocumented — a private overload, a passkey
|
||||
type, or a compile-time gate. Reject an unresolved wildcard payload at Blueprint
|
||||
compile time.
|
||||
|
||||
---
|
||||
|
||||
### MB-04 - Payload pointer retained after the callback returns
|
||||
|
||||
**Mechanism.** The bus passes a pointer to the sender's stack object rather than
|
||||
copying. A listener stores that pointer.
|
||||
|
||||
**Why it is silent.** The memory is still mapped and, for a short while, still
|
||||
holds the right bytes. Reading it immediately after the broadcast usually works.
|
||||
The failure requires the stack to be reused, which depends on what runs next.
|
||||
|
||||
**Why the obvious check misses it.** The callback signature hands you a reference,
|
||||
and references are normally safe to keep. Nothing in the type expresses "valid
|
||||
until this function returns". The one place the constraint is written down is the
|
||||
bus's own documentation, which the person writing the listener has no reason to
|
||||
open.
|
||||
|
||||
**Symptom.** Corrupted payload fields read at a later tick — values that are
|
||||
plausible, occasionally correct, and change with unrelated code. This is one of
|
||||
the few entries in this file that produces a crash, and the crash is far from the
|
||||
cause.
|
||||
|
||||
**Detect.** Find listeners that store rather than consume:
|
||||
|
||||
```bash
|
||||
rg -n -A8 "void .*\(FGameplayTag\s+\w+,\s*const\s+F\w+&\s*\w+\)" --glob "*.cpp" . \
|
||||
| rg "=\s*&|Ptr\s*=|AddRaw|CopyRef"
|
||||
```
|
||||
|
||||
Any assignment of the payload's address into a member is the finding. A copy of
|
||||
the payload by value is fine.
|
||||
|
||||
**Guardrail.** State the lifetime in the callback's own documentation, and copy at
|
||||
the boundary. Where the bus crosses into a scripting layer, copy by value — the
|
||||
measured reference does exactly this for its Blueprint path, and nulls the stored
|
||||
pointer immediately after the delegate returns.
|
||||
|
||||
---
|
||||
|
||||
## Ownership and lifetime
|
||||
|
||||
### MB-05 - The convenient overload cannot express the match type
|
||||
|
||||
**Mechanism.** Three registration overloads exist. The one that binds a UObject
|
||||
member weakly — the safest and by far the most used — delegates to the raw form
|
||||
without forwarding the match type, so it is always exact.
|
||||
|
||||
**Why it is silent.** Exact matching is the right default, so every listener
|
||||
registered this way behaves correctly. The feature that is unavailable is one
|
||||
nobody is currently trying to use, and its absence produces no error because you
|
||||
cannot pass the argument at all.
|
||||
|
||||
**Why the obvious check misses it.** The feature is fully implemented, documented
|
||||
in the enum, exposed in the Blueprint node, and covered by the broadcast walk. A
|
||||
review that asks "do we support hierarchical matching?" finds all of that and
|
||||
answers yes. The gap is one unforwarded default argument in a header.
|
||||
|
||||
**Symptom.** A project ships a complete hierarchical matching implementation that
|
||||
nothing in its own C++ ever uses, and nobody notices, because using it requires
|
||||
switching to a less convenient overload for reasons that are never stated.
|
||||
|
||||
**Detect.** Count the feature's declaration sites against its use sites, and
|
||||
exclude the implementation itself:
|
||||
|
||||
```bash
|
||||
rg -n "PartialMatch" --glob "*.h" --glob "*.cpp" . # declared
|
||||
rg -n "PartialMatch" --glob "*.cpp" . | rg -v "MessageRouter|MessageSubsystem"
|
||||
```
|
||||
|
||||
An empty second result with a rich first result is the finding. In the measured
|
||||
reference the game code used partial matching zero times; the only similarly
|
||||
named symbols elsewhere belonged to two unrelated subsystems that had copied the
|
||||
pattern.
|
||||
|
||||
**Guardrail.** Every overload of a registration API forwards every option, or the
|
||||
option does not exist. If an overload deliberately restricts, name it so —
|
||||
`RegisterExactListener` — rather than silently dropping an argument.
|
||||
|
||||
---
|
||||
|
||||
### MB-06 - Parameter struct with no bound callback returns an invalid handle silently
|
||||
|
||||
**Mechanism.** The options-struct registration form checks that a callback is
|
||||
bound and, if not, returns a default-constructed invalid handle without
|
||||
registering and without logging.
|
||||
|
||||
**Why it is silent.** The caller receives a handle-shaped value. Storing it,
|
||||
passing it around and eventually unregistering it all work. The only difference
|
||||
is that nothing was ever registered.
|
||||
|
||||
**Why the obvious check misses it.** The return type is not optional and not an
|
||||
error code; it is a handle, and handles look like success. Checking validity
|
||||
requires knowing that this overload can fail, which is stated nowhere at the call
|
||||
site.
|
||||
|
||||
**Symptom.** A listener that never receives anything, in code that reads as fully
|
||||
wired. Because unregistration of an invalid handle is also silent, no part of the
|
||||
lifecycle complains.
|
||||
|
||||
**Detect.** Find registrations whose returned handle is never validity-checked:
|
||||
|
||||
```bash
|
||||
rg -n "RegisterListener\(" --glob "*.cpp" . -A3 | rg -v "IsValid|ensure|check"
|
||||
```
|
||||
|
||||
**Guardrail.** Return an explicit failure, or log at warning level, or assert in
|
||||
development builds. A silent failure path on a registration API is a listener
|
||||
that cannot be debugged.
|
||||
|
||||
---
|
||||
|
||||
### MB-07 - Async listener outlives its owner until the next message
|
||||
|
||||
**Mechanism.** An async listen action detects its owner's death by noticing that
|
||||
its dynamic delegate is no longer bound — which it can only do *while handling a
|
||||
message*. Until the next message arrives, the action and its registration stay
|
||||
alive.
|
||||
|
||||
**Why it is silent.** The cleanup does eventually happen, and it happens before
|
||||
anything observable goes wrong. On a busy channel the window is a frame.
|
||||
|
||||
**Why the obvious check misses it.** There *is* cleanup, and reading it shows
|
||||
correct logic. The gap is a timing property — cleanup is driven by traffic rather
|
||||
than by the owner's destruction — and timing properties are invisible in a code
|
||||
review that asks "is cleanup implemented?".
|
||||
|
||||
**Symptom.** On a quiet channel, listener records accumulate for destroyed
|
||||
owners. On a shutdown path, the last message of a session may be delivered into
|
||||
an object graph that is already tearing down.
|
||||
|
||||
**Detect.** Find cleanup that is triggered by message handling rather than by
|
||||
destruction:
|
||||
|
||||
```bash
|
||||
rg -n -B6 "SetReadyToDestroy" --glob "*.cpp" . | rg -i "IsBound|HandleMessage"
|
||||
rg -n -i "//\s*@?TODO.*(proactive|cleanup)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
The second command matters: in the measured reference the authors had recorded
|
||||
exactly this limitation as a TODO with a tracker reference, which is the
|
||||
strongest possible confirmation that a finding is real and known.
|
||||
|
||||
**Guardrail.** Tie the async action's lifetime to its owner's destruction
|
||||
directly, not to the arrival of the next message.
|
||||
|
||||
---
|
||||
|
||||
### MB-08 - Listener map is not cleared on map travel
|
||||
|
||||
**Mechanism.** A GameInstance-scoped bus is cleared only at GameInstance
|
||||
teardown. Listeners registered by objects belonging to a world that has been
|
||||
travelled away from remain in the map.
|
||||
|
||||
**Why it is silent.** Weak binding means the dead objects are never called. The
|
||||
records are inert: they cost memory and a little iteration time, and produce no
|
||||
wrong behaviour.
|
||||
|
||||
**Why the obvious check misses it.** Deinitialisation exists and is correct for
|
||||
the scope it was written for. "Does the bus clean up?" is answered yes. The
|
||||
question that finds this is narrower — *at which lifecycle event?* — and the
|
||||
answer is one that only matters in a project that travels between maps, which a
|
||||
single-map test never does.
|
||||
|
||||
**Symptom.** Listener counts that grow monotonically across a session of map
|
||||
changes. Diagnosis is difficult because the leak is bounded by session length and
|
||||
the objects themselves are correctly garbage collected.
|
||||
|
||||
**Detect.** Find the scope and the reset point, and compare them:
|
||||
|
||||
```bash
|
||||
rg -n "public UGameInstanceSubsystem|public UWorldSubsystem" --glob "*.h" .
|
||||
rg -n -A4 "::Deinitialize\(\)" --glob "*.cpp" . | rg "Reset\(|Empty\(|Clear\("
|
||||
```
|
||||
|
||||
A GameInstance scope whose only reset is in `Deinitialize` leaks across travel by
|
||||
construction. Confirm at runtime: log the listener count, travel, log again.
|
||||
|
||||
**Guardrail.** Either scope the bus to the world, or hook map change and drop
|
||||
records whose weak owner has expired. A weak callback is not an unsubscription.
|
||||
|
||||
---
|
||||
|
||||
### MB-09 - Teardown calls an accessor that asserts
|
||||
|
||||
**Mechanism.** The static accessor resolves the world with an assert-on-failure
|
||||
mode and then asserts on both the world and the subsystem. Teardown code calls it
|
||||
to unregister.
|
||||
|
||||
**Why it is silent.** In every ordinary shutdown the world is still present when
|
||||
components tear down, so the asserts pass. The failure needs an unusual teardown
|
||||
order — a subsystem destroyed first, PIE cancelled mid-initialisation, a level
|
||||
unloaded during travel.
|
||||
|
||||
**Why the obvious check misses it.** Unregistering in `EndPlay` is the correct
|
||||
pattern and is what a reviewer is looking for. That the *accessor* is the
|
||||
strictest of the several available ways to reach the subsystem is a property of a
|
||||
different file.
|
||||
|
||||
**Symptom.** An assert during shutdown, in a build configuration and teardown
|
||||
order that nobody can reproduce on demand, at a call site that is doing the right
|
||||
thing.
|
||||
|
||||
**Detect.** Find teardown paths that use the asserting accessor:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A2 "::EndPlay|NativeDestruct|BeginDestroy" --glob "*.cpp" . \
|
||||
| rg "Subsystem::Get\(|::Get\(this\)"
|
||||
rg -n -A6 "static .*& Get\(" --glob "*.h" . | rg "Assert|check\("
|
||||
```
|
||||
|
||||
The pairing is the finding: an asserting accessor called from a destruction path.
|
||||
In the measured reference this pattern appeared in three separate consumer
|
||||
classes, all of which could have unregistered through the handle instead.
|
||||
|
||||
**Guardrail.** Unregister through the handle, which resolves the subsystem weakly
|
||||
and no-ops when it is gone. Provide a non-asserting accessor and use it in every
|
||||
teardown path.
|
||||
|
||||
---
|
||||
|
||||
## Contract and vocabulary
|
||||
|
||||
### MB-10 - The channel string is the contract, declared in several places
|
||||
|
||||
**Mechanism.** Publisher and each subscriber declare their own file-local tag
|
||||
constant for the same channel string. The engine deduplicates by string, so it
|
||||
works.
|
||||
|
||||
**Why it is silent.** It is correct. Every site resolves to the same tag, the join
|
||||
succeeds, and the arrangement has a real benefit: modules do not depend on each
|
||||
other's headers.
|
||||
|
||||
**Why the obvious check misses it.** Each declaration is locally idiomatic and
|
||||
each file reads well. Nothing links them. A rename in one file compiles, links,
|
||||
and passes review — and the reviewer of that file has no way to see the other
|
||||
three.
|
||||
|
||||
**Symptom.** Editing one declaration silently detaches one participant from the
|
||||
channel. This is the symmetric-typo failure from tag governance, arriving through
|
||||
a different door: correctness means all sites agree, not that any one is right.
|
||||
|
||||
**Detect.** Count independent declarations per channel string:
|
||||
|
||||
```bash
|
||||
rg -o -N 'UE_DEFINE_GAMEPLAY_TAG[_A-Z]*\([^,]+,\s*"([^"]+)"' -r '$1' \
|
||||
--glob "*.cpp" . | sort | uniq -c | sort -rn | head
|
||||
```
|
||||
|
||||
Any string with a count above one is a contract with several owners. In the
|
||||
measured reference one elimination channel was declared independently at four
|
||||
sites — the publisher and three separate processors in a feature plugin.
|
||||
|
||||
**Guardrail.** Shared channels live in one header owned by the publisher. Keep
|
||||
file-local declarations only for channels that never leave their file.
|
||||
|
||||
---
|
||||
|
||||
### MB-11 - Listener order assumed without a contract
|
||||
|
||||
**Mechanism.** Nothing guarantees callback order, and removal by swap-with-last
|
||||
actively reorders the array as a side effect of unrelated unsubscriptions.
|
||||
|
||||
**Why it is silent.** An order exists on every run and is stable as long as
|
||||
nothing unsubscribes. Code that depends on it works, and keeps working, until an
|
||||
unrelated feature adds a listener or removes one.
|
||||
|
||||
**Why the obvious check misses it.** The dependency is never written down —
|
||||
that is what makes it an assumption. No search finds "this listener must run
|
||||
first"; the knowledge lives in the fact that it currently does.
|
||||
|
||||
**Symptom.** Ordering-dependent behaviour that changes when an unrelated system
|
||||
subscribes, in a different module, in a different release.
|
||||
|
||||
**Detect.** Find the removal strategy and the documented guarantee, and check for
|
||||
listeners that mutate shared state:
|
||||
|
||||
```bash
|
||||
rg -n "RemoveAtSwap|RemoveSwap" --glob "*.cpp" .
|
||||
rg -n -i "order.*not guaranteed|call order" --glob "*.h" .
|
||||
```
|
||||
|
||||
If the second command finds an explicit warning from the bus authors, the
|
||||
guarantee does not exist and any dependence on order is yours to remove.
|
||||
|
||||
**Guardrail.** If ordering matters, it is not a bus concern — sequence the work in
|
||||
one owner and broadcast the outcome. Never introduce priority through
|
||||
registration order.
|
||||
|
||||
---
|
||||
|
||||
### MB-12 - Reentrancy with no depth limit
|
||||
|
||||
**Mechanism.** Broadcasting from inside a listener callback is legal and
|
||||
supported. Nothing bounds the recursion.
|
||||
|
||||
**Why it is silent.** It is a legitimate and useful pattern — an aggregator that
|
||||
consumes one channel and publishes a derived fact on another does exactly this.
|
||||
The dangerous case differs only in that the graph of channels contains a cycle.
|
||||
|
||||
**Why the obvious check misses it.** Each listener is individually correct and
|
||||
each broadcast is individually justified. The cycle exists in the graph formed by
|
||||
several modules, which no file shows, and which a feature plugin can complete by
|
||||
adding one processor.
|
||||
|
||||
**Symptom.** Stack overflow, or a hang, with a call stack that is one repeating
|
||||
pattern thousands of frames deep and no obvious origin.
|
||||
|
||||
**Detect.** Map the graph: which listeners publish, and on which channels?
|
||||
|
||||
```bash
|
||||
rg -n -B10 "BroadcastMessage\(" --glob "*.cpp" . \
|
||||
| rg -B2 "OnMessageReceived|::Handle\w*Message"
|
||||
```
|
||||
|
||||
Every hit is an edge from an input channel to an output channel. Draw them and
|
||||
look for a cycle. In the measured reference this pattern was present and
|
||||
deliberate — a processor consuming eliminations and publishing assists — with no
|
||||
cycle, which is the correct state and worth confirming rather than assuming.
|
||||
|
||||
**Guardrail.** Keep a broadcast-depth counter in development builds and assert
|
||||
past a small bound. Document, per processor, which channels it reads and writes.
|
||||
|
||||
---
|
||||
|
||||
## Scope and misuse
|
||||
|
||||
### MB-13 - A message used where state is required
|
||||
|
||||
**Mechanism.** A fact is published as an event only. A listener that subscribes
|
||||
after the event has no way to learn it.
|
||||
|
||||
**Why it is silent.** Every listener that existed at publication time is correct.
|
||||
Subscription usually happens during initialisation, before anything interesting
|
||||
has been published, so the whole system is correct for the entire development
|
||||
period.
|
||||
|
||||
**Why the obvious check misses it.** Testing subscribes early by construction —
|
||||
you start the game, the widget is created, then things happen. Reproducing
|
||||
requires a listener created *after* the fact, which means mid-match widget
|
||||
creation, a late-joining client, or a feature activated at runtime.
|
||||
|
||||
**Symptom.** A widget created mid-match shows a default value forever. A late
|
||||
joiner's UI is empty while everyone else's is correct.
|
||||
|
||||
**Detect.** For each channel, ask whether any subscriber can be created late:
|
||||
|
||||
```bash
|
||||
rg -n "RegisterListener\(" --glob "*.cpp" . -B6 | rg -i "NativeConstruct|BeginPlay|OnActivated"
|
||||
```
|
||||
|
||||
Widget construction and feature activation are the late-creation paths. For each,
|
||||
apply the test from the skill: if a listener subscribing one second late must know
|
||||
the current value, the fact is state and the bus is the wrong carrier.
|
||||
|
||||
**Guardrail.** Replicate or store the state; use the message purely as a change
|
||||
notification. On subscription, read the current value directly from its owner —
|
||||
which requires that the value have an owner, and that is the real design output.
|
||||
|
||||
---
|
||||
|
||||
### MB-14 - A command disguised as an event
|
||||
|
||||
**Mechanism.** A message is published to make something happen, rather than to
|
||||
report that it has happened.
|
||||
|
||||
**Why it is silent.** With exactly one listener it behaves identically to a
|
||||
function call. It works, and it looks decoupled.
|
||||
|
||||
**Why the obvious check misses it.** The mechanism is indistinguishable from
|
||||
correct use — same API, same payload shape. Only the *name* reveals it, and names
|
||||
in the imperative mood ("give", "apply", "spawn") pass review because they
|
||||
describe what the sender wants.
|
||||
|
||||
**Symptom.** Zero listeners means the command is silently dropped with no failure
|
||||
result. Two listeners means it happens twice. Neither is reported, because a bus
|
||||
has no concept of a required handler.
|
||||
|
||||
**Detect.** Grep the channel vocabulary for the imperative mood:
|
||||
|
||||
```bash
|
||||
rg -o -N '"([A-Za-z]+\.[A-Za-z.]*Message[A-Za-z.]*)"' -r '$1' --glob "*.cpp" . \
|
||||
| sort -u | rg -i "^\w+\.(Add|Give|Set|Apply|Spawn|Request|Please|Do)"
|
||||
```
|
||||
|
||||
A channel named for an action rather than for a completed fact is the finding.
|
||||
Past-tense names — "equipped", "eliminated", "changed" — are the correct shape.
|
||||
|
||||
**Guardrail.** Commands go to one accountable handler that can fail and say so.
|
||||
Events report facts in the past tense and tolerate zero listeners.
|
||||
|
||||
---
|
||||
|
||||
### MB-15 - Broadcast treated as replication
|
||||
|
||||
**Mechanism.** A message is published on the server and expected to reach
|
||||
clients. The bus is local to one process and has no networking at all.
|
||||
|
||||
**Why it is silent.** In a single-process test — the editor's default multiplayer
|
||||
mode — the host and its client share one process and, depending on the bus's
|
||||
scope, may share one instance. The message appears to cross the network.
|
||||
|
||||
**Why the obvious check misses it.** Everything about the code reads as
|
||||
transport: a channel, a payload, publish and subscribe. The absence of networking
|
||||
is a property of a module's dependency list, which is not where anyone looks when
|
||||
reasoning about a message flow.
|
||||
|
||||
**Symptom.** A feature that works in the editor and does nothing on a dedicated
|
||||
server. Alternatively, a listen server processes both the server-side publication
|
||||
and the client-side rebroadcast and everything happens twice.
|
||||
|
||||
**Detect.** Prove the absence, then find who assumed otherwise:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*(Server|Client|NetMulticast)|NetSerialize|Replicated" \
|
||||
--glob "*.cpp" --glob "*.h" <bus-module-path>/ # expect: nothing
|
||||
rg -n -B4 "BroadcastMessage\(" --glob "*.cpp" . | rg "HasAuthority|WITH_SERVER_CODE|NM_"
|
||||
```
|
||||
|
||||
The second command finds the correct pattern as well as the incorrect one: a
|
||||
net-mode check before a rebroadcast is the guard that prevents double processing
|
||||
on a listen server. Its absence next to a rebroadcast is the finding.
|
||||
|
||||
**Guardrail.** Replicate state or send an RPC; rebroadcast locally on arrival;
|
||||
guard the rebroadcast with a net-mode check so a listen server does not process
|
||||
both paths.
|
||||
|
||||
---
|
||||
|
||||
## Cost
|
||||
|
||||
### MB-16 - Per-broadcast copy of the listener array
|
||||
|
||||
**Mechanism.** For mutation safety the listener array of each non-empty bucket is
|
||||
copied before iteration. The copied element type holds a callable object, whose
|
||||
copy may allocate.
|
||||
|
||||
**Why it is silent.** It is correct, and correctness was the reason for it. The
|
||||
cost is proportional to broadcast frequency times listener count times tag depth,
|
||||
and all three are small early in a project.
|
||||
|
||||
**Why the obvious check misses it.** The line is a defensive copy with a comment
|
||||
explaining why it is needed, which is exactly what a reviewer wants to see. The
|
||||
cost is invisible unless you know that the element type is expensive to copy —
|
||||
which requires reading a different struct in a different header.
|
||||
|
||||
**Symptom.** Allocator churn proportional to message traffic, appearing in
|
||||
profiles as generic allocation cost rather than as bus cost. Discovered when a
|
||||
high-frequency channel — damage in a shooter — is added late.
|
||||
|
||||
**Detect.** Find the copy and the depth walk, then measure:
|
||||
|
||||
```bash
|
||||
rg -n -B2 "TArray<\w*ListenerData>\s*\w+\(" --glob "*.cpp" .
|
||||
rg -n "RequestDirectParent" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Both present means cost per broadcast scales with tag depth even when no
|
||||
hierarchical listener exists. Confirm with an allocation profile while
|
||||
broadcasting on the highest-frequency channel.
|
||||
|
||||
**Guardrail.** Copy handles rather than callables, or iterate an index with a
|
||||
generation counter, or skip the ancestor walk when the bucket count for parent
|
||||
tags is zero. Measure before optimising — this is a real cost, not a large one.
|
||||
|
||||
---
|
||||
|
||||
### MB-17 - A declared field nobody uses, and the search that hides it
|
||||
|
||||
**Mechanism.** A member is declared in a public struct, never written and never
|
||||
read. It survives because removing it feels risky.
|
||||
|
||||
**Why it is silent.** An unused field costs a few bytes and nothing else. It
|
||||
serialises, copies and compiles like any other member.
|
||||
|
||||
**Why the obvious check misses it.** This entry is here less for the dead field
|
||||
than for the search that fails to prove it dead. A naive recursive search of a
|
||||
built project matches compiled debug symbols, so a field with exactly one source
|
||||
occurrence returns several hits and reads as "in use". The instrument reports the
|
||||
build output as if it were code.
|
||||
|
||||
**Symptom.** Dead API surface that persists across refactors, plus — more
|
||||
expensively — a general loss of trust in the searches used to make deletion
|
||||
decisions.
|
||||
|
||||
**Detect.** Search source only, and say so explicitly:
|
||||
|
||||
```bash
|
||||
F='StateClearedHandle'
|
||||
rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" . # source truth
|
||||
rg -n "\b$F\b" . # unfiltered, for contrast
|
||||
```
|
||||
|
||||
In the measured reference the filtered search returned exactly one line — the
|
||||
declaration — while the unfiltered one returned four, three of them debug-symbol
|
||||
files. Always compare the two before concluding either "dead" or "in use".
|
||||
|
||||
**Guardrail.** Restrict every deletion-decision search by file type. Delete dead
|
||||
fields in the same change that proves them dead, and record the search you ran.
|
||||
@@ -0,0 +1,300 @@
|
||||
# Patterns: a message bus read line by line
|
||||
|
||||
A worked reading of one real tag-addressed message bus and its use in the product
|
||||
that ships it. The runtime module is about a thousand lines and was read in full,
|
||||
which is why several claims below are **negative results stated with confidence**
|
||||
rather than hedged — "there is no networking in this module" is a conclusion from
|
||||
reading every line of it, not from a search that found nothing.
|
||||
|
||||
Individual defects are in [failure-modes.md](failure-modes.md), one entry each,
|
||||
with a recipe. This file is about the shapes: what the bus is, where its costs
|
||||
sit, and how the surrounding product worked around what it does not do.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not answerable from source, and left open.
|
||||
|
||||
---
|
||||
|
||||
## 1. What the mechanism actually is
|
||||
|
||||
A publish/subscribe bus where the **channel is a tag** and the **message is an
|
||||
arbitrary struct**. Publisher and subscriber never reference each other; the
|
||||
entire contract is the pair (channel tag, struct type).
|
||||
|
||||
The implementation is one flat map from tag to listener list **[measured]**.
|
||||
There is no tag tree in the data structure — hierarchy is walked at broadcast
|
||||
time by asking each tag for its direct parent until the chain ends.
|
||||
|
||||
Two properties fall out of that design and explain most of this file:
|
||||
|
||||
1. **The join is a string and a reflection pointer.** Neither is checked by the
|
||||
compiler. Everything in the failure-mode list is a consequence.
|
||||
2. **Delivery is synchronous.** Callbacks run inside the broadcast call before it
|
||||
returns **[measured]**. Good for debugging — the stack is continuous — and the
|
||||
reason reentrancy is possible at all.
|
||||
|
||||
The module's dependencies are the engine core, the engine, and the tag system
|
||||
**[measured]**. Nothing from the game, nothing from the ability system. **[derived]**
|
||||
That self-containment is why this kind of plugin is worth lifting wholesale, and
|
||||
why its defects travel with it.
|
||||
|
||||
---
|
||||
|
||||
## 2. Type safety is real, and has exactly one hole
|
||||
|
||||
The payload crosses the bus as an untyped pointer plus a reflected struct type.
|
||||
The gate is one line **[measured]**: deliver when the listener never had a valid
|
||||
type, **or** when the sent struct derives from the listener's expected struct.
|
||||
|
||||
That gives three branches, and they are not equally safe:
|
||||
|
||||
| Branch | Behaviour | Assessment |
|
||||
|---|---|---|
|
||||
| Sent type derives from expected type | delivered | correct, and makes exact equality a special case |
|
||||
| Types unrelated | logged as an error, listener skipped, broadcast continues | correct and loud |
|
||||
| Listener registered with no expected type | **check bypassed entirely** | documented "for internal use" — and reachable from a Blueprint node with an empty type pin **[measured]** |
|
||||
|
||||
The third row is the hole. What makes it interesting is not that it exists — an
|
||||
internal escape hatch is a reasonable thing to have — but that the comment
|
||||
asserting it is internal is the only thing protecting it **[derived]**. In the
|
||||
Blueprint path a second, stricter check downstream limits the damage to a failed
|
||||
payload extraction. In C++ the same registration combined with a templated
|
||||
callback would reinterpret unrelated bytes with no diagnostic at all. Recipe:
|
||||
MB-03.
|
||||
|
||||
There is also a defensive feature worth copying: the expected struct type is held
|
||||
weakly, with a flag recording that it was once valid, so a listener whose struct
|
||||
was garbage collected is detected and removed at broadcast time with a warning
|
||||
**[measured]**. Epic's own comment on those two fields says they were added in
|
||||
response to real problems. **[derived]** That is a project that was bitten and
|
||||
responded structurally — and the removal code it added is where MB-01 lives.
|
||||
|
||||
---
|
||||
|
||||
## 3. The defect worth reading twice
|
||||
|
||||
During the parent walk, a stale listener is removed using the **original
|
||||
broadcast channel** rather than the ancestor tag currently being visited
|
||||
**[measured]**. Handle IDs are unique only within a channel.
|
||||
|
||||
This entry earns its place because of *where* the mistake is, not what it does:
|
||||
|
||||
- the loop header three lines up rebinds what "the channel" means on every
|
||||
iteration;
|
||||
- the error-log statement a few lines **below** the bug uses the loop variable
|
||||
correctly **[measured]**;
|
||||
- so the same function contains both the right and the wrong idiom, and the wrong
|
||||
one is the one that mutates state.
|
||||
|
||||
**[derived]** Two outcomes, both quiet. Either nothing is found and the warning
|
||||
repeats on every subsequent broadcast forever, or an unrelated listener holding
|
||||
the same channel-local ID in the broadcast channel's bucket is removed instead,
|
||||
and its owner silently stops receiving.
|
||||
|
||||
It requires hierarchical matching plus an expired struct type simultaneously, so
|
||||
the audited project never hits it — its own usage has neither. **That is exactly
|
||||
why it matters to someone lifting the plugin**: the conditions that make it
|
||||
harmless are properties of the original product, not of the code. Recipe: MB-01.
|
||||
|
||||
---
|
||||
|
||||
## 4. A fully built feature that nothing uses
|
||||
|
||||
Hierarchical matching is implemented in the broadcast walk, exposed in the enum,
|
||||
and selectable in the Blueprint node **[measured]**. Game-code usage: **zero**
|
||||
**[measured]**.
|
||||
|
||||
The cause is one unforwarded default argument. The overload that binds a UObject
|
||||
member weakly — the safest form, and the one used almost everywhere — delegates to
|
||||
the raw form without passing the match type, so it is always exact **[measured]**.
|
||||
Reaching hierarchical matching from C++ requires switching to a lambda or an
|
||||
options struct, for reasons the API never states.
|
||||
|
||||
**[derived]** The general shape is worth more than the instance: **a feature can
|
||||
be complete, correct, tested and documented, and still be unreachable from the
|
||||
path everyone actually uses.** Auditing for "is it implemented?" finds it.
|
||||
Auditing for "who calls it?" finds the truth. Recipe: MB-05.
|
||||
|
||||
Two neighbouring subsystems in the same product had independently copied the
|
||||
channel-tag-plus-match-type pattern **[measured]** — evidence that the idea was
|
||||
considered good enough to reproduce, in a project where the original was never
|
||||
switched on.
|
||||
|
||||
---
|
||||
|
||||
## 5. Blueprint parity is where the asymmetries live
|
||||
|
||||
The scripting layer is a separate module, editor-only, with its own node
|
||||
**[measured]**. It does several things well: a wildcard output pin retyped from
|
||||
the selected payload struct, a compile-time error when a wildcard payload pin is
|
||||
connected but untyped, and a copy of the payload by value at the boundary so
|
||||
scripting never holds the sender's stack pointer **[measured]**.
|
||||
|
||||
And then the two paths disagree about types **[measured]**:
|
||||
|
||||
| | C++ | Blueprint |
|
||||
|---|---|---|
|
||||
| Type gate | derived accepted | strict equality |
|
||||
| Mismatch | logged as an error | **silent skip** |
|
||||
|
||||
**[derived]** A Blueprint listener on a base struct silently misses derived
|
||||
payloads that its C++ equivalent receives, and reports nothing. The observable is
|
||||
"the node never fires", which everyone investigates as a wrong tag first, because
|
||||
that hypothesis is cheap to test.
|
||||
|
||||
Neither behaviour is wrong on its own. They were written for different consumers
|
||||
and each is correct against its own local reasoning; nothing in either file
|
||||
mentions the other. **The rule that would have caught it is procedural, not
|
||||
technical: one test matrix, run through both APIs.** Recipe: MB-02.
|
||||
|
||||
---
|
||||
|
||||
## 6. Lifetime: three separate ways to leak
|
||||
|
||||
| Shape | Measured | Consequence **[derived]** |
|
||||
|---|---|---|
|
||||
| Listener map cleared only at GameInstance teardown | reset appears in `Deinitialize` and nowhere else | Records from a travelled-away world persist for the session. Weak binding stops the calls, not the accumulation. MB-08 |
|
||||
| Async node detects owner death only while handling a message | cleanup keyed on the delegate being unbound, inside the message handler; authors' own TODO records the limitation with a tracker ID | On a quiet channel, cleanup waits for traffic. MB-07 |
|
||||
| Options-struct registration with no bound callback | returns a default invalid handle, no log | A listener that never fires, in code that reads as wired. MB-06 |
|
||||
|
||||
The first row generalises past this plugin: **a weak callback is not an
|
||||
unsubscription.** It prevents a call into a dead object and leaves the record.
|
||||
Any bus that binds weakly and never sweeps will accumulate, and the accumulation
|
||||
is invisible precisely because weak binding made it harmless.
|
||||
|
||||
The second row is a good example of a finding you should *want* to discover:
|
||||
the authors documented the limitation themselves, in a comment, with a reference.
|
||||
A TODO that names the problem is the strongest confirmation available from
|
||||
reading source.
|
||||
|
||||
### And a teardown hazard on the consumer side
|
||||
|
||||
The static accessor resolves the world in assert-on-failure mode and then asserts
|
||||
on both the world and the subsystem **[measured]**. Three consumer classes in the
|
||||
product call it from destruction paths **[measured]** — where the world may already
|
||||
be gone — when unregistering through the handle would have resolved the subsystem
|
||||
weakly and no-opped. **[derived]** Correct in every ordinary shutdown; an assert
|
||||
in the unusual orders that nobody can reproduce on demand. Recipe: MB-09.
|
||||
|
||||
---
|
||||
|
||||
## 7. How the product worked around no networking
|
||||
|
||||
The bus has no networking at all: no server, client or multicast functions, no
|
||||
custom serialisation, no replicated properties **[measured]**, in a module small
|
||||
enough to have been read in full. The surrounding product needed networked facts
|
||||
anyway, and solved it three ways — worth comparing, because they are not equally
|
||||
good.
|
||||
|
||||
**1. RPC wrappers on game state and player state.** A multicast or client-targeted
|
||||
RPC receives the message struct and, on arrival, rebroadcasts it locally
|
||||
**[measured]**. The rebroadcast is guarded by a net-mode check so a listen server
|
||||
does not process both the server-side publication and the client-side
|
||||
rebroadcast **[measured]**. That guard is the transferable part; without it every
|
||||
listen-server host sees everything twice. The headers say plainly that these are
|
||||
for notifications that can be lost.
|
||||
|
||||
**2. A replicated fast array that rebroadcasts on arrival.** Fully implemented and
|
||||
**used nowhere in the project** — no member of that type exists anywhere
|
||||
**[measured]**. It is a worked example rather than working code, and it carries an
|
||||
empty removal callback (see the network-authority skill for why that matters at
|
||||
adoption time).
|
||||
|
||||
**3. Replicated state plus a local broadcast on the replication callback.** The
|
||||
cleanest and the most used **[measured]**. State replicates; the message is purely
|
||||
a local notification that it changed.
|
||||
|
||||
**[derived]** The third pattern is the one to copy, and the reason is the
|
||||
state-versus-event test: only in that arrangement is a late joiner correct,
|
||||
because the truth is in the replicated state and the message carries none of it.
|
||||
|
||||
---
|
||||
|
||||
## 8. What the bus does not do, stated plainly
|
||||
|
||||
Read as a list of things you will otherwise discover one at a time:
|
||||
|
||||
- **No sticky or replay messages.** A subscriber cannot ask for the last value.
|
||||
The product compensates by having new widgets read the owning component
|
||||
directly on construction **[measured]** — which only works because the state has
|
||||
an owner. MB-13.
|
||||
- **No ordering guarantee.** The authors say so in a header comment, and removal
|
||||
by swap-with-last actively reorders the array **[measured]**. MB-11.
|
||||
- **No per-player filtering.** Consumers filter inside the callback, comparing a
|
||||
target field against their own owner **[measured]**.
|
||||
- **No cancellation.** Every matching listener receives every message.
|
||||
- **No recursion guard.** Broadcasting from inside a callback is supported and
|
||||
used deliberately by an aggregator in the product **[measured]**; nothing bounds
|
||||
the depth if the channel graph ever contains a cycle. MB-12.
|
||||
- **No thread safety.** Game thread only **[measured]**.
|
||||
- **No networking.** §7.
|
||||
|
||||
---
|
||||
|
||||
## 9. The architecture it made possible
|
||||
|
||||
Worth stating, because the failure-mode list above is not an argument against the
|
||||
pattern:
|
||||
|
||||
```text
|
||||
ability system --damage / elimination-->
|
||||
processors in a feature plugin --assist / chain / streak-->
|
||||
scripted relay --notification-->
|
||||
UI widget
|
||||
```
|
||||
|
||||
No layer references the next **[measured]**. The processors live in a feature
|
||||
plugin that the base module has never heard of, and they are added and removed
|
||||
with the feature. **[derived]** That is the whole case for a bus: it let
|
||||
game-mode-specific logic move into a shippable plugin without creating a single
|
||||
reference from the base module.
|
||||
|
||||
Note also the vocabulary discipline that came with it: channels named so bus
|
||||
traffic is distinguishable from other tags, tags declared next to the publisher,
|
||||
payload fields initialised at declaration because a payload never passes through
|
||||
serialisation and gets no free defaults **[measured]**.
|
||||
|
||||
And the cost of that discipline, measured: one elimination channel string was
|
||||
declared independently at **four** sites — the publisher and three processors
|
||||
**[measured]**. It works because the engine deduplicates by string. **[derived]**
|
||||
It means the string is the real contract and a single-site edit detaches one
|
||||
participant silently. Recipe: MB-10.
|
||||
|
||||
---
|
||||
|
||||
## 10. What source reading could not settle
|
||||
|
||||
Stated because it bounds several claims above:
|
||||
|
||||
- **Blueprint call sites are invisible to text search.** Three RPC wrappers and
|
||||
one notification channel have zero C++ callers or publishers, and all are
|
||||
script-callable **[measured]**. Whether they are used, and by what, is
|
||||
**[open]**. They are recorded as open questions, not as dead code.
|
||||
- **The extent of scripted bus usage is unknown.** A search for the listen node
|
||||
across all source returns nothing **[measured]**, which means the entire
|
||||
scripted traffic of the bus lives in binary assets.
|
||||
- **The channel-key defect was found by reading, not by running.** It has not been
|
||||
reproduced at runtime **[open]**, and its trigger conditions do not occur in the
|
||||
audited project.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against the gameplay message router plugin and its consumers in Epic's
|
||||
Lyra Starter Game on Unreal Engine 5.6, read as source in a single workspace. The
|
||||
runtime module was read in full, line by line; the negative results in §7 and §8
|
||||
rest on that rather than on a search.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; each
|
||||
`MB-` identifier resolves back to the audited location there, so any specific
|
||||
claim above can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One plugin, one engine version, one workspace. These are examples and failure
|
||||
evidence, not guarantees about other versions. Binary assets were not read, so
|
||||
every statement about scripted *usage* is bounded by what source could show and is
|
||||
marked open where it is open. Re-run the recipes in
|
||||
[failure-modes.md](failure-modes.md) against your own tree before acting on
|
||||
anything here.
|
||||
@@ -0,0 +1,386 @@
|
||||
---
|
||||
name: ue-gameplay-tag-governance
|
||||
description: >-
|
||||
Govern GameplayTags in Unreal Engine as a typed vocabulary rather than loose
|
||||
strings: choosing declaration source (native macro, project ini, plugin ini,
|
||||
DataTable, platform ini, runtime), namespace design, constraining tag inputs
|
||||
with meta=(Categories=...), exact vs hierarchical matching, redirects and
|
||||
renaming, replication of tag sets, tag-driven counters, and auditing an
|
||||
existing tag inventory. Use when adding a tag namespace, wiring systems
|
||||
through tags, reviewing tag usage, or debugging a tag-driven feature that
|
||||
silently does nothing.
|
||||
---
|
||||
|
||||
# UE gameplay tag governance
|
||||
|
||||
The invariant:
|
||||
|
||||
> A GameplayTag is a foreign key with no database behind it. Every tag join must
|
||||
> be validated on both sides by something other than the compiler.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Eighteen detection recipes for silent tag failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting or reviewing a tag-driven design. Each
|
||||
entry names the mechanism, why it stays silent, why the obvious check misses it,
|
||||
and a command you can run against your own tree.
|
||||
|
||||
Related skills: `ue-gameplay-messaging`, `ue-gas-architecture`,
|
||||
`ue-modular-gameplay`, `ue-input-architecture`, `ue-architecture-guardrails`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why tags need governance at all
|
||||
|
||||
A tag replaces a compile-time reference with a runtime string comparison. That is
|
||||
the whole point — and the whole cost. Once a system is joined by a tag:
|
||||
|
||||
- the compiler no longer sees the dependency;
|
||||
- "find all references" no longer finds it;
|
||||
- a typo is not an error, it is an empty result;
|
||||
- deleting the producer does not break the consumer's build.
|
||||
|
||||
The characteristic failure is therefore never a crash. It is **nothing happens**,
|
||||
reported weeks later by a designer, in a build nobody can bisect.
|
||||
|
||||
Governance means deliberately re-adding the checks the compiler stopped doing.
|
||||
|
||||
---
|
||||
|
||||
## 2. Decide the declaration source before the first tag
|
||||
|
||||
Unreal accepts tags from at least six independent sources. They do not behave
|
||||
the same and they are not equally visible to search.
|
||||
|
||||
| Source | Form | Compile-time symbol | Found by grep pattern |
|
||||
|---|---|---|---|
|
||||
| Native, registered | `UE_DEFINE_GAMEPLAY_TAG_COMMENT` + `_EXTERN` in a header | yes | symbol name |
|
||||
| Native, file-local | `UE_DEFINE_GAMEPLAY_TAG_STATIC` | file-local only | symbol name |
|
||||
| Project config | `+GameplayTagList=(Tag="...")` in the project tag ini | no | `+GameplayTagList=` |
|
||||
| Plugin config | `GameplayTagList=(Tag="...")` — **no leading `+`**, different section | no | `GameplayTagList=(Tag=` |
|
||||
| DataTable | binary asset referenced by `+GameplayTagTableList` | no | not text-searchable |
|
||||
| Runtime | `AddNativeGameplayTag` during module startup | no | call site only |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Default to native, registered tags** for anything C++ reads or writes. They
|
||||
give Ctrl+Click, rename support, and a link error instead of a silent miss.
|
||||
2. **Use ini only for namespaces owned by data**, where designers add entries
|
||||
without engineers — cosmetic variants, surface types, platform traits.
|
||||
3. **Never split one namespace across sources.** A namespace half in C++ and half
|
||||
in ini has no single place to read its vocabulary.
|
||||
4. **Record the source of every namespace** in one document. This is the only
|
||||
defence against source #6, which no static inventory can see.
|
||||
5. **Never trust a single grep pattern** when auditing. Plugin ini syntax differs
|
||||
from project ini syntax by exactly one character, and a pattern anchored on
|
||||
`+` returns a confident, wrong zero. Recipe: TG-05.
|
||||
|
||||
---
|
||||
|
||||
## 3. Design the namespace as a tree that means something
|
||||
|
||||
Hierarchy is not decoration. `MatchesTag` and partial matching turn parentage into
|
||||
executable logic: a system asking "is this an action?" will answer yes for every
|
||||
descendant, forever, including ones added next year by someone who never read that
|
||||
call site.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **A parent tag is a promise.** Adding a child enrols it in every hierarchical
|
||||
check that already matches the parent. Only create a child if that enrolment
|
||||
is intended.
|
||||
2. **Do not encode two dimensions in one path.** `Weapon.Rifle.Silenced` conflates
|
||||
class and modifier; the modifier cannot then be queried independently. Use two
|
||||
tags.
|
||||
3. **Depth is coupling.** Every additional level is another place a hierarchical
|
||||
query can accidentally match.
|
||||
4. **Reserve a leaf-only convention** for tags used as identities (channels, cue
|
||||
names, state nodes). If a tag is an identity, nothing should ever be its child.
|
||||
5. **Document each parent's contract** where the parent is declared — what a
|
||||
descendant is required to satisfy.
|
||||
|
||||
### Naming
|
||||
|
||||
Pick one scheme and enforce it in review:
|
||||
|
||||
```text
|
||||
<Domain>.<Category>.<Specific>
|
||||
|
||||
Ability.Type.Action.Reload
|
||||
InitState.DataAvailable
|
||||
Event.Combat.Elimination
|
||||
Platform.Trait.Input.PrimaryController
|
||||
```
|
||||
|
||||
Do not abbreviate inconsistently. Do not mix singular and plural at the same
|
||||
level. These are unenforceable by tooling and therefore purely a review duty.
|
||||
|
||||
---
|
||||
|
||||
## 4. Constrain every tag input
|
||||
|
||||
An unconstrained `FGameplayTag` UPROPERTY offers the designer the entire project
|
||||
vocabulary — thousands of entries, of which a handful are legal. This is the
|
||||
single highest-leverage guardrail available, and it is one line:
|
||||
|
||||
```cpp
|
||||
UPROPERTY(EditAnywhere, meta = (Categories = "Ability.Type.Action"))
|
||||
FGameplayTag ActionTag;
|
||||
|
||||
UPROPERTY(EditAnywhere, meta = (Categories = "Event.Combat"))
|
||||
FGameplayTagContainer ObservedChannels;
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Every editable tag property gets `Categories`.** No exceptions for "it's
|
||||
obvious" — obvious is exactly what a typo defeats.
|
||||
2. **`UPARAM(meta=(Categories=...))`** for `BlueprintCallable` tag arguments.
|
||||
3. **Verify the constraint actually resolves.** `Categories` pointing at a
|
||||
non-existent namespace produces an *empty* dropdown, not an error. It looks
|
||||
like a working constraint and is worse than none, because it silently blocks
|
||||
the designer while passing review. Recipe: TG-03.
|
||||
4. **Re-verify `Categories` after any namespace rename.** It is a string; the
|
||||
rename tool does not touch it.
|
||||
5. **Measure coverage**: count tag properties versus properties carrying
|
||||
`Categories`. Anything below full coverage is a list of known holes, and
|
||||
should be tracked as such. Recipe: TG-04.
|
||||
|
||||
### Verification procedure for a constraint
|
||||
|
||||
Open the property in the editor. If the dropdown is empty, the namespace does not
|
||||
exist. This takes five seconds and is the only check that exists.
|
||||
|
||||
---
|
||||
|
||||
## 5. Choose exact versus hierarchical matching per call site
|
||||
|
||||
Three different operations look alike and mean different things:
|
||||
|
||||
| Call | Semantics |
|
||||
|---|---|
|
||||
| `A == B` | exact identity |
|
||||
| `A.MatchesTagExact(B)` | exact identity, explicit |
|
||||
| `A.MatchesTag(B)` | A is B **or a descendant of** B |
|
||||
| `Container.HasTag(B)` | any element matches B or descends from it |
|
||||
| `Container.HasTagExact(B)` | any element is exactly B |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **State the intent at every call site.** Prefer the explicit `...Exact` form
|
||||
even when `==` would do, so the next reader does not have to infer.
|
||||
2. **Never mix the two forms across one feature's read and write paths.** If a
|
||||
rule is stored hierarchically and queried exactly (or the reverse), the feature
|
||||
works for the tags that were tested and fails for the rest — the worst possible
|
||||
failure distribution. Recipe: TG-07.
|
||||
3. **Hierarchical matching is a public API.** Once a system matches on a parent,
|
||||
the parent's whole subtree is part of that system's contract.
|
||||
4. **Default to exact** for identities; use hierarchical only where "any kind of
|
||||
X" is genuinely the rule.
|
||||
|
||||
---
|
||||
|
||||
## 6. Renaming: redirects are not optional
|
||||
|
||||
A tag lives in binary assets, ini files, DataTables, and blueprint graphs. A
|
||||
rename without a redirect silently zeroes every one of those references.
|
||||
|
||||
```ini
|
||||
[/Script/GameplayTags.GameplayTagsSettings]
|
||||
+GameplayTagRedirects=(OldTagName="Ability.Type.Action.Fire",NewTagName="Ability.Type.Action.WeaponFire")
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Add the redirect in the same commit as the rename.** Not after; a build
|
||||
between the two loses data permanently in any asset that gets resaved.
|
||||
2. **Never delete a redirect** until every asset in every branch has been resaved
|
||||
and verified. Redirects are cheap; lost designer work is not.
|
||||
3. **A project with zero redirects and a non-trivial tag history has already lost
|
||||
references** — or has never renamed a tag, which is itself worth confirming.
|
||||
Recipe: TG-18.
|
||||
4. **Typos are renames.** Fixing a misspelled tag requires a redirect exactly like
|
||||
any other rename.
|
||||
|
||||
### The symmetric-typo trap
|
||||
|
||||
A misspelled tag that is misspelled *identically* in every location works
|
||||
perfectly. It will keep working until one person fixes one of the spellings —
|
||||
at which point the join breaks with no error at either end. Treat a discovered
|
||||
typo as a coordinated multi-file change with a redirect, never as a local fix.
|
||||
Recipe: TG-02.
|
||||
|
||||
---
|
||||
|
||||
## 7. Replication of tag vocabulary
|
||||
|
||||
Tags replicate as indices, not strings, when fast replication is enabled. That
|
||||
requires **client and server to have byte-identical tag registries**.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Any source of tags that can differ between builds breaks fast replication.**
|
||||
That includes plugin-supplied tags when the plugin loads on one side only, and
|
||||
any runtime tag registration.
|
||||
2. **Register tags before networking starts**, unconditionally on both sides.
|
||||
3. **Watch the replicated container size cap.** It limits how many tags fit in one
|
||||
replicated container; exceeding it is a runtime failure, not a compile error.
|
||||
Recipe: TG-11.
|
||||
4. **Do not gate tag registration on `IsRunningDedicatedServer()` or platform.**
|
||||
Gate the *content* those tags refer to, never the vocabulary.
|
||||
|
||||
---
|
||||
|
||||
## 8. Tag-driven counters and state
|
||||
|
||||
Tag-to-count containers (stacks) are a common pattern for ammo, charges,
|
||||
resources. They are arithmetic wearing a tag as a label, and inherit arithmetic's
|
||||
problems.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Validate the amount, not only the tag.** A zero or negative delta that
|
||||
silently returns is a lost write that reports as "the ability did nothing".
|
||||
Recipe: TG-12.
|
||||
2. **Guard the accumulator.** Unchecked `int32` addition on a stack count is an
|
||||
overflow waiting for an idle-game or a modifier loop. Recipe: TG-13.
|
||||
3. **Decide the contract for removing more than exists**, then implement it —
|
||||
clamp to zero, fail loudly, or allow negative. Silence is not a contract.
|
||||
4. **Keep the map and the array in sync in every replication callback** if the
|
||||
container maintains both a replicated array and a lookup map. Every
|
||||
`PreReplicatedRemove` / `PostReplicatedAdd` / `PostReplicatedChange` must
|
||||
update the map, or clients read stale counts.
|
||||
|
||||
---
|
||||
|
||||
## 9. String-to-tag lookup is a debug tool
|
||||
|
||||
Resolving a tag from a runtime string — especially with partial/substring
|
||||
matching — returns the first registry match in registry order. That order is not
|
||||
stable against adding unrelated tags.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Confine partial matching to cheat/console/debug paths.**
|
||||
2. **Do not export a partial-match helper with a permissive default** across a
|
||||
module boundary. The first production caller who passes `true` gets a
|
||||
non-deterministic tag and no warning. Recipe: TG-09.
|
||||
3. **Exact lookup with explicit failure handling** is the only form allowed in
|
||||
gameplay code.
|
||||
|
||||
---
|
||||
|
||||
## 10. Auditing an existing tag inventory
|
||||
|
||||
Run this before trusting any tag-driven system you did not write. Each step maps
|
||||
to a recipe in the failure modes.
|
||||
|
||||
```text
|
||||
1. Count declarations per source (all six). Reconcile against the documented
|
||||
list. Count with two patterns, not one. TG-05, TG-06
|
||||
2. Count editable tag properties; count those with Categories.
|
||||
Report the ratio. TG-04
|
||||
3. Resolve every Categories string against the actual namespace list.
|
||||
Empty dropdown = broken constraint. TG-03
|
||||
4. List every parent tag used in a hierarchical match; enumerate its
|
||||
current descendants and confirm each one belongs. TG-08
|
||||
5. Search for tags declared but never read, and read but never declared. TG-16
|
||||
6. Count redirects. Zero redirects plus renamed tags in history =
|
||||
data loss. TG-18
|
||||
7. Diff the tag registry between a client build and a server build. TG-10
|
||||
```
|
||||
|
||||
Steps 1 and 3 find most real defects. Step 4 finds the ones that will appear
|
||||
later, when someone adds a child.
|
||||
|
||||
---
|
||||
|
||||
## 11. Review checklist
|
||||
|
||||
- [ ] Namespace has one documented declaration source.
|
||||
- [ ] Tags read by C++ are native, not ini-only strings.
|
||||
- [ ] Every editable tag property/parameter carries `Categories`.
|
||||
- [ ] Every `Categories` string resolves to a namespace that exists.
|
||||
- [ ] Exact vs hierarchical matching is explicit and consistent per feature.
|
||||
- [ ] Parent tags used in hierarchical matches have a documented contract.
|
||||
- [ ] Renames ship with redirects in the same commit.
|
||||
- [ ] Tag registration is unconditional and identical on client and server.
|
||||
- [ ] Tag counters validate amount and guard overflow.
|
||||
- [ ] Partial string lookup exists only in debug paths.
|
||||
- [ ] A tag-driven feature has a test that fails when the tag is wrong.
|
||||
|
||||
---
|
||||
|
||||
## 12. Tests
|
||||
|
||||
The purpose of these tests is to convert "nothing happened" into a red build.
|
||||
|
||||
### Vocabulary
|
||||
|
||||
- every tag referenced in C++ resolves at startup;
|
||||
- every tag in ini is referenced somewhere, or is explicitly marked data-owned;
|
||||
- registry contents identical between client and server targets.
|
||||
|
||||
### Matching
|
||||
|
||||
- exact query does not match a descendant;
|
||||
- hierarchical query matches every current descendant;
|
||||
- adding a new child does not change any exact query's result.
|
||||
|
||||
### Constraints
|
||||
|
||||
- each `Categories` string yields a non-empty candidate set;
|
||||
- a property rejects a tag outside its category.
|
||||
|
||||
### Counters
|
||||
|
||||
- add zero, add negative, remove more than present, remove from empty;
|
||||
- replication callbacks leave map and array consistent;
|
||||
- large accumulations do not overflow.
|
||||
|
||||
### Migration
|
||||
|
||||
- renaming a tag with a redirect preserves asset references;
|
||||
- loading an asset saved with the old name resolves to the new tag.
|
||||
|
||||
---
|
||||
|
||||
## 13. When not to use a tag
|
||||
|
||||
Use a stronger construct when:
|
||||
|
||||
- the set of values is closed and owned by code → **enum**;
|
||||
- the value selects an implementation → **class reference or DataAsset**;
|
||||
- the join must be checked at compile time → **interface or typed pointer**;
|
||||
- exactly one consumer exists and it is in the same module → **direct call**.
|
||||
|
||||
A tag earns its cost when the vocabulary is open, the consumers are unknown to the
|
||||
producer, and the join genuinely crosses a module or data boundary. Used anywhere
|
||||
else, it converts a compile error into a silent runtime no-op — which is the only
|
||||
thing it reliably does.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
Epic's Lyra Starter Game on Unreal Engine 5.6, read rather than run. Source
|
||||
addresses stay in the research archive that produced this skill; what ships is the
|
||||
detection recipe, because an address in someone else's tree is not something you
|
||||
can act on and a recipe is.
|
||||
|
||||
Each entry carries a stable identifier (`TG-01`, `TG-02`, …) that resolves back to
|
||||
the audited location in that archive. If you need the original address to settle a
|
||||
dispute, it exists and can be produced.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one specific reference project on one engine
|
||||
version in one workspace. It is evidence of failure modes, not a guarantee about
|
||||
other engine versions or other samples. Re-run the recipes against your own tree
|
||||
before trusting any specific claim — including the counts, which are properties of
|
||||
that project on that day and are quoted only to show the shape of the contrast a
|
||||
recipe produces.
|
||||
+670
@@ -0,0 +1,670 @@
|
||||
# Failure modes: gameplay tag governance
|
||||
|
||||
Eighteen ways a tag-driven system stops working without reporting anything.
|
||||
|
||||
The common shape of every entry below: **the system does not fail, it declines to
|
||||
act.** Tag defects do not produce crashes, log errors or failed builds. A missing
|
||||
join and a correctly-evaluated "no" are the same observable. Design the gates
|
||||
accordingly — the only reliable detection is a test that asserts the *positive*
|
||||
behaviour, or a recipe that compares two counts.
|
||||
|
||||
Recipes use `rg` and are written to be run from a project root. Most of them are
|
||||
deliberately a **pair** of searches: the class of defect here is "the thing is
|
||||
present, and the thing that would give it meaning is absent", and a single search
|
||||
can only ever show you the first half.
|
||||
|
||||
---
|
||||
|
||||
## Vocabulary and declaration
|
||||
|
||||
### TG-01 - Typo in a string-declared tag
|
||||
|
||||
**Mechanism.** A tag declared in ini, or authored into an asset by hand, is
|
||||
misspelled. No compiler, linker or loader rejects it, because a tag is a string
|
||||
until the moment something tries to match it.
|
||||
|
||||
**Why it is silent.** The lookup succeeds in the sense that matters to the engine:
|
||||
it returns "no match", which is a legal answer that the calling system is required
|
||||
to handle gracefully. A misspelled tag is indistinguishable from a condition that
|
||||
is legitimately not met right now.
|
||||
|
||||
**Why the obvious check misses it.** Review reads the tag and it looks right —
|
||||
misspellings survive review precisely because they are near-misses. Compile,
|
||||
cook and startup all pass. The only artefact that would reveal it is a runtime
|
||||
comparison that nobody logs, because logging every failed tag match would drown
|
||||
the log.
|
||||
|
||||
**Symptom.** The feature does nothing. No log line, no error. Usually reported by
|
||||
a designer, weeks later, as "this never worked" — by which time the change that
|
||||
introduced it is far outside any bisect window.
|
||||
|
||||
**Detect.** Resolve the declared vocabulary against the referenced one:
|
||||
|
||||
```bash
|
||||
# every tag string declared in config
|
||||
rg -o -N 'Tag="([^"]+)"' -r '$1' Config/ Plugins/*/Config/ | sort -u > /tmp/declared
|
||||
# every tag string mentioned in source
|
||||
rg -o -N '"([A-Z][A-Za-z]+(\.[A-Za-z][A-Za-z0-9]*)+)"' -r '$1' Source/ | sort -u > /tmp/used
|
||||
comm -13 /tmp/declared /tmp/used # referenced but never declared
|
||||
```
|
||||
|
||||
The third command is the finding: a tag string that code compares against and
|
||||
that no source declares.
|
||||
|
||||
**Guardrail.** Native tag macros for anything C++ reads, so a typo becomes a link
|
||||
error. `Categories` constraints on every editable tag property so authoring uses a
|
||||
dropdown rather than a keyboard.
|
||||
|
||||
---
|
||||
|
||||
### TG-02 - Symmetric typo
|
||||
|
||||
**Mechanism.** The same misspelling exists at both ends of a join. The feature
|
||||
works. A later cleanup corrects one side.
|
||||
|
||||
**Why it is silent.** It is not merely silent — it is *correct*. Both sides agree,
|
||||
the join resolves, the feature ships and works for years. Nothing is wrong until
|
||||
someone makes it right.
|
||||
|
||||
**Why the obvious check misses it.** Every check that exists passes, because the
|
||||
system is working. The defect has no runtime signature at all; it is latent in the
|
||||
spelling, and it converts into a failure only through a well-intentioned edit
|
||||
whose diff looks obviously correct — which is why it also survives review of the
|
||||
fix.
|
||||
|
||||
**Symptom.** A previously working feature breaks, and the commit that broke it is
|
||||
a one-word spelling correction that every reviewer approves.
|
||||
|
||||
**Detect.** Before correcting any tag spelling, count the sites:
|
||||
|
||||
```bash
|
||||
T='Primarly' # the suspected misspelling
|
||||
rg -n --no-heading "$T" Config/ Source/ Plugins/ | tee /tmp/sites | wc -l
|
||||
```
|
||||
|
||||
On the measured reference this returned six sites across four files: two project
|
||||
tag declarations, two platform-specific config files, and a native tag macro plus
|
||||
its use in one widget class. Correcting any one of them would have broken the
|
||||
join with no error at either end.
|
||||
|
||||
**Guardrail.** Treat a discovered typo as a coordinated multi-site rename plus a
|
||||
redirect, never as a local edit. The rule generalises: **when a string literal is
|
||||
the join, correctness means "all sides agree", not "the string is right".**
|
||||
|
||||
---
|
||||
|
||||
### TG-03 - `Categories` filter pointing at a namespace that does not exist
|
||||
|
||||
**Mechanism.** The `Categories` metadata string is not validated against the tag
|
||||
tree. A stale or mistyped namespace yields an empty candidate set.
|
||||
|
||||
**Why it is silent.** An empty dropdown is a legitimate state — it is what a
|
||||
correctly-configured filter shows before anyone has created tags under that
|
||||
namespace. The editor has no way to distinguish "this filter is wrong" from "this
|
||||
namespace is empty today".
|
||||
|
||||
**Why the obvious check misses it.** The property has a constraint, which is the
|
||||
thing reviewers check for, and it is spelled plausibly. Confirming it requires
|
||||
opening the property in the editor and looking at the dropdown — a five-second
|
||||
check that nothing prompts anyone to perform, because the code reads as complete.
|
||||
|
||||
**Symptom.** Designers see an empty dropdown, conclude the tags are not ready yet,
|
||||
and leave the field blank. The feature is then wired to an unset tag, which is the
|
||||
same failure as TG-01 arriving by a different route.
|
||||
|
||||
**Detect.** Extract every filter string and confirm each resolves:
|
||||
|
||||
```bash
|
||||
rg -o -N 'Categories\s*=\s*"([^"]+)"' -r '$1' Source/ | sort -u > /tmp/filters
|
||||
while read -r ns; do
|
||||
n=$(rg -c "Tag=\"$ns" Config/ Plugins/*/Config/ 2>/dev/null | awk -F: '{s+=$2} END {print s+0}')
|
||||
[ "$n" = "0" ] && echo "EMPTY FILTER: $ns"
|
||||
done < /tmp/filters
|
||||
```
|
||||
|
||||
On the measured reference one filter named a namespace that appeared nowhere else
|
||||
in the project, while the real tags lived under a differently-named root. That
|
||||
filter was one of only fourteen constrained properties in the entire codebase —
|
||||
one of the few places tag input was validated at all, and it was misconfigured.
|
||||
|
||||
**Guardrail.** Open every constrained property once and confirm the dropdown is
|
||||
non-empty. Re-check after any namespace rename: `Categories` is a string, and no
|
||||
rename tool touches it.
|
||||
|
||||
---
|
||||
|
||||
### TG-04 - Unconstrained tag property
|
||||
|
||||
**Mechanism.** A tag field without `Categories` offers the entire project
|
||||
vocabulary — thousands of entries, of which a handful are legal here.
|
||||
|
||||
**Why it is silent.** Any tag the designer picks is a valid tag. It serialises, it
|
||||
loads, it compares. The system that reads it simply never matches, which is the
|
||||
TG-01 observable again.
|
||||
|
||||
**Why the obvious check misses it.** There is nothing to see. The absence of one
|
||||
optional metadata specifier does not appear in a diff as a change, does not warn,
|
||||
and does not look like an omission in review — most properties do not have it
|
||||
either, so it reads as normal.
|
||||
|
||||
**Symptom.** Wrong-namespace tags assigned in assets; the system ignores them
|
||||
silently. The designer's mental model — "I set the tag, so it is wired" — is
|
||||
reasonable and wrong.
|
||||
|
||||
**Detect.** Report the coverage ratio, which is the only number that matters:
|
||||
|
||||
```bash
|
||||
total=$(rg -c "FGameplayTag(Container)?\s+\w+;" Source/ | awk -F: '{s+=$2} END {print s+0}')
|
||||
constrained=$(rg -c "meta\s*=\s*\([^)]*Categories" Source/ | awk -F: '{s+=$2} END {print s+0}')
|
||||
echo "constrained $constrained of $total"
|
||||
```
|
||||
|
||||
On the measured reference this was fourteen of sixty-four — under sixteen per cent
|
||||
— and one of the fourteen was broken per TG-03. A low ratio is not itself a bug;
|
||||
it is a list of places where a typo cannot be caught.
|
||||
|
||||
**Guardrail.** `meta=(Categories=...)` on every editable tag property and `UPARAM`
|
||||
on every Blueprint-exposed tag argument. Track the ratio as a number that only
|
||||
goes up.
|
||||
|
||||
---
|
||||
|
||||
### TG-05 - A grep pattern that misses a declaration source
|
||||
|
||||
**Mechanism.** Plugin tag config uses `GameplayTagList=` without the leading `+`
|
||||
and under a different section header than project config. A pattern anchored on
|
||||
`+` never sees it.
|
||||
|
||||
**Why it is silent.** The audit succeeds. It produces a number, the number is
|
||||
plausible, and nothing indicates it is partial. A confident wrong answer is
|
||||
produced by a working tool.
|
||||
|
||||
**Why the obvious check misses it.** The obvious check *is* the grep, and greps
|
||||
do not report what they could not have matched. This is the general hazard of
|
||||
proving absence by search: **an empty or small result proves the pattern, not the
|
||||
codebase.**
|
||||
|
||||
**Symptom.** An audit reports a small tag vocabulary and concludes the project is
|
||||
safer than it is. Renames and migrations are then planned against an inventory
|
||||
that is missing most of its entries.
|
||||
|
||||
**Detect.** Count with two patterns and require them to agree:
|
||||
|
||||
```bash
|
||||
rg -c '^\+GameplayTagList=' Config/ # project form
|
||||
rg -c 'GameplayTagList=\(Tag=' Config/ Plugins/*/Config/ # both forms
|
||||
```
|
||||
|
||||
On the measured reference the first pattern found 99 tags in one file; the second
|
||||
found 212 across five files, four of which were plugin configs. The difference is
|
||||
not an edge case — it is the majority of the vocabulary.
|
||||
|
||||
**Guardrail.** Reconcile every count against a documented per-namespace source
|
||||
list. Never conclude "absent" from one pattern; widen it, then widen the path.
|
||||
|
||||
---
|
||||
|
||||
### TG-06 - Tag declared in a source no static search can see
|
||||
|
||||
**Mechanism.** DataTable-sourced tags live in binary assets. Runtime registration
|
||||
leaves no declaration text at all.
|
||||
|
||||
**Why it is silent.** These tags work perfectly. They resolve, match and replicate
|
||||
like any other. They are invisible only to *tooling*, and tooling is not what
|
||||
exercises them.
|
||||
|
||||
**Why the obvious check misses it.** Every text-based inventory is structurally
|
||||
incapable of seeing them. Widening the grep does not help, because there is no
|
||||
text to find — the failure is in the choice of instrument, not its configuration.
|
||||
|
||||
**Symptom.** Tags exist at runtime that no inventory, document or search reports.
|
||||
A rename sweep misses them; a "delete unused tags" pass deletes something that is
|
||||
in use.
|
||||
|
||||
**Detect.** Ask the runtime, not the filesystem:
|
||||
|
||||
```bash
|
||||
rg -n "GameplayTagTableList|AddNativeGameplayTag|AddTagTableRow" Config/ Source/
|
||||
```
|
||||
|
||||
Any hit means your static inventory is incomplete by construction. Close the gap
|
||||
by dumping the tag registry from a running build and diffing it against the static
|
||||
inventory; the difference is exactly the set this entry describes.
|
||||
|
||||
**Guardrail.** Document every namespace's declaration source in one place. Forbid
|
||||
runtime registration outside module startup, so at least the timing is knowable.
|
||||
|
||||
---
|
||||
|
||||
## Matching semantics
|
||||
|
||||
### TG-07 - Exact/hierarchical mismatch inside one feature
|
||||
|
||||
**Mechanism.** One code path uses hierarchical matching and another uses exact
|
||||
comparison, on the same data, sometimes in the same class.
|
||||
|
||||
**Why it is silent.** Both calls are correct in isolation and both are correct for
|
||||
the tags used during development — typically leaf tags, where hierarchical and
|
||||
exact matching give identical answers. The two forms only diverge for a tag that
|
||||
has descendants, which is a case that arrives later.
|
||||
|
||||
**Why the obvious check misses it.** Reading either call site shows a defensible
|
||||
choice. The defect exists only in the relationship between two call sites that may
|
||||
be hundreds of lines apart, and no tool reports "these two comparisons on the same
|
||||
field disagree about semantics".
|
||||
|
||||
**Symptom.** The rule works for the tags used in testing and fails for others —
|
||||
the failure distribution that is hardest to reproduce, because the reporter's
|
||||
repro steps look identical to the ones that pass.
|
||||
|
||||
**Detect.** List both forms per file and look for files that contain both:
|
||||
|
||||
```bash
|
||||
rg -l "MatchesTag\(" Source/ > /tmp/hier
|
||||
rg -l "MatchesTagExact\(|HasTagExact\(" Source/ > /tmp/exact
|
||||
comm -12 <(sort /tmp/hier) <(sort /tmp/exact)
|
||||
```
|
||||
|
||||
Files in the intersection are candidates. On the measured reference one mapping
|
||||
class used hierarchical matching in two methods and exact matching in a third, all
|
||||
against the same tag container.
|
||||
|
||||
**Guardrail.** State matching intent explicitly at each call site — prefer the
|
||||
`...Exact` form even where `==` would do. Review the read and write paths of a
|
||||
rule together, never separately.
|
||||
|
||||
---
|
||||
|
||||
### TG-08 - Unintended enrolment via a new child tag
|
||||
|
||||
**Mechanism.** A system matches on a parent tag. Adding a child anywhere in the
|
||||
project enrols that child in the system automatically.
|
||||
|
||||
**Why it is silent.** The enrolment is the feature working as designed. Nothing
|
||||
malfunctions; a new tag simply acquires behaviour its author did not know about,
|
||||
and that behaviour is usually plausible.
|
||||
|
||||
**Why the obvious check misses it.** The person adding the tag is looking at the
|
||||
tag list, not at every hierarchical match in the codebase. The person who wrote
|
||||
the match is not present. There is no declaration anywhere that says "this parent
|
||||
is load-bearing", so the coupling exists only in the reader's memory.
|
||||
|
||||
**Symptom.** A newly added, unrelated tag activates behaviour nobody wired, in a
|
||||
subsystem nobody touched, in the same release as dozens of other changes.
|
||||
|
||||
**Detect.** Enumerate the parents that carry behaviour, then their descendants:
|
||||
|
||||
```bash
|
||||
rg -o -N 'MatchesTag\(([A-Za-z_:]+)\)' -r '$1' Source/ | sort -u # load-bearing parents
|
||||
# for each parent P, list current descendants:
|
||||
rg -o -N 'Tag="(P\.[^"]+)"' -r '$1' Config/ Plugins/*/Config/
|
||||
```
|
||||
|
||||
Read the descendant list and confirm every entry belongs in that subsystem. Do
|
||||
this at the moment you add a hierarchical match, and record the result next to
|
||||
the parent declaration.
|
||||
|
||||
**Guardrail.** Document each parent's contract at its declaration and treat parent
|
||||
tags used in hierarchical matches as public API. Reserve a leaf-only convention
|
||||
for tags that are identities.
|
||||
|
||||
---
|
||||
|
||||
### TG-09 - Partial string lookup in gameplay code
|
||||
|
||||
**Mechanism.** A helper resolves a tag from a substring, returning the first match
|
||||
in registry order.
|
||||
|
||||
**Why it is silent.** It returns a real, valid tag. Everything downstream works
|
||||
normally with it. The answer is only wrong in the sense that it is not the answer
|
||||
you would have chosen, and there is no way to notice that from the outside.
|
||||
|
||||
**Why the obvious check misses it.** The function has a plausible name and a
|
||||
`bool` parameter that defaults to the safe behaviour. Reviewing the helper shows a
|
||||
correct implementation of what it says it does. The hazard is in the *export* —
|
||||
it crosses a module boundary with a permissive mode available — and export
|
||||
surfaces are reviewed for compilation, not for misuse potential.
|
||||
|
||||
**Symptom.** The resolved tag changes when an unrelated tag is added elsewhere in
|
||||
the project. Non-deterministic behaviour, no error, and a repro that depends on
|
||||
the contents of the tag registry rather than on anything in the feature.
|
||||
|
||||
**Detect.** Find partial-match helpers and check who calls them with matching
|
||||
enabled:
|
||||
|
||||
```bash
|
||||
rg -n "PartialMatch|bMatchPartialString|FindTagByString" Source/
|
||||
rg -n "FindTagByString\([^)]*true" Source/ # callers opting into partial
|
||||
```
|
||||
|
||||
On the measured reference the helper was exported from the game module with the
|
||||
permissive parameter defaulting to off, and its only current callers were cheat
|
||||
commands passing it on — safe today, one production caller away from
|
||||
non-deterministic.
|
||||
|
||||
**Guardrail.** Confine partial matching to cheat and console paths. Do not export
|
||||
such a helper across a module boundary at all; if you must, do not give it a
|
||||
permissive mode.
|
||||
|
||||
---
|
||||
|
||||
## Replication
|
||||
|
||||
### TG-10 - Divergent tag registries under fast replication
|
||||
|
||||
**Mechanism.** Fast replication sends tags by registry index. Client and server
|
||||
must hold identical registries for an index to mean the same thing on both ends.
|
||||
|
||||
**Why it is silent.** Indices are always valid. A mismatched index does not fail
|
||||
to deserialise — it resolves to a different, perfectly real tag. The receiving
|
||||
side then acts correctly on the wrong fact.
|
||||
|
||||
**Why the obvious check misses it.** PIE shares one process and therefore one
|
||||
registry, so the divergence cannot exist in the environment where testing happens.
|
||||
Any check performed in the editor is performed on a configuration where the bug is
|
||||
impossible by construction.
|
||||
|
||||
**Symptom.** Wrong tags received, or replication failures, only in networked
|
||||
builds with asymmetric plugin loading — and the symptom is a wrong behaviour, not
|
||||
an error.
|
||||
|
||||
**Detect.** Find every conditional path into tag registration:
|
||||
|
||||
```bash
|
||||
rg -n "FastReplication" Config/
|
||||
rg -n -B4 "AddNativeGameplayTag|RegisterGameplayTag" Source/ \
|
||||
| rg "IsRunningDedicatedServer|WITH_EDITOR|#if|IsClient|bIsServer"
|
||||
```
|
||||
|
||||
Any conditional wrapping tag registration is a divergence risk. Confirm by
|
||||
dumping the registry from a client build and a dedicated-server build and diffing
|
||||
them; that diff is the only authoritative answer.
|
||||
|
||||
**Guardrail.** Register vocabulary unconditionally on both sides. Gate the
|
||||
*content* a tag refers to, never the vocabulary itself.
|
||||
|
||||
---
|
||||
|
||||
### TG-11 - Replicated container size cap
|
||||
|
||||
**Mechanism.** A setting fixes the number of bits used for a replicated tag
|
||||
container's element count, which caps how many tags fit in one container.
|
||||
|
||||
**Why it is silent.** Under the cap, everything is correct. Over it, the excess
|
||||
entries are simply not transmitted — a truncation, not an error. The receiving
|
||||
side sees a valid container with fewer tags.
|
||||
|
||||
**Why the obvious check misses it.** The cap lives in a settings file as a bit
|
||||
count, not as a tag count, so its practical meaning requires a mental exponent.
|
||||
Nothing connects it to the container that will eventually exceed it, and the
|
||||
container that exceeds it does so only under gameplay conditions that accumulate
|
||||
tags.
|
||||
|
||||
**Symptom.** Tags beyond the cap are lost in transit with no compile-time or
|
||||
runtime signal. Debugging starts at the gameplay logic, which is correct.
|
||||
|
||||
**Detect.** Read the cap, convert it, and compare against your worst case:
|
||||
|
||||
```bash
|
||||
rg -n "NumBitsForContainerSize|FastReplication" Config/
|
||||
```
|
||||
|
||||
On the measured reference the value was 6 bits — a hard limit of 63 entries per
|
||||
replicated container. Then find containers that grow without bound:
|
||||
|
||||
```bash
|
||||
rg -n "AddTag\(|AppendTags\(" Source/ | rg -i "replicat|net"
|
||||
```
|
||||
|
||||
**Guardrail.** Know the cap as a tag count, not a bit count, and write it next to
|
||||
any container that grows dynamically. Assert the size where it grows.
|
||||
|
||||
---
|
||||
|
||||
## Counters
|
||||
|
||||
### TG-12 - Tag counter swallowing zero or negative deltas
|
||||
|
||||
**Mechanism.** A stack-modifying function guards on `if (Amount > 0)` with no
|
||||
`else`, so a zero or negative amount returns silently.
|
||||
|
||||
**Why it is silent.** Returning without acting is a legal response to a degenerate
|
||||
input, and the caller has no return value to check. The write is lost with the
|
||||
same observable as a write that was never attempted.
|
||||
|
||||
**Why the obvious check misses it.** The function *does* validate — there is a
|
||||
warning path for an invalid tag right beside the silent path for an invalid
|
||||
amount. Seeing the tag validation satisfies the reviewer's question "is this
|
||||
input checked?", and the asymmetry between the two checks is invisible unless you
|
||||
compare them deliberately.
|
||||
|
||||
**Symptom.** A lost write reported as "the ability did nothing". Because the tag
|
||||
was valid, the tag-level warning never fires, so the log is clean.
|
||||
|
||||
**Detect.** Compare validation coverage between the two parameters:
|
||||
|
||||
```bash
|
||||
rg -n -A6 "void .*Stack(Count|Tag)?\(|AddStack|RemoveStack" Source/ \
|
||||
| rg "if\s*\(.*Amount|ensure|UE_LOG"
|
||||
rg -n "//\s*@?TODO" Source/ | rg -i "stack|amount|negative|zero"
|
||||
```
|
||||
|
||||
On the measured reference the invalid-tag path carried warnings at both the add
|
||||
and remove sites while the invalid-amount path returned silently at both, and an
|
||||
unresolved TODO next to the remove path asked exactly the question this entry
|
||||
describes.
|
||||
|
||||
**Guardrail.** Validate the amount as strictly as the tag. Define and implement
|
||||
the contract for removing more than exists — clamp, fail loudly, or allow
|
||||
negative. Silence is not a contract.
|
||||
|
||||
---
|
||||
|
||||
### TG-13 - Unguarded accumulation on a stack count
|
||||
|
||||
**Mechanism.** Unchecked `int32` addition on a tag count, with no upper bound.
|
||||
|
||||
**Why it is silent.** Every individual addition is correct. The wrap requires an
|
||||
accumulation that no test performs, and after it wraps the value is still a valid
|
||||
integer that every subsequent comparison evaluates without complaint.
|
||||
|
||||
**Why the obvious check misses it.** Overflow guards are absent from most
|
||||
arithmetic in most codebases, so their absence here does not stand out. The
|
||||
addition is one line in a function whose logic is otherwise about tags, and the
|
||||
reviewer's attention is on the tag semantics.
|
||||
|
||||
**Symptom.** A count wraps to negative under a modifier loop or a long session.
|
||||
Every `>= required` check then fails, so the ability with the most resources
|
||||
behaves as if it had none.
|
||||
|
||||
**Detect.** Find unbounded accumulation on stack counts:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A2 "StackCount\s*\+=|Count\s*\+=\s*\w+" Source/ \
|
||||
| rg -v "Clamp|Min\(|check\(|ensure"
|
||||
```
|
||||
|
||||
Any addition with no clamp in the surrounding lines is the finding.
|
||||
|
||||
**Guardrail.** Clamp against a documented maximum on every add, and state the
|
||||
maximum where the container is declared.
|
||||
|
||||
---
|
||||
|
||||
## Stubs and documentation
|
||||
|
||||
### TG-14 - Tag-shaped property with no reader
|
||||
|
||||
**Mechanism.** A tag container is exposed for authoring while the code that would
|
||||
consume it is stubbed or deferred.
|
||||
|
||||
**Why it is silent.** An unread property is not an error at any level of the
|
||||
stack. It serialises, it appears in the editor, and it survives every cook. The
|
||||
authored value simply goes nowhere.
|
||||
|
||||
**Why the obvious check misses it.** The property exists, is documented by its
|
||||
name, and appears in the class next to properties that do work. "Is this
|
||||
implemented?" is answered by the presence of the field. The absent half is the
|
||||
*reader*, and no default tooling asks who reads a property.
|
||||
|
||||
**Symptom.** Designers configure the field; nothing happens. The feature appears
|
||||
implemented in the editor and in review, so the investigation starts from "the
|
||||
data must be wrong".
|
||||
|
||||
**Detect.** For each editable tag property, search for a reader:
|
||||
|
||||
```bash
|
||||
P='HiddenByTags' # substitute your property
|
||||
rg -n "\b$P\b" Source/ # declaration plus any uses
|
||||
rg -n "\b$P\b" Source/ | rg -v "UPROPERTY|^\s*FGameplayTag" # actual reads
|
||||
```
|
||||
|
||||
A declaration with no read is a stub. On the measured reference one widget base
|
||||
class exposed such a container while both of its consumer sites hardcoded the
|
||||
result of the check to a constant, with the real expression parked in a comment.
|
||||
|
||||
**Guardrail.** Do not ship an editable property whose consumer does not exist. If
|
||||
it must exist for serialisation reasons, mark it non-editable or log when it is
|
||||
non-empty.
|
||||
|
||||
---
|
||||
|
||||
### TG-15 - Restoring a stubbed tag check without restoring its inputs
|
||||
|
||||
**Mechanism.** The stub from TG-14 coexists with other damage: the values the
|
||||
restored check would select between are themselves overwritten at runtime.
|
||||
|
||||
**Why it is silent.** Both halves are individually invisible. The stub is silent
|
||||
per TG-14; the overwrite is a normal-looking assignment in an initialisation path.
|
||||
Neither produces a symptom while the feature is switched off.
|
||||
|
||||
**Why the obvious check misses it.** The obvious check is reading the stub, which
|
||||
tells you exactly what to implement. It does not tell you that the inputs to the
|
||||
thing you are about to implement have been clobbered somewhere else, because that
|
||||
assignment lives in a different method with a different purpose.
|
||||
|
||||
**Symptom.** The obvious one-line fix produces wrong behaviour instead of correct
|
||||
behaviour, and the fix gets blamed and reverted — burying the original defect
|
||||
under a failed attempt.
|
||||
|
||||
**Detect.** Before re-enabling any stub, verify every input it reads is still
|
||||
authored where it appears to be authored:
|
||||
|
||||
```bash
|
||||
V='ShownVisibility' # a field the stub would select between
|
||||
rg -n "\b$V\b\s*=" Source/ # who writes it, and when
|
||||
```
|
||||
|
||||
On the measured reference the designer-authored visibility fields were overwritten
|
||||
by a setter called during construction, so implementing the tag check alone would
|
||||
have selected between two values that no longer held the authored data.
|
||||
|
||||
**Guardrail.** Treat "implement the stub" as a task that starts with tracing the
|
||||
stub's inputs, not with writing the missing expression.
|
||||
|
||||
---
|
||||
|
||||
### TG-16 - Vocabulary without documentation
|
||||
|
||||
**Mechanism.** The per-tag comment field is optional and tag names are terse.
|
||||
|
||||
**Why it is silent.** An undocumented tag works exactly as well as a documented
|
||||
one. The cost is paid by people, later, and never by the build.
|
||||
|
||||
**Why the obvious check misses it.** There is no check. Nothing measures
|
||||
documentation coverage, and the field is empty by default, so an empty one looks
|
||||
like every other one.
|
||||
|
||||
**Symptom.** Duplicate tags with different spellings for the same concept; tags
|
||||
nobody dares delete because nobody knows what reads them. The vocabulary grows
|
||||
monotonically because deletion is unsafe.
|
||||
|
||||
**Detect.** Measure the ratio:
|
||||
|
||||
```bash
|
||||
total=$(rg -c 'GameplayTagList=\(Tag=' Config/ | awk -F: '{s+=$2} END {print s+0}')
|
||||
empty=$(rg -c 'DevComment=""' Config/ | awk -F: '{s+=$2} END {print s+0}')
|
||||
echo "undocumented $empty of $total"
|
||||
```
|
||||
|
||||
On the measured reference this was 68 of 99 in the project config — sixty-nine per
|
||||
cent undocumented, in a project otherwise notable for its care.
|
||||
|
||||
**Guardrail.** Require a comment in review for every declared tag, and run a
|
||||
periodic declared-but-unread / read-but-undeclared diff so the vocabulary can
|
||||
shrink as well as grow.
|
||||
|
||||
---
|
||||
|
||||
### TG-17 - Vocabulary with no central registry
|
||||
|
||||
**Mechanism.** File-local native tag macros scatter definitions across the module,
|
||||
each valid, none discoverable from the others.
|
||||
|
||||
**Why it is silent.** Every definition works. Scattering costs nothing at runtime
|
||||
and nothing at build time; it costs only the ability to answer a question.
|
||||
|
||||
**Why the obvious check misses it.** Looking at the central registry file shows a
|
||||
well-organised list, and it is genuinely well organised — it is simply a minority
|
||||
of the vocabulary. The check confirms what it can see, and what it cannot see does
|
||||
not announce itself.
|
||||
|
||||
**Symptom.** No single place answers "what tags does this project have". Audits
|
||||
and renames are incomplete by construction, and each incomplete audit produces a
|
||||
number that looks authoritative.
|
||||
|
||||
**Detect.** Compare the central file against the whole tree, and include plugins:
|
||||
|
||||
```bash
|
||||
rg -c "UE_DEFINE_GAMEPLAY_TAG" Source/ Plugins/ | awk -F: '{s+=$2} END {print "total", s}'
|
||||
rg -l "UE_DEFINE_GAMEPLAY_TAG" Source/ Plugins/ | wc -l
|
||||
rg -c "UE_DEFINE_GAMEPLAY_TAG" Source/*/GameplayTags.cpp 2>/dev/null
|
||||
```
|
||||
|
||||
On the measured reference this gave 94 definitions across 31 files, of which 35
|
||||
sat in the central registry — so roughly two thirds of the vocabulary lived
|
||||
outside the file whose name promises to hold it. Note that restricting the search
|
||||
to the main source directory undercounts by 18, which is TG-05 arriving again in a
|
||||
different disguise.
|
||||
|
||||
**Guardrail.** A central registry header for shared tags; file-local macros only
|
||||
for tags used in exactly one file and never authored in data. Measure the central
|
||||
share and treat a falling ratio as drift.
|
||||
|
||||
---
|
||||
|
||||
## Migration
|
||||
|
||||
### TG-18 - Rename without a redirect
|
||||
|
||||
**Mechanism.** Tag references in assets and configs are strings resolved at load
|
||||
time. A rename leaves them unresolvable.
|
||||
|
||||
**Why it is silent.** An unresolvable tag reference loads as an empty tag. The
|
||||
asset opens, the editor shows a blank field, and the system reads "no tag set" —
|
||||
which is a legal state that means "not configured yet".
|
||||
|
||||
**Why the obvious check misses it.** The rename itself is a code or config change
|
||||
that compiles and passes review. The damage is in binary assets that nobody opens
|
||||
during the change, and it becomes permanent only when one of them is resaved for
|
||||
an unrelated reason — at which point the old string is gone from disk entirely.
|
||||
|
||||
**Symptom.** Assets silently lose the tag. If they are resaved before anyone
|
||||
notices, the reference cannot be recovered from the asset; it has to be
|
||||
reconstructed from memory or version control.
|
||||
|
||||
**Detect.** Count redirects against rename history:
|
||||
|
||||
```bash
|
||||
rg -n "GameplayTagRedirects" Config/ Plugins/*/Config/ # declared redirects
|
||||
git log --oneline -S 'GameplayTagList' -- Config/ | head # config tag churn
|
||||
```
|
||||
|
||||
On the measured reference the first command returned nothing at all, confirmed
|
||||
with two independent patterns. Zero redirects means either that no tag has ever
|
||||
been renamed, or that every rename was destructive — and the second command tells
|
||||
you which.
|
||||
|
||||
**Guardrail.** Ship the redirect in the same commit as the rename. Never remove a
|
||||
redirect until every branch has been resaved and verified. Redirects are cheap;
|
||||
reconstructing a designer's work is not.
|
||||
@@ -0,0 +1,205 @@
|
||||
# Patterns: tag vocabulary in a measured reference product
|
||||
|
||||
A worked example of the governance rules from `SKILL.md`, against one specific
|
||||
reference project read as source. It exists so the failure modes in
|
||||
[failure-modes.md](failure-modes.md) can be read as *shapes that occur together*
|
||||
rather than as a checklist of independent risks.
|
||||
|
||||
The project audited below is a careful, well-organised codebase. That is the
|
||||
point. Every finding here survived review by competent engineers, because none of
|
||||
them has a runtime signature.
|
||||
|
||||
Markers: **[measured]** — read in source or config; **[derived]** — conclusion
|
||||
from measured facts; **[open]** — not answerable without the editor or a
|
||||
networked build.
|
||||
|
||||
---
|
||||
|
||||
## 1. Six declaration sources, and only one of them is easy to find
|
||||
|
||||
The project declares tags from five distinct sources simultaneously
|
||||
**[measured]**:
|
||||
|
||||
| Source | Count | Visible to a naive text search |
|
||||
|---|---:|---|
|
||||
| Project config, `+`-prefixed | 99 | yes |
|
||||
| Plugin configs, no `+`, different section | 113 across 4 files | **no** |
|
||||
| Native macros, central registry file | 35 | by symbol |
|
||||
| Native macros, elsewhere in the tree | 59 across 30 files | by symbol |
|
||||
| DataTable rows in binary assets | **[open]** | **no** |
|
||||
|
||||
Two numbers in that table are the lesson.
|
||||
|
||||
**212 versus 99.** A pattern anchored on the project-config form finds 99 tags and
|
||||
reports them confidently. The real config vocabulary is 212 across five files
|
||||
**[measured]**. The pattern is not subtly wrong; it misses the majority. Recipe:
|
||||
TG-05.
|
||||
|
||||
**94 across 31 files, 35 central.** Native definitions split three ways — 24 plain
|
||||
macros, 35 with comments, 35 file-local statics **[measured]**. The file whose
|
||||
name promises to be the registry holds roughly a third of them. Restricting the
|
||||
search to the primary source directory instead of including plugins undercounts by
|
||||
18 **[measured]** — the same failure as the config case, one directory level up.
|
||||
Recipe: TG-17.
|
||||
|
||||
**[derived]** There is no single artefact in this project that answers "what tags
|
||||
exist". Any audit, rename or deletion built on one is incomplete by construction,
|
||||
and will produce a number that looks authoritative.
|
||||
|
||||
---
|
||||
|
||||
## 2. Validation coverage: fourteen of sixty-four, one of them broken
|
||||
|
||||
Editable tag properties: 64. Properties carrying a `Categories` constraint: 14
|
||||
**[measured]** — under sixteen per cent.
|
||||
|
||||
That ratio is a list of places where a typo cannot be caught. But the sharper
|
||||
finding is inside the fourteen: one constraint named a namespace that appears
|
||||
nowhere else in the project, while the tags it was meant to offer live under a
|
||||
differently-named root **[measured]**.
|
||||
|
||||
**[derived]** The filter does not narrow the picker, it empties it. And it does so
|
||||
in one of the few places where tag input was validated at all — so the single most
|
||||
protective mechanism in the file was the one that was misconfigured.
|
||||
|
||||
The general form is worth stating: **a constraint is only a guardrail if something
|
||||
proves it resolves.** An empty dropdown is indistinguishable from a namespace
|
||||
nobody has populated yet, so the editor cannot tell you, and neither can the
|
||||
compiler. Recipes: TG-03, TG-04.
|
||||
|
||||
---
|
||||
|
||||
## 3. A misspelling that works because every side repeats it
|
||||
|
||||
The token `Primarly` — for `Primarily` — appears at six sites across four files
|
||||
**[measured]**: two project tag declarations, two platform-specific config files,
|
||||
and a native tag macro plus its use in a widget class.
|
||||
|
||||
**[derived]** It functions precisely because every side misspells it identically.
|
||||
Fixing any one site breaks the join, with no error at either end and a diff that
|
||||
every reviewer would approve.
|
||||
|
||||
This is the cleanest available demonstration of the invariant at the top of the
|
||||
skill: **when a string literal is the join, correctness means "all sides agree",
|
||||
not "the string is right".** Recipe: TG-02.
|
||||
|
||||
The project's exposure to this class is total in a specific way: **zero
|
||||
`GameplayTagRedirects` are declared**, confirmed with two independent search
|
||||
patterns **[measured]**. Every rename in this project is therefore destructive,
|
||||
including the corrective rename that fixing `Primarly` would require. Recipe:
|
||||
TG-18.
|
||||
|
||||
---
|
||||
|
||||
## 4. Replication settings that carry an unstated limit
|
||||
|
||||
| Setting | Value **[measured]** | What it means in tags |
|
||||
|---|---|---|
|
||||
| Fast replication | enabled | Tags travel as registry indices; client and server registries must match exactly |
|
||||
| Bits for container size | 6 | A hard cap of 63 entries per replicated tag container |
|
||||
|
||||
**[derived]** The first row makes plugin-asymmetric tag loading a correctness
|
||||
problem rather than a cosmetic one, and PIE cannot reproduce it: one process, one
|
||||
registry. The second row is a gameplay limit expressed as a bit count, sitting in
|
||||
a settings file, with nothing connecting it to the containers that will grow.
|
||||
|
||||
**[open]** Whether client and server registries actually match under asymmetric
|
||||
feature-plugin loading is unanswerable without a dedicated-server build, and is
|
||||
left open. Recipes: TG-10, TG-11.
|
||||
|
||||
---
|
||||
|
||||
## 5. Counters: the tag is validated, the amount is not
|
||||
|
||||
A tag-to-count container in the project validates its tag argument at both the add
|
||||
and the remove site, with a warning on failure. Neither site validates the
|
||||
*amount*: a guard of the form "if the amount is positive" returns silently
|
||||
otherwise, and the accumulator has no clamp **[measured]**. An unresolved TODO
|
||||
next to the remove path asks exactly what should happen when more is removed than
|
||||
exists **[measured]**.
|
||||
|
||||
**[derived]** This asymmetry is the entire finding, and it is invisible for a
|
||||
structural reason: seeing the tag validation satisfies the reviewer's question
|
||||
"are the inputs checked here?". One checked parameter reads as a checked function.
|
||||
|
||||
Worth noting for balance: the same container implements all three replication
|
||||
callbacks correctly, keeping its array and lookup map consistent **[measured]** —
|
||||
which is exactly the discipline that a different structure in the same project
|
||||
fails to apply. Recipes: TG-12, TG-13.
|
||||
|
||||
---
|
||||
|
||||
## 6. A stub with a second layer under it
|
||||
|
||||
One widget base class exposes an editable tag container for "hide when these tags
|
||||
are present". Both consumer sites hardcode the check result to a constant, with
|
||||
the real expression parked in a comment **[measured]**. That alone is an ordinary
|
||||
stub: the property is authored, read nowhere, and does nothing.
|
||||
|
||||
The second layer is what makes it worth including here. A setter in the same class
|
||||
overwrites the designer-authored visibility values during construction
|
||||
**[measured]**.
|
||||
|
||||
**[derived]** So the obvious one-line fix — implement the hardcoded check —
|
||||
produces *wrong* behaviour rather than correct behaviour, because it would select
|
||||
between two values that no longer hold what the designer authored. The fix then
|
||||
gets blamed and reverted, burying the original defect under a failed attempt.
|
||||
|
||||
**A stub can be deeper than its TODO admits.** Before implementing one, trace its
|
||||
inputs. Recipes: TG-14, TG-15.
|
||||
|
||||
---
|
||||
|
||||
## 7. Documentation coverage
|
||||
|
||||
68 of 99 project-config tags carry an empty comment field — sixty-nine per cent
|
||||
undocumented **[measured]**.
|
||||
|
||||
**[derived]** This is why vocabularies only grow. A tag nobody can explain is a tag
|
||||
nobody can safely delete, so the safe action is always to add another one. Over a
|
||||
project's life that produces duplicate tags for one concept, which is the
|
||||
condition that makes hierarchical matching dangerous (TG-08) and renames necessary
|
||||
(TG-18) — both of which this project is unprotected against.
|
||||
|
||||
One tag's own comment in the project admits that the tag sits in the wrong
|
||||
namespace and needs to move **[measured]**. With zero redirects declared, that move
|
||||
cannot be made without breaking every asset that references it. The comment is an
|
||||
accurate description of a debt that the configuration makes unpayable.
|
||||
|
||||
---
|
||||
|
||||
## 8. What the reference got right
|
||||
|
||||
1. **A central registry file exists** and is well organised — the problem is its
|
||||
share, not its quality **[measured]**.
|
||||
2. **Replication callbacks on the tag-stack container are complete**, keeping the
|
||||
replicated array and the lookup map consistent in all three **[measured]**.
|
||||
3. **Tag validation with explicit warnings** at the stack add/remove sites — the
|
||||
discipline is present, it is simply applied to one parameter of two
|
||||
**[measured]**.
|
||||
4. **Constraints exist at all.** Fourteen constrained properties is a low ratio,
|
||||
and it is fourteen more than most projects have.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source and
|
||||
config in a single workspace. Source addresses stay in the research archive that
|
||||
produced this skill; each `TG-` identifier resolves back to the audited location
|
||||
there. Numbers without an author are worth less than no numbers, so: this is where
|
||||
these came from, and they can be produced on request.
|
||||
|
||||
One correction is recorded here rather than hidden: an earlier form of this
|
||||
material stated that 59 native definitions were spread over "40+ files". The
|
||||
re-measurement performed while writing this file gave **94 definitions across 31
|
||||
files**, and that is the number above. The original figure was never verified; the
|
||||
recipe in TG-17 is what produced the corrected one.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Every claim above refers to one project on one engine version. These are examples
|
||||
and failure evidence, not guarantees about other engine versions or other samples.
|
||||
The counts are properties of that project on that day and are quoted only to show
|
||||
the shape of the contrast a recipe produces — re-run the recipes against your own
|
||||
tree rather than importing the numbers.
|
||||
@@ -0,0 +1,452 @@
|
||||
---
|
||||
name: ue-gas-architecture
|
||||
description: >-
|
||||
Design or review production Gameplay Ability System architecture in Unreal
|
||||
Engine: ability packages and revocation receipts, buffered ability input on
|
||||
the ability system component, activation policies and exclusivity groups,
|
||||
attribute set responsibility boundaries, gameplay effect contexts, tag
|
||||
relationship policy, match phases as abilities, and the health/death seam.
|
||||
Use when adding combat, abilities, equipment-granted powers, phases or rounds,
|
||||
damage and healing, or diagnosing ability activation and lifecycle bugs.
|
||||
---
|
||||
|
||||
# UE GAS architecture
|
||||
|
||||
The invariant:
|
||||
|
||||
> GAS owns capabilities and effects; domain components own durable gameplay
|
||||
> state; data packages compose capabilities; every temporary grant keeps a
|
||||
> receipt.
|
||||
|
||||
Worked analysis of a production GAS layer, including the parts worth copying
|
||||
verbatim: [patterns](references/patterns.md).
|
||||
Detection recipes for silent GAS failures: [failure modes](references/failure-modes.md).
|
||||
|
||||
Related skills: `ue-data-driven-architecture`, `ue-input-architecture`,
|
||||
`ue-modular-gameplay`, `ue-cosmetics-and-teams`, `ue-multiplayer-authority`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Package grants as ability sets
|
||||
|
||||
Do not scatter calls to grant abilities, apply effects and add attribute sets
|
||||
across pawns, equipment and modes. Define one immutable capability package:
|
||||
|
||||
```text
|
||||
AbilitySet
|
||||
├─ abilities: class + level + optional semantic input tag
|
||||
├─ gameplay effects: class + level
|
||||
└─ attribute sets: class
|
||||
```
|
||||
|
||||
The same package can be granted by pawn data at spawn, by an equipment
|
||||
definition while equipped, by a feature action for a mode, by a pickup, or by a
|
||||
test fixture.
|
||||
|
||||
### Grant order is not cosmetic
|
||||
|
||||
Attributes, then abilities, then effects. Effects in the third group may modify
|
||||
attributes created by the first; the reverse order silently applies modifiers to
|
||||
attributes that do not exist yet.
|
||||
|
||||
### Grant only on authority
|
||||
|
||||
The package's grant function must reject non-authority mutation at its first
|
||||
line. Clients receive replicated specs, effects and attributes; they do not
|
||||
author grants.
|
||||
|
||||
### Return a grant receipt
|
||||
|
||||
```text
|
||||
GrantedHandles
|
||||
├─ AbilitySpecHandles
|
||||
├─ ActiveGameplayEffectHandles
|
||||
└─ GrantedAttributeSet pointers
|
||||
```
|
||||
|
||||
The revoke function removes exactly those resources. This is what prevents one
|
||||
equipment item from revoking another owner's grants — and it is the single most
|
||||
reusable piece of a GAS layer.
|
||||
|
||||
### Decide permanence explicitly
|
||||
|
||||
Passing no receipt means "grant for the lifetime of the ability system
|
||||
component". That is a legitimate design choice, and it must be a choice rather
|
||||
than an omission. Review every caller:
|
||||
|
||||
| Caller | Receipt required? |
|
||||
|---|---|
|
||||
| pawn baseline from pawn data | optional, permanent by design |
|
||||
| equipment | yes |
|
||||
| feature action | yes, per actor and per activation context |
|
||||
| temporary pickup | yes, or rely on effect duration handles |
|
||||
| test or cheat command | yes |
|
||||
|
||||
A single reference implementation demonstrates both: the pawn baseline passes
|
||||
nothing because it is meant to last as long as the player state, and equipment
|
||||
passes a receipt stored in its own applied-item entry.
|
||||
|
||||
---
|
||||
|
||||
## 2. Treat input as a buffered concern of the ability system component
|
||||
|
||||
Input events should not call activation directly. Send a semantic input tag to
|
||||
the ability system component and process once per frame after player input:
|
||||
|
||||
```text
|
||||
InputAction
|
||||
→ InputTag pressed/released
|
||||
→ ASC records spec handles
|
||||
→ PlayerController.PostProcessInput
|
||||
→ ASC.ProcessAbilityInput
|
||||
→ activation policy chooses the activation set
|
||||
→ TryActivateAbility
|
||||
```
|
||||
|
||||
What this buys:
|
||||
|
||||
- simultaneous inputs resolve together instead of in binding order;
|
||||
- held input is distinguishable from newly pressed input;
|
||||
- activation can be suppressed globally for a frame;
|
||||
- activation policy belongs to the ability rather than to each binding site.
|
||||
|
||||
### Required buffers
|
||||
|
||||
Held spec handles, pressed this frame, released this frame. Clear the transient
|
||||
buffers after processing, and drop stale handles when specs are removed.
|
||||
|
||||
### Order within the frame matters
|
||||
|
||||
Collect held abilities first, then pressed, then activate everything in one pass,
|
||||
and only then process releases. Activating as you iterate means a held input can
|
||||
activate an ability and then immediately deliver it a press event it should never
|
||||
have seen.
|
||||
|
||||
### Add a global input block tag
|
||||
|
||||
When a project-defined "ability input blocked" tag is present on the component:
|
||||
clear the buffers, activate nothing, and define explicitly whether already-active
|
||||
held abilities receive a release or a cancel. One tag is cheaper and far more
|
||||
reviewable than disabling dozens of input actions during UI, death or stun.
|
||||
|
||||
---
|
||||
|
||||
## 3. Put activation rules on the ability
|
||||
|
||||
### Activation policy — when
|
||||
|
||||
- `OnInputTriggered` — activate once when pressed;
|
||||
- `WhileInputActive` — activate and retry while held;
|
||||
- `OnSpawn` — activate when granted and the avatar is ready.
|
||||
|
||||
Do not infer these from the input event type at the binding site. The ability
|
||||
knows its own lifecycle; the binding does not.
|
||||
|
||||
`OnSpawn` needs two entry points, not one: the ability may be granted to a live
|
||||
avatar, or the avatar may appear for an already-granted ability. Handle both, and
|
||||
skip activation when the avatar is being torn down.
|
||||
|
||||
### Activation group — concurrency
|
||||
|
||||
- `Independent` — unaffected by others;
|
||||
- `Exclusive_Replaceable` — exclusive, but another exclusive may displace it;
|
||||
- `Exclusive_Blocking` — exclusive and prevents other exclusives from starting.
|
||||
|
||||
The component tracks active counts per group and enforces transitions. This
|
||||
replaces an N×N web of per-ability blocking tags for broad concurrency rules.
|
||||
|
||||
Rules:
|
||||
|
||||
- an ability may enter a group only if not blocked;
|
||||
- entering an exclusive group cancels other replaceable exclusives;
|
||||
- group counts must balance across activate and end;
|
||||
- an ability in a replaceable group must not be able to refuse cancellation —
|
||||
that combination is a contradiction and should log rather than obey;
|
||||
- changing group at runtime is explicit and validated.
|
||||
|
||||
Death is the clearest worked example: on activation it cancels everything not
|
||||
explicitly marked as surviving death, refuses cancellation, and moves itself into
|
||||
the blocking group for the duration.
|
||||
|
||||
---
|
||||
|
||||
## 4. Move tag relationships into a policy asset
|
||||
|
||||
Abilities carry their intrinsic tags. A pawn or archetype selects a relationship
|
||||
map:
|
||||
|
||||
```text
|
||||
AbilityTag
|
||||
├─ tags to block
|
||||
├─ tags to cancel
|
||||
├─ activation required tags
|
||||
└─ activation blocked tags
|
||||
```
|
||||
|
||||
This removes repeated policy from N ability defaults and lets the same ability
|
||||
classes run under different archetype rules. It matters most when abilities
|
||||
arrive from separate feature plugins that must not know about each other.
|
||||
|
||||
### Relationship validation
|
||||
|
||||
- [ ] the key tag is valid and appropriately scoped;
|
||||
- [ ] parent-tag matching is intentional, not incidental;
|
||||
- [ ] block and cancel are not confused with each other;
|
||||
- [ ] required and blocked sets are not contradictory;
|
||||
- [ ] cycles are documented and tested;
|
||||
- [ ] every archetype that needs policy actually selects a map.
|
||||
|
||||
A linear scan is acceptable for a small matrix. Profile before assuming it scales:
|
||||
in the measured reference the lookup is a linear array walk performed on every
|
||||
activation and on every block/cancel application, and its authors marked it as
|
||||
provisional in three separate places.
|
||||
|
||||
---
|
||||
|
||||
## 5. Split attribute sets by direction of responsibility
|
||||
|
||||
Do not create one character-attributes dumping ground.
|
||||
|
||||
**Receiving set.** Health, max health, the meta attributes damage and healing,
|
||||
clamping, and the out-of-health signal. Damage and healing are transient inputs
|
||||
to execution, not durable replicated state: they arrive, convert into a health
|
||||
change, and reset to zero in the same call.
|
||||
|
||||
**Source set.** Base damage, base heal, outgoing coefficients — the values
|
||||
captured *from the instigator*.
|
||||
|
||||
Execution calculations capture the source set from the instigator and the
|
||||
receiving set from the target. The split makes "who contributes this value"
|
||||
structural rather than a comment.
|
||||
|
||||
### Gates on the receiving set
|
||||
|
||||
- clamp in the correct hook, and in all of them that can change the value;
|
||||
- guard against duplicate out-of-health events with an explicit flag that resets
|
||||
when the value recovers;
|
||||
- record pre-execution values when messages need a before-and-after;
|
||||
- notify domain components through delegates;
|
||||
- keep game-flow state — death started, death finished — out of the attribute set
|
||||
entirely.
|
||||
|
||||
### Mark what modifiers must not touch
|
||||
|
||||
Attributes that may only be changed by an execution should say so in metadata.
|
||||
This is the difference between "we agreed not to modify health directly" and a
|
||||
rule the editor enforces.
|
||||
|
||||
---
|
||||
|
||||
## 6. Extend the gameplay effect context for combat metadata
|
||||
|
||||
When damage needs source or target data beyond the stock context, derive a
|
||||
project context and implement allocation, script-struct identity, duplication and
|
||||
network serialization.
|
||||
|
||||
Store only what must survive execution and replication. A context is not a place
|
||||
to hang a mutable combat object graph.
|
||||
|
||||
Validation:
|
||||
|
||||
- [ ] the context survives duplication with its custom fields intact;
|
||||
- [ ] custom fields serialize deterministically;
|
||||
- [ ] object references are network-safe, and fields that are deliberately not
|
||||
replicated are documented as server-only at their declaration;
|
||||
- [ ] execution calculations handle absent hit data;
|
||||
- [ ] script exposure does not permit invalid mutation.
|
||||
|
||||
---
|
||||
|
||||
## 7. Separate numeric death detection from the death lifecycle
|
||||
|
||||
```text
|
||||
effect execution changes Health
|
||||
→ receiving set emits out-of-health
|
||||
→ health component sends a death gameplay event
|
||||
→ death ability starts
|
||||
→ health component enters DeathStarted, then DeathFinished
|
||||
→ game mode handles respawn
|
||||
```
|
||||
|
||||
Responsibilities: GAS owns the *number*; the health component owns the *state*
|
||||
and its replication; the death ability owns the *sequence*; the game mode owns
|
||||
the player lifecycle.
|
||||
|
||||
Do not represent death only as "health is zero". Consumers need a monotonic state
|
||||
they can observe over the network, and a predicted transition needs somewhere to
|
||||
be corrected.
|
||||
|
||||
### The death trigger must be unambiguous
|
||||
|
||||
Configure it in the ability's class defaults so exactly one path can start it, and
|
||||
have the ability's end path force the final transition — insurance against a
|
||||
designer-authored sequence that does not reach the end.
|
||||
|
||||
---
|
||||
|
||||
## 8. Model match phases as abilities when the lifecycle fits
|
||||
|
||||
A phase ability runs on the ability system component of the game state:
|
||||
|
||||
```text
|
||||
Game.Playing
|
||||
├─ Game.Playing.Warmup
|
||||
├─ Game.Playing.SuddenDeath
|
||||
└─ Game.PostGame
|
||||
```
|
||||
|
||||
Starting a phase grants and activates the ability, registers its tag, cancels
|
||||
incompatible siblings, and notifies observers. Ending the ability ends the phase.
|
||||
|
||||
Hierarchical tags give the rule for free: **parents and children coexist, siblings
|
||||
do not.** Starting `Game.Playing.SuddenDeath` while `Game.Playing` is active
|
||||
leaves the parent running; starting `Game.PostGame` ends the whole subtree.
|
||||
|
||||
### Use this pattern when
|
||||
|
||||
- the phase has a real enter/active/exit lifecycle;
|
||||
- ability tasks, effects and tags are natural fits for its logic;
|
||||
- the server owns transitions;
|
||||
- observers want tag-based matching rather than an enum switch.
|
||||
|
||||
### Do not use it when
|
||||
|
||||
- a replicated enum fully describes the state;
|
||||
- the phase must never be cancellable;
|
||||
- clients must query authoritative phase state and no replication path has been
|
||||
designed for it. This is the trap: if the phase abilities are server-only, a
|
||||
client-side query of "is this phase active" answers false forever, and it
|
||||
answers it without an error.
|
||||
|
||||
### Observer handles are mandatory
|
||||
|
||||
Registration returns a handle and supports unsubscribe. An array of callbacks
|
||||
with no handles grows until the world resets, and every entry pins whatever it
|
||||
captured.
|
||||
|
||||
---
|
||||
|
||||
## 9. Failure feedback belongs at the ability system seam
|
||||
|
||||
Activation failure should produce structured tags, not a log line:
|
||||
|
||||
```text
|
||||
Ability.ActivateFail.Cooldown
|
||||
Ability.ActivateFail.Cost
|
||||
Ability.ActivateFail.TagsBlocked
|
||||
Ability.ActivateFail.IsDead
|
||||
```
|
||||
|
||||
For a locally controlled avatar, route them straight to feedback. For a
|
||||
server-side rejection of a locally predicted action, notify the owning client with
|
||||
the ability and the failure tags. Map tags to user-facing text and to animations
|
||||
in data, and deliver them over a message bus so the UI never learns about GAS.
|
||||
|
||||
Never parse log strings in UI.
|
||||
|
||||
### Compose costs rather than accepting one
|
||||
|
||||
A single cost effect is rarely enough. A list of inline cost objects — each with
|
||||
its own check, its own application, its own failure tag, and an optional
|
||||
"only charge on hit" rule — turns cost into data. The hit determination should be
|
||||
computed once and cached, on authority only.
|
||||
|
||||
---
|
||||
|
||||
## 10. Equipment and feature integration
|
||||
|
||||
Equipment grants ability sets when equipped and retains the receipt in its
|
||||
applied-item entry. Unequip reverses exactly that receipt.
|
||||
|
||||
A feature action that grants abilities retains, per actor: direct ability handles,
|
||||
attribute instances, ability-set receipts, and the extension request handle.
|
||||
|
||||
Every integration answers:
|
||||
|
||||
- Who owns the ability system component — the pawn or the player state?
|
||||
- Which object is the source object for the spec?
|
||||
- When is avatar info initialized?
|
||||
- Is the grant permanent or reversible?
|
||||
- What happens if the actor is destroyed before the feature deactivates?
|
||||
|
||||
See `ue-modular-gameplay` for actor-extension timing and rollback.
|
||||
|
||||
---
|
||||
|
||||
## 11. Review checklist
|
||||
|
||||
### Ability set
|
||||
|
||||
- [ ] authority-only grant, checked at the top of the function;
|
||||
- [ ] grant order is attributes, abilities, effects;
|
||||
- [ ] every class reference non-null and of the expected type;
|
||||
- [ ] semantic input tag valid and constrained;
|
||||
- [ ] receipt retained for every temporary grant;
|
||||
- [ ] revoke removes only owned resources;
|
||||
- [ ] source object deliberately chosen.
|
||||
|
||||
### Ability system component
|
||||
|
||||
- [ ] input buffers processed once per frame, after player input;
|
||||
- [ ] blocked-input behaviour explicit for held abilities;
|
||||
- [ ] activation group counters balance;
|
||||
- [ ] avatar change activates spawn-policy abilities exactly once;
|
||||
- [ ] relationship map selected per archetype;
|
||||
- [ ] failure tags reach local feedback.
|
||||
|
||||
### Attributes and damage
|
||||
|
||||
- [ ] source and target responsibilities separated into different sets;
|
||||
- [ ] meta attributes are not treated as durable state;
|
||||
- [ ] clamping occurs in every path that can change the value;
|
||||
- [ ] team and friendly-fire rules applied in one authoritative calculation;
|
||||
- [ ] out-of-health cannot fire twice for one death;
|
||||
- [ ] custom effect context serializes.
|
||||
|
||||
### Phases and death
|
||||
|
||||
- [ ] the server owns transitions;
|
||||
- [ ] observers can unregister;
|
||||
- [ ] sibling and parent tag semantics are tested, not assumed;
|
||||
- [ ] clients have a defined observation path that does not depend on a
|
||||
server-only subsystem;
|
||||
- [ ] death start and finish are durable states, not health comparisons.
|
||||
|
||||
---
|
||||
|
||||
## 12. Decision summary
|
||||
|
||||
Use GAS for a concern when at least one is true: it is a grantable and revocable
|
||||
capability; prediction or networked activation matters; it is naturally expressed
|
||||
through effects, tags, cost and cooldown; it benefits from task-based lifecycle
|
||||
with cancel and end.
|
||||
|
||||
Keep it out of GAS when the concern is durable domain state better served by a
|
||||
replicated component, presentation only, static definition composition, or a
|
||||
global lifecycle with no ability semantics.
|
||||
|
||||
A production GAS architecture is not "everything is an ability". It is a clean
|
||||
seam between capability lifecycle, numeric effects, domain state and data-driven
|
||||
composition.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The patterns and failure modes in `references/` come from a line-by-line audit of
|
||||
the GAS layer of Epic's Lyra Starter Game on Unreal Engine 5.6 — roughly 4 900
|
||||
lines across 51 files — read as source rather than run.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; what ships
|
||||
is the detection recipe. Each entry carries a stable identifier (`GA-01` and up)
|
||||
that resolves back to the audited location, so any specific claim can be produced
|
||||
on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The architecture described here
|
||||
is transferable; the specific defects are evidence, not guarantees about other
|
||||
versions. Re-run the detection recipes against your own tree before acting on any
|
||||
specific claim.
|
||||
@@ -0,0 +1,502 @@
|
||||
# Failure modes: GAS architecture
|
||||
|
||||
Fourteen ways a Gameplay Ability System layer misbehaves without reporting
|
||||
anything.
|
||||
|
||||
The shared property, stated once: **GAS failures present as absence.** An ability
|
||||
that does not activate, a phase query that answers false, a grant that is never
|
||||
revoked, a cost that is never charged. The framework is built to tolerate a
|
||||
refused activation, because a refused activation is a normal gameplay state. So
|
||||
the diagnostic burden is entirely yours, and the useful question is never "did it
|
||||
error" but "which side of the seam owns this fact".
|
||||
|
||||
Each entry gives the mechanism, why it stays silent, why the obvious check misses
|
||||
it, the symptom a human reports, a `Detect` recipe, and the guardrail.
|
||||
Identifiers (`GA-01` and up) are stable and resolve back to the audited source in
|
||||
the research archive.
|
||||
|
||||
Recipes use `rg` and run from a project's source root. They were executed against
|
||||
the audited project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## Ownership and receipts
|
||||
|
||||
### GA-01 - Grant with no receipt, permanent by omission
|
||||
|
||||
**Mechanism.** An ability set is granted with the out-parameter for handles left
|
||||
null. The grant then lasts as long as the ability system component.
|
||||
|
||||
**Why it is silent.** Permanent is a legitimate lifetime, and the API is designed
|
||||
to offer it. Nothing distinguishes "permanent because we decided" from "permanent
|
||||
because the argument was easy to omit".
|
||||
|
||||
**Why the obvious check misses it.** The call is correct, the API is used as
|
||||
documented, and the reviewer sees a grant that works. The defect only exists
|
||||
relative to an intention that lives in nobody's head at review time.
|
||||
|
||||
**Symptom.** A capability that should have been removed with the feature, the
|
||||
equipment or the round survives it. Usually noticed after a respawn, when the
|
||||
player has two of something.
|
||||
|
||||
**Detect.** List every grant site and classify it by whether it keeps a receipt:
|
||||
|
||||
```bash
|
||||
rg -n -B3 "GiveToAbilitySystem" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Any call passing `nullptr` for the handles parameter is a permanent grant. In the
|
||||
audited project this is exactly one site — the pawn baseline on the player state —
|
||||
and the equipment path beside it passes a receipt stored in its applied-item
|
||||
entry. One permanent grant with a stated reason is a design; several are a leak.
|
||||
|
||||
**Guardrail.** Make permanence explicit at the call site, in a comment naming the
|
||||
lifetime the grant is tied to. Review new grant sites against the table of
|
||||
allowed permanent grants rather than against the API signature.
|
||||
|
||||
---
|
||||
|
||||
### GA-02 - Revocation that removes more than it granted
|
||||
|
||||
**Mechanism.** Teardown clears abilities by class, by tag, or by clearing all
|
||||
abilities on the component, instead of by the handles the grant returned.
|
||||
|
||||
**Why it is silent.** The abilities the caller wanted removed *are* removed. The
|
||||
extra removals affect capabilities another owner granted, and that owner is not
|
||||
watching.
|
||||
|
||||
**Why the obvious check misses it.** The teardown function looks thorough. "Remove
|
||||
everything matching this class" reads as more robust than "remove these three
|
||||
handles", not less.
|
||||
|
||||
**Symptom.** Unequipping one item removes an ability granted by a different
|
||||
feature. Reproduces only when two owners granted overlapping capabilities, which
|
||||
is rare in a test and common in a shipped game.
|
||||
|
||||
**Detect.** Find removals that are not keyed on stored handles:
|
||||
|
||||
```bash
|
||||
rg -n "ClearAbility\(|ClearAllAbilities|RemoveActiveGameplayEffect" --glob "*.cpp" . -B5 \
|
||||
| rg -n "(GetAssetTags|StaticClass|ForEach|\.Class)"
|
||||
```
|
||||
|
||||
Removal keyed on anything other than a handle produced by the matching grant is
|
||||
the finding.
|
||||
|
||||
**Guardrail.** Every revoke path takes a receipt and removes exactly its contents.
|
||||
If a receipt was not kept, the correct fix is to keep one, not to widen the
|
||||
removal.
|
||||
|
||||
---
|
||||
|
||||
### GA-03 - Grant order that applies effects before their attributes exist
|
||||
|
||||
**Mechanism.** An ability set grants effects before it grants attribute sets, so
|
||||
an effect modifies an attribute that has not been added yet.
|
||||
|
||||
**Why it is silent.** Applying a modifier to a missing attribute is not an error
|
||||
in the framework — the modifier simply finds nothing to modify. The effect is
|
||||
applied, the handle is valid, the log is clean.
|
||||
|
||||
**Why the obvious check misses it.** Each block of the grant function is correct
|
||||
in isolation, and the whole function reads as a straightforward three-part loop.
|
||||
Order dependencies between the parts are invisible unless you know the effects
|
||||
reference the attributes.
|
||||
|
||||
**Symptom.** An attribute sits at its default value despite an effect that should
|
||||
have initialized it. Investigated as a bad effect asset.
|
||||
|
||||
**Detect.** Read the grant function and confirm the sequence:
|
||||
|
||||
```bash
|
||||
rg -n -A40 "::GiveToAbilitySystem" --glob "*.cpp" . \
|
||||
| rg -n "(AddAttributeSetSubobject|GiveAbility|ApplyGameplayEffect)"
|
||||
```
|
||||
|
||||
The line numbers must run attributes, then abilities, then effects. The audited
|
||||
implementation gets this right and the ordering is worth copying verbatim.
|
||||
|
||||
**Guardrail.** State the order as a comment in the grant function, and add a test
|
||||
that grants a set whose effect initializes an attribute the same set provides.
|
||||
|
||||
---
|
||||
|
||||
## State that lies
|
||||
|
||||
### GA-04 - Server-only subsystem queried from clients
|
||||
|
||||
**Mechanism.** Phase abilities are configured server-only, and the subsystem that
|
||||
tracks them is created on every client anyway. Its map of active phases is
|
||||
therefore always empty on a client, and the "is this phase active" query always
|
||||
returns false there.
|
||||
|
||||
**Why it is silent.** False is a valid answer to that question. The client is not
|
||||
told that it asked a question it cannot answer; it is told "no".
|
||||
|
||||
**Why the obvious check misses it.** The subsystem exists on the client, is
|
||||
correctly initialized, and answers immediately. Every structural check passes.
|
||||
The creation gate that would have prevented it exists too — commented out, with
|
||||
the unconditional `return true` left in place. Reading that function tells you
|
||||
someone thought about it, not what they concluded.
|
||||
|
||||
**Symptom.** Client-side UI that reacts to match phase never reacts. Blamed on
|
||||
replication, on the widget, on the tag — rarely on the query itself.
|
||||
|
||||
**Detect.** Find subsystem creation gates that were considered and abandoned:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "ShouldCreateSubsystem" --glob "*.cpp" . | rg -n "^\s*//.*(return|check|World)"
|
||||
```
|
||||
|
||||
Then, for every server-authoritative subsystem, check whether any client code
|
||||
path queries its state:
|
||||
|
||||
```bash
|
||||
rg -n "IsPhaseActive|IsRoundActive|GetCurrentPhase" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
**Guardrail.** A subsystem whose state is authority-only either refuses to be
|
||||
created without authority, or its queries return a tri-state that distinguishes
|
||||
"no" from "I cannot know". Clients observe phase through replicated tags,
|
||||
replicated state or messages — never through the authoritative tracker.
|
||||
|
||||
---
|
||||
|
||||
### GA-05 - Observer registration with no handle
|
||||
|
||||
**Mechanism.** A subscription API appends a callback to an array and returns
|
||||
nothing. There is no way to unsubscribe.
|
||||
|
||||
**Why it is silent.** Subscriptions work. The list grows, and a growing list has
|
||||
no symptom until it is large or until a captured object needed to die.
|
||||
|
||||
**Why the obvious check misses it.** The API is complete from the caller's
|
||||
perspective: subscribe, get called back. The missing half is a function that does
|
||||
not exist, and reviews do not notice absent functions.
|
||||
|
||||
**Symptom.** Callbacks firing on objects that logically ended, and a list that
|
||||
grows until the world resets. In the audited project the authors documented this
|
||||
themselves, in two consecutive comments above the function, and shipped it —
|
||||
including the honest second thought that a handle would not help if callers
|
||||
ignored it.
|
||||
|
||||
**Detect.** Find subscription functions that return void:
|
||||
|
||||
```bash
|
||||
rg -n "void \w+::When\w+|void \w+::(Register|Subscribe|Observe)\w*\(" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Any registration whose return type is void is a subscription that cannot be
|
||||
undone.
|
||||
|
||||
**Guardrail.** Registration returns a handle; the handle unregisters. If callers
|
||||
cannot be trusted to hold handles, bind weakly *and* sweep expired entries — but
|
||||
the sweep is an addition to the handle, not a substitute for it.
|
||||
|
||||
---
|
||||
|
||||
### GA-06 - Write-only field
|
||||
|
||||
**Mechanism.** A member is declared and assigned in the constructor, and read
|
||||
nowhere.
|
||||
|
||||
**Why it is silent.** It is a correctly initialized member of a working class. It
|
||||
costs nothing at runtime and breaks nothing.
|
||||
|
||||
**Why the obvious check misses it.** Searching for the name finds two hits — a
|
||||
declaration and an assignment — which reads as "declared and used". An
|
||||
unused-variable warning does not fire, because the write *is* a use.
|
||||
|
||||
**Symptom.** No runtime symptom. The cost is comprehension: the field's name
|
||||
promises behaviour, and a future engineer will implement against a flag that
|
||||
nothing consumes. In the audited project the field is a cancellation-logging flag
|
||||
whose own comment describes it as temporary, added while tracking a bug that has
|
||||
since closed.
|
||||
|
||||
**Detect.** Compare writes against reads for suspicious members:
|
||||
|
||||
```bash
|
||||
F='bLogCancelation'
|
||||
rg -n "\b$F\b" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Two hits — declaration and assignment — with no comparison, no branch and no
|
||||
pass-by-value is the finding. This is the mirror image of the missing-writer
|
||||
check: here the writer exists and the reader does not.
|
||||
|
||||
**Guardrail.** Delete it. A flag with no consumer is a comment that pretends to be
|
||||
code, and it will eventually be believed.
|
||||
|
||||
---
|
||||
|
||||
## Performance and correctness on the activation path
|
||||
|
||||
### GA-07 - Function-static containers on a per-activation path
|
||||
|
||||
**Mechanism.** A function called during every ability activation declares
|
||||
function-local `static` containers as scratch space to avoid reallocation.
|
||||
|
||||
**Why it is silent.** It is correct on a single thread, and every activation
|
||||
happens on the game thread today. The reuse is invisible because each call clears
|
||||
before use.
|
||||
|
||||
**Why the obvious check misses it.** The pattern looks like a deliberate
|
||||
optimization, and it is one. Nothing at the call site says "this function is not
|
||||
reentrant", and reentrancy through an ability that activates another ability is
|
||||
plausible but not obvious.
|
||||
|
||||
**Symptom.** Nothing, until the function is called from a second thread or
|
||||
reentrantly — at which point tag requirements are computed from another call's
|
||||
scratch data and an ability activates when it should not.
|
||||
|
||||
**Detect.** Find statics in functions on the activation path:
|
||||
|
||||
```bash
|
||||
rg -n "static F\w+Container|static TArray" --glob "*.cpp" . -B10 \
|
||||
| rg -n "(CanActivate|SatisfyTagRequirements|ActivateAbility)"
|
||||
```
|
||||
|
||||
The audited project has three such statics in one activation-path function.
|
||||
|
||||
**Guardrail.** Use a member scratch buffer on an instanced object, or accept the
|
||||
allocation. If the static stays, document the single-thread and
|
||||
non-reentrancy assumption at the declaration, where the next reader will see it.
|
||||
|
||||
---
|
||||
|
||||
### GA-08 - Linear scan marked provisional, on the hot path
|
||||
|
||||
**Mechanism.** A relationship or lookup table is walked linearly on every
|
||||
activation, with a comment acknowledging that a real index would be needed at
|
||||
scale.
|
||||
|
||||
**Why it is silent.** It is correct at every size. Only the cost changes, and cost
|
||||
does not report itself.
|
||||
|
||||
**Why the obvious check misses it.** The comment is the check, and it is
|
||||
addressed to a future that has not arrived. Reviewers read "for now" as "someone
|
||||
is tracking this".
|
||||
|
||||
**Symptom.** Activation cost grows with the size of a designer-authored data
|
||||
asset, discovered during a late-project profiling pass when the table has grown
|
||||
by two orders of magnitude.
|
||||
|
||||
**Detect.** Find provisional comments on lookup paths:
|
||||
|
||||
```bash
|
||||
rg -n "Simple iteration|for now|O\(n\)|linear" --glob "*.cpp" . -A4 \
|
||||
| rg -n "(for\s*\(|ForEach)"
|
||||
```
|
||||
|
||||
In the audited project the same provisional comment appears three times in one
|
||||
mapping class, and the mapping is consulted on every activation and on every
|
||||
block-and-cancel application.
|
||||
|
||||
**Guardrail.** Record the size at which the structure must change, and add an
|
||||
assertion or a test that fails when the data crosses it. "For now" without a
|
||||
number never ends.
|
||||
|
||||
---
|
||||
|
||||
## Attributes and damage
|
||||
|
||||
### GA-09 - Meta attribute treated as durable state
|
||||
|
||||
**Mechanism.** An incoming-damage or incoming-healing attribute — designed as a
|
||||
one-shot channel that converts to a health change and resets to zero — is read,
|
||||
replicated or stored as if it held a value between executions.
|
||||
|
||||
**Why it is silent.** The attribute exists and always has a value. Reading it
|
||||
outside an execution returns zero, which is a plausible number.
|
||||
|
||||
**Why the obvious check misses it.** Nothing in the type distinguishes a meta
|
||||
attribute from a state attribute. The distinction lives in a comment and in the
|
||||
post-execution code that zeroes it.
|
||||
|
||||
**Symptom.** UI or logic that reads "current damage" and always sees zero, or a
|
||||
replicated attribute that costs bandwidth and carries nothing.
|
||||
|
||||
**Detect.** Confirm each meta attribute is zeroed after conversion and excluded
|
||||
from replication:
|
||||
|
||||
```bash
|
||||
rg -n "SetDamage\(0|SetHealing\(0" --glob "*.cpp" .
|
||||
rg -n "DOREPLIFETIME.*\b(Damage|Healing)\b" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
A meta attribute that appears in the replication list, or one that is never
|
||||
zeroed, is the finding.
|
||||
|
||||
**Guardrail.** Group meta attributes under an explicit comment banner in the
|
||||
header, keep them out of the replication list, and mark state attributes that
|
||||
executions alone may change so the editor enforces it rather than the team
|
||||
remembering it.
|
||||
|
||||
---
|
||||
|
||||
### GA-10 - Out-of-health fired more than once for one death
|
||||
|
||||
**Mechanism.** The zero-health signal is emitted from a change handler that can
|
||||
run several times as clamping, max-health changes and replication settle.
|
||||
|
||||
**Why it is silent.** Each individual emission is correct. Listeners are expected
|
||||
to handle a broadcast; nothing says "at most once per life".
|
||||
|
||||
**Why the obvious check misses it.** The emitting code is a single line in a
|
||||
single place. The multiplicity comes from how many paths reach it, which is
|
||||
visible only by enumerating callers of the enclosing handler.
|
||||
|
||||
**Symptom.** Two death sequences, doubled elimination messages, score credited
|
||||
twice. Often first noticed in the kill feed rather than in gameplay.
|
||||
|
||||
**Detect.** Find the emission and check for a latch:
|
||||
|
||||
```bash
|
||||
rg -n -B12 "OnOutOfHealth.Broadcast" --glob "*.cpp" . | rg -n "bOutOfHealth|bAlready|if\s*\(!"
|
||||
```
|
||||
|
||||
An emission with no guarding flag, or a flag that is never reset when health
|
||||
recovers, is the finding. The audited implementation has both the latch and the
|
||||
reset — copy that shape.
|
||||
|
||||
**Guardrail.** Latch the signal, reset the latch when the value recovers, and
|
||||
assert in development that the death transition runs once per life.
|
||||
|
||||
---
|
||||
|
||||
### GA-11 - Placeholder attribute initialization
|
||||
|
||||
**Mechanism.** Attributes are initialized by assigning one to another in code —
|
||||
health set to max health — with a comment saying a data-driven initialization
|
||||
will replace it.
|
||||
|
||||
**Why it is silent.** The result is a live, plausible character. Nothing about a
|
||||
full-health spawn looks provisional.
|
||||
|
||||
**Why the obvious check misses it.** The data-driven mechanism usually *exists*
|
||||
alongside it — an initialization-data field on the grant structure, a curve table
|
||||
setting — so a reviewer looking for "is initialization data-driven?" finds the
|
||||
machinery and stops.
|
||||
|
||||
**Symptom.** Designer-authored initialization values have no effect, because the
|
||||
code path that would consume them is bypassed by the placeholder.
|
||||
|
||||
**Detect.** Find initialization marked as temporary, then check whether the
|
||||
data-driven path has any consumer:
|
||||
|
||||
```bash
|
||||
rg -n "TEMP|placeholder|Eventually this will" --glob "*.cpp" . -A3 \
|
||||
| rg -n "(SetNumericAttributeBase|InitFromMetaDataTable|InitStats)"
|
||||
rg -n "GlobalCurveTableName" Config/
|
||||
```
|
||||
|
||||
In the audited project the placeholder is present, the global curve table is set
|
||||
to none, and the initialization-data field on the grant structure has no traced
|
||||
consumer.
|
||||
|
||||
**Guardrail.** A placeholder initialization is acceptable only with the
|
||||
data-driven path deleted or disabled, so nobody can author data that silently
|
||||
does nothing.
|
||||
|
||||
---
|
||||
|
||||
## Lifecycle
|
||||
|
||||
### GA-12 - Activation group counters that do not balance
|
||||
|
||||
**Mechanism.** A per-group active count is incremented on activation and
|
||||
decremented on end. Any path that ends an ability without notifying the component
|
||||
leaks a count.
|
||||
|
||||
**Why it is silent.** A leaked count does not error. It makes the group look
|
||||
occupied, so later activations are refused — which is indistinguishable from
|
||||
correct exclusivity.
|
||||
|
||||
**Why the obvious check misses it.** Both the increment and the decrement exist
|
||||
and are correctly paired in the normal path. The leak comes from an abnormal end:
|
||||
a destroyed avatar, a cancelled ability that refuses cancellation, an early return.
|
||||
|
||||
**Symptom.** After some sequence involving death or a cancelled ability, exclusive
|
||||
abilities stop activating for that pawn. Restarting the match fixes it, which
|
||||
makes it look like corruption rather than arithmetic.
|
||||
|
||||
**Detect.** Find increments and decrements and confirm they are in paired
|
||||
notification hooks:
|
||||
|
||||
```bash
|
||||
rg -n "ActivationGroupCounts" --glob "*.cpp" . -B4
|
||||
```
|
||||
|
||||
Any mutation outside the activated/ended notification pair is a leak candidate.
|
||||
Then assert the invariant directly: the exclusive counts must never exceed one in
|
||||
total.
|
||||
|
||||
**Guardrail.** Assert the invariant in development builds at every mutation, and
|
||||
test the abnormal ends explicitly — avatar destroyed mid-ability, cancel refused,
|
||||
ability ended from a task callback.
|
||||
|
||||
---
|
||||
|
||||
### GA-13 - Death represented only as a health comparison
|
||||
|
||||
**Mechanism.** Consumers ask "is health at or below zero" instead of reading a
|
||||
durable death state.
|
||||
|
||||
**Why it is silent.** The comparison is true at the right moments. It is also true
|
||||
during the window before the death sequence starts, and true again if health is
|
||||
restored and re-zeroed within one sequence.
|
||||
|
||||
**Why the obvious check misses it.** The comparison is simple, local and obviously
|
||||
correct. The state machine it replaces lives in another component, and using it
|
||||
requires knowing it exists.
|
||||
|
||||
**Symptom.** Systems that disagree about whether a character is dead, especially
|
||||
across the network, and a predicted death that cannot be corrected because there
|
||||
is no state to roll back.
|
||||
|
||||
**Detect.** Find health comparisons used as death tests:
|
||||
|
||||
```bash
|
||||
rg -n "GetHealth\(\)\s*<=?\s*0|Health\s*<=\s*0\.?0?f?" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Every hit outside the attribute set itself should be reading a death state
|
||||
instead.
|
||||
|
||||
**Guardrail.** One replicated, monotonic death state with an explicit
|
||||
start and finish, owned by a domain component, with the ability system supplying
|
||||
only the numeric signal that triggers it.
|
||||
|
||||
---
|
||||
|
||||
### GA-14 - Failure reason discarded at the seam
|
||||
|
||||
**Mechanism.** Activation fails, the framework records a reason, and the project
|
||||
does not forward it — or forwards it only on the server, where no UI exists.
|
||||
|
||||
**Why it is silent.** A refused activation is normal. The player sees nothing,
|
||||
which is exactly what a refused activation looks like when the reason is missing.
|
||||
|
||||
**Why the obvious check misses it.** The failure path exists and is exercised
|
||||
constantly. What is absent is the delivery of the reason to the machine that can
|
||||
render it, and absence of delivery has no call site to review.
|
||||
|
||||
**Symptom.** The player presses a button and nothing happens, with no feedback
|
||||
distinguishing "on cooldown" from "out of ammo" from "you are dead". Designers
|
||||
compensate with generic click sounds.
|
||||
|
||||
**Detect.** Check that failure tags are both produced and delivered to the owning
|
||||
client:
|
||||
|
||||
```bash
|
||||
rg -n "OptionalRelevantTags|NotifyAbilityFailed|ActivateFail" --glob "*.cpp" . -A6 \
|
||||
| rg -n "(Client|Broadcast|Message)"
|
||||
```
|
||||
|
||||
Failure tags produced with no client-notification path is the finding. The audited
|
||||
project does this correctly: a client notification carries the tags, the ability
|
||||
maps them to text and animation in data, and delivery goes over a message bus so
|
||||
the UI never references GAS.
|
||||
|
||||
**Guardrail.** Every failure reason reaches the locally controlled player as a
|
||||
tag, never as a log line, and the mapping from tag to feedback lives in data.
|
||||
@@ -0,0 +1,340 @@
|
||||
# Patterns: a production GAS layer read end to end
|
||||
|
||||
A worked reading of one real Gameplay Ability System layer — roughly 4 900 lines
|
||||
across 51 files — audited as source.
|
||||
|
||||
This file is different in balance from the other reference documents in this
|
||||
bundle. Most of them are about what a codebase got wrong. **This layer is mostly
|
||||
worth copying**, and the useful output is a list of the specific decisions that
|
||||
earn their keep, each with the reason it was made. The defects are in
|
||||
[failure-modes.md](failure-modes.md) and are fewer than the patterns.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not answerable from source, and left open.
|
||||
|
||||
---
|
||||
|
||||
## 1. The shape of the layer
|
||||
|
||||
Two ability system components, on different owners, for different reasons
|
||||
**[measured]**:
|
||||
|
||||
- one on the **player state**, with the pawn as avatar — so capabilities survive
|
||||
the pawn's death;
|
||||
- one on the **game state**, with itself as avatar — which is what makes match
|
||||
phases possible at all (§5).
|
||||
|
||||
The only global hooks are three config lines **[measured]**: a replacement
|
||||
globals class, a cue manager class, and an explicitly empty global curve table.
|
||||
**[derived]** That is a remarkably small integration surface for a layer this
|
||||
size, and the reason is that everything else is composed from data assets rather
|
||||
than registered globally.
|
||||
|
||||
---
|
||||
|
||||
## 2. The ability set, and why the receipt is the whole idea
|
||||
|
||||
An ability set is a primary data asset holding three arrays: abilities, effects,
|
||||
attribute sets **[measured]**. Grants happen in the order **attributes,
|
||||
abilities, effects** **[measured]** — not cosmetic, since effects in the third
|
||||
group may modify attributes created by the first.
|
||||
|
||||
Two details make it reusable rather than merely tidy:
|
||||
|
||||
**The input tag lives in the grant data, not in the ability** **[measured]**. The
|
||||
same ability class can be bound to different semantic inputs by different sets,
|
||||
and the ability itself hides the stock input category entirely **[measured]**.
|
||||
|
||||
**The revocation receipt.** The grant function takes an optional out-parameter —
|
||||
three parallel arrays of handles — and the revoke function removes exactly those
|
||||
**[measured]**:
|
||||
|
||||
```text
|
||||
GrantedHandles
|
||||
├─ AbilitySpecHandles -> ClearAbility
|
||||
├─ GameplayEffectHandles -> RemoveActiveGameplayEffect
|
||||
└─ GrantedAttributeSets -> RemoveSpawnedAttribute
|
||||
```
|
||||
|
||||
**[derived]** This is a cloakroom ticket. Whoever granted holds the ticket and can
|
||||
undo exactly their own grant without touching anyone else's. It is the answer to
|
||||
the question that breaks most modular gameplay implementations — *who removes
|
||||
this, and how do they know it was theirs?*
|
||||
|
||||
Three consumers use it, and their differences are the design **[measured]**:
|
||||
|
||||
| Consumer | Receipt | Why |
|
||||
|---|---|---|
|
||||
| equipment | stored in the applied-item entry | must reverse on unequip |
|
||||
| feature action | array per actor | must reverse on deactivation |
|
||||
| pawn baseline on player state | **passes null** | intended to last as long as the player state |
|
||||
|
||||
**[derived]** The optional out-parameter is therefore an API that encodes a
|
||||
lifetime decision. Passing null means "permanent, deliberately". One such call in
|
||||
a project is a design; several are GA-01.
|
||||
|
||||
Authority is gated at the first line of both grant and revoke **[measured]**, and
|
||||
invalid entries log with the asset name and index rather than crashing or
|
||||
silently skipping **[measured]** — a small thing that turns a content error into a
|
||||
searchable message.
|
||||
|
||||
---
|
||||
|
||||
## 3. Input as a buffered concern
|
||||
|
||||
Enhanced input does not call activation. It sends a semantic tag, the component
|
||||
records spec handles, and one call per frame from the controller's post-process
|
||||
input step does the work **[measured]**.
|
||||
|
||||
The ordering inside that call is the part to copy, and the reason is in the
|
||||
authors' own comment **[measured]**: collect held abilities with the
|
||||
while-input-active policy, then pressed abilities with the on-triggered policy,
|
||||
**then activate everything in one pass**, and only then process releases. Activate
|
||||
as you iterate and a held input first activates an ability and then delivers it a
|
||||
press event it should never have seen.
|
||||
|
||||
Two more decisions worth taking:
|
||||
|
||||
- **A single global block tag** clears every buffer and returns early
|
||||
**[measured]**. Suppressing all ability input during UI, death or stun is one
|
||||
tag rather than dozens of disabled input actions.
|
||||
- **Input events are replicated as events rather than directly** **[measured]**,
|
||||
with comments explaining that the wait-for-input ability tasks depend on it. A
|
||||
subtle choice that is easy to reverse by accident when copying.
|
||||
|
||||
---
|
||||
|
||||
## 4. Activation policy and exclusivity
|
||||
|
||||
Two small enums replace two large problems.
|
||||
|
||||
**Policy — when to activate** **[measured]**: on input triggered, while input
|
||||
active, on spawn. **[derived]** Automatic fire, single fire and passives become one
|
||||
enum in class defaults rather than three different implementations.
|
||||
|
||||
The spawn policy needs two entry points, not one — the ability granted to a live
|
||||
avatar, and the avatar appearing for an already-granted ability — and the audited
|
||||
implementation wires both **[measured]**, while skipping activation when the
|
||||
avatar is being torn down **[measured]**.
|
||||
|
||||
**Group — how to relate to others** **[measured]**: independent, exclusive
|
||||
replaceable, exclusive blocking. The component keeps per-group counts, cancels
|
||||
other replaceable exclusives when an exclusive starts, and asserts that no more
|
||||
than one exclusive is ever active **[measured]**.
|
||||
|
||||
**[derived]** Without this you express "firing cancels reloading, but the finisher
|
||||
cancels nothing" as an N×N matrix of blocking tags on every ability. With it, one
|
||||
enum plus one behaviour tag.
|
||||
|
||||
One invariant is enforced rather than documented: an ability in the replaceable
|
||||
group **cannot** refuse cancellation — the call logs an error and does nothing
|
||||
**[measured]**. **[derived]** That combination is a contradiction, and catching it
|
||||
in code rather than in review is the difference between a rule and a convention.
|
||||
|
||||
The death ability is the worked example **[measured]**: on activation it cancels
|
||||
everything not tagged as surviving death, refuses cancellation, and *moves itself*
|
||||
into the blocking group for the duration.
|
||||
|
||||
---
|
||||
|
||||
## 5. Match phases as abilities
|
||||
|
||||
The highest return per line in the layer: about 330 lines of C++ for a
|
||||
hierarchical match state machine **[measured]**.
|
||||
|
||||
A phase is an ability running on the game state's ability system component. While
|
||||
the ability is active, the phase is active. The ability's defaults differ from the
|
||||
base in exactly the two ways that matter **[measured]**: server-initiated
|
||||
execution and server-only security. A client cannot start a phase.
|
||||
|
||||
**Hierarchy comes from tag nesting, and that is the entire rule** **[measured]**.
|
||||
On starting a phase, every active phase whose tag does not match the incoming tag
|
||||
is cancelled:
|
||||
|
||||
| Active | Starting | Result |
|
||||
|---|---|---|
|
||||
| `Game.Playing` | `Game.Playing.SuddenDeath` | parent stays, child starts |
|
||||
| `Game.Playing`, `Game.Playing.SuddenDeath` | `Game.Playing.Overtime` | sibling ends, parent stays |
|
||||
| `Game.Playing.*` | `Game.PostGame` | whole subtree ends |
|
||||
| `Game.Playing` | second ability, same tag | both run |
|
||||
|
||||
**Parents and children coexist; siblings do not.**
|
||||
|
||||
What the pattern gets for free, because a phase is an ordinary ability
|
||||
**[derived]**: ability tasks as phase timers and triggers, effects applied for the
|
||||
phase's duration, cancellation by the same rules as anything else, and all of the
|
||||
per-mode logic authored in data rather than in new C++.
|
||||
|
||||
One observer detail is worth copying deliberately: subscribing to a phase that is
|
||||
**already active** fires the callback immediately **[measured]**. That closes the
|
||||
classic race where a system subscribes after the phase started and waits forever.
|
||||
The end-observer path has no equivalent **[measured]**, which is consistent —
|
||||
there is no "already ended" to catch up on.
|
||||
|
||||
### And two limitations the authors recorded themselves
|
||||
|
||||
- **Observers return no handle** **[measured]**, with two consecutive comments
|
||||
above the function saying so — including the honest second thought that a handle
|
||||
would not help if callers ignored it. GA-05.
|
||||
- **The subsystem's creation gate is commented out** **[measured]**, leaving an
|
||||
unconditional `return true`. So the subsystem exists on clients, where the
|
||||
active-phase map is always empty and the phase query always answers false.
|
||||
**[derived]** Client code must learn about phases through replicated tags,
|
||||
effects or messages — never through this subsystem. GA-04.
|
||||
|
||||
---
|
||||
|
||||
## 6. Attribute sets split by direction
|
||||
|
||||
Three classes on a clean axis **[measured]**:
|
||||
|
||||
- a base providing accessor macros and a six-parameter change delegate whose
|
||||
comment honestly warns that some parameters are null on clients;
|
||||
- a **receiving** set: health, max health, and the meta attributes damage and
|
||||
healing;
|
||||
- a **source** set: base damage, base heal.
|
||||
|
||||
The executions make the split structural rather than nominal **[measured]**: they
|
||||
capture the source set from the instigator and write into the receiving set on the
|
||||
target. **[derived]** The same actor usually owns both, but their roles in a
|
||||
calculation are different, and that difference lives in the class layout instead
|
||||
of a comment.
|
||||
|
||||
### The meta-attribute pattern
|
||||
|
||||
Incoming damage arrives in a non-replicated attribute, converts to a health change
|
||||
in post-execution, and resets to zero in the same call **[measured]**.
|
||||
|
||||
**[derived]** Damage and healing become one-shot channels rather than state. They
|
||||
need no replication, no manual reset, and no "current damage" that a reader might
|
||||
mistake for something durable. Health and damage are additionally marked as hidden
|
||||
from modifiers **[measured]** — so only an execution can change them, enforced by
|
||||
the editor rather than by agreement.
|
||||
|
||||
### Three interception levels, each doing one job
|
||||
|
||||
**[measured]**: a pre-execution veto that can cancel the whole effect (where
|
||||
damage immunity and a development-only god mode live, both respecting a
|
||||
self-destruct exception so suicide still works); a post-execution step that
|
||||
converts meta attributes, broadcasts change delegates and publishes a damage
|
||||
message; and clamping applied in both change hooks with a shared helper.
|
||||
|
||||
A latch flag prevents a duplicate out-of-health broadcast and resets when health
|
||||
recovers **[measured]** — the correct shape for GA-10.
|
||||
|
||||
### Where the damage number comes from
|
||||
|
||||
Four multiplications **[measured]**: base damage, distance attenuation, physical
|
||||
material attenuation, and a team multiplier that is **0 or 1 rather than a
|
||||
branch**. The two attenuation curves come from an interface implemented by the
|
||||
damage *source* — the weapon — rather than from the execution. **[derived]** Falloff
|
||||
therefore lives in weapon data where a designer owns it, and the execution stays
|
||||
generic.
|
||||
|
||||
---
|
||||
|
||||
## 7. The death seam: three responsibilities, three classes
|
||||
|
||||
```text
|
||||
effect execution changes Health
|
||||
→ receiving set broadcasts out-of-health
|
||||
→ health component sends a death gameplay event
|
||||
→ death ability runs the sequence
|
||||
→ health component enters DeathStarted, then DeathFinished
|
||||
```
|
||||
|
||||
**[measured]** throughout. The division is the point:
|
||||
|
||||
- **GAS owns the number** — how much health, and the signal that it reached zero;
|
||||
- **the health component owns the state** — a replicated enum with its own
|
||||
replication callback that can roll back a predicted transition and logs invalid
|
||||
ones;
|
||||
- **the death ability owns the sequence** — what plays, when control returns.
|
||||
|
||||
**[derived]** None of the three knows the others' internals; they communicate
|
||||
through delegates and a gameplay event.
|
||||
|
||||
Three details worth taking:
|
||||
|
||||
- the death trigger is configured in the ability's **class defaults**
|
||||
**[measured]**, so exactly one path can start it;
|
||||
- the ability's end path **always** calls the finish transition **[measured]** —
|
||||
insurance against a designer-authored sequence that does not reach the end;
|
||||
- death state is neither an attribute nor a tag but a replicated enum
|
||||
**[measured]**; the tags exist too, set loosely, because the state is already
|
||||
replicated by other means.
|
||||
|
||||
Alongside, a damage message and an elimination message go out over a message bus
|
||||
**[measured]**, so kill feeds, statistics and accolades never reference GAS.
|
||||
|
||||
---
|
||||
|
||||
## 8. Two more things worth copying
|
||||
|
||||
**Composable costs.** The stock system offers one cost effect. Here, an ability
|
||||
holds a list of inline cost objects, each with its own check, its own application,
|
||||
its own failure tag, and an optional "only charge on hit" rule whose hit
|
||||
determination is computed once, cached, and only on authority **[measured]**.
|
||||
**[derived]** Ammunition, inventory items and player resources become three small
|
||||
classes rather than three special cases.
|
||||
|
||||
**Failure reasons that reach the player.** The stock system swallows why an
|
||||
activation failed. Here the component notifies the owning client when the avatar
|
||||
is not locally controlled, the ability maps failure tags to user-facing text and
|
||||
to animations through two data maps, and delivery goes over a message bus
|
||||
**[measured]**. **[derived]** "Out of ammo" on the HUD costs no bespoke RPC, and
|
||||
the UI never learns that GAS exists. GA-14.
|
||||
|
||||
---
|
||||
|
||||
## 9. Copy carefully
|
||||
|
||||
Two places where the audited implementation is explicitly provisional, and says so:
|
||||
|
||||
- **The relationship mapping is a linear array walk**, consulted on every
|
||||
activation and on every block-and-cancel application, with the same "simple
|
||||
iteration for now" comment in three separate functions **[measured]**.
|
||||
**[derived]** Fine at the size measured; profile before assuming it scales.
|
||||
GA-08.
|
||||
- **Three function-static containers** are used as scratch space in a function on
|
||||
the activation path **[measured]**. **[derived]** Correct on one thread, not
|
||||
reentrant, and nothing at the call site says so. GA-07.
|
||||
|
||||
And one thing that reads as finished and is not: attribute initialization is a
|
||||
placeholder that assigns health from max health in code, with a comment saying a
|
||||
data-driven source will replace it **[measured]**. The global curve table is
|
||||
configured to none **[measured]**, and the initialization-data field on the grant
|
||||
structure has no traced consumer **[open]**. GA-11.
|
||||
|
||||
---
|
||||
|
||||
## 10. What source reading could not settle
|
||||
|
||||
- **Binary assets were not read.** The concrete phase abilities, the ability sets,
|
||||
the contents of the relationship mapping for each archetype, and the activation
|
||||
groups configured on real abilities are all **[open]**.
|
||||
- **Who starts the first phase is unknown** **[open]**. A search for the start
|
||||
call finds only the subsystem itself, so the caller is in a scripted asset.
|
||||
- **One field could not be classified from this layer alone**: a
|
||||
cancellation-logging flag is declared and assigned in the constructor and read
|
||||
nowhere in the audited directories **[measured]**. Whether something outside them
|
||||
reads it is **[open]** — though its own comment describes it as temporary,
|
||||
added while tracking a bug. GA-06.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against the GAS layer of Epic's Lyra Starter Game on Unreal Engine 5.6,
|
||||
read as source in a single workspace. Source addresses stay in the research
|
||||
archive that produced this skill; each `GA-` identifier resolves back to the
|
||||
audited location there, so any specific claim above can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. Binary assets were not read, so
|
||||
every statement about authored content is marked open rather than concluded. The
|
||||
architecture is transferable; the specific defects are evidence, not guarantees
|
||||
about other versions. Re-run the recipes in
|
||||
[failure-modes.md](failure-modes.md) against your own tree before acting on
|
||||
anything here.
|
||||
@@ -0,0 +1,387 @@
|
||||
---
|
||||
name: ue-input-architecture
|
||||
description: >-
|
||||
Design or review Enhanced Input architecture in Unreal Engine: physical
|
||||
mapping contexts, input actions, semantic gameplay tags, input config data,
|
||||
native versus ability input paths, per-frame buffering on the ability system
|
||||
component, player-mappable settings, device switching, and runtime injection
|
||||
and removal of input by feature plugins. Use when adding controls, rebinding,
|
||||
gamepad or touch support, ability input, mode-specific input, or debugging
|
||||
actions that bind correctly and never fire.
|
||||
---
|
||||
|
||||
# UE input architecture
|
||||
|
||||
The invariant:
|
||||
|
||||
> Physical controls, semantic commands and gameplay implementations are three
|
||||
> separate layers joined by validated data.
|
||||
|
||||
```text
|
||||
key / gamepad / touch
|
||||
→ input mapping context
|
||||
→ input action
|
||||
→ input config: input action → input tag
|
||||
→ native handler OR ability system component
|
||||
→ ability set: input tag → ability
|
||||
```
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes for silent input failures: [failure modes](references/failure-modes.md).
|
||||
|
||||
**The characteristic failure of this architecture is a button that does nothing
|
||||
and logs nothing.** Read the failure modes before adopting it — most of them are
|
||||
variations on that one observable.
|
||||
|
||||
Related skills: `ue-gameplay-tag-governance`, `ue-gas-architecture`,
|
||||
`ue-modular-gameplay`, `ue-ui-architecture`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Keep the three layers independent
|
||||
|
||||
### Physical — the mapping context
|
||||
|
||||
Owns key, button and axis to input action; modifiers and triggers; priority
|
||||
relative to other contexts; player-mappable metadata.
|
||||
|
||||
Changing a keyboard layout or a gamepad binding should stop here.
|
||||
|
||||
### Semantic — the input config
|
||||
|
||||
An immutable asset owning pairs:
|
||||
|
||||
```text
|
||||
IA_Move → InputTag.Move
|
||||
IA_WeaponFire → InputTag.Weapon.Fire
|
||||
IA_Dash → InputTag.Ability.Dash
|
||||
```
|
||||
|
||||
Gameplay depends on tags, not on asset paths or physical keys.
|
||||
|
||||
### Capability — handlers and ability sets
|
||||
|
||||
Native tags bind to explicit handlers. Ability tags travel to the ability system
|
||||
component as payload. Granted ability specs carry the matching semantic tag.
|
||||
|
||||
Changing an ability implementation should not require touching a mapping context.
|
||||
|
||||
---
|
||||
|
||||
## 2. Split native and ability actions
|
||||
|
||||
Make the distinction explicit in the config, with two separate lists.
|
||||
|
||||
**Native** — deterministic controller and pawn functions: move, look, crouch,
|
||||
autorun. Each costs one explicit binding line, and the tag is a **lookup key at
|
||||
bind time only**. The handler receives an input value; the tag does not exist at
|
||||
runtime.
|
||||
|
||||
**Ability** — bound through one pressed/released pair with the row's tag passed
|
||||
as a payload argument. One pair of functions serves any number of abilities, and
|
||||
the count of new abilities you can add without writing code is unbounded.
|
||||
|
||||
### Classification test
|
||||
|
||||
Native when the action controls continuous motion, must execute immediately and
|
||||
deterministically, and needs no ability lifecycle.
|
||||
|
||||
Ability when the action activates a grantable capability, can be blocked or
|
||||
cancelled by tags, has cost or cooldown or prediction, or changes with equipment
|
||||
or mode.
|
||||
|
||||
### Watch what the ability path fixes at bind time
|
||||
|
||||
If ability binding is hardcoded to two trigger events, then hold, tap and
|
||||
double-tap variants must be expressed through triggers and modifiers in the
|
||||
mapping context rather than through trigger-event selection. That is a
|
||||
reasonable constraint — but it is a constraint, and it is invisible until
|
||||
someone needs a third event.
|
||||
|
||||
---
|
||||
|
||||
## 3. Join semantic tags at grant time
|
||||
|
||||
When an ability set grants an ability, add the configured input tag to the
|
||||
spec's dynamic source tags. Lookup then becomes:
|
||||
|
||||
```text
|
||||
pressed input tag
|
||||
→ scan activatable specs
|
||||
→ exact tag match
|
||||
→ buffer the spec handle
|
||||
```
|
||||
|
||||
Do not store an input action pointer on the ability. That couples a capability to
|
||||
one content asset and blocks per-mode and per-device remapping.
|
||||
|
||||
### The match is exact, not hierarchical
|
||||
|
||||
A spec tagged `InputTag.Weapon` is **not** matched by a press of
|
||||
`InputTag.Weapon.Fire`. The tag hierarchy organises the vocabulary and filters
|
||||
the editor picker; it does not participate in lookup. Assuming otherwise
|
||||
produces a binding that is correct in every visible respect and never fires.
|
||||
Recipe: IN-01.
|
||||
|
||||
### Validate the join
|
||||
|
||||
The input config and the ability set are independent assets. Their tag equality
|
||||
is a foreign key with no database behind it, so validation is the only compiler
|
||||
it will ever have:
|
||||
|
||||
- every input-driven ability has a non-empty tag;
|
||||
- every intended ability tag is reachable from at least one active config;
|
||||
- every config ability tag has a granted consumer in the intended composition;
|
||||
- duplicates are intentional;
|
||||
- category constraints hold on both sides.
|
||||
|
||||
---
|
||||
|
||||
## 4. Buffer ability input, process once per frame
|
||||
|
||||
Forwarding a tag must not activate immediately. Record held handles, pressed this
|
||||
frame, released this frame; process from the controller's post-input step in
|
||||
three passes — held with the while-active policy, then pressed with the
|
||||
on-triggered policy, then activate everything in one batch, then releases.
|
||||
|
||||
This removes event-order dependence and makes simultaneous input resolvable.
|
||||
See `ue-gas-architecture` for the activation policies and the global block tag.
|
||||
|
||||
**Accept the consequence:** the ability path is one input-processing step behind
|
||||
the native path within the same frame. For most games that is invisible. It is
|
||||
still a real difference between the two paths, and it belongs in the
|
||||
classification decision rather than being discovered later.
|
||||
|
||||
---
|
||||
|
||||
## 5. Initialize input at a named readiness boundary
|
||||
|
||||
Binding depends on several independently arriving pieces: a local controller, the
|
||||
input subsystem, the pawn data and its config, the project's input component
|
||||
subclass, and any contexts added by features.
|
||||
|
||||
Do not treat `BeginPlay` as ready. Use init states and emit a semantic event —
|
||||
"bind inputs now" — once the dependencies hold.
|
||||
|
||||
Feature actions must respond to **both** the generic extension-added event and
|
||||
the semantic ready event, because either can arrive first. Set the readiness flag
|
||||
*before* broadcasting, or a handler that checks it during the broadcast will see
|
||||
false.
|
||||
|
||||
### Base initialization checklist
|
||||
|
||||
- [ ] the component is the expected input subclass, checked with a diagnostic
|
||||
that names the fix;
|
||||
- [ ] clearing existing mappings is deliberate and its consequences are handled;
|
||||
- [ ] default contexts added to the local player subsystem;
|
||||
- [ ] settings registration separated from runtime activation (§6);
|
||||
- [ ] native actions bound explicitly;
|
||||
- [ ] ability actions bound **and their handles retained** (§7);
|
||||
- [ ] the ready event is sent exactly once per initialization;
|
||||
- [ ] late feature handlers can bind safely.
|
||||
|
||||
### Clearing all mappings is a broadcast dependency
|
||||
|
||||
If initialization clears every mapping on the subsystem, it also removes contexts
|
||||
added by already-active features. Recovery then depends on those features hearing
|
||||
the ready event and re-adding their contexts — which makes correctness a property
|
||||
of event ordering rather than of stored state. It works; it is fragile; and it
|
||||
must be a documented decision rather than a side effect. Recipe: IN-08.
|
||||
|
||||
---
|
||||
|
||||
## 6. Separate "register with settings" from "activate now"
|
||||
|
||||
These are two different operations:
|
||||
|
||||
```text
|
||||
RegisterInputMappingContext(IMC) // discoverable and rebindable
|
||||
AddMappingContext(IMC, priority) // actually active
|
||||
```
|
||||
|
||||
A context may be active but not remappable, or registered but not currently
|
||||
active. All four combinations are meaningful.
|
||||
|
||||
**One flag must not gate both.** If it does, unchecking "expose this for
|
||||
rebinding" also silently removes the input — a designer changes a settings
|
||||
checkbox and the character stops responding, with nothing in the log. Recipe:
|
||||
IN-06.
|
||||
|
||||
---
|
||||
|
||||
## 7. Modular input needs ownership receipts
|
||||
|
||||
### Adding a mapping context
|
||||
|
||||
Data: context, priority, register-with-settings flag. Retain per activation
|
||||
context: the extension and delegate handles, the exact players touched, and the
|
||||
exact contexts and priorities activated.
|
||||
|
||||
### Adding ability bindings
|
||||
|
||||
Data: a list of input configs. Retain per pawn: **every returned bind handle**,
|
||||
which config produced it, and the readiness subscription.
|
||||
|
||||
### Symmetric removal
|
||||
|
||||
```text
|
||||
remove the exact bind handles this config created
|
||||
remove the exact contexts this action added
|
||||
unregister from settings only if this action registered them
|
||||
release extension and delegate handles
|
||||
```
|
||||
|
||||
**Never implement removal by clearing all bindings or all mappings on the pawn.**
|
||||
That destroys other features' ownership, and it will look like it works.
|
||||
|
||||
The measured reference composes this well and tears it down badly: bind handles
|
||||
are declared as locals and discarded, so the removal function *cannot* be
|
||||
implemented without changing the component that binds. Copy the composition;
|
||||
write the teardown yourself. Recipes: IN-03, IN-04.
|
||||
|
||||
---
|
||||
|
||||
## 8. Device switching belongs above gameplay semantics
|
||||
|
||||
Gameplay should not branch on keyboard versus gamepad for the same command. The
|
||||
UI layer owns the active input type, glyphs, controller brand and style,
|
||||
change notifications, and disconnected-device handling.
|
||||
|
||||
Gameplay may still use **distinct actions and tags** where the semantics genuinely
|
||||
differ — mouse delta and analog stick look need different modifiers and
|
||||
sensitivity, so:
|
||||
|
||||
```text
|
||||
InputTag.Look.Mouse
|
||||
InputTag.Look.Stick
|
||||
```
|
||||
|
||||
That is semantic distinction, not device sniffing. The test is whether the two
|
||||
paths do different arithmetic; if they do, they are different commands.
|
||||
|
||||
Per-player sensitivity, dead zones and inversion belong in input modifiers that
|
||||
read settings — not in handler code. Note that a modifier which resolves a setting
|
||||
by property **name** through reflection will break silently on a rename, with no
|
||||
compiler help. Recipe: IN-10.
|
||||
|
||||
---
|
||||
|
||||
## 9. Context priority and consumption
|
||||
|
||||
A handler can be perfectly bound and never fire because an earlier context
|
||||
consumes the key.
|
||||
|
||||
Before adding a mapping: list every active context in priority order; find all
|
||||
mappings for the physical key; inspect triggers, modifiers and consume-input
|
||||
behaviour; establish whether the contexts can be active simultaneously; and
|
||||
confirm no raw key path bypasses the system.
|
||||
|
||||
Priority is meaningful only between concurrently active contexts. Document why a
|
||||
feature needs to outrank the baseline.
|
||||
|
||||
### Avoid raw key fallback
|
||||
|
||||
Raw key handling bypasses rebinding, device parity, triggers and modifiers,
|
||||
context priority, and input profiles. Use it only for editor and debug controls,
|
||||
with a stated reason.
|
||||
|
||||
---
|
||||
|
||||
## 10. The config-file boundary
|
||||
|
||||
With Enhanced Input, the project input config file should contain the
|
||||
player-input and input-component classes, user-settings class registration, debug
|
||||
bindings, UI input defaults, and legacy axis properties only where the engine
|
||||
still requires them.
|
||||
|
||||
Gameplay mappings live in mapping context assets; semantic mappings live in input
|
||||
config assets. **Legacy axis configuration that sits beside Enhanced Input is
|
||||
actively misleading** — it looks like the place sensitivity is set, and it is not.
|
||||
Recipe: IN-09.
|
||||
|
||||
---
|
||||
|
||||
## 11. Test matrix
|
||||
|
||||
### Composition
|
||||
|
||||
Baseline pawn data and config; a mode action set adding a second config;
|
||||
equipment granting and removing matching abilities; a feature activating both
|
||||
before and after pawn readiness; two features adding distinct configs at once.
|
||||
|
||||
### Lifecycle
|
||||
|
||||
```text
|
||||
spawn → bind → activate feature → press / hold / release
|
||||
→ deactivate feature → verify no response
|
||||
→ reactivate → verify exactly one response
|
||||
```
|
||||
|
||||
### Devices
|
||||
|
||||
Keyboard and mouse; gamepad; touch; hot-swap while UI is open; a saved and
|
||||
reloaded remap; two local players with separate profiles.
|
||||
|
||||
### Network and abilities
|
||||
|
||||
Owning-client predicted activation; server rejection with failure feedback; input
|
||||
held across a possession or avatar change; the block tag clearing buffered state;
|
||||
an ability removed while its input is held.
|
||||
|
||||
---
|
||||
|
||||
## 12. Review checklist
|
||||
|
||||
- [ ] Physical key knowledge stops at the mapping context.
|
||||
- [ ] Gameplay uses semantic tags, not asset references.
|
||||
- [ ] Native versus ability classification is intentional.
|
||||
- [ ] Ability input is buffered, not activated inside the event callback.
|
||||
- [ ] Config and ability-set tags are validated as a join, in both directions.
|
||||
- [ ] Tag matching semantics are known to be exact, and the data suits that.
|
||||
- [ ] Binding waits for a named readiness state, not for begin play.
|
||||
- [ ] Base and feature bindings both retain handles.
|
||||
- [ ] Runtime activation and settings registration are independent.
|
||||
- [ ] Deactivation removes only what this owner added.
|
||||
- [ ] Priority and consumption conflicts have been audited for shared keys.
|
||||
- [ ] Device icons and method switching stay in the UI layer.
|
||||
- [ ] Activate, deactivate, reactivate produces no duplicate callbacks.
|
||||
|
||||
---
|
||||
|
||||
## 13. When this architecture is too much
|
||||
|
||||
For a single-mode project with one device family, no dynamic abilities and no
|
||||
rebinding, direct bindings are honest and this indirection is overhead.
|
||||
|
||||
Introduce semantic tags when at least one becomes true: abilities are granted
|
||||
dynamically; several pawn archetypes reuse controls; features add input at
|
||||
runtime; rebinding and device parity matter; gameplay must be independent of
|
||||
content asset paths.
|
||||
|
||||
The architecture is paid for by **modularity of content** and **many archetypes
|
||||
with different action sets**. Without either, the cost is real and the benefit is
|
||||
not: to answer "what does the left mouse button do?" a reader must traverse
|
||||
mapping context, action, config, tag, pawn data, ability set and spec tags —
|
||||
seven assets instead of one call stack.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
the input layer of Epic's Lyra Starter Game on Unreal Engine 5.6 — about 870
|
||||
lines across seven file pairs, plus its integration points in the hero component
|
||||
and two feature actions — read as source rather than run.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; what
|
||||
ships is the detection recipe. Each entry carries a stable identifier (`IN-06`
|
||||
and up) that resolves back to the audited location.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. Mapping contexts and input
|
||||
actions are binary assets and were not read, so every claim about which keys map
|
||||
to which actions is out of scope rather than concluded. Re-run the recipes
|
||||
against your own tree before acting on any specific claim.
|
||||
@@ -0,0 +1,405 @@
|
||||
# Failure modes: input architecture
|
||||
|
||||
Eleven ways a tag-addressed input layer binds correctly and never fires.
|
||||
|
||||
The shared property is unusually literal here: **the ability input path has no
|
||||
error case.** A press resolves to a tag, the tag is compared against granted
|
||||
ability specs, and no match means no activation. That is the same code path as
|
||||
"you do not currently have this ability", which is a normal state. Nothing logs,
|
||||
because from the framework's view nothing went wrong.
|
||||
|
||||
The native path is louder — an unresolved tag produces an error at bind time —
|
||||
which is worth knowing when triaging: **if the button is native and silent, the
|
||||
problem is upstream of the config; if it is an ability and silent, the problem
|
||||
could be anywhere along seven assets.**
|
||||
|
||||
Recipes use `rg` from a project's source root, and were executed against the
|
||||
audited project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## The join
|
||||
|
||||
### IN-01 - Tag lookup is exact, and the hierarchy suggests otherwise
|
||||
|
||||
**Mechanism.** The ability spec lookup compares the pressed tag against the
|
||||
spec's dynamic source tags with an exact match. A spec tagged with a parent is
|
||||
not found by a press of its child.
|
||||
|
||||
**Why it is silent.** No match means no activation, which is indistinguishable
|
||||
from not having been granted the ability. There is no "close match" concept to
|
||||
warn about.
|
||||
|
||||
**Why the obvious check misses it.** Everything else about the tag layer is
|
||||
hierarchical — the editor picker filters by category, the vocabulary is a tree,
|
||||
and neighbouring systems in the same codebase do use hierarchical matching. The
|
||||
reasonable assumption is wrong here, and nothing at either the config or the
|
||||
ability-set end states the matching semantics.
|
||||
|
||||
**Symptom.** An ability that is granted, bound, and correct in every inspectable
|
||||
respect simply does not activate. Designers re-author the tag, re-save the
|
||||
assets, and eventually route around it by duplicating the leaf.
|
||||
|
||||
**Detect.** Read the comparison and confirm the semantics, then check whether any
|
||||
authored tag relies on the other one:
|
||||
|
||||
```bash
|
||||
rg -n "HasTagExact|MatchesTag" --glob "*.cpp" . | rg -i "input|spec"
|
||||
```
|
||||
|
||||
Then list every input tag used on an ability spec and every tag used in a config,
|
||||
and confirm the sets are equal rather than merely overlapping. In the audited
|
||||
project both lookup sites use exact matching, on adjacent lines.
|
||||
|
||||
**Guardrail.** State the matching semantics in a comment at the lookup and in the
|
||||
tag namespace's own documentation. If hierarchical matching is wanted, implement
|
||||
it deliberately — do not let readers infer it from the tag shape.
|
||||
|
||||
---
|
||||
|
||||
### IN-02 - Two tag sources for one namespace
|
||||
|
||||
**Mechanism.** Native input tags are declared in C++ as compile-time symbols;
|
||||
ability input tags are declared in a config file as strings. Both live under the
|
||||
same namespace root.
|
||||
|
||||
**Why it is silent.** Both work. The engine resolves both to the same registry,
|
||||
and a tag from either source behaves identically at runtime.
|
||||
|
||||
**Why the obvious check misses it.** Looking at either source shows a complete,
|
||||
well-organised list. There is no artefact anywhere that shows both, so the
|
||||
vocabulary appears smaller than it is and its split appears not to exist.
|
||||
|
||||
**Symptom.** An audit of the input vocabulary misses half of it. A rename sweep
|
||||
covers one source. A developer looking for a tag symbol in C++ concludes a tag
|
||||
does not exist when it is declared in config.
|
||||
|
||||
**Detect.** Count both sources and compare against the namespace:
|
||||
|
||||
```bash
|
||||
rg -n 'UE_DEFINE_GAMEPLAY_TAG\w*\([^,]+,\s*"InputTag\.' --glob "*.cpp" .
|
||||
rg -n 'GameplayTagList=\(Tag="InputTag\.' Config/ Plugins/*/Config/
|
||||
```
|
||||
|
||||
Both non-empty is the finding. In the audited project the native movement tags
|
||||
come from C++ and the ability tags from config, and the two sets do not overlap
|
||||
at all — which is a defensible split, undocumented.
|
||||
|
||||
**Guardrail.** One namespace, one declaration source, documented. If a split is
|
||||
deliberate, draw the line at a sub-namespace boundary and write it down where
|
||||
both halves are visible.
|
||||
|
||||
---
|
||||
|
||||
## Ownership and teardown
|
||||
|
||||
### IN-03 - Bind handles discarded, making removal impossible
|
||||
|
||||
**Mechanism.** The binding helper returns handles through an out-parameter. Both
|
||||
callers declare that array as a local and let it fall out of scope.
|
||||
|
||||
**Why it is silent.** The bindings work perfectly. Handles are needed only to
|
||||
undo, and nothing undoes during a normal session — pawns are recreated between
|
||||
matches, which disposes of the input component and hides the leak.
|
||||
|
||||
**Why the obvious check misses it.** The out-parameter is passed correctly, the
|
||||
call is idiomatic, and a removal function exists elsewhere in the class. The
|
||||
defect is the *storage duration* of a local variable, which no review of API
|
||||
usage examines.
|
||||
|
||||
**Symptom.** Deactivating a feature leaves its ability bindings on a live pawn.
|
||||
Harmless while the ability is also revoked — the tag reaches the component and
|
||||
finds no spec — but a re-grant resurrects phantom input.
|
||||
|
||||
**Detect.** Find handle arrays declared at the call site rather than as members:
|
||||
|
||||
```bash
|
||||
rg -n -B3 "BindAbilityActions|BindAction\(" --glob "*.cpp" . \
|
||||
| rg "TArray<uint32>\s+\w+;|TArray<\w*Handle>\s+\w+;"
|
||||
```
|
||||
|
||||
In the audited project this appears at two sites in one component: the initial
|
||||
binding and the feature-driven addition.
|
||||
|
||||
**Guardrail.** Handles live on the object whose lifetime governs the binding,
|
||||
stored in the same change that creates them. If there is nowhere to store them,
|
||||
the input is not removable and that must be stated before it ships.
|
||||
|
||||
---
|
||||
|
||||
### IN-04 - Removal API that exists and does nothing
|
||||
|
||||
**Mechanism.** The class exposes a removal function whose body is a comment. A
|
||||
feature action calls it during teardown.
|
||||
|
||||
**Why it is silent.** The call compiles and returns. Nothing distinguishes
|
||||
"removed the bindings" from "did nothing", because neither returns a result.
|
||||
|
||||
**Why the obvious check misses it.** The teardown path *is* implemented in the
|
||||
feature action — it iterates its tracked pawns and calls removal on each — so a
|
||||
review of the action finds a complete, symmetric implementation. The empty body
|
||||
is one call away in a different subsystem.
|
||||
|
||||
**Symptom.** As IN-03, from the other end. A team scopes "implement the missing
|
||||
removal" from the empty function, and finds the estimate was an order of
|
||||
magnitude low because the handles no longer exist to remove.
|
||||
|
||||
**Detect.** Find removal functions with no statements, and dead removal helpers:
|
||||
|
||||
```bash
|
||||
rg -n -A4 "void \w+::Remove\w+\(" --glob "*.cpp" . | rg -B2 "@?TODO"
|
||||
rg -c "RemoveBinds|RemoveBindingByHandle" --glob "*.cpp" --glob "*.h" .
|
||||
```
|
||||
|
||||
The second command matters: in the audited project a working removal helper
|
||||
exists on the input component and has **zero callers** — the mechanism is
|
||||
present, and the data it needs was thrown away by IN-03.
|
||||
|
||||
**Guardrail.** A removal function with no body must not compile quietly: make it
|
||||
pure virtual, assert, or delete the caller. Then check that its helper has at
|
||||
least one caller — a correct utility nobody can call is the same as no utility.
|
||||
|
||||
---
|
||||
|
||||
### IN-05 - Tracking container that is never written
|
||||
|
||||
**Mechanism.** The context-adding action keeps a list of the controllers it
|
||||
touched, removes from it, and iterates it during reset — but the add path takes
|
||||
the per-context data as a parameter and never writes to it.
|
||||
|
||||
**Why it is silent.** The reset loop finds the list empty and completes. Contexts
|
||||
still come off, but only reactively — when the receiver is reported as going
|
||||
away — so behaviour depends on teardown order.
|
||||
|
||||
**Why the obvious check misses it.** Four of five operations on the container
|
||||
exist, including an assertion at activation that it starts empty. Any check of
|
||||
"is this container used?" answers yes emphatically. Only the producer is missing.
|
||||
|
||||
**Symptom.** Mapping contexts survive feature deactivation when the receiving
|
||||
controller outlives the feature. Intermittent, ordering-dependent.
|
||||
|
||||
**Detect.** Compare reads against writes per container:
|
||||
|
||||
```bash
|
||||
C='ControllersAddedTo'
|
||||
rg -n "\b$C\b" --glob "*.cpp" --glob "*.h" .
|
||||
rg -n "\b$C\b\s*\.\s*(Add|AddUnique|Emplace|Push)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
In the audited project the second search is empty, the authors left a comment on
|
||||
the reset path noting exactly that, and a sibling action in the same directory
|
||||
does the equivalent correctly — so the two files diff against each other.
|
||||
|
||||
**Guardrail.** A container whose only mutations are removals has no producer.
|
||||
Assert non-empty after the add path runs.
|
||||
|
||||
---
|
||||
|
||||
## Configuration traps
|
||||
|
||||
### IN-06 - One flag secretly gating two operations
|
||||
|
||||
**Mechanism.** Registering a context for rebinding and activating it are separate
|
||||
operations. In one code path the activation call is nested **inside** the
|
||||
`if` that tests the register-with-settings flag.
|
||||
|
||||
**Why it is silent.** Both operations are legitimate and both are performed when
|
||||
the flag is set — which it is by default. Nothing about the default case is
|
||||
wrong, so the coupling never manifests during development.
|
||||
|
||||
**Why the obvious check misses it.** The flag's name describes exactly one of the
|
||||
two things it controls. A designer reading the property, and a reviewer reading
|
||||
the property, both understand it correctly and neither is looking at the brace
|
||||
structure twenty lines away. Worse: the *same* project has a second
|
||||
implementation of the same feature where the flag correctly controls only
|
||||
registration, so a reader who checks one path and generalises is wrong.
|
||||
|
||||
**Symptom.** Unchecking "expose this for rebinding" makes the character stop
|
||||
responding entirely, with nothing in the log. Diagnosed as an asset problem.
|
||||
|
||||
**Detect.** Find activation calls nested inside a settings-registration branch:
|
||||
|
||||
```bash
|
||||
rg -n -B8 "AddMappingContext\(" --glob "*.cpp" . | rg "bRegisterWithSettings|RegisterInputMappingContext"
|
||||
```
|
||||
|
||||
Any hit is the finding. Then check every other implementation of the same
|
||||
operation in the project and compare — the discrepancy between two paths is
|
||||
stronger evidence than either path alone.
|
||||
|
||||
**Guardrail.** Registration and activation are separate statements at the same
|
||||
brace level, each guarded by its own condition. Test all four combinations of the
|
||||
two flags, because all four are meaningful.
|
||||
|
||||
---
|
||||
|
||||
### IN-07 - Nested condition that disables an unrelated feature
|
||||
|
||||
**Mechanism.** The loop that applies default mapping contexts sits inside an
|
||||
`if` that tests for a semantic config. A pawn with no config therefore gets no
|
||||
contexts either.
|
||||
|
||||
**Why it is silent.** Every shipped pawn has a config, so the guard is always
|
||||
true and the nesting is never exercised. The two things are unrelated in concept
|
||||
and adjacent in code.
|
||||
|
||||
**Why the obvious check misses it.** The guard is correct for the block it was
|
||||
written for. The extra scope it captures is a brace position, not a statement,
|
||||
and nothing names the coupling.
|
||||
|
||||
**Symptom.** A new pawn archetype authored without a config is completely
|
||||
unresponsive — not "no abilities", but no movement, no look, nothing. Read as a
|
||||
much deeper problem than a missing asset reference.
|
||||
|
||||
**Detect.** Read the initialization function's brace structure:
|
||||
|
||||
```bash
|
||||
rg -n -A25 "::InitializePlayerInput" --glob "*.cpp" . | rg -n "if \(|for \(|AddMappingContext"
|
||||
```
|
||||
|
||||
A mapping-context loop indented inside a config guard is the finding.
|
||||
|
||||
**Guardrail.** Contexts and configs are independent inputs to initialization and
|
||||
belong in sibling blocks. Validate that a pawn without a config is a deliberate
|
||||
configuration rather than a broken one.
|
||||
|
||||
---
|
||||
|
||||
### IN-08 - Clearing all mappings, then relying on a broadcast to restore them
|
||||
|
||||
**Mechanism.** Pawn input initialization clears every mapping on the subsystem
|
||||
before adding its own — removing contexts that active features added earlier. A
|
||||
readiness broadcast afterwards prompts those features to re-add theirs.
|
||||
|
||||
**Why it is silent.** It works, because the broadcast happens and the features
|
||||
listen. Correctness is real; it is just a property of event ordering rather than
|
||||
of stored state.
|
||||
|
||||
**Why the obvious check misses it.** The clear is a defensible line with an
|
||||
obvious intent — start from a known state. The features' re-add is in a different
|
||||
file, reached by an event. Nothing at the clear site says what it is destroying,
|
||||
and nothing at the re-add site says why it is necessary.
|
||||
|
||||
**Symptom.** A feature's input disappears when a pawn re-initializes, in any
|
||||
ordering where that feature does not receive or does not act on the ready event —
|
||||
a feature activated during the broadcast, one whose handler early-returns, or one
|
||||
whose receiver is not yet registered.
|
||||
|
||||
**Detect.** Find the clear and confirm the recovery path:
|
||||
|
||||
```bash
|
||||
rg -n -A3 "ClearAllMappings" --glob "*.cpp" .
|
||||
rg -n "SendGameFrameworkComponentExtensionEvent" --glob "*.cpp" . -B4
|
||||
```
|
||||
|
||||
A clear with no explicit re-application, relying on an event broadcast elsewhere,
|
||||
is the finding. In the audited project the flag that gates readiness is set
|
||||
*before* the broadcast — the correct order, and worth verifying rather than
|
||||
assuming.
|
||||
|
||||
**Guardrail.** Either do not clear what you do not own, or re-apply from a stored
|
||||
list rather than from a broadcast. If the broadcast is the mechanism, comment it
|
||||
at the clear site so the dependency is visible from both ends.
|
||||
|
||||
---
|
||||
|
||||
## Cost and misleading surfaces
|
||||
|
||||
### IN-09 - Legacy axis configuration beside Enhanced Input
|
||||
|
||||
**Mechanism.** The project input config retains legacy axis entries — dead zones,
|
||||
sensitivities — while actual sensitivity and dead zone are applied by input
|
||||
modifiers.
|
||||
|
||||
**Why it is silent.** The legacy values are read by the legacy path, which is not
|
||||
in use. Changing them does nothing, and doing nothing produces no error.
|
||||
|
||||
**Why the obvious check misses it.** They are exactly where a developer would
|
||||
look for sensitivity, with exactly the names they would search for. The file
|
||||
answers the question asked; the answer is stale.
|
||||
|
||||
**Symptom.** Time lost tuning values that have no effect, and a plausible but
|
||||
wrong belief about where input tuning lives — which then propagates into
|
||||
documentation.
|
||||
|
||||
**Detect.** Compare the legacy entries against the modifiers that really apply:
|
||||
|
||||
```bash
|
||||
rg -n "AxisConfig|Sensitivity=|DeadZone" Config/*.ini
|
||||
rg -n "class \w*InputModifier" --glob "*.h" .
|
||||
```
|
||||
|
||||
Both non-empty is the finding. In the audited project eleven legacy axis entries
|
||||
coexist with four modifiers that read player settings, and the modifiers win.
|
||||
|
||||
**Guardrail.** Delete legacy entries the engine does not require, or annotate them
|
||||
in place as inert with a pointer to the modifiers. A misleading config is more
|
||||
expensive than a missing one.
|
||||
|
||||
---
|
||||
|
||||
### IN-10 - Setting resolved by property name through reflection
|
||||
|
||||
**Mechanism.** An input modifier scales a value by looking up a settings property
|
||||
**by name** through reflection, with a cache.
|
||||
|
||||
**Why it is silent.** A rename produces no compile error, because the name is a
|
||||
string in data. The lookup fails at runtime and the modifier returns unscaled
|
||||
input — a plausible value.
|
||||
|
||||
**Why the obvious check misses it.** Renaming a property is exactly the operation
|
||||
IDEs make safe. The developer performs a rename refactor, everything compiles,
|
||||
every reference updates, and this one does not because it is not a reference.
|
||||
|
||||
**Symptom.** Sensitivity, inversion or dead zone silently stops respecting the
|
||||
player's setting after an unrelated refactor. Reported as "the setting does
|
||||
nothing" long after the rename.
|
||||
|
||||
**Detect.** Find name-based property resolution:
|
||||
|
||||
```bash
|
||||
rg -n "FindPropertyByName|FindFProperty|GetPropertyByName" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Each hit is a link the compiler does not check. Add a startup assertion for each.
|
||||
|
||||
**Guardrail.** Resolve settings through typed accessors. Where reflection is
|
||||
unavoidable, assert at startup that every name resolves, so a rename fails loudly
|
||||
at launch rather than quietly in gameplay.
|
||||
|
||||
---
|
||||
|
||||
## Loading
|
||||
|
||||
### IN-11 - Blocking load on the pawn initialization path
|
||||
|
||||
**Mechanism.** A mapping context soft reference is resolved with a synchronous
|
||||
load during pawn input initialization.
|
||||
|
||||
**Why it is silent.** It succeeds. The asset arrives, input works, and the cost is
|
||||
a stall measured in milliseconds on a warm cache.
|
||||
|
||||
**Why the obvious check misses it.** A synchronous load is the simplest correct
|
||||
code, and correctness is what review checks. The cost appears only on a cold
|
||||
cache, on slower storage, or at a moment — player spawn — that nobody profiles
|
||||
because it happens before the part they were measuring.
|
||||
|
||||
**Symptom.** A hitch at spawn that is attributed to the level, the character mesh
|
||||
or the ability system, because the input layer is not where anyone looks for a
|
||||
frame spike.
|
||||
|
||||
**Detect.** Find synchronous loads on initialization paths, and contrast them with
|
||||
the bare accessors elsewhere:
|
||||
|
||||
```bash
|
||||
rg -n "LoadSynchronous\(\)" --glob "*.cpp" . -B6 | rg "Init|BeginPlay|Setup"
|
||||
rg -n "\.Get\(\)" --glob "*.cpp" . -A2 | rg -v "ensure|check|if\s*\(|UE_LOG"
|
||||
```
|
||||
|
||||
The pairing is the interesting part: the audited project loads synchronously on
|
||||
the base path and uses bare, unguarded accessors on the feature paths — so one
|
||||
route hitches and the other silently skips when a bundle did not load.
|
||||
|
||||
**Guardrail.** Preload input assets with the bundle that owns them and resolve
|
||||
with a guarded accessor that logs on failure. Neither blocking nor silently
|
||||
skipping is acceptable on a path this visible.
|
||||
@@ -0,0 +1,225 @@
|
||||
# Patterns: tag-addressed input in a measured reference product
|
||||
|
||||
A worked example of the three-layer architecture from `SKILL.md`, measured in one
|
||||
reference project. It exists to answer the question the norms cannot: **what does
|
||||
this architecture actually cost, and what does it actually buy?**
|
||||
|
||||
The short version, and it is the reason this file is worth reading: the
|
||||
architecture works, the composition benefits are real and measurable, and the
|
||||
teardown half was never finished. Those three facts are not in tension. They are
|
||||
what a reference project looks like.
|
||||
|
||||
Markers: **[measured]** — read in source or config; **[derived]** — conclusion
|
||||
from measured facts; **[open]** — not answerable from source alone.
|
||||
|
||||
---
|
||||
|
||||
## 1. The shape, in one chain
|
||||
|
||||
```text
|
||||
key
|
||||
→ mapping context
|
||||
→ input action
|
||||
→ [payload: input tag] handler on the hero component
|
||||
→ ability system component buffers the spec handle
|
||||
~~~ end of the input-processing step ~~~
|
||||
→ controller post-process input
|
||||
→ activation policy decides
|
||||
→ activate
|
||||
```
|
||||
|
||||
Seven hops, five of them across assets **[measured]**. That count is the whole
|
||||
trade: it is why a feature plugin can add controls without touching the base
|
||||
module, and it is why a broken binding takes an afternoon to diagnose.
|
||||
|
||||
---
|
||||
|
||||
## 2. What the split actually buys, with numbers
|
||||
|
||||
| Benefit | Measured evidence |
|
||||
|---|---|
|
||||
| Code cost per new ability control | **zero lines** — one loop binds every row of the config **[measured]** |
|
||||
| Code cost per new native control | one explicit line; the project has exactly five and they have not changed **[measured]** |
|
||||
| Feature-plugin controls without base-module edits | one plugin ships 11 input actions and its own config, and the base module contains no reference to them **[measured]** |
|
||||
| One switch to suppress all ability input | a single tag on the ability system component clears the buffers and returns early **[measured]** |
|
||||
| Same action, different ability per character | the join is a string in two assets; no code path is per-character **[derived]** |
|
||||
|
||||
The fourth row is the one most often underestimated. Without a semantic layer,
|
||||
suppressing input during a menu means either a flag consulted in every handler or
|
||||
unbinding and rebinding at runtime. Here it is one tag, checked once per frame,
|
||||
in one place.
|
||||
|
||||
---
|
||||
|
||||
## 3. What it costs, also with numbers
|
||||
|
||||
**Compile-time errors become runtime silence.** The native path logs an error
|
||||
when a tag cannot be resolved at bind time. The ability path has **no error case
|
||||
at all** — the pressed tag is compared against granted specs, and no match is
|
||||
indistinguishable from not having the ability **[measured]**. This asymmetry is
|
||||
worth internalising as a triage rule: a silent native control is a config
|
||||
problem; a silent ability control could be anywhere along the chain in §1.
|
||||
|
||||
**The lookup is exact, and everything around it looks hierarchical.** Both lookup
|
||||
sites use exact tag matching **[measured]**, while the editor picker filters by
|
||||
category and the vocabulary is a tree. A spec tagged with a parent is never found
|
||||
by a press of its child. Nothing at either end of the join states this. Recipe:
|
||||
IN-01.
|
||||
|
||||
**Two declaration sources for one namespace.** Native tags are C++ symbols;
|
||||
ability tags are config strings; the two sets do not overlap **[measured]**. Both
|
||||
are defensible individually. Together they mean no single artefact lists the input
|
||||
vocabulary. Recipe: IN-02.
|
||||
|
||||
**Linear scans on every press.** The lookup iterates all activatable specs
|
||||
**[measured]**. At the scale of this project that is free; it is proportional to
|
||||
ability count times presses per frame, and neither number is bounded by the
|
||||
design.
|
||||
|
||||
---
|
||||
|
||||
## 4. Teardown: the half that was not finished
|
||||
|
||||
This is the cluster worth studying, because the four findings compound into one
|
||||
outcome and each looks minor alone.
|
||||
|
||||
| Finding **[measured]** | Alone | In combination |
|
||||
|---|---|---|
|
||||
| Bind handles declared as locals at two call sites and discarded | a missed `TArray` member | removal has no data to work with |
|
||||
| The removal function's body is a single TODO comment | an unimplemented method | the feature action's teardown calls into it and returns successfully |
|
||||
| A working `RemoveBinds` helper exists with **zero callers** | dead code | the mechanism was built and could never be reached |
|
||||
| The context action's tracking list is read, iterated and removed from — and never written | an empty loop | contexts come off only reactively, so behaviour depends on teardown order |
|
||||
|
||||
**[derived]** Deactivating a feature leaves its ability bindings on a live pawn.
|
||||
The practical damage is limited because pawns are recreated between matches and
|
||||
because an orphan binding sends a tag that matches no spec — but the leak is real,
|
||||
and a later re-grant of the same ability resurrects phantom input.
|
||||
|
||||
The instructive part is the third row. Someone wrote the removal helper. It is
|
||||
correct. It has never been called, because the handles it needs were thrown away
|
||||
one function earlier. **A correct utility that nothing can call is
|
||||
indistinguishable from a utility that does not exist** — and it is worse, because
|
||||
its presence answers "is removal supported?" with a yes.
|
||||
|
||||
The original authors had also left a comment on the broken tracking line noting
|
||||
that the container is never modified **[measured]**. That is a known defect that
|
||||
shipped, which is the most honest possible evidence for the general rule: *finding
|
||||
a defect and fixing a defect are different projects, and reference code shows you
|
||||
the first.*
|
||||
|
||||
Recipes: IN-03, IN-04, IN-05.
|
||||
|
||||
---
|
||||
|
||||
## 5. One flag, two operations — the sharpest single finding
|
||||
|
||||
In the base initialization path, the call that makes a mapping context **active**
|
||||
sits *inside* the branch that checks whether the context should be registered with
|
||||
the player's settings **[measured]**.
|
||||
|
||||
In the feature-plugin path for the same data structure, the two are separate: the
|
||||
settings registration is skipped by the flag, and the activation call runs
|
||||
unconditionally **[measured]**.
|
||||
|
||||
**[derived]** So the same checkbox means "expose this for rebinding" in one place
|
||||
and "expose this for rebinding **and also make it work at all**" in another. Clear
|
||||
the flag on a base mapping and the controls go silent, with no error and no
|
||||
obvious connection between cause and effect.
|
||||
|
||||
Two properties make this the best example in the file:
|
||||
|
||||
1. **It is verifiable in ten seconds** by reading the two files side by side, and
|
||||
invisible in any amount of reading of either file alone.
|
||||
2. **The correct version exists in the same codebase**, which removes any argument
|
||||
about intent. This is not a design decision applied inconsistently; one of the
|
||||
two is wrong.
|
||||
|
||||
Recipe: IN-06. The same initialization path carries a second nesting defect: the
|
||||
loop over default mapping contexts is nested inside a check for the pawn's input
|
||||
config, so a pawn with no config gets no contexts either and is completely mute
|
||||
**[measured]** — two unrelated features disabled by one condition. Recipe: IN-07.
|
||||
|
||||
---
|
||||
|
||||
## 6. Initialization order as a broadcast dependency
|
||||
|
||||
Initialization clears **every** mapping on the subsystem before adding its own
|
||||
**[measured]**. That includes contexts added by features that activated earlier.
|
||||
|
||||
Recovery works, and it works by broadcast: the component sets its readiness flag
|
||||
and then sends a semantic "bind inputs now" event, which both feature actions
|
||||
listen for alongside the generic extension event **[measured]**. Features re-add
|
||||
what was just removed.
|
||||
|
||||
**[derived]** This is correct and it is fragile in a specific way: correctness is a
|
||||
property of *event ordering* rather than of stored state. Nothing reconciles what
|
||||
was cleared against what was restored, and a feature that is active but not
|
||||
listening for the ready event loses its contexts permanently.
|
||||
|
||||
Worth copying from this: the readiness flag is set **before** the broadcast, not
|
||||
after **[measured]**. A handler that checks readiness during the broadcast sees
|
||||
true. That ordering is one line and it is the difference between the pattern
|
||||
working and not. Recipe: IN-08.
|
||||
|
||||
---
|
||||
|
||||
## 7. Sediment: things that exist and do nothing
|
||||
|
||||
A reference project accumulates these, and each one costs a reader time:
|
||||
|
||||
- two mapping-management functions whose bodies contain only a check and a comment,
|
||||
called from the initialization path, doing nothing **[measured]** — the remains of
|
||||
a superseded settings API;
|
||||
- a user-settings class and a key-profile class registered in config, both of which
|
||||
are empty overrides **[measured]** — extension points presented as features;
|
||||
- a legacy axis-configuration block sitting beside the modern input system, whose
|
||||
deadzone and sensitivity values are overridden by modifiers that read player
|
||||
settings **[measured]**. A reader tuning the legacy values changes nothing.
|
||||
Recipe: IN-09;
|
||||
- a settings-driven modifier that resolves a property **by name through
|
||||
reflection** **[measured]**, so renaming the underlying setting breaks it
|
||||
silently with no compiler involvement. Recipe: IN-10.
|
||||
|
||||
**[derived]** None of these is a bug. Collectively they are the reason "read the
|
||||
reference and copy what it does" is a bad adoption strategy: a meaningful share of
|
||||
what it does is nothing.
|
||||
|
||||
---
|
||||
|
||||
## 8. What the reference got right
|
||||
|
||||
1. **The payload trick.** Passing the semantic tag as a bound-delegate payload is
|
||||
what makes one function pair serve unlimited abilities **[measured]**. It is
|
||||
the single most portable idea in this file.
|
||||
2. **Buffer then activate in one batch.** Collecting handles and activating after
|
||||
all input is processed removes event-order dependence, and the authors left a
|
||||
comment explaining precisely why — held input would otherwise activate an
|
||||
ability and then deliver a press event to it **[measured]**.
|
||||
3. **Setting readiness before broadcasting it** (§6).
|
||||
4. **Separate actions for mouse look and stick look**, with different tags and
|
||||
different handlers, because the stick needs frame-time scaling and the mouse
|
||||
does not **[measured]**. This is semantic distinction done correctly — not
|
||||
device sniffing in gameplay code, but two genuinely different commands.
|
||||
5. **Device switching kept entirely out of the input layer.** The input module
|
||||
contains no reference to input-device types at all **[measured]**; device
|
||||
changes affect icons, haptics and UI style, and never gameplay bindings.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source and
|
||||
config in a single workspace: roughly 870 lines across the input module, plus the
|
||||
hero component, two feature actions and the input configuration file. Blueprint
|
||||
assets — the mapping contexts and input actions themselves — are binary and were
|
||||
not read, so every statement here about which keys map to what is **[open]**.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; each
|
||||
`IN-` identifier resolves back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
These are examples and failure evidence from one project on one engine version,
|
||||
not guarantees about others. The counts are properties of that project on that
|
||||
day and are quoted to show the shape of a contrast, not to be imported. Re-run the
|
||||
recipes against your own tree.
|
||||
@@ -0,0 +1,425 @@
|
||||
---
|
||||
name: ue-modular-gameplay
|
||||
description: >-
|
||||
Design or review modular gameplay in Unreal Engine: experience definitions,
|
||||
feature plugins, reusable action sets, feature actions, actor extension
|
||||
handlers, component init-state chains, primary asset bundles, and runtime
|
||||
activation and deactivation. Use when adding game modes, seasonal or
|
||||
downloadable features, modular components, abilities, input or UI, or when
|
||||
fixing initialization-order and feature teardown bugs.
|
||||
---
|
||||
|
||||
# UE modular gameplay
|
||||
|
||||
The invariant:
|
||||
|
||||
> A feature may select and attach behaviour without modifying the receiving base
|
||||
> class, but it owns every resource it adds and must remove exactly that
|
||||
> resource.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Eighteen detection recipes for silent modularity failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting this architecture. Nearly all of them are
|
||||
teardown defects, and teardown defects share one property: **activation is what
|
||||
gets tested, so the half that is broken is the half nobody exercises.**
|
||||
|
||||
Related skills: `ue-data-driven-architecture`, `ue-input-architecture`,
|
||||
`ue-ui-architecture`, `ue-gas-architecture`, `ue-multiplayer-authority`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Separate the four layers
|
||||
|
||||
```text
|
||||
Playlist / session entry
|
||||
└─ Experience definition
|
||||
├─ feature plugins to activate
|
||||
├─ reusable action sets
|
||||
├─ mode-specific actions
|
||||
└─ default pawn data
|
||||
|
||||
Feature data asset
|
||||
└─ lifecycle and asset scanning of one plugin
|
||||
|
||||
Action set
|
||||
└─ reusable composition slice
|
||||
|
||||
Feature action
|
||||
└─ one apply/rollback operation
|
||||
```
|
||||
|
||||
Do not merge them:
|
||||
|
||||
- the playlist owns map, session and menu metadata, not gameplay composition;
|
||||
- the experience owns one runtime composition, not plugin lifecycle policy;
|
||||
- the feature data asset owns plugin registration and scanning, not a game mode;
|
||||
- an action set expresses reuse; it is not an inherited base experience;
|
||||
- an action performs one focused mutation and owns its rollback state.
|
||||
|
||||
### Prefer composition over experience inheritance
|
||||
|
||||
A mode should read as a sum:
|
||||
|
||||
```text
|
||||
SharedInput + StandardComponents + StandardHUD + ModeDelta
|
||||
```
|
||||
|
||||
not as a hidden defaults chain:
|
||||
|
||||
```text
|
||||
Base -> Shooter -> Team -> TeamDeathmatch
|
||||
```
|
||||
|
||||
Flat action sets keep dependencies, asset loading and teardown auditable. A
|
||||
reference product measured for this skill goes further and *forbids* the second
|
||||
form in validation, pointing authors at composition instead.
|
||||
|
||||
---
|
||||
|
||||
## 2. Experience loading is a state machine
|
||||
|
||||
```text
|
||||
Unloaded
|
||||
→ LoadingAssets
|
||||
→ LoadingGameFeatures
|
||||
→ ExecutingActions
|
||||
→ Loaded
|
||||
→ Deactivating
|
||||
→ Unloaded
|
||||
```
|
||||
|
||||
Each transition needs an explicit completion criterion. Never track this with a
|
||||
handful of booleans — `bAssetsReady && bPluginsReady && bActionsDone` cannot
|
||||
express "failed", and failure is the state you will need most.
|
||||
|
||||
### The server selects; every peer realises locally
|
||||
|
||||
Replicate the selected definition, never the result of loading it:
|
||||
|
||||
```text
|
||||
server selects experience ID
|
||||
→ replicated definition reaches the client
|
||||
→ server and client run the same load pipeline
|
||||
→ role-specific bundles and actions decide their local work
|
||||
```
|
||||
|
||||
This is "replicate intent, realise presentation locally" applied to composition.
|
||||
|
||||
### Document selection precedence
|
||||
|
||||
If an experience can come from matchmaking, a URL, developer settings, the
|
||||
command line, world settings and a default, define the exact precedence, log the
|
||||
winning source, validate the asset ID before loading, and provide an explicit
|
||||
fallback.
|
||||
|
||||
### Delay only for a named dependency
|
||||
|
||||
Deferring selection by one frame so that startup settings can initialise is
|
||||
legitimate **because the dependency has a name**. A next-tick timer with no
|
||||
identified prerequisite is initialization-order debt with a timer wrapped around
|
||||
it. Recipe: MG-12.
|
||||
|
||||
---
|
||||
|
||||
## 3. Bundle and role boundaries
|
||||
|
||||
Collect the experience plus every action set as primary assets, then request
|
||||
named bundles: a shared bundle always, the client bundle when a client exists,
|
||||
the server bundle when authority exists.
|
||||
|
||||
Mark soft references at their declaration:
|
||||
|
||||
```cpp
|
||||
meta=(AssetBundles="Client")
|
||||
meta=(AssetBundles="Client,Server")
|
||||
```
|
||||
|
||||
| Action target | Expected role |
|
||||
|---|---|
|
||||
| HUD, layout, input icons | Client |
|
||||
| scoring authority, bots, team assignment | Server |
|
||||
| replicated scoring component | Client + Server |
|
||||
| ability classes and effects needed for prediction | usually both |
|
||||
|
||||
**A client/server flag on an action does not guarantee correct cook bundles.**
|
||||
Validate both, separately.
|
||||
|
||||
### Editor bundle loading is not shipping
|
||||
|
||||
A policy that loads client data when not a dedicated server and server data when
|
||||
not a client-only build loads **both** in the editor. Every memory and load-time
|
||||
figure taken in-editor therefore describes a configuration that ships to nobody.
|
||||
Use it for direction of change only. Recipe: MG-18.
|
||||
|
||||
---
|
||||
|
||||
## 4. Extend actors through receivers and extension handlers
|
||||
|
||||
Do not scan the world for actors and patch them. Use the component manager's
|
||||
extension protocol:
|
||||
|
||||
```text
|
||||
receiver actor registers itself
|
||||
action registers an extension handler for a target class
|
||||
manager invokes the handler for existing AND future receivers
|
||||
returned request handle owns the subscription
|
||||
```
|
||||
|
||||
The receiver lifecycle is symmetric and belongs in the base class:
|
||||
|
||||
```text
|
||||
PreInitializeComponents : add receiver
|
||||
BeginPlay : send ready events
|
||||
EndPlay : remove receiver
|
||||
```
|
||||
|
||||
Note what this buys: base gameplay classes register as receivers and **know
|
||||
nothing about features**. If a class does not inherit the modular base, those
|
||||
three calls are written by hand — and forgetting the third is a leak with no
|
||||
symptom until the second activation.
|
||||
|
||||
### Use semantic readiness events
|
||||
|
||||
A generic "extension added" event means only that the actor is known to the
|
||||
manager. It may not yet have its pawn data, ability system, controller or input.
|
||||
Emit named events at the real dependency boundary — abilities ready, bind inputs
|
||||
now, actor ready — and have handlers accept both the generic event and the
|
||||
semantic one, so activation order stops mattering. Recipe: MG-13.
|
||||
|
||||
### Retain request handles
|
||||
|
||||
Handler registration and component requests return ownership handles that are
|
||||
reference counted. **Dropping the handle is the unsubscribe operation**; there is
|
||||
no explicit removal call. Storing them per activation context is therefore not
|
||||
bookkeeping, it is the only teardown mechanism that exists.
|
||||
|
||||
---
|
||||
|
||||
## 5. Use an init-state chain
|
||||
|
||||
Networked actor composition has non-deterministic arrival order. Replace
|
||||
scattered checks in `BeginPlay`, possession callbacks and replication callbacks
|
||||
with a monotonic chain:
|
||||
|
||||
```text
|
||||
Spawned → DataAvailable → DataInitialized → GameplayReady
|
||||
```
|
||||
|
||||
- **Spawned** — the actor and component exist.
|
||||
- **DataAvailable** — required replicated and config data exist.
|
||||
- **DataInitialized** — every participating feature reached DataAvailable and
|
||||
cross-feature initialisation ran.
|
||||
- **GameplayReady** — safe for active gameplay.
|
||||
|
||||
The load-bearing barrier is "have all features on this actor reached
|
||||
DataAvailable", which synchronises components without any of them naming the
|
||||
others.
|
||||
|
||||
### Rules
|
||||
|
||||
- states only advance; they reset with the actor's lifecycle, not by backtracking;
|
||||
- every transition has an observable prerequisite;
|
||||
- replication callbacks re-enter the same check rather than duplicating init;
|
||||
- feature-specific work runs once, at a named transition;
|
||||
- test server, owning client and simulated proxy separately — a proxy never
|
||||
acquires a controller, and a gate that waits for one deadlocks it while working
|
||||
perfectly in a single-process test.
|
||||
|
||||
---
|
||||
|
||||
## 6. Action design contract
|
||||
|
||||
An action class implements behaviour; its instance stores targets and
|
||||
configuration. A world-aware base action should filter to game worlds, handle
|
||||
worlds that already exist at activation, subscribe for future ones, and release
|
||||
its world delegates on deactivation — leaving resource tracking to the concrete
|
||||
action.
|
||||
|
||||
### State must be keyed by activation context
|
||||
|
||||
One action object can be active in several contexts at once — client and server
|
||||
in a single-process test are the common case. **State in plain member fields is a
|
||||
cross-context leak.** Key it:
|
||||
|
||||
```cpp
|
||||
struct FPerContextData { /* ... */ };
|
||||
TMap<FGameFeatureStateChangeContext, FPerContextData> ContextData;
|
||||
```
|
||||
|
||||
Recipe: MG-11.
|
||||
|
||||
### One responsibility per action
|
||||
|
||||
Good: add components; add abilities; add input binding; add input mapping
|
||||
context; add widgets; add cue path.
|
||||
|
||||
Bad: "set up shooter mode".
|
||||
|
||||
A focused action has a testable rollback and can be reused in an action set.
|
||||
|
||||
### The mandatory ownership ledger
|
||||
|
||||
Write this table **before** implementing any action:
|
||||
|
||||
| Added resource | Stored ownership | Removal |
|
||||
|---|---|---|
|
||||
| extension callback | request handle | release the handle |
|
||||
| component | component request | remove the owned component |
|
||||
| ability, effect, attribute set | granted handles | revoke the handles |
|
||||
| input bind | bind handles per pawn | remove those exact binds |
|
||||
| input mapping context | player + context | remove from the same subsystem |
|
||||
| widget extension | extension handle | unregister |
|
||||
| layout | widget or layout handle | deactivate and remove |
|
||||
| global delegate | delegate handle | unbind |
|
||||
|
||||
**If the middle column is blank, reject the action.** Every teardown failure in
|
||||
the measured reference is a blank middle column that shipped. Recipes: MG-01
|
||||
through MG-05.
|
||||
|
||||
---
|
||||
|
||||
## 7. Deactivation is not optional
|
||||
|
||||
"Modular" is not proven by successful activation. The round trip is the test:
|
||||
|
||||
```text
|
||||
inactive baseline
|
||||
→ activate → verify additions exactly once
|
||||
→ deactivate → verify exact baseline
|
||||
→ reactivate → verify no duplicates and no stale state
|
||||
```
|
||||
|
||||
Run it with: actors that existed before activation; actors spawned after it; an
|
||||
actor destroyed before the feature deactivates; two worlds in one process; client
|
||||
and server roles; a partial load failure; and a plugin required by two
|
||||
experiences at once.
|
||||
|
||||
### Removal must be idempotent
|
||||
|
||||
Resources come off by two independent paths: the explicit reset at deactivation,
|
||||
and reactively when the actor dies first and the manager reports it. Both must be
|
||||
safe to run, in either order, and must operate on "is there a record?" rather than
|
||||
assuming one. Recipe: MG-08.
|
||||
|
||||
### Reference-count shared plugin activation
|
||||
|
||||
If two experiences need one plugin, count activations. One owner deactivating
|
||||
must not unload a plugin the other still requires. **Check that the count is not
|
||||
editor-only** — a reference count compiled out of shipping builds is a reference
|
||||
count that protects development and nothing else. Recipe: MG-16.
|
||||
|
||||
### Diff requirements on experience change
|
||||
|
||||
```text
|
||||
old - new = deactivate and unload
|
||||
new - old = load and activate
|
||||
intersection = keep active
|
||||
```
|
||||
|
||||
Do not unload everything and reload it immediately. If this diff is not
|
||||
implemented, define experience changes as map-travel-only and say so in writing —
|
||||
that is a legitimate scope decision, and leaving it undefined is not.
|
||||
|
||||
---
|
||||
|
||||
## 8. Failure handling
|
||||
|
||||
A production loader must not assert its way through configuration failure.
|
||||
Represent failure as a state:
|
||||
|
||||
```text
|
||||
Loaded
|
||||
Failed(reason, failing asset / plugin / action)
|
||||
```
|
||||
|
||||
Handle: an invalid asset ID; a missing plugin URL; a cancelled load; plugin
|
||||
activation failure; action validation failure; timeout or partial completion; and
|
||||
a deactivation pauser that never completes.
|
||||
|
||||
The loading screen consumes loader state and reason. **A permanent loading screen
|
||||
is not error handling** — it is the absence of it, rendered.
|
||||
|
||||
---
|
||||
|
||||
## 9. Validation checklist
|
||||
|
||||
### Experience
|
||||
|
||||
- [ ] Asset ID registered and resolves.
|
||||
- [ ] Default pawn data valid, or explicitly no pawn.
|
||||
- [ ] Plugin names resolve to URLs.
|
||||
- [ ] Action sets non-null and duplicate-free.
|
||||
- [ ] Reusable behaviour in action sets; only the mode delta inline.
|
||||
- [ ] Client and server bundles cover every soft reference.
|
||||
|
||||
### Action
|
||||
|
||||
- [ ] Target actor class valid; addition list non-empty.
|
||||
- [ ] Existing and future actors both handled.
|
||||
- [ ] Semantic readiness event used where the generic one is too early.
|
||||
- [ ] Every request, delegate, grant, widget and input handle retained.
|
||||
- [ ] Deactivation reverses activation exactly.
|
||||
- [ ] Repeated activate/deactivate is idempotent.
|
||||
- [ ] Per-context state cannot leak into another world.
|
||||
- [ ] Index or key passed to a deferred callback matches the entry it names.
|
||||
|
||||
### Init states
|
||||
|
||||
- [ ] Every prerequisite is explicit.
|
||||
- [ ] Every participating feature can reach every state.
|
||||
- [ ] No feature waits on a signal produced only by a path it disabled.
|
||||
- [ ] Authority and local-control branches tested.
|
||||
- [ ] Late replication re-enters the same check.
|
||||
|
||||
---
|
||||
|
||||
## 10. Review questions
|
||||
|
||||
Before approving a modular feature:
|
||||
|
||||
1. Which definition selects it?
|
||||
2. Which plugin owns its code and content?
|
||||
3. Which action applies each effect?
|
||||
4. Which readiness state makes that effect safe?
|
||||
5. Which handle proves ownership?
|
||||
6. What is the exact inverse operation?
|
||||
7. What is loaded on client, on server, in the editor?
|
||||
8. Can another feature share the same plugin or resource?
|
||||
9. What happens on partial failure?
|
||||
10. Does activate → deactivate → activate return to the same observable state?
|
||||
|
||||
**If questions 4 to 6 have no precise answer, the feature is dynamically attached
|
||||
but not modular.** That distinction is the entire subject of this skill: attaching
|
||||
is easy and demonstrable, detaching is neither, and only one of them gets
|
||||
demonstrated.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
the feature-action layer of Epic's Lyra Starter Game on Unreal Engine 5.6, read
|
||||
rather than run. The engine-side component manager and feature subsystem are not
|
||||
part of that project, so every statement about them is inferred from call sites
|
||||
in the audited code and is marked as such rather than presented as measured.
|
||||
|
||||
That audit found the architecture sound and its teardown half incomplete in five
|
||||
separate actions — which is why this skill treats the ownership ledger and the
|
||||
round-trip test as mandatory rather than advisory. Several of the gaps are
|
||||
recorded by the original authors in their own comments, which is the strongest
|
||||
confirmation a source audit can produce.
|
||||
|
||||
Source addresses stay in the research archive; what ships is the detection
|
||||
recipe. Each entry carries a stable identifier (`MG-01`, `MG-02`, …) that
|
||||
resolves back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. Binary assets were not read, so
|
||||
which actions the shipped features actually configure is unknown and is recorded
|
||||
as open rather than guessed. Re-run the recipes against your own tree before
|
||||
trusting any specific claim.
|
||||
@@ -0,0 +1,701 @@
|
||||
# Failure modes: modular gameplay
|
||||
|
||||
Nineteen ways a feature attaches correctly and detaches incompletely.
|
||||
|
||||
The shared property is stated once, because it explains the whole list:
|
||||
**activation is what gets demonstrated.** A feature is switched on, someone
|
||||
checks the ability appeared, the widget rendered, the input worked, and the
|
||||
review is done. Deactivation is exercised by nobody during development, because
|
||||
in development you restart the editor instead. So the broken half is
|
||||
systematically the half nobody looks at, and its symptoms — a duplicate on the
|
||||
second activation, a bind that survives its feature — appear later and are
|
||||
attributed elsewhere.
|
||||
|
||||
A second property applies to most entries: the reference implementations of these
|
||||
actions **look symmetric**. There is an `Add` and there is a `Remove`, a reset
|
||||
loop and a reactive removal path. The defect is almost never a missing function;
|
||||
it is a function that removes something other than what was added, or that reads
|
||||
a container nobody fills.
|
||||
|
||||
Recipes use `rg` and are written to be run from a project's source root. Most are
|
||||
a pair: one search finds the addition, the other looks for the matching removal.
|
||||
|
||||
---
|
||||
|
||||
## Teardown
|
||||
|
||||
### MG-01 - Rollback function with an empty body
|
||||
|
||||
**Mechanism.** An action's removal path calls a method on a game component, and
|
||||
that method's body is a comment saying it is not implemented yet.
|
||||
|
||||
**Why it is silent.** Every caller succeeds. The call compiles, executes and
|
||||
returns; nothing distinguishes "removed the bindings" from "did nothing" at the
|
||||
call site, because neither returns a result.
|
||||
|
||||
**Why the obvious check misses it.** The action's teardown *is* implemented and
|
||||
reads correctly — it iterates its tracked pawns and calls a removal method on
|
||||
each. Reviewing the action finds a complete, symmetric implementation. The empty
|
||||
body is one call away, in a different file owned by a different subsystem, and it
|
||||
is the game component's problem rather than the feature's.
|
||||
|
||||
**Symptom.** Ability input bindings survive the deactivation of the feature that
|
||||
added them. The second activation adds a second set. Nothing errors.
|
||||
|
||||
**Detect.** Find removal functions whose body is only a comment:
|
||||
|
||||
```bash
|
||||
rg -n -A4 "void \w+::Remove\w+\(" --glob "*.cpp" . \
|
||||
| rg -B2 "^\s*//\s*@?TODO|^\s*\}"
|
||||
```
|
||||
|
||||
Read each hit: a `Remove` whose body contains no statements is the finding. In
|
||||
the measured reference the removal of an additional input config was exactly this
|
||||
— a three-line function containing one TODO comment.
|
||||
|
||||
**Guardrail.** A removal function with no body must not compile silently. Make it
|
||||
pure virtual, assert inside it, or delete the caller. An unimplemented rollback
|
||||
that is *called* is worse than one that does not exist, because the call site
|
||||
proves the intent was there.
|
||||
|
||||
---
|
||||
|
||||
### MG-02 - Bind handles discarded at the moment they are created
|
||||
|
||||
**Mechanism.** A binding call takes an out-parameter array of handles. The caller
|
||||
declares that array as a local variable and lets it go out of scope.
|
||||
|
||||
**Why it is silent.** The bindings work. Handles are only needed to undo, and
|
||||
nothing undoes during a normal session.
|
||||
|
||||
**Why the obvious check misses it.** The out-parameter is passed correctly and
|
||||
the call is idiomatic — some codebases even name the argument in a comment at the
|
||||
call site, which makes it look deliberate. The defect is the *storage duration*
|
||||
of a local variable, which no search for "is the API used correctly?" examines.
|
||||
This is also the root cause of MG-01: the removal function cannot be implemented,
|
||||
because the information it would need was thrown away one function earlier.
|
||||
|
||||
**Symptom.** Rollback is not merely unimplemented but *impossible* without
|
||||
changing the component that binds. An estimate of "implement the missing removal"
|
||||
comes out an order of magnitude low.
|
||||
|
||||
**Detect.** Find handle arrays declared as locals at a binding call:
|
||||
|
||||
```bash
|
||||
rg -n -B2 "BindAbilityActions|BindAction\(" --glob "*.cpp" . \
|
||||
| rg "TArray<uint32>\s+\w+;|TArray<FInputBindingHandle>\s+\w+;"
|
||||
```
|
||||
|
||||
Any handle container declared in the same scope as the call, rather than as a
|
||||
member, is discarded. In the measured reference this appeared at two separate
|
||||
call sites in one component.
|
||||
|
||||
**Guardrail.** Handles returned by a binding API are stored by the object whose
|
||||
lifetime governs the binding, in the same change that creates the binding. If
|
||||
there is nowhere to store them, the feature is not removable and that must be
|
||||
stated before it ships.
|
||||
|
||||
---
|
||||
|
||||
### MG-03 - Tracking container that is read but never written
|
||||
|
||||
**Mechanism.** An action declares a container of the actors it has affected,
|
||||
removes from it in the reactive path, iterates it during reset — and never adds
|
||||
to it. The function that would add takes the per-context data as a parameter and
|
||||
ignores it.
|
||||
|
||||
**Why it is silent.** The reset loop runs, finds the container empty, and
|
||||
completes successfully. An empty loop is a legal outcome; it looks exactly like a
|
||||
feature that affected nothing because no receivers existed.
|
||||
|
||||
**Why the obvious check misses it.** Every symptom of correctness is present: the
|
||||
container is declared, an `ensure` at activation asserts it starts empty, the
|
||||
reset loop pops from it correctly, and the reactive path removes from it. Four of
|
||||
the five operations exist. Only the write is missing, and "is this container
|
||||
used?" answers yes at four sites.
|
||||
|
||||
**Symptom.** Input mapping contexts are never removed at deactivation. They come
|
||||
off only if the receiving actor happens to die first and the reactive path fires.
|
||||
The result depends on teardown order, which makes it intermittent.
|
||||
|
||||
**Detect.** For every tracking container, compare reads against writes:
|
||||
|
||||
```bash
|
||||
C='ControllersAddedTo'
|
||||
rg -n "\b$C\b" --glob "*.cpp" --glob "*.h" . # all uses
|
||||
rg -n "\b$C\b\s*\.\s*(Add|AddUnique|Emplace|Push)" --glob "*.cpp" . # writes only
|
||||
```
|
||||
|
||||
Uses without writes is the finding. The measured reference is unusually clear
|
||||
here: a sibling action in the same directory tracks its pawns with
|
||||
`AddUnique` in exactly the place the broken one omits it, so the two files can be
|
||||
diffed against each other. The original authors had also left a comment on the
|
||||
broken line noting that the container is never modified — a known defect that
|
||||
shipped.
|
||||
|
||||
**Guardrail.** A container whose only writes are removals is a container with no
|
||||
producer. Assert non-empty after the add path runs, or derive the collection from
|
||||
the resource itself rather than maintaining a parallel list.
|
||||
|
||||
---
|
||||
|
||||
### MG-04 - Reset path narrower than the reactive removal path
|
||||
|
||||
**Mechanism.** An action has two removal routes: a reset invoked at deactivation,
|
||||
and a per-actor removal invoked when the manager reports a receiver going away.
|
||||
The per-actor route cleans up more than the reset does.
|
||||
|
||||
**Why it is silent.** The reset completes and clears its own bookkeeping, so the
|
||||
action's internal state is consistent afterwards. What is left behind lives in
|
||||
another subsystem — a pushed layout widget still on its layer — and that
|
||||
subsystem was never told.
|
||||
|
||||
**Why the obvious check misses it.** Both functions exist and both are correct
|
||||
for what they enumerate. Comparing them requires reading two loops in different
|
||||
parts of a file and noticing that one iterates two collections and the other
|
||||
iterates one. Nothing names the missing collection.
|
||||
|
||||
**Symptom.** A layout widget pushed by a feature stays on screen after the
|
||||
feature is deactivated. On reactivation there are two.
|
||||
|
||||
**Detect.** Enumerate what each removal path touches and diff the sets:
|
||||
|
||||
```bash
|
||||
rg -n -A20 "::Reset\(" --glob "*.cpp" . | rg "Empty\(|Unregister|Deactivate|Remove"
|
||||
rg -n -A20 "::Remove\w+\(" --glob "*.cpp" . | rg "Empty\(|Unregister|Deactivate|Remove"
|
||||
```
|
||||
|
||||
A collection handled in the second output and absent from the first is the
|
||||
finding.
|
||||
|
||||
**Guardrail.** One removal implementation, called by both routes. If they must
|
||||
differ, the reset calls the per-actor removal in a loop rather than reimplementing
|
||||
a subset of it.
|
||||
|
||||
---
|
||||
|
||||
### MG-05 - Two removal semantics for one logical resource
|
||||
|
||||
**Mechanism.** The same kind of resource is revoked differently depending on which
|
||||
field granted it: one path clears immediately, the other defers removal until the
|
||||
resource finishes naturally.
|
||||
|
||||
**Why it is silent.** Both are legitimate APIs with legitimate uses. Deferred
|
||||
removal is correct when an ability is mid-activation; immediate removal is
|
||||
correct when tearing down. Neither errors.
|
||||
|
||||
**Why the obvious check misses it.** Each call site is individually defensible,
|
||||
and they are in different functions. The inconsistency is only visible if you ask
|
||||
"what happens to a granted ability at deactivation?" and discover the answer is
|
||||
"it depends which field you used", which is not a question the code invites.
|
||||
|
||||
**Symptom.** Deactivation behaviour differs between two configurations that a
|
||||
designer considers equivalent. Timing-dependent, so it reproduces inconsistently.
|
||||
|
||||
**Detect.** List every revocation call for one resource kind and compare:
|
||||
|
||||
```bash
|
||||
rg -n "ClearAbility|SetRemoveAbilityOnEnd|RemoveActiveGameplayEffect|RemoveSpawnedAttribute" \
|
||||
--glob "*.cpp" .
|
||||
```
|
||||
|
||||
Two different revocation calls for the same resource kind in one teardown path is
|
||||
the finding.
|
||||
|
||||
**Guardrail.** Pick one semantic per resource kind and document it at the grant
|
||||
site. If both are genuinely needed, the choice belongs in the data, named, not in
|
||||
which code path happened to grant it.
|
||||
|
||||
---
|
||||
|
||||
### MG-06 - Removal that is not idempotent
|
||||
|
||||
**Mechanism.** Resources come off by two independent routes — explicit reset at
|
||||
deactivation, and reactive removal when the actor dies first. If removal assumes a
|
||||
record exists, running it twice or in the wrong order misbehaves.
|
||||
|
||||
**Why it is silent.** In the common ordering — feature deactivates while actors
|
||||
are alive — only one route runs. The double-run needs an actor destroyed during
|
||||
teardown, which is a shutdown-order accident.
|
||||
|
||||
**Why the obvious check misses it.** Each route is correct in isolation. Nothing
|
||||
in either function says "this may also be reached from the other path"; the
|
||||
coupling is a property of the manager's event dispatch, which is engine code.
|
||||
|
||||
**Symptom.** An assert or a stale entry during shutdown, in an ordering nobody can
|
||||
reproduce on demand.
|
||||
|
||||
**Detect.** Check that removal is guarded by record existence, and that reset
|
||||
loops tolerate mutation:
|
||||
|
||||
```bash
|
||||
rg -n -A8 "::Remove\w+\(" --glob "*.cpp" . | rg "Find\(|Contains\(|if\s*\("
|
||||
rg -n -A6 "while\s*\(!\w+\.IsEmpty\(\)\)" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
The second pattern is the correct shape: a reset written as "while not empty,
|
||||
take the top and remove it" survives a removal that mutates the same collection.
|
||||
A plain range-for over a collection that removal mutates does not.
|
||||
|
||||
**Guardrail.** Removal operates on the presence of a record, never on the
|
||||
assumption of one. Write reset loops as drain loops.
|
||||
|
||||
---
|
||||
|
||||
## Correctness of what gets attached
|
||||
|
||||
### MG-07 - Index incremented inside the guard it is guarding
|
||||
|
||||
**Mechanism.** A loop builds one deferred callback per configuration entry,
|
||||
passing the entry's index as payload. The increment sits inside the branch that
|
||||
skips invalid entries, so after any skip the payload index no longer matches the
|
||||
list.
|
||||
|
||||
**Why it is silent.** Every index remains in range and resolves to a real entry.
|
||||
The callback applies a valid, well-formed configuration — just the wrong one.
|
||||
There is no out-of-bounds access and nothing to assert on.
|
||||
|
||||
**Why the obvious check misses it.** Reading the loop shows an increment that
|
||||
appears to advance per iteration; it takes reading the brace structure to see it
|
||||
is inside the guard. Editor-time validation flags the invalid entry that triggers
|
||||
it, so the condition is believed impossible — but that validation is an editor
|
||||
check, not a runtime guarantee, and it does not run on cooked data.
|
||||
|
||||
**Symptom.** Actors receive another entry's abilities. Because both sets are
|
||||
valid, this presents as a design or data error and gets investigated in the
|
||||
content.
|
||||
|
||||
**Detect.** Find counters incremented inside a conditional within a loop:
|
||||
|
||||
```bash
|
||||
rg -n -B8 "\b\w*Index\+\+|\+\+\w*Index" --glob "*.cpp" . | rg "if\s*\(|continue"
|
||||
```
|
||||
|
||||
Read each hit's brace structure. In the measured reference the increment sat
|
||||
inside a null-class guard, six lines below the `if` that opened it.
|
||||
|
||||
**Guardrail.** Derive the index from the loop itself rather than maintaining a
|
||||
parallel counter. If a counter is needed, increment it in the loop header where
|
||||
its relationship to the iteration is structural.
|
||||
|
||||
---
|
||||
|
||||
### MG-08 - Soft reference resolved with a bare accessor and no fallback
|
||||
|
||||
**Mechanism.** An action resolves a soft class or asset pointer with a
|
||||
non-loading accessor. If the bundle did not load it, the accessor returns null
|
||||
and the entry is skipped.
|
||||
|
||||
**Why it is silent.** The skip is unconditional and unlogged. A feature that
|
||||
grants nothing looks identical to a feature whose grant list is empty, and both
|
||||
are legal.
|
||||
|
||||
**Why the obvious check misses it.** Every call site is correct given the
|
||||
precondition "bundles are loaded", which is true in the editor, where everything
|
||||
is loaded. The precondition is established in a completely different subsystem —
|
||||
asset bundle configuration — and the code that depends on it does not state that
|
||||
it does.
|
||||
|
||||
**Symptom.** A packaged build silently omits abilities, widgets or mappings that
|
||||
work in the editor. Bisecting finds nothing, because no code changed.
|
||||
|
||||
**Detect.** Compare loading resolutions against non-loading ones, and check for a
|
||||
fallback:
|
||||
|
||||
```bash
|
||||
rg -c "LoadSynchronous" --glob "*.cpp" .
|
||||
rg -n "\.Get\(\)" --glob "*.cpp" . -A2 | rg -v "ensure|check|if\s*\(|UE_LOG"
|
||||
```
|
||||
|
||||
In the measured reference the feature-action directory used three synchronous
|
||||
loads against ten bare accessors, none of the latter guarded or logged. Both
|
||||
numbers matter: the loads are a runtime hitch, the bare accessors are a silent
|
||||
omission, and they appear in the same files.
|
||||
|
||||
**Guardrail.** Either load, or assert, or log — never skip silently. A soft
|
||||
reference that is expected to be bundle-loaded should assert that it was.
|
||||
|
||||
---
|
||||
|
||||
### MG-09 - Unreachable code after a return
|
||||
|
||||
**Mechanism.** A validation function returns its result and is followed by a
|
||||
second, unreachable return.
|
||||
|
||||
**Why it is silent.** It is genuinely harmless. The compiler may not even warn.
|
||||
|
||||
**Why the obvious check misses it.** It is invisible to behaviour-driven review
|
||||
because it has no behaviour. It is included here for one reason: **it is evidence
|
||||
about the file rather than a defect in it.** Dead code after a return means
|
||||
someone edited this validation and did not read what followed. Treat it as a
|
||||
marker for where attention lapsed, and read the surrounding function more
|
||||
carefully than you otherwise would.
|
||||
|
||||
**Symptom.** None. Its value is diagnostic.
|
||||
|
||||
**Detect.**
|
||||
|
||||
```bash
|
||||
rg -n -A3 "^\s*return \w+;\s*$" --glob "*.cpp" . | rg -A1 "^\s*$" | rg "return"
|
||||
```
|
||||
|
||||
**Guardrail.** Enable and honour unreachable-code warnings. Where one appears,
|
||||
re-read the function rather than only deleting the line.
|
||||
|
||||
---
|
||||
|
||||
## Context and scope
|
||||
|
||||
### MG-10 - Action state in plain members rather than per-context
|
||||
|
||||
**Mechanism.** One action object can be active in several state-change contexts
|
||||
simultaneously — client and server in a single process is the ordinary case.
|
||||
State kept in plain member fields is shared between them.
|
||||
|
||||
**Why it is silent.** With one context, which is every standalone run, the
|
||||
behaviour is correct. The corruption requires two contexts, and the second
|
||||
context's writes look like legitimate updates.
|
||||
|
||||
**Why the obvious check misses it.** Member fields on an action object are the
|
||||
obvious place to keep state, and nothing in the base class signature suggests
|
||||
otherwise. The requirement is transmitted only by convention — a context
|
||||
parameter that every lifecycle method receives and that a single-context
|
||||
implementation can safely ignore.
|
||||
|
||||
**Symptom.** One process's client and server tread on each other's tracking data.
|
||||
Teardown removes the wrong context's resources, or removes them twice.
|
||||
|
||||
**Detect.** Find action state that is not keyed by context:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "class \w*GameFeatureAction\w*" --glob "*.h" . \
|
||||
| rg "^\s*(TArray|TMap|TSet|TSharedPtr|bool|int32)\s+\w+" \
|
||||
| rg -v "FGameFeatureStateChangeContext"
|
||||
```
|
||||
|
||||
Any mutable member that is not inside a per-context map is a candidate. The
|
||||
correct shape — a per-context struct in a map keyed by the change context —
|
||||
appears in four separate actions in the measured reference and is worth copying
|
||||
verbatim.
|
||||
|
||||
**Guardrail.** All action state lives in a per-context map. Assert at activation
|
||||
that the context's slot is empty, and reset it if not — that assertion is what
|
||||
turns a silent leak into a visible one.
|
||||
|
||||
---
|
||||
|
||||
### MG-11 - Activation reference count compiled out of shipping
|
||||
|
||||
**Mechanism.** A count that prevents one owner from deactivating a plugin another
|
||||
still needs is wrapped in an editor-only guard, and additionally checked against
|
||||
an editor runtime flag inside.
|
||||
|
||||
**Why it is silent.** In the editor, where multiple worlds share a process, the
|
||||
count is present and correct. In a shipping build there is one world per process,
|
||||
so the case it protects against does not arise — until it does.
|
||||
|
||||
**Why the obvious check misses it.** The mechanism exists, is correct, and is
|
||||
tested every time anyone uses multi-world play-in-editor. Discovering that it does
|
||||
not ship means reading the preprocessor guard around it, which is at the top of a
|
||||
file nobody opens while reasoning about deactivation.
|
||||
|
||||
**Symptom.** No symptom in a single-world build. A latent hazard the moment any
|
||||
build hosts two experiences, which is exactly when a project starts caring about
|
||||
modularity.
|
||||
|
||||
**Detect.**
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A10 "RequestToDeactivate|NotifyOfPluginActivation|ActivationCount" \
|
||||
--glob "*.cpp" . | rg "#if|GIsEditor|WITH_EDITOR"
|
||||
```
|
||||
|
||||
A reference count inside an editor guard, in code that ships, is the finding.
|
||||
|
||||
**Guardrail.** Reference counting is a correctness mechanism, not a development
|
||||
convenience. If it only exists in the editor, say so where it is declared and
|
||||
state what protects shipping instead.
|
||||
|
||||
---
|
||||
|
||||
### MG-12 - Plugins leak enabled between experiences
|
||||
|
||||
**Mechanism.** Experience teardown deactivates the plugins it listed, but a switch
|
||||
between experiences is not implemented as a diff — so plugins required by the old
|
||||
experience and not the new one may stay enabled.
|
||||
|
||||
**Why it is silent.** An enabled plugin that nothing uses has no symptom. Its
|
||||
content stays resident and its actions have already been deactivated, so the
|
||||
observable state is correct.
|
||||
|
||||
**Why the obvious check misses it.** Deactivation exists and works for the case it
|
||||
was written for: ending play. Experience-to-experience switching is a different
|
||||
lifecycle that reuses the same teardown, and nothing in that teardown announces
|
||||
its scope.
|
||||
|
||||
**Symptom.** Memory that does not return to baseline across mode changes, and a
|
||||
second experience that inherits capabilities from the first.
|
||||
|
||||
**Detect.** Look for the diff, and for the authors' own admissions:
|
||||
|
||||
```bash
|
||||
rg -n -i "//\s*@?TODO.*(leak|deactivat|unload|experience)" --glob "*.cpp" .
|
||||
rg -n "GameFeaturePluginURLs" --glob "*.cpp" . -B4 -A8 | rg "Difference|Intersect|Diff"
|
||||
```
|
||||
|
||||
The first search matters more than the second. In the measured reference the
|
||||
experience manager carries an eight-item block of the authors' own TODOs, one of
|
||||
which states plainly that plugins are leaked enabled and that diffing requirements
|
||||
is the intended fix. **A defect the authors documented is the cheapest one you
|
||||
will ever find** — read those blocks first in any codebase you are adopting.
|
||||
|
||||
**Guardrail.** Implement the requirement diff, or define experience changes as
|
||||
map-travel-only and write that decision down. An undefined middle is what leaks.
|
||||
|
||||
---
|
||||
|
||||
### MG-13 - Editor loads both role bundles
|
||||
|
||||
**Mechanism.** The loading policy loads client data when not a dedicated server
|
||||
and server data when not a client-only build. In the editor neither condition
|
||||
excludes, so both load.
|
||||
|
||||
**Why it is silent.** It is intentional and it is correct — the editor must be
|
||||
able to run either role. Nothing is wrong; the number it produces simply describes
|
||||
no shipping configuration.
|
||||
|
||||
**Why the obvious check misses it.** Profiling produces a figure, and figures from
|
||||
a profiler carry an authority their provenance does not. The condition that
|
||||
inflates it is two boolean expressions in a policy class, in a different
|
||||
subsystem from the one being measured.
|
||||
|
||||
**Symptom.** Memory and load-time budgets built on editor measurements, wrong in a
|
||||
direction that feels safe until a platform limit is real.
|
||||
|
||||
**Detect.**
|
||||
|
||||
```bash
|
||||
rg -n -A6 "GetGameFeatureLoadingMode|bLoadClientData|bLoadServerData" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
If neither flag excludes in the editor, both bundle sets load. In the measured
|
||||
reference the authors noted the resulting editor hitching in a comment beside it.
|
||||
|
||||
**Guardrail.** Editor asset measurements are directional only. Any absolute budget
|
||||
requires a packaged build of the target role, and any budget quoted without one
|
||||
carries that caveat in writing.
|
||||
|
||||
---
|
||||
|
||||
## Timing and readiness
|
||||
|
||||
### MG-14 - A delay with no named dependency
|
||||
|
||||
**Mechanism.** Selection or initialisation is deferred by a tick to let something
|
||||
else finish first.
|
||||
|
||||
**Why it is silent.** It works. One frame is enough today, on this machine, for
|
||||
this content set.
|
||||
|
||||
**Why the obvious check misses it.** A next-tick deferral is a common, accepted
|
||||
idiom, and the code around it is correct. Whether it is legitimate depends
|
||||
entirely on whether the thing being waited for is *named* — and an unnamed wait
|
||||
looks identical to a named one.
|
||||
|
||||
**Symptom.** Initialisation order breaks on a slower machine, a larger content
|
||||
set, or after an unrelated subsystem changes its own timing. The failure appears
|
||||
far from the deferral.
|
||||
|
||||
**Detect.** Find deferrals and check for a stated prerequisite:
|
||||
|
||||
```bash
|
||||
rg -n -B6 "SetTimerForNextTick|GetTimerManager\(\)\.SetTimer" --glob "*.cpp" . \
|
||||
| rg -i "//|comment|wait|until|so that"
|
||||
```
|
||||
|
||||
A deferral whose surrounding comment names what it waits for is debt with a
|
||||
receipt. One with no comment is debt with a timer wrapped around it. In the
|
||||
measured reference the one-frame delay before experience selection *is*
|
||||
documented — it waits for startup settings — which is the correct form.
|
||||
|
||||
**Guardrail.** Every delay names its prerequisite in a comment at the call site,
|
||||
and is replaced by an explicit readiness signal when one becomes available.
|
||||
|
||||
---
|
||||
|
||||
### MG-15 - Handler attached to the generic extension event
|
||||
|
||||
**Mechanism.** An action registers for "extension added", which fires when the
|
||||
manager first learns of the actor — before its pawn data, ability system,
|
||||
controller or input exist.
|
||||
|
||||
**Why it is silent.** The handler runs, finds what it needs missing, and returns
|
||||
without acting. Returning early on a missing prerequisite is defensive
|
||||
programming, and looks like it.
|
||||
|
||||
**Why the obvious check misses it.** The registration is correct, the handler is
|
||||
correct, and the early return is correct. Nothing is wrong at any single point.
|
||||
The feature simply never applies, because the one event that would have carried it
|
||||
arrived too early and no later event re-triggers it.
|
||||
|
||||
**Symptom.** A feature that works when activated after the actor exists and does
|
||||
nothing when activated before, or the reverse — order-dependent, and the order
|
||||
differs between the editor and a packaged run.
|
||||
|
||||
**Detect.** List which events each handler responds to:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "HandleActorExtension|FExtensionHandlerDelegate" --glob "*.cpp" . \
|
||||
| rg "EventName ==|NAME_"
|
||||
```
|
||||
|
||||
A handler that responds only to the generic added/removed events, with no
|
||||
semantic ready event, is the finding. The correct shape in the measured reference
|
||||
dispatches on both: the generic event *and* a named readiness event such as
|
||||
"abilities ready" or "bind inputs now", so whichever arrives second does the work.
|
||||
|
||||
**Guardrail.** Handlers accept both the generic event and a semantic readiness
|
||||
event, making activation order irrelevant. Emitting a new readiness event costs a
|
||||
name and one call — it requires no change to the feature infrastructure.
|
||||
|
||||
---
|
||||
|
||||
### MG-16 - Receiver registration written by hand and left incomplete
|
||||
|
||||
**Mechanism.** Actors register with the component manager in three places: add
|
||||
receiver early, send a ready event at begin play, remove receiver at end play.
|
||||
Classes inheriting a modular base get all three. Classes that do not inherit it
|
||||
have them written by hand.
|
||||
|
||||
**Why it is silent.** Missing the removal leaks a registration whose owner is
|
||||
gone. Weak references prevent a crash; the record remains. Missing the ready
|
||||
event means features silently never attach to that actor class.
|
||||
|
||||
**Why the obvious check misses it.** The class works. It renders, it ticks, it
|
||||
does its job. The two or three lines that make it visible to features are
|
||||
infrastructure boilerplate, and boilerplate absent from a file looks like
|
||||
boilerplate that was not needed there.
|
||||
|
||||
**Symptom.** A feature attaches to every actor class except one, with no error —
|
||||
or registrations that accumulate across level transitions.
|
||||
|
||||
**Detect.** Compare the three calls per class:
|
||||
|
||||
```bash
|
||||
rg -l "AddGameFrameworkComponentReceiver" --glob "*.cpp" . > /tmp/add
|
||||
rg -l "RemoveGameFrameworkComponentReceiver" --glob "*.cpp" . > /tmp/rem
|
||||
comm -23 <(sort /tmp/add) <(sort /tmp/rem)
|
||||
```
|
||||
|
||||
Any file that adds and never removes is the finding. In the measured reference
|
||||
one HUD class inherits a non-modular engine base and therefore carries all three
|
||||
calls by hand — correctly, but by hand, which is the fragile arrangement.
|
||||
|
||||
**Guardrail.** Put the three calls in a base class and inherit it. Where that is
|
||||
impossible, add a test that asserts the receiver count returns to baseline after a
|
||||
level transition.
|
||||
|
||||
---
|
||||
|
||||
## Composition
|
||||
|
||||
### MG-17 - Handle dropped without realising it was the unsubscribe
|
||||
|
||||
**Mechanism.** Extension handler and component requests return reference-counted
|
||||
handles. There is no explicit removal API — releasing the last reference *is* the
|
||||
removal.
|
||||
|
||||
**Why it is silent.** Keeping a handle costs nothing visible, and dropping one
|
||||
produces no event. Both directions of the mistake are quiet: holding too long
|
||||
leaks a subscription, releasing too early silently detaches a live feature.
|
||||
|
||||
**Why the obvious check misses it.** Nothing in the calling code says
|
||||
"unsubscribe". Reviewing teardown for removal calls finds none — correctly, because
|
||||
there are none — and it takes knowing the ownership model to recognise that
|
||||
emptying an array of handles is the teardown.
|
||||
|
||||
**Symptom.** Either features that cannot be removed, or features that detach when
|
||||
an unrelated container is cleared.
|
||||
|
||||
**Detect.** Confirm every handle-producing call has a stored destination:
|
||||
|
||||
```bash
|
||||
rg -n "AddExtensionHandler|AddComponentRequest" --glob "*.cpp" . -A3 \
|
||||
| rg -v "\.Add\(|=\s*\w+\.Add"
|
||||
```
|
||||
|
||||
A call whose result is not stored is a subscription that ends immediately.
|
||||
|
||||
**Guardrail.** Name the container for what it does — a handle array is an
|
||||
ownership ledger, not a cache — and comment at its declaration that clearing it is
|
||||
the unsubscribe operation.
|
||||
|
||||
---
|
||||
|
||||
### MG-18 - Plugin activation order assumed
|
||||
|
||||
**Mechanism.** Plugin URLs from an experience and from each of its action sets
|
||||
merge into one list and activate in parallel. Action *execution* order is
|
||||
deterministic; plugin *activation* order is not.
|
||||
|
||||
**Why it is silent.** Parallel activation usually completes before anything
|
||||
depends on the result, and when it does not, the dependent code has its own
|
||||
readiness check that eventually passes.
|
||||
|
||||
**Why the obvious check misses it.** Action ordering *is* deterministic and
|
||||
documented — experience actions first, then action sets in order. Having verified
|
||||
that, it is natural to assume the same of plugins, and nothing contradicts it.
|
||||
|
||||
**Symptom.** A feature that depends on another plugin's registration works
|
||||
consistently until content or load timing changes.
|
||||
|
||||
**Detect.** Check whether the plugin list is loaded as a batch:
|
||||
|
||||
```bash
|
||||
rg -n -A10 "GameFeaturePluginURLs" --glob "*.cpp" . | rg "for\s*\(|LoadGameFeaturePlugin"
|
||||
```
|
||||
|
||||
A loop issuing loads without sequencing between them means order is not
|
||||
guaranteed.
|
||||
|
||||
**Guardrail.** Express cross-plugin dependencies as readiness states rather than
|
||||
activation order. If ordering is genuinely required, sequence the loads and say
|
||||
why in the code.
|
||||
|
||||
---
|
||||
|
||||
## Failure handling
|
||||
|
||||
### MG-19 - Asynchronous deactivation counted but not honoured
|
||||
|
||||
**Mechanism.** The deactivation context supports pausers for asynchronous
|
||||
teardown. The implementation counts them and, if any are outstanding, logs an
|
||||
error and proceeds.
|
||||
|
||||
**Why it is silent.** It is not entirely silent — it logs. But it logs and
|
||||
*continues*, so teardown completes and the sequence appears successful. A log line
|
||||
during shutdown competes with everything else logged during shutdown.
|
||||
|
||||
**Why the obvious check misses it.** The support looks present: the context is
|
||||
constructed with a pauser callback, the count is tracked, the branch exists.
|
||||
Reviewing for "is async deactivation handled?" finds all the machinery. Reading
|
||||
what the branch *does* is a separate step.
|
||||
|
||||
**Symptom.** An action with asynchronous teardown is cut short. Whatever it was
|
||||
waiting to release stays held, and the failure is attributed to whichever
|
||||
subsystem later notices the leak.
|
||||
|
||||
**Detect.** Find pauser handling and read the branch:
|
||||
|
||||
```bash
|
||||
rg -n -A8 "NumExpectedPausers|DeactivatingContext" --glob "*.cpp" . \
|
||||
| rg "UE_LOG|Error|return|check"
|
||||
```
|
||||
|
||||
An error log with no wait and no abort is the finding. The measured reference
|
||||
states its own limitation in the log text — that asynchronous deactivation is not
|
||||
fully supported — which makes it a documented gap rather than an unknown one.
|
||||
|
||||
**Guardrail.** Either wait for pausers before completing teardown, or reject
|
||||
actions that request asynchronous deactivation at validation time. Logging and
|
||||
continuing converts a design gap into a runtime one.
|
||||
@@ -0,0 +1,302 @@
|
||||
# Patterns: a feature-action layer read end to end
|
||||
|
||||
A worked reading of one real modular gameplay implementation: how a feature
|
||||
attaches behaviour to actors it does not own, and what happens when it detaches.
|
||||
|
||||
Individual defects are in [failure-modes.md](failure-modes.md), one entry each,
|
||||
with a recipe. This file is about the shapes worth copying, and about one
|
||||
finding that only becomes visible when you count.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not answerable from the available source, and left open.
|
||||
|
||||
**A scope caveat that shapes everything below.** The engine-side component
|
||||
manager, request handle, feature action base and feature subsystem are not part
|
||||
of the audited project — they live in engine plugins that were not in the tree.
|
||||
Every statement about them here is inferred from call sites **[derived]**, and is
|
||||
marked so. This is the honest version of a common situation: you can audit how a
|
||||
project *uses* an engine subsystem far more cheaply than you can audit the
|
||||
subsystem, and the usage is where your defects will be anyway.
|
||||
|
||||
---
|
||||
|
||||
## 1. The core idea: the action does not look for actors
|
||||
|
||||
The mechanism that makes this architecture work is one inversion.
|
||||
|
||||
A feature action does **not** scan the world for actors to modify. It registers
|
||||
an extension handler for a target *class* with a component manager, and the
|
||||
manager invokes that handler for every matching receiver — the ones that already
|
||||
exist and the ones that appear later **[measured]**.
|
||||
|
||||
```text
|
||||
receiver actor registers itself with the manager
|
||||
action registers an extension handler for a class
|
||||
manager invokes the handler for existing AND future receivers
|
||||
returned request handle owns the subscription
|
||||
```
|
||||
|
||||
The receiver side is three calls, and in the audited project they live in modular
|
||||
base classes rather than in gameplay code **[measured]**: add receiver in
|
||||
pre-initialize, send a ready event at begin play, remove receiver at end play.
|
||||
|
||||
**[derived]** The consequence is the whole point of the architecture: base
|
||||
gameplay classes register as receivers and **know nothing about features at all**.
|
||||
Nothing in the character, player state or game mode references the feature system.
|
||||
A feature plugin can be added to the project without touching any of them.
|
||||
|
||||
The audited project contains one instructive exception. A HUD class inherits a
|
||||
non-modular engine base, so those three calls are written out by hand in it
|
||||
**[measured]**. Correct — and fragile in a way the inherited version is not,
|
||||
because the third call is easy to omit and its omission has no symptom until
|
||||
something counts registrations. Recipe: MG-16.
|
||||
|
||||
---
|
||||
|
||||
## 2. Ownership is a reference count, and nothing says so
|
||||
|
||||
Handler registration and component requests return shared handles **[measured]**.
|
||||
There is no explicit unregister call anywhere in the audited code
|
||||
**[measured]** — the only way to detach is to release the last reference.
|
||||
|
||||
Every reset path therefore begins by emptying an array of handles **[measured]**,
|
||||
and that line *is* the unsubscribe operation even though it reads as ordinary
|
||||
container cleanup.
|
||||
|
||||
**[derived]** Two failure directions follow, and both are quiet. Hold a handle too
|
||||
long and the subscription leaks. Clear the array early and a live feature detaches
|
||||
with no event. Reviewing teardown for "removal calls" finds none — correctly,
|
||||
because there are none — so the reviewer must already know the ownership model to
|
||||
see that teardown is present. Recipe: MG-17.
|
||||
|
||||
One detail in the audited code confirms the model from the inside: when a
|
||||
component already exists, the action still issues a request for it in some cases,
|
||||
with a comment explaining that requests are reference counted **[measured]**.
|
||||
**[derived]** That is a project telling you, in its own code, how the engine
|
||||
subsystem it depends on behaves — which is the best evidence available when the
|
||||
subsystem's source is not in the tree.
|
||||
|
||||
---
|
||||
|
||||
## 3. Per-context state, and why it is not optional
|
||||
|
||||
One action object can be active in several state-change contexts at once — client
|
||||
and server in a single process being the ordinary case. Every lifecycle method
|
||||
receives a context, and four separate actions in the audited project key their
|
||||
state by it **[measured]**:
|
||||
|
||||
```cpp
|
||||
struct FPerContextData { /* tracked resources */ };
|
||||
TMap<FGameFeatureStateChangeContext, FPerContextData> ContextData;
|
||||
```
|
||||
|
||||
This is the pattern to copy verbatim. **[derived]** State in plain member fields
|
||||
works perfectly with one context and corrupts silently with two, and the second
|
||||
context's writes look like legitimate updates.
|
||||
|
||||
Two supporting habits appear consistently and are worth copying with it
|
||||
**[measured]**:
|
||||
|
||||
- at activation, `ensure` that this context's slot is empty and force a reset if
|
||||
not, **before** calling the base implementation — an assertion that converts a
|
||||
leak into a visible failure;
|
||||
- at deactivation, call the base implementation **first**, then reset the
|
||||
context's data.
|
||||
|
||||
The base class itself keeps its delegate handles in a map keyed the same way
|
||||
**[measured]**, which is how it supports being active in two contexts at once.
|
||||
|
||||
Recipe: MG-10.
|
||||
|
||||
---
|
||||
|
||||
## 4. Semantic readiness events
|
||||
|
||||
A generic "extension added" event means only that the manager knows about the
|
||||
actor. **[derived]** It does not mean the actor has its pawn data, ability system,
|
||||
controller, or initialised input — so a handler attached to that event alone will
|
||||
find its prerequisites missing, return early, and never run again.
|
||||
|
||||
The audited project solves this by emitting its own named events at the real
|
||||
dependency boundaries — abilities ready, bind inputs now **[measured]** — and by
|
||||
having every handler dispatch on **both** the generic event and the semantic one
|
||||
**[measured]**:
|
||||
|
||||
```text
|
||||
extension removed / receiver removed -> remove
|
||||
extension added / <semantic event> -> add
|
||||
```
|
||||
|
||||
**[derived]** Whichever arrives second does the work, so activation order stops
|
||||
mattering. This is the single most transferable idea in the file, and it costs a
|
||||
`static const FName` plus one call to emit — no change to the feature
|
||||
infrastructure is required. Recipe: MG-15.
|
||||
|
||||
---
|
||||
|
||||
## 5. The finding that only appears when you count
|
||||
|
||||
Eight actions were audited. Five have incomplete teardown **[measured]**:
|
||||
|
||||
| Action | Attaches | Detaches |
|
||||
|---|---|---|
|
||||
| add abilities | grants abilities, attribute sets, ability sets; stores every handle | mostly correct — but two revocation semantics for one resource kind, depending on which field granted it |
|
||||
| add widgets | pushes layouts, registers extensions, stores handles | reset unregisters extensions and does **not** deactivate layouts; the per-actor path does both |
|
||||
| add input binding | delegates to a game component | the component's removal method has an empty body |
|
||||
| add input mapping context | adds contexts to the input subsystem | tracks affected controllers in a container **nothing ever writes to** |
|
||||
| splitscreen config | votes to disable | correct — decrements the vote |
|
||||
|
||||
**[derived]** Read one at a time, each looks like an isolated oversight. Read
|
||||
together, they are one pattern: **the add path is complete in all eight, and the
|
||||
remove path is complete in three.** That ratio is the argument for the ownership
|
||||
ledger in the skill — not a rule someone invented, but the shape the evidence
|
||||
takes.
|
||||
|
||||
Two of these deserve their detail, because the detail is what makes them
|
||||
recognisable elsewhere.
|
||||
|
||||
### The container nobody fills
|
||||
|
||||
An action declares a container of the controllers it has affected. An `ensure` at
|
||||
activation asserts it starts empty. The reset loop drains it correctly. The
|
||||
reactive path removes from it. **Nothing adds to it** — the function that would
|
||||
takes the per-context data as a parameter and ignores it **[measured]**.
|
||||
|
||||
Four of five operations exist, so any search for "is this container used?"
|
||||
answers yes emphatically. The one missing operation is the producer.
|
||||
|
||||
Two things make this the clearest entry in the file. First, a **sibling action in
|
||||
the same directory** does it correctly, with `AddUnique` in exactly the place the
|
||||
broken one omits **[measured]** — so the two files diff against each other.
|
||||
Second, the original authors left a comment on the broken line stating that the
|
||||
container is never modified **[measured]**. A known defect that shipped. Recipe:
|
||||
MG-03.
|
||||
|
||||
### The rollback that cannot be written
|
||||
|
||||
The input-binding action's teardown iterates its tracked pawns and calls a removal
|
||||
method on the game component. That method's body is a single TODO comment
|
||||
**[measured]**.
|
||||
|
||||
But the cause is one level further down: the function that creates the bindings
|
||||
declares the handle array as a **local variable** and lets it go out of scope
|
||||
**[measured]**, at two separate call sites.
|
||||
|
||||
**[derived]** So the removal is not merely unimplemented, it is impossible to
|
||||
implement without changing the component — the information it would need was
|
||||
discarded at creation. An estimate of "implement the missing rollback" made from
|
||||
reading the empty function comes out an order of magnitude low. Recipes: MG-01,
|
||||
MG-02.
|
||||
|
||||
---
|
||||
|
||||
## 6. What the reference got right
|
||||
|
||||
Stated deliberately, because the section above is not an argument against the
|
||||
architecture:
|
||||
|
||||
1. **Actions do not search for actors.** The extension handler protocol handles
|
||||
existing and future receivers uniformly **[measured]**.
|
||||
2. **Per-context state in all four attaching actions** **[measured]**, with the
|
||||
activation-time emptiness assertion.
|
||||
3. **Reset loops written as drain loops** — "while not empty, take the top,
|
||||
remove it" **[measured]** — which survives a removal that mutates the same
|
||||
collection. A range-for would not.
|
||||
4. **Two removal routes, both idempotent**: explicit reset at deactivation, and
|
||||
reactive removal when the actor dies first, keyed on whether a record exists
|
||||
**[measured]**.
|
||||
5. **Composition enforced over inheritance.** Validation rejects deriving one
|
||||
scripted experience from another and points the author at action sets
|
||||
**[measured]**. That is a design rule with a compiler behind it.
|
||||
6. **Action sets are separate primary assets** with their own bundle data
|
||||
**[measured]**, which is what allows them to be loaded alongside the experience
|
||||
rather than through it.
|
||||
7. **A one-frame deferral with a named reason** **[measured]** — it waits for
|
||||
startup settings. A delay whose prerequisite is named is debt with a receipt;
|
||||
an unnamed one is debt with a timer around it. Recipe: MG-14.
|
||||
8. **A policy class as the place for pre-activation work.** Logic that must run
|
||||
before activation, is process-global rather than per-world, or needs to see all
|
||||
of a plugin's actions at once, belongs in a lifecycle observer rather than in an
|
||||
action **[measured]**. One action in the audited set is pure data with a getter,
|
||||
executed entirely by such an observer **[measured]** — a clean separation worth
|
||||
copying.
|
||||
|
||||
---
|
||||
|
||||
## 7. The authors' own TODO block
|
||||
|
||||
The experience manager carries a block of eight TODOs from its authors
|
||||
**[measured]**, covering: asynchronous experience loading; explicit error
|
||||
handling instead of assertions; running actions in phases rather than all at
|
||||
once; support for deactivating an experience; and — stated plainly — that plugins
|
||||
are **leaked enabled** between experiences, with diffing requirements named as the
|
||||
intended fix.
|
||||
|
||||
**[derived]** This is the cheapest finding in any audit, and the most reliable.
|
||||
The authors know; they wrote it down; nobody adopting the code reads it, because
|
||||
adoption reads the class list rather than the comments.
|
||||
|
||||
Two practical consequences:
|
||||
|
||||
- **Read the TODO blocks first** in any codebase you are adopting. They are a
|
||||
defect list written by the people best placed to write one.
|
||||
- Treat "the authors documented this gap" as *stronger* evidence than a static
|
||||
finding of your own, not weaker. It removes the possibility that you have
|
||||
misread the intent.
|
||||
|
||||
Recipe: MG-12. The same reading also surfaces asynchronous deactivation: pausers
|
||||
are counted, and when any are outstanding the code logs an error and proceeds
|
||||
**[measured]**. The limitation is stated in the log text itself. Recipe: MG-19.
|
||||
|
||||
---
|
||||
|
||||
## 8. Editor loading is not shipping loading
|
||||
|
||||
The loading policy loads client data when not a dedicated server, and server data
|
||||
when not a client-only build **[measured]**. **[derived]** In the editor neither
|
||||
condition excludes, so both bundle sets load — and the authors noted the resulting
|
||||
hitching in a comment beside it **[measured]**.
|
||||
|
||||
Every in-editor memory and load-time figure therefore describes a configuration
|
||||
that ships to nobody. Directionally useful, absolutely wrong. Recipe: MG-13.
|
||||
|
||||
---
|
||||
|
||||
## 9. What source reading could not settle
|
||||
|
||||
- **The engine subsystems were not in the tree.** Component manager, request
|
||||
handle, feature action base, feature subsystem, and the add-components action
|
||||
**[open]**. Everything above about them is inferred from call sites.
|
||||
- **Binary assets were not read.** Which actions the shipped feature plugins
|
||||
actually configure, and with what data, is **[open]**. What *is* measured: those
|
||||
plugins contain no feature-action C++ of their own, so they use only actions
|
||||
from the base module and the engine **[measured]**.
|
||||
- **One teardown question is genuinely undecidable from this tree.** Whether
|
||||
clearing the handle array causes the manager to dispatch removal events — which
|
||||
would make the narrower reset path in MG-04 harmless — depends on engine code
|
||||
that was not available **[open]**. It is recorded as an open question rather than
|
||||
resolved in either direction, because a plausible answer either way would be a
|
||||
guess.
|
||||
|
||||
That last item is worth its space. It would have been easy to state MG-04 as a
|
||||
confirmed defect; the honest form is "this is a defect unless an engine behaviour
|
||||
we could not read compensates for it", and a reader adopting the code can settle
|
||||
it in ten minutes with the engine source that they have and the audit did not.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against the feature-action layer of Epic's Lyra Starter Game on Unreal
|
||||
Engine 5.6, read as source in a single workspace. Source addresses stay in the
|
||||
research archive that produced this skill; each `MG-` identifier resolves back to
|
||||
the audited location there, so any specific claim above can be produced on
|
||||
request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace, with the engine-side subsystems
|
||||
outside the readable tree. These are examples and failure evidence, not
|
||||
guarantees about other versions. Claims marked open stay open. Re-run the recipes
|
||||
in [failure-modes.md](failure-modes.md) against your own tree before acting on
|
||||
anything here.
|
||||
@@ -0,0 +1,390 @@
|
||||
---
|
||||
name: ue-multiplayer-authority
|
||||
description: >-
|
||||
Design and review the authority model of a networked Unreal Engine product:
|
||||
which side owns each fact, what is predicted versus replicated versus
|
||||
validated, where responsibility boundaries run between GameState, PlayerState,
|
||||
Controller, Pawn and components, RPC and validation discipline, hit
|
||||
registration and lag compensation, replication modes, relevancy and cost, and
|
||||
the cheat surface. Use when starting a multiplayer feature, porting
|
||||
single-player code to network, reviewing trust boundaries, or investigating
|
||||
desync, rubber-banding, "works in PIE but not on dedicated server", or
|
||||
suspected cheating.
|
||||
---
|
||||
|
||||
# UE multiplayer authority
|
||||
|
||||
The invariant:
|
||||
|
||||
> Every fact in a networked game has exactly one owner. Authority is not a
|
||||
> property of code that says `HasAuthority` — it is the guarantee that no other
|
||||
> machine can produce that fact.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Twenty detection recipes for silent trust failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting or reviewing a networked design. Each
|
||||
entry names the mechanism, why it stays silent, why the obvious check misses it,
|
||||
and a command you can run against your own tree.
|
||||
|
||||
Related skills: `ue-gas-architecture`, `ue-gameplay-messaging`,
|
||||
`ue-modular-gameplay`, `ue-cosmetics-and-teams`, `ue-architecture-guardrails`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Start from one question
|
||||
|
||||
Before any code: **for each fact in the game, who is allowed to be wrong about
|
||||
it?**
|
||||
|
||||
- If a client being wrong is a cosmetic glitch → the client may own it.
|
||||
- If a client being wrong is an advantage → the server must own it, and must be
|
||||
able to produce the value independently.
|
||||
- If a client being wrong is unrecoverable → the server must own it *and*
|
||||
persist it.
|
||||
|
||||
This question, answered per fact and written down, is the authority model.
|
||||
Everything else in this document is enforcement of that table.
|
||||
|
||||
### The three states of a networked fact
|
||||
|
||||
| State | Meaning | Who can be wrong |
|
||||
|---|---|---|
|
||||
| **Predicted** | Client computes it immediately for responsiveness; server may overrule | Client, temporarily, visibly |
|
||||
| **Replicated** | Server computes it, clients receive it | Nobody — clients only observe |
|
||||
| **Validated** | Client proposes it, server independently verifies before accepting | Client may lie; server must detect |
|
||||
|
||||
**The third is the expensive one, and it is the one that is usually missing.**
|
||||
"Replicated" is often mistaken for "validated": a value that travelled from the
|
||||
client to the server and was then replicated back out to everyone is *client
|
||||
authored*, no matter how authoritative it looks on arrival.
|
||||
|
||||
---
|
||||
|
||||
## 2. What actually enforces authority
|
||||
|
||||
Unreal offers several mechanisms that look similar and enforce very different
|
||||
amounts. Know which is which before relying on one.
|
||||
|
||||
| Mechanism | Enforces | Does not enforce |
|
||||
|---|---|---|
|
||||
| `HasAuthority()` guard | that this code path runs on the server | anything about callers that skip the guard |
|
||||
| `UFUNCTION(Server)` | routing to the server | that the arguments are truthful |
|
||||
| `WithValidation` | that a `_Validate` function runs and can disconnect | anything, if `_Validate` returns `true` unconditionally |
|
||||
| `BlueprintAuthorityOnly` | that **Blueprint** cannot call this | nothing at all for C++ callers |
|
||||
| replicated property | one-way flow server → client | that the server computed the value itself |
|
||||
| `GetLifetimeReplicatedProps` conditions | who receives it | who may change it |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **A public mutator on a replicated fact must guard itself**, not rely on its
|
||||
callers. `if (!HasAuthority()) return;` at the top of the function, not at the
|
||||
call site — call sites multiply.
|
||||
2. **`WithValidation` with an empty `_Validate` is worse than none**: it passes
|
||||
review as validation while enforcing nothing.
|
||||
3. **`BlueprintAuthorityOnly` is a Blueprint editor affordance.** Never cite it as
|
||||
a security property.
|
||||
4. **Write the guard even when the only current caller is the server.** The guard
|
||||
is documentation the compiler enforces; the caller list is not stable.
|
||||
|
||||
---
|
||||
|
||||
## 3. Responsibility boundaries by layer
|
||||
|
||||
Put each fact in exactly one place. The most common source of networked bugs is a
|
||||
fact that lives in two.
|
||||
|
||||
| Layer | Exists on | Owns | Must not own |
|
||||
|---|---|---|---|
|
||||
| **GameMode** | server only | match rules, spawning, scoring decisions | anything a client must read |
|
||||
| **GameState** | server + all clients | match-wide replicated facts: phase, timer, team scores | per-player secrets |
|
||||
| **PlayerState** | server + all clients | per-player facts everyone may see: name, team, score, abilities | input, camera, UI state |
|
||||
| **PlayerController** | server + owning client | input, camera, client-directed commands, UI ownership | facts other players need |
|
||||
| **Pawn / Character** | server + all clients | embodiment: transform, movement, physical state | player identity, persistent progression |
|
||||
| **Components** | follow their owner | one cohesive slice of the owner's facts | facts belonging to a different layer |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **`GameMode` exists only on the server.** Any client code path referencing it
|
||||
is a bug that PIE will hide.
|
||||
2. **Persistent player facts belong to `PlayerState`, not the Pawn.** Pawns die.
|
||||
3. **Input and camera belong to the Controller.** A Pawn reading input directly
|
||||
cannot be possessed by AI or spectated.
|
||||
4. **A component's authority is its owner's authority.** A component attached to a
|
||||
client-owned actor cannot be authoritative regardless of its internal guards.
|
||||
5. **Simulated proxies have no controller.** Any logic gated on "has a controller"
|
||||
silently excludes them; any logic that assumes one will null-deref on them.
|
||||
|
||||
---
|
||||
|
||||
## 4. RPC discipline
|
||||
|
||||
### Choosing the direction
|
||||
|
||||
```text
|
||||
Client -> Server UFUNCTION(Server, Reliable/Unreliable, WithValidation)
|
||||
Server -> owner UFUNCTION(Client, ...)
|
||||
Server -> everyone UFUNCTION(NetMulticast, ...)
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Every `Server` RPC carries `WithValidation` and a `_Validate` that actually
|
||||
checks.** Range, ownership, state legality, rate. An RPC without validation is
|
||||
a client-side API into your server.
|
||||
2. **Reliable is a budget, not a default.** Reliable RPCs are ordered and queued;
|
||||
overflowing the queue disconnects the client. Anything cosmetic, frequent or
|
||||
idempotent should be unreliable.
|
||||
3. **Never multicast what a replicated property already carries.** The property
|
||||
handles joiners and relevancy; a multicast does not.
|
||||
4. **Multicasts do not reach late joiners or newly relevant clients.** Anything a
|
||||
late joiner must know is state, not an event.
|
||||
5. **Validate on the server that the sender owns the thing it is acting on.**
|
||||
Routing to the server proves nothing about the *subject* of the call.
|
||||
6. **Rate-limit anything a client can call in a loop.**
|
||||
|
||||
### Debug and cheat commands
|
||||
|
||||
Cheat entry points reachable from a client are the highest-value target in the
|
||||
codebase. They must be compiled out of shipping builds — not merely hidden,
|
||||
disabled by a flag, or guarded by UI. A `Server` RPC that executes an arbitrary
|
||||
cheat string, present in a shipping binary, is a remote console.
|
||||
|
||||
---
|
||||
|
||||
## 5. Prediction
|
||||
|
||||
Prediction exists to hide latency, not to save server work. The server must still
|
||||
compute everything it accepts.
|
||||
|
||||
### Safe to predict
|
||||
|
||||
- movement through the engine's own prediction (correction path already exists);
|
||||
- animation, audio, VFX, camera;
|
||||
- UI affordances (cooldown wheels, ammo counters) that a correction can restore;
|
||||
- ability activation under a prediction key, where the server can reject.
|
||||
|
||||
### Never predict
|
||||
|
||||
- damage application;
|
||||
- score, currency, progression;
|
||||
- inventory grants;
|
||||
- death and elimination;
|
||||
- anything persisted.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Every predicted value needs a defined correction path** — what happens
|
||||
visually when the server disagrees. Undefined correction means the client
|
||||
keeps a wrong value forever.
|
||||
2. **Prediction is a client-side optimisation of a server-side truth.** If the
|
||||
server cannot independently produce the value, the feature is not predicted,
|
||||
it is client-authoritative.
|
||||
3. **A prediction key is not validation.** It reconciles ordering; it does not
|
||||
check whether the client was entitled to act.
|
||||
|
||||
---
|
||||
|
||||
## 6. Hit registration — the decision that shapes the whole product
|
||||
|
||||
This is where authority is usually lost, and it is a design decision that must be
|
||||
made explicitly and early, because retrofitting is expensive.
|
||||
|
||||
There are three positions:
|
||||
|
||||
| Model | Server does | Cost | Cheat exposure |
|
||||
|---|---|---|---|
|
||||
| **Server-authoritative trace** | traces from its own state at the current tick | cheap | none, but shots feel like they miss under latency |
|
||||
| **Lag-compensated rewind** | stores per-tick historical positions, rewinds to the client's timestamp, re-traces | expensive: memory, CPU, complexity | low; the standard for competitive shooters |
|
||||
| **Client-reported hit** | accepts the client's hit result | free | total |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Choose consciously and write the choice down.** Most projects arrive at the
|
||||
third by default, because the first feels bad and the second is work.
|
||||
2. **If you accept client-reported hits, know exactly what a malicious client
|
||||
gains** — typically arbitrary damage, arbitrary range, arbitrary target,
|
||||
arbitrary surface multipliers.
|
||||
3. **The server cannot validate what it cannot reproduce.** Without stored
|
||||
historical positions there is nothing to validate a client's hit *against*,
|
||||
so bolting on "validation" later is not a patch, it is building the rewind
|
||||
subsystem.
|
||||
4. **Sanity checks are not validation, but they are not worthless.** Maximum
|
||||
range, line-of-sight from the shooter's replicated position, fire-rate limits,
|
||||
team check, and "was the target ever near there recently" raise the cost of
|
||||
trivial cheats even without full rewind.
|
||||
5. **Never let the client choose damage-modifying inputs.** The surface/material
|
||||
that multiplies damage, the distance used for falloff, and the trace origin
|
||||
must all come from server-known state, even in a client-reported-hit model.
|
||||
These are the cheapest things to move server-side and the most valuable.
|
||||
|
||||
### Recognising a stub
|
||||
|
||||
Watch for parameters threaded through many function signatures but never branched
|
||||
on, and for booleans initialised to a constant with a `//@TODO`. These are the
|
||||
attachment points of a rewind subsystem that was designed and never built. They
|
||||
make the code *look* like it validates. Recipes: NA-06, NA-07 in the failure
|
||||
modes.
|
||||
|
||||
---
|
||||
|
||||
## 7. Replication cost and relevancy
|
||||
|
||||
Authority decisions determine correctness; relevancy decisions determine whether
|
||||
the product ships.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Pick an ability-system replication mode deliberately.** `Full` replicates all
|
||||
gameplay effects to all clients; `Mixed` replicates effects to the owner and
|
||||
tags/cues to everyone; `Minimal` replicates no effects to anyone.
|
||||
Player-controlled actors want `Mixed`; AI want `Minimal`.
|
||||
2. **Net cull distance is a gameplay decision, not a performance knob.** Setting
|
||||
it beyond the map size disables relevancy culling entirely — every actor
|
||||
replicates to every client, which is correct for a small arena and fatal for a
|
||||
large map.
|
||||
3. **Prefer replicated state over event spam.** State is idempotent, survives
|
||||
packet loss, and reaches late joiners.
|
||||
4. **FastArraySerializer for collections that change incrementally**, with the
|
||||
`PreReplicatedRemove` / `PostReplicatedAdd` / `PostReplicatedChange` callbacks
|
||||
all implemented — an empty callback is a client-side desync with no symptom on
|
||||
the server.
|
||||
5. **Replication graph before you need it, or never.** Enabling it late changes
|
||||
which actors clients see, and every relevancy assumption in gameplay code has
|
||||
to be re-verified. If it ships disabled, treat routing configuration as
|
||||
untested.
|
||||
6. **Cosmetic state should not replicate.** Replicate intent (which skin, which
|
||||
team), realise presentation locally.
|
||||
|
||||
---
|
||||
|
||||
## 8. Lifecycle across the network
|
||||
|
||||
1. **Readiness is a state chain, not a timing assumption.** Actor spawned,
|
||||
controller assigned, PlayerState replicated, data initialised, gameplay ready —
|
||||
each an explicit state with explicit transitions.
|
||||
2. **Client and server reach each state in a different order.** Any code that
|
||||
assumes an order will work for the listen-server host and fail for remote
|
||||
clients.
|
||||
3. **A simulated proxy passes through the chain without ever acquiring a
|
||||
controller.** Gates must handle that case explicitly rather than by omission.
|
||||
4. **Every dynamic addition needs its exact removal**: granted abilities, applied
|
||||
effects, registered listeners, spawned cues, added tag paths. Test the removal
|
||||
on both server and client, in the same session, more than once.
|
||||
5. **Death is not destruction.** Define the full sequence — death started, effects
|
||||
removed, input released, cues cleaned, actor destroyed or respawned — and give
|
||||
each step one owner.
|
||||
|
||||
---
|
||||
|
||||
## 9. The cheat surface review
|
||||
|
||||
Enumerate, for the shipping build, every path a malicious client can reach:
|
||||
|
||||
```text
|
||||
1. Every Server RPC -> does _Validate independently verify?
|
||||
2. Every RPC argument -> is any of it trusted without recomputation?
|
||||
3. Every cheat/debug entry point -> is it compiled out, not just disabled?
|
||||
4. Every damage input -> which come from the client?
|
||||
5. Every fail-open error path -> what does the code do when a check errors out?
|
||||
6. Every client-side gate -> is there a matching server-side one?
|
||||
```
|
||||
|
||||
Item 5 deserves attention: a permission check that returns "allowed" when it
|
||||
cannot determine the answer converts an unknown state into an exploit. Failure to
|
||||
determine must deny.
|
||||
|
||||
---
|
||||
|
||||
## 10. Test matrix
|
||||
|
||||
**PIE is not a network test.** A listen-server host shares memory with its client;
|
||||
half of the failure modes here cannot occur in it.
|
||||
|
||||
### Mandatory configurations
|
||||
|
||||
- dedicated server + two remote clients;
|
||||
- one client with 150 ms latency and 2% packet loss;
|
||||
- a client joining mid-match;
|
||||
- a client disconnecting mid-action;
|
||||
- an actor becoming relevant and then irrelevant.
|
||||
|
||||
### Per feature
|
||||
|
||||
- does it work for a remote client, not just the host?
|
||||
- does a late joiner reach the correct state?
|
||||
- what does a simulated proxy see?
|
||||
- what happens if the client sends the RPC twice, out of order, or with absurd
|
||||
arguments?
|
||||
- is the state correct after death and respawn, twice?
|
||||
|
||||
### Authority regression
|
||||
|
||||
- for each fact in the authority table, assert on the server that the client
|
||||
cannot change it — with a test client that tries.
|
||||
|
||||
---
|
||||
|
||||
## 11. Review checklist
|
||||
|
||||
- [ ] Every fact has a named owner in a written authority table.
|
||||
- [ ] Each fact is classified predicted / replicated / validated.
|
||||
- [ ] Public mutators of replicated state guard themselves.
|
||||
- [ ] Every `Server` RPC has `WithValidation` with a real `_Validate`.
|
||||
- [ ] No trust in client-supplied damage inputs: origin, distance, surface, target.
|
||||
- [ ] Hit registration model is chosen and documented.
|
||||
- [ ] Cheat entry points are compiled out of shipping.
|
||||
- [ ] Error paths deny rather than allow.
|
||||
- [ ] Replication mode and net cull distance chosen per class, not inherited.
|
||||
- [ ] FastArray callbacks all implemented.
|
||||
- [ ] Every addition has a verified removal.
|
||||
- [ ] Feature tested on a dedicated server with a remote client under latency.
|
||||
- [ ] `BlueprintAuthorityOnly` is not cited as enforcement anywhere.
|
||||
|
||||
---
|
||||
|
||||
## 12. When client authority is the right answer
|
||||
|
||||
Deliberate client authority is legitimate when:
|
||||
|
||||
- the fact is purely presentational;
|
||||
- the game is cooperative or single-player-with-friends and the threat model is
|
||||
"no adversary";
|
||||
- the cost of server validation exceeds the value of the protected fact;
|
||||
- an anti-cheat layer outside the game handles the threat.
|
||||
|
||||
In every one of these cases, the decision is written down with its threat model.
|
||||
The failure worth avoiding is not "we trusted the client" — it is **"we did not
|
||||
know we trusted the client."**
|
||||
|
||||
That distinction is what this skill exists to preserve. A codebase can be an
|
||||
excellent architectural reference and simultaneously an unsafe network reference;
|
||||
those are separate axes, and copying the first without noticing the second is the
|
||||
common accident.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
Epic's Lyra Starter Game on Unreal Engine 5.6, read rather than run. That project
|
||||
is an outstanding architectural reference and a poor network-security reference;
|
||||
those are independent axes, and the audit exists to keep them apart. Source
|
||||
addresses stay in the research archive that produced this skill; what ships is the
|
||||
detection recipe, because an address in someone else's tree is not something you
|
||||
can act on and a recipe is.
|
||||
|
||||
Each entry carries a stable identifier (`NA-01`, `NA-02`, …) that resolves back to
|
||||
the audited location in that archive. If you need the original address to settle a
|
||||
dispute, it exists and can be produced.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one specific reference project on one engine
|
||||
version in one workspace. It is evidence of failure modes, not a guarantee about
|
||||
other engine versions or other samples. Several claims in it are explicitly marked
|
||||
as unanswerable without a dedicated-server build, and they stay unanswered. Re-run
|
||||
the recipes against your own tree before trusting any specific claim.
|
||||
@@ -0,0 +1,764 @@
|
||||
# Failure modes: multiplayer authority
|
||||
|
||||
Twenty ways a networked Unreal project loses authority without anything going
|
||||
wrong on screen. Each entry names the mechanism, why it stays silent, why the
|
||||
obvious check misses it, the symptom, a recipe you can run against your own tree,
|
||||
and the guardrail.
|
||||
|
||||
Two properties recur across this list and are worth stating once, because they
|
||||
explain why these survive review:
|
||||
|
||||
- **PIE cannot reproduce most of them.** A listen-server host shares memory with
|
||||
its client, so the client/server split that produces the bug does not exist
|
||||
during the test that would have caught it. Entries NA-01, NA-04, NA-11, NA-13,
|
||||
NA-17 and NA-19 are invisible in PIE by construction.
|
||||
- **Several of them look correct in review.** A `WithValidation` marker, a
|
||||
`BlueprintAuthorityOnly` specifier, a `bIsSimulated` parameter threaded through
|
||||
signatures — each reads as evidence of a server-side path that is not there.
|
||||
The specifier is real; the enforcement is not.
|
||||
|
||||
Recipes below use `rg`. They are written to be run from the root of a project's
|
||||
source tree and to be read as a *comparison of two numbers*, not as a single
|
||||
search: the whole class of defect here is "the thing is present, and the thing
|
||||
that would make it mean something is absent".
|
||||
|
||||
---
|
||||
|
||||
## Trust boundaries
|
||||
|
||||
### NA-01 - Replicated is mistaken for validated
|
||||
|
||||
**Mechanism.** A value travels client to server and is then replicated back out to
|
||||
everyone. On arrival it lives in an authoritative-looking replicated property, and
|
||||
every later reader treats it as a server fact.
|
||||
|
||||
**Why it is silent.** Replication works perfectly. The value is consistent on every
|
||||
machine, arrives in order, survives relevancy changes and late joins. Consistency
|
||||
is not authorship, but consistency is what everyone observes and tests.
|
||||
|
||||
**Why the obvious check misses it.** Review asks "is this replicated?" and the
|
||||
answer is yes. There is no compiler concept, specifier or naming convention that
|
||||
distinguishes a server-computed replicated value from a server-relayed one. Both
|
||||
compile to the same declaration.
|
||||
|
||||
**Symptom.** Nothing at all, until someone modifies a client. Then the value is
|
||||
whatever that client chose, and it is authoritative across the whole session.
|
||||
|
||||
**Detect.** This one is a table, not a grep. List every replicated property and
|
||||
name, for each, the server-side computation that produced it:
|
||||
|
||||
```bash
|
||||
rg -n "UPROPERTY\([^)]*Replicated" Source/ | wc -l # facts to account for
|
||||
rg -n "GetLifetimeReplicatedProps" Source/ # where they are declared
|
||||
```
|
||||
|
||||
For each result, answer in writing: *what did the server compute to get this?* If
|
||||
the answer is "the client sent it", the fact is client-authored. A property whose
|
||||
only writer is an RPC handler that assigns the parameter directly is the shape to
|
||||
look for.
|
||||
|
||||
**Guardrail.** Classify every fact as predicted / replicated / validated in a
|
||||
written table before the feature is built. "Validated" requires the server to
|
||||
independently reproduce the value, not merely to receive it.
|
||||
|
||||
---
|
||||
|
||||
### NA-02 - `WithValidation` with an empty `_Validate`
|
||||
|
||||
**Mechanism.** The RPC carries the `WithValidation` specifier and the paired
|
||||
`_Validate` function exists, so the engine calls it. Its body returns `true`
|
||||
unconditionally.
|
||||
|
||||
**Why it is silent.** The mechanism is fully wired and working. The engine really
|
||||
does call the validator and really would disconnect a client that failed it. There
|
||||
is simply no input it rejects, and a validator that accepts everything is
|
||||
indistinguishable at runtime from one that has nothing to reject yet.
|
||||
|
||||
**Why the obvious check misses it.** The specifier is what reviewers grep for, and
|
||||
it is present. Declarations and implementations live in different files, so a
|
||||
header review — the natural place to audit an RPC surface — never sees the body.
|
||||
|
||||
**Symptom.** The RPC accepts anything. Review signs off on the subsystem as
|
||||
validated, and the estimate for "add validation" is later scoped as a patch when
|
||||
it is a design.
|
||||
|
||||
**Detect.** Read every validator body, not the declarations:
|
||||
|
||||
```bash
|
||||
rg -n --stats "WithValidation" Source/ # declared validators
|
||||
rg -n -A3 "_Validate\(" Source/ | rg -B1 "^\s*return true;" # bodies that accept all
|
||||
```
|
||||
|
||||
Compare the counts. On the measured reference the second command returned two
|
||||
sites, both belonging to cheat RPCs. Any validator whose body is a bare
|
||||
`return true;` is an unimplemented validator regardless of the specifier.
|
||||
|
||||
**Guardrail.** Require every `_Validate` body to check range, ownership, state
|
||||
legality and rate. Treat a one-line `return true;` in a validator as a review
|
||||
blocker, the same way an empty `catch` block is.
|
||||
|
||||
---
|
||||
|
||||
### NA-03 - `BlueprintAuthorityOnly` cited as enforcement
|
||||
|
||||
**Mechanism.** A function is marked `BlueprintAuthorityOnly`. The specifier stops
|
||||
Blueprint graphs from calling it on a non-authoritative actor. C++ callers are
|
||||
entirely unaffected.
|
||||
|
||||
**Why it is silent.** In a project where the function is only ever called from
|
||||
Blueprint, the specifier does exactly what everyone believes it does. The gap
|
||||
opens the first time a C++ caller appears, and that caller works, so nothing draws
|
||||
attention to it.
|
||||
|
||||
**Why the obvious check misses it.** The word "authority" is in the specifier.
|
||||
Grepping for authority-related terms finds it and it looks like a guard. It is an
|
||||
editor affordance whose name describes intent rather than enforcement.
|
||||
|
||||
**Symptom.** A function documented and reviewed as authority-guarded is reachable
|
||||
from any C++ path, including client code, and mutates replicated state from there.
|
||||
|
||||
**Detect.** Enumerate the sites and check each for a real guard in the body:
|
||||
|
||||
```bash
|
||||
rg -n "BlueprintAuthorityOnly" Source/ # claimed guards
|
||||
rg -n -A6 "BlueprintAuthorityOnly" Source/ \
|
||||
| rg "HasAuthority|GetLocalRole" # actual guards
|
||||
```
|
||||
|
||||
A large gap between the two counts is the finding. On the measured reference the
|
||||
specifier appeared at five sites across three subsystems, with no `HasAuthority`
|
||||
in the guarded bodies.
|
||||
|
||||
**Guardrail.** Never cite `BlueprintAuthorityOnly` as a security property in a
|
||||
review, a comment or a design document. Add an explicit `HasAuthority()` guard
|
||||
inside the function body; keep the specifier only as an editor hint.
|
||||
|
||||
---
|
||||
|
||||
### NA-04 - Guard placed at the call site rather than in the mutator
|
||||
|
||||
**Mechanism.** A public mutator of replicated state relies on its callers to check
|
||||
authority. Today every caller does.
|
||||
|
||||
**Why it is silent.** The invariant holds for as long as the caller list is what it
|
||||
was when the code was written. Nothing degrades gradually; the code is correct
|
||||
until one day it is not, and the transition is a new caller, not an edit to the
|
||||
function.
|
||||
|
||||
**Why the obvious check misses it.** Reading the function shows nothing wrong —
|
||||
there is no missing line, because the check was never meant to be there. Reading
|
||||
the call sites shows a check at each one. Both halves pass review independently.
|
||||
The defect only exists in the relationship between them, which no single file
|
||||
review inspects.
|
||||
|
||||
**Symptom.** Works until a new caller appears — commonly a subclass in a feature
|
||||
plugin, a Blueprint node, or a second caller added by someone who reasonably
|
||||
assumed the mutator guards itself. Then a client writes replicated state directly.
|
||||
|
||||
**Detect.** Find public or virtual writers of replicated state whose own body has
|
||||
no guard:
|
||||
|
||||
```bash
|
||||
# 1. names of replicated properties
|
||||
rg -n -o "UPROPERTY\([^)]*Replicated[^)]*\)\s*\n\s*\w+\s+(\w+)" -r '$1' Source/
|
||||
# 2. for each name, find assignments and check the enclosing function
|
||||
rg -n -B12 "\bDeathState\s*=" Source/ | rg "HasAuthority|void |virtual"
|
||||
```
|
||||
|
||||
Substitute your own property name in step 2. The finding is an assignment whose
|
||||
preceding function signature is public or virtual and whose body contains no
|
||||
authority check.
|
||||
|
||||
**Guardrail.** `if (!HasAuthority()) return;` as the first line of the mutator.
|
||||
Write it even when the only current caller is the server: the guard is
|
||||
documentation the compiler enforces, and the caller list is not a stable property.
|
||||
|
||||
---
|
||||
|
||||
### NA-05 - Client-supplied damage inputs
|
||||
|
||||
**Mechanism.** Distance, trace origin, surface material or target identity arrive
|
||||
in the RPC payload and feed the damage calculation directly.
|
||||
|
||||
**Why it is silent.** Every value is in a plausible range, because a normal client
|
||||
sends plausible values. The damage numbers that result look exactly like correct
|
||||
damage numbers. There is no error state and no log line, because nothing failed.
|
||||
|
||||
**Why the obvious check misses it.** The damage code is usually reviewed for
|
||||
correctness of the formula, and the formula is correct. The question that matters
|
||||
is not "does this compute the right number?" but "who chose the inputs?", and the
|
||||
inputs arrive as ordinary function parameters several layers away from where they
|
||||
entered the process.
|
||||
|
||||
**Symptom.** Arbitrary damage multipliers from a modified client, invisible in
|
||||
logs. In a competitive game this surfaces as player reports that a specific
|
||||
weapon or surface "feels wrong", long before anyone suspects the trust boundary.
|
||||
|
||||
**Detect.** Trace each damage-modifying input back to its origin:
|
||||
|
||||
```bash
|
||||
rg -n "PhysicalMaterial|PhysMat|TraceStart|HitLocation|Distance" \
|
||||
Source/ --glob "*Damage*" --glob "*Execution*"
|
||||
```
|
||||
|
||||
For each hit, answer: did the server compute this, or did it arrive from the
|
||||
client? On the measured reference the falloff distance and the surface multiplier
|
||||
both came from client-supplied fields, while the team check beside them was
|
||||
genuinely server-side — the model was incomplete, not absent, which is the harder
|
||||
case to spot.
|
||||
|
||||
**Guardrail.** Recompute every damage-modifying input from server-known state. The
|
||||
shooter's replicated position and a server-side collision query cover most of it
|
||||
and are cheap. Client-supplied values may inform presentation, never arithmetic
|
||||
that changes an outcome.
|
||||
|
||||
---
|
||||
|
||||
### NA-06 - Validation flag initialised to a constant
|
||||
|
||||
**Mechanism.** A boolean local is initialised to a literal — `const bool bIsValid
|
||||
= true;` — and then consumed by branches that read as validation. There is no
|
||||
computation between the initialisation and the use.
|
||||
|
||||
**Why it is silent.** A permanently-valid result is a legal runtime state. Every
|
||||
branch takes the path it would take on a successful validation, which is also the
|
||||
path taken during all normal play, so the code behaves exactly as intended for
|
||||
every non-malicious input.
|
||||
|
||||
**Why the obvious check misses it.** The flag has a meaningful name, appears in
|
||||
conditionals, and is used more than once. Any search for "is validation present?"
|
||||
finds a variable called `bIsTargetDataValid` being consumed by a gate and stops
|
||||
there. The tell is the *initialiser*, which is three lines above and looks like
|
||||
boilerplate.
|
||||
|
||||
**Symptom.** The code path reads as validated in review and in estimates. Nothing
|
||||
is observable at runtime until a modified client sends implausible data and it is
|
||||
accepted.
|
||||
|
||||
**Detect.** Look for boolean locals initialised to a literal and later consumed as
|
||||
gates, restricted to network-relevant code:
|
||||
|
||||
```bash
|
||||
rg -n "const bool b\w+\s*=\s*(true|false)\s*;" Source/
|
||||
```
|
||||
|
||||
Then, for each hit, check whether the name promises a computation
|
||||
(`bIsValid`, `bCanApply`, `bIsAuthorized`) and whether the variable is read by a
|
||||
branch. On the measured reference this returned eight sites; most were legitimate
|
||||
call-parameter constants, and exactly one — a target-data validity flag inside a
|
||||
weapon ability — was a validation stub. The recipe's value is that it reduces the
|
||||
whole codebase to a list short enough to read.
|
||||
|
||||
**Guardrail.** Treat any boolean local in a network path whose initialiser is a
|
||||
literal and whose name asserts a property as unimplemented until proven otherwise.
|
||||
If validation is not written yet, name the variable for what it is.
|
||||
|
||||
---
|
||||
|
||||
### NA-07 - Dead parameters that imply a server path
|
||||
|
||||
**Mechanism.** A parameter such as `bIsSimulated` is threaded through many function
|
||||
signatures, passed a literal at the only call site, and never branched on
|
||||
anywhere.
|
||||
|
||||
**Why it is silent.** The parameter is passed and received correctly at every
|
||||
level. Nothing is unused in the compiler's view, so no warning fires. The code
|
||||
does what it does today with complete consistency.
|
||||
|
||||
**Why the obvious check misses it.** Signature review is exactly the wrong
|
||||
instrument here, and it is the instrument reviewers reach for when auditing a
|
||||
subsystem's shape. A parameter appearing in ten signatures is strong evidence to a
|
||||
human that a feature exists. Confirming it requires following the value to a
|
||||
branch, which no IDE affordance does for you.
|
||||
|
||||
**Symptom.** A reviewer concludes a simulated or authoritative path exists.
|
||||
Estimates for "just add server validation" come out an order of magnitude low,
|
||||
because the attachment points were mistaken for the subsystem.
|
||||
|
||||
**Detect.** Count mentions, then count branches:
|
||||
|
||||
```bash
|
||||
P='bIsSimulated'
|
||||
rg -n "\b$P\b" Source/ | wc -l # mentions: many
|
||||
rg -n "if\s*\(\s*!?\s*$P\b|\?\s*.*:\s*.*\b$P\b" Source/ # branches: ?
|
||||
```
|
||||
|
||||
On the measured reference the parameter appeared at ten sites and was branched on
|
||||
at zero. A high first number with an empty second is a mounting point, not a
|
||||
feature.
|
||||
|
||||
**Guardrail.** Trace every network-relevant parameter to a branch. No branch means
|
||||
no feature: either implement it or delete the parameter, because its presence is
|
||||
an active lie to the next reader.
|
||||
|
||||
---
|
||||
|
||||
### NA-08 - Retrofitting hit validation without rewind
|
||||
|
||||
**Mechanism.** The server holds no historical positions, so there is nothing to
|
||||
validate a timestamped client hit against. Validation is scheduled as a task on
|
||||
top of the existing weapon code.
|
||||
|
||||
**Why it is silent.** The absence is an absence. Nothing in the codebase points at
|
||||
the missing subsystem; there is no stub, no TODO, no failing test. The weapon
|
||||
system works, feels good, and passes playtests, because every client is honest
|
||||
during a playtest.
|
||||
|
||||
**Why the obvious check misses it.** The question "do we validate hits?" gets
|
||||
answered by looking at the hit path, which is full of plausible machinery. The
|
||||
right question is one level down and is about a subsystem nobody expects to find
|
||||
by reading weapon code: does the server store per-tick positions at all?
|
||||
|
||||
**Symptom.** "Add server validation" is scoped as a patch and turns out to be a
|
||||
subsystem. Partial attempts produce shots that visibly miss under latency, which
|
||||
is then treated as a regression and reverted.
|
||||
|
||||
**Detect.** Ask whether the machinery exists at all, before reading any weapon
|
||||
code:
|
||||
|
||||
```bash
|
||||
rg -n "FSavedMove|ServerMove|CompressedFlags|SavedMoves|PositionHistory" Source/
|
||||
```
|
||||
|
||||
Empty output means there is no lag compensation and no custom movement
|
||||
prediction, and therefore no basis for hit validation in principle. On the
|
||||
measured reference this returned nothing at all, which explains every other
|
||||
finding in this file about client trust: they are consequences of one absent
|
||||
subsystem, not independent oversights.
|
||||
|
||||
**Guardrail.** Decide the hit-registration model — server trace, lag-compensated
|
||||
rewind, or client-reported — before the weapon system exists, and record the
|
||||
decision with its threat model. Interim sanity checks (max range, line of sight
|
||||
from the replicated position, fire rate, team) raise cheat cost without full
|
||||
rewind and are worth having in every model.
|
||||
|
||||
---
|
||||
|
||||
### NA-09 - Fail-open permission check
|
||||
|
||||
**Mechanism.** A check returns "allowed" on an indeterminate or error result. The
|
||||
error branch was written to avoid blocking legitimate play during development.
|
||||
|
||||
**Why it is silent.** The indeterminate case does not arise in normal play. Both
|
||||
participants are always fully initialised, always have a PlayerState, always have
|
||||
a resolved team. The permissive branch is dead code for the entire duration of
|
||||
every test.
|
||||
|
||||
**Why the obvious check misses it.** The function has a check, the check is
|
||||
correct in the determinate cases, and unit tests cover those cases. Reviewing the
|
||||
error branch requires asking "what if this returns the error value?", and the
|
||||
answer during development is "it doesn't".
|
||||
|
||||
**Symptom.** Rules bypassed exactly in the edge cases nobody tested — early
|
||||
initialisation, missing PlayerState, a disconnecting client, an actor spawned
|
||||
before its owner. For a friendly-fire check this means damage applied where team
|
||||
rules should have prevented it.
|
||||
|
||||
**Detect.** Find the error branches of permission checks and read their return
|
||||
value:
|
||||
|
||||
```bash
|
||||
rg -n -B3 -A3 "InvalidArgument|Indeterminate|Unknown" Source/ \
|
||||
| rg "return true|return .*Allow|bCanDamage\s*=\s*true"
|
||||
rg -n "//\s*@?TODO.*temporar" -i Source/
|
||||
```
|
||||
|
||||
The second command is worth running on its own: a permissive branch marked
|
||||
temporary is the highest-confidence instance of this class.
|
||||
|
||||
**Guardrail.** Failure to determine must deny. Make the error branch explicit and
|
||||
restrictive, and if a permissive default is genuinely needed during development,
|
||||
gate it on a build configuration so it cannot ship.
|
||||
|
||||
---
|
||||
|
||||
### NA-10 - Cheat RPC surface in shipping
|
||||
|
||||
**Mechanism.** `Server` RPCs taking arbitrary command strings are declared outside
|
||||
any build-configuration guard. The implementation may be neutralised; the surface
|
||||
is not.
|
||||
|
||||
**Why it is silent.** The commands are useful, work as intended, and are reached
|
||||
through a UI that only appears in development. Hiding the entry point in the UI
|
||||
feels like removing it, and the difference never shows up in play.
|
||||
|
||||
**Why the obvious check misses it.** The guard that reviewers look for is usually
|
||||
present *somewhere* — around the cheat manager, around the console UI, around the
|
||||
implementation body. Confirming the RPC declaration itself is unguarded means
|
||||
checking a header for the absence of a preprocessor directive, which is not
|
||||
something a code search naturally expresses.
|
||||
|
||||
**Symptom.** A remote console in the shipped binary. Any client that can encode
|
||||
the RPC can execute arbitrary cheat strings on the server.
|
||||
|
||||
**Detect.** Check whether the declarations sit inside a build guard:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*Server[^)]*\)" -A2 Source/ | rg -i "cheat|exec|command"
|
||||
rg -n -B8 "void Server(Cheat|Exec)" Source/ | rg "#if|UE_BUILD_SHIPPING"
|
||||
```
|
||||
|
||||
An empty second result with a non-empty first is the finding. Note the
|
||||
distinction the recipe forces: whether the *implementation* is neutralised in a
|
||||
shipping configuration is a separate question that only a packaged build answers,
|
||||
and it should stay marked as open until someone answers it.
|
||||
|
||||
**Guardrail.** Wrap declaration *and* implementation in `#if !UE_BUILD_SHIPPING`
|
||||
or an equivalent cheat-manager guard. Hiding the UI is not removal.
|
||||
|
||||
---
|
||||
|
||||
## Replication shape
|
||||
|
||||
### NA-11 - Multicast used where state belongs
|
||||
|
||||
**Mechanism.** A one-shot multicast conveys a fact that a late joiner needs. It
|
||||
reaches everyone connected and relevant at the moment it fires.
|
||||
|
||||
**Why it is silent.** For everyone present, it works. The fact arrives, the visual
|
||||
plays, the state is right. The failure is scoped precisely to players who were not
|
||||
there, and those players see a plausible world — just the wrong one.
|
||||
|
||||
**Why the obvious check misses it.** Testing is done by players who join before
|
||||
the match starts, which is the configuration in which the bug cannot occur.
|
||||
Joining mid-match is a separate deliberate test that has to be designed.
|
||||
|
||||
**Symptom.** Players who join, or become relevant, after the event see the wrong
|
||||
state permanently. It never self-corrects, because nothing re-sends an event.
|
||||
|
||||
**Detect.** Enumerate multicasts and ask what each one carries:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*NetMulticast" -A3 Source/
|
||||
```
|
||||
|
||||
For each, answer: if a player joined one second after this fired, what would they
|
||||
see? Then run the test — join mid-round and inspect the state of every feature.
|
||||
This is a runtime check by nature; the grep only produces the list to test.
|
||||
|
||||
**Guardrail.** Anything a late joiner must know is replicated state. Multicasts
|
||||
carry transient presentation only — a sound, an impact, a one-frame effect whose
|
||||
absence is invisible a second later.
|
||||
|
||||
---
|
||||
|
||||
### NA-12 - Reliable RPC overflow
|
||||
|
||||
**Mechanism.** Frequent or per-frame reliable RPCs fill the reliable queue.
|
||||
|
||||
**Why it is silent.** The queue is generous. At two players on a local network it
|
||||
never fills. The cost scales with player count, latency and packet loss
|
||||
simultaneously, so it stays invisible across the entire range of configurations
|
||||
used during development.
|
||||
|
||||
**Why the obvious check misses it.** Each individual call site is defensible —
|
||||
one reliable RPC is nothing. The failure is a sum over call sites and time, and no
|
||||
review reads all of them together.
|
||||
|
||||
**Symptom.** Clients disconnect under load, typically first observed in a
|
||||
playtest at full player count, and usually attributed to the network rather than
|
||||
to a specific call site.
|
||||
|
||||
**Detect.** Count reliable RPCs and look for the ones a client can trigger
|
||||
repeatedly:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*\bReliable\b" Source/ | wc -l
|
||||
rg -n "UFUNCTION\([^)]*\bReliable\b" -A3 Source/ \
|
||||
| rg -i "tick|update|move|aim|look|input|cosmetic"
|
||||
```
|
||||
|
||||
The second command finds reliable calls on names that imply frequency. Confirm
|
||||
with a load test at maximum player count and 2% packet loss, watching for
|
||||
queue-overflow disconnects.
|
||||
|
||||
**Guardrail.** Reliable is a budget, not a default. Cosmetic, frequent or
|
||||
idempotent calls are unreliable; rate-limit anything a client can call in a loop.
|
||||
|
||||
---
|
||||
|
||||
### NA-13 - Incomplete FastArray callbacks
|
||||
|
||||
**Mechanism.** `PreReplicatedRemove`, `PostReplicatedAdd` and
|
||||
`PostReplicatedChange` are partially implemented. A derived cache or local
|
||||
listener list is updated in some of them and not others.
|
||||
|
||||
**Why it is silent.** The server is always right, because the server never runs
|
||||
the callbacks that are missing — it mutates the array directly. Only clients
|
||||
apply deltas, so only clients drift, and they drift into a state that is
|
||||
internally consistent.
|
||||
|
||||
**Why the obvious check misses it.** All three callbacks exist as overrides, so a
|
||||
structural review finds a complete-looking implementation. The empty one has a
|
||||
valid signature and compiles. Server-side tests pass because the defect cannot
|
||||
exist on the server.
|
||||
|
||||
**Symptom.** Client-side desync with no symptom on the server — the hardest class
|
||||
of bug to attribute, because every server-side log and dump shows correct state.
|
||||
|
||||
**Detect.** Compare the three callbacks per FastArray type:
|
||||
|
||||
```bash
|
||||
rg -n "FFastArraySerializer" -l Source/ | while read -r f; do
|
||||
echo "== $f"
|
||||
rg -n -A4 "Pre(ReplicatedRemove)|Post(ReplicatedAdd|ReplicatedChange)" "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Read the bodies. An empty body next to two implemented ones is the finding.
|
||||
On the measured reference one replicated structure had an empty removal hook and
|
||||
no callers at all — harmless while unused, a silent desync the moment it is
|
||||
adopted, which is the worst possible time to discover it.
|
||||
|
||||
**Guardrail.** Implement all three or none. Assert cache/array consistency in
|
||||
development builds so the drift becomes an error rather than a mystery.
|
||||
|
||||
---
|
||||
|
||||
### NA-14 - Relevancy disabled by an inherited constant
|
||||
|
||||
**Mechanism.** A net cull distance chosen for a small arena is inherited by a
|
||||
large map. Beyond the map's own size, the setting disables relevancy culling
|
||||
entirely.
|
||||
|
||||
**Why it is silent.** It is not a bug on the map it was chosen for — it is the
|
||||
correct setting there, and it makes gameplay better. The value carries forward
|
||||
because it is in a base class, and base-class defaults are not re-examined when a
|
||||
new map is added.
|
||||
|
||||
**Why the obvious check misses it.** The number is a squared distance, which
|
||||
makes it unreadable at a glance: nine hundred million does not announce itself as
|
||||
three hundred metres. Reviewers see a tuned-looking constant and move on.
|
||||
|
||||
**Symptom.** Bandwidth and server CPU scale with total actor count instead of
|
||||
local density. Discovered at player-count scale-up, late, when the fix requires
|
||||
re-verifying every gameplay assumption about who can see whom.
|
||||
|
||||
**Detect.** Find the value and convert it:
|
||||
|
||||
```bash
|
||||
rg -n "SetNetCullDistanceSquared|NetCullDistanceSquared\s*=" Source/ Config/
|
||||
```
|
||||
|
||||
Take the square root of each result and compare it to the largest map's diagonal.
|
||||
On the measured reference the value corresponded to a 300 m radius on arena maps —
|
||||
correct there, fatal if inherited into an open world.
|
||||
|
||||
**Guardrail.** Treat net cull distance as a per-class gameplay decision with a
|
||||
stated map-size assumption written next to it, not as an inherited performance
|
||||
knob.
|
||||
|
||||
---
|
||||
|
||||
### NA-15 - Replication graph enabled late
|
||||
|
||||
**Mechanism.** The graph implementation and its routing configuration exist, and
|
||||
the graph ships disabled. The routing is therefore never exercised.
|
||||
|
||||
**Why it is silent.** Disabled means the default replication path runs, and the
|
||||
default path is correct. The configuration file looks maintained. Nothing
|
||||
degrades, because nothing is using it.
|
||||
|
||||
**Why the obvious check misses it.** The presence of a configured routing table is
|
||||
read as evidence that the feature is in use. Confirming otherwise means finding
|
||||
the one boolean that disables the whole thing, which lives in a settings class
|
||||
rather than next to the routing.
|
||||
|
||||
**Symptom.** Enabling it changes which actors clients see, breaking gameplay code
|
||||
that had silently assumed universal relevancy. The breakage appears far from the
|
||||
change, in features nobody associated with replication.
|
||||
|
||||
**Detect.** Check the enable flag before trusting the routing:
|
||||
|
||||
```bash
|
||||
rg -n "bDisableReplicationGraph|bEnableReplicationGraph" Source/ Config/
|
||||
rg -n "ClassNodeMapping|RepGraph" Config/
|
||||
```
|
||||
|
||||
A disabled flag beside a populated routing table means the routing is untested
|
||||
configuration, not working configuration.
|
||||
|
||||
**Guardrail.** Enable it early or treat it as unavailable. Re-verify every
|
||||
relevancy assumption at the moment it is turned on, and treat that as a project
|
||||
milestone rather than a settings change.
|
||||
|
||||
---
|
||||
|
||||
### NA-16 - Wrong ability-system replication mode for AI
|
||||
|
||||
**Mechanism.** `Mixed` is inherited for AI-controlled actors. `Mixed` replicates
|
||||
gameplay effects to the owning client — and an AI has no owning client to
|
||||
receive them.
|
||||
|
||||
**Why it is silent.** It is correct for player-controlled actors, which is where
|
||||
it was chosen, and it is not wrong for AI either — merely wasteful. Waste has no
|
||||
symptom until bandwidth is the constraint.
|
||||
|
||||
**Why the obvious check misses it.** All three modes are one-line calls in
|
||||
constructors spread across different classes. There is no place where the choice
|
||||
is expressed as a policy, so nobody reviews the set of choices together.
|
||||
|
||||
**Symptom.** Bandwidth spent replicating gameplay effects nobody reads. Shows up
|
||||
as an unexplained baseline in profiling, attributed to player count.
|
||||
|
||||
**Detect.** List every mode choice in one place:
|
||||
|
||||
```bash
|
||||
rg -n "SetReplicationMode|EGameplayEffectReplicationMode::" Source/ \
|
||||
| sed 's/.*EGameplayEffectReplicationMode:://' | sort | uniq -c
|
||||
```
|
||||
|
||||
On the measured reference all three ability-system components used `Mixed`, with
|
||||
`Minimal` and `Full` absent from the module entirely — a default applied
|
||||
uniformly rather than a decision made three times.
|
||||
|
||||
**Guardrail.** `Mixed` for player-controlled, `Minimal` for AI, `Full` only with a
|
||||
written reason. Put the three choices in one comment block so they are reviewed
|
||||
as a policy.
|
||||
|
||||
---
|
||||
|
||||
## Lifecycle
|
||||
|
||||
### NA-17 - Readiness gate that assumes a controller
|
||||
|
||||
**Mechanism.** An initialisation chain waits for a controller. Simulated proxies
|
||||
on remote clients never receive one.
|
||||
|
||||
**Why it is silent.** The host is fine. The listen-server's own pawn has a
|
||||
controller, and so does every pawn in a single-process test. The stall happens
|
||||
only on machines that are not running the authority, which is the configuration
|
||||
least often exercised.
|
||||
|
||||
**Why the obvious check misses it.** The gate reads as obviously correct — of
|
||||
course initialisation waits for the controller. The missing knowledge is not in
|
||||
the code, it is that simulated proxies exist and never satisfy the condition. A
|
||||
reviewer who knows that spots it instantly; one who does not cannot find it by
|
||||
reading more carefully.
|
||||
|
||||
**Symptom.** Remote clients see uninitialised characters forever. Reproduces only
|
||||
with a dedicated server, so it is often first reported by QA and not reproducible
|
||||
by the developer.
|
||||
|
||||
**Detect.** Find readiness conditions that mention a controller:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A4 "GetController\(\)|Controller\s*!=\s*nullptr|IsPlayerControlled" \
|
||||
Source/ --glob "*Init*" --glob "*Extension*" --glob "*Ready*"
|
||||
```
|
||||
|
||||
For each, ask what a simulated proxy does at that line. The measured reference
|
||||
handles this correctly — its init-state gate permits progression without a
|
||||
controller — and is worth reading precisely because the naive version of the same
|
||||
gate deadlocks every simulated proxy while working perfectly in PIE.
|
||||
|
||||
**Guardrail.** An explicit named state chain, with the no-controller case handled
|
||||
deliberately rather than by omission. Write down which states a simulated proxy
|
||||
reaches.
|
||||
|
||||
---
|
||||
|
||||
### NA-18 - Addition without verified removal
|
||||
|
||||
**Mechanism.** Granted abilities, applied effects, registered listeners and
|
||||
spawned cues accumulate across death/respawn or feature activation cycles,
|
||||
because the addition path was built and the removal path was assumed.
|
||||
|
||||
**Why it is silent.** One cycle is fine. Two cycles produce a duplicate, and a
|
||||
duplicate of an idempotent-looking effect is usually invisible. The cost grows
|
||||
linearly with session length, and development sessions are short.
|
||||
|
||||
**Why the obvious check misses it.** Both halves exist in most cases — there is a
|
||||
removal function, it is called, and it removes *something*. Whether it removes
|
||||
exactly what was added requires holding the handle, and a review that sees
|
||||
`Remove` next to `Add` reasonably concludes the pair is complete.
|
||||
|
||||
**Symptom.** Duplicated effects after the second respawn; leaks that only appear
|
||||
in long sessions; a second feature activation that adds a second copy of every
|
||||
widget.
|
||||
|
||||
**Detect.** Compare addition and removal sites per resource type, then test:
|
||||
|
||||
```bash
|
||||
rg -n "GiveAbility|ApplyGameplayEffect|AddLoadedGameplayCuePath|RegisterListener" \
|
||||
Source/ | wc -l
|
||||
rg -n "ClearAbility|RemoveActiveGameplayEffect|RemoveLoadedGameplayCuePath|Unregister" \
|
||||
Source/ | wc -l
|
||||
```
|
||||
|
||||
Unequal counts are a starting point, not a verdict. The verdict is runtime:
|
||||
baseline, activate, deactivate, compare to baseline, reactivate — on both server
|
||||
and client, twice in one session.
|
||||
|
||||
**Guardrail.** Every dynamic addition retains an ownership handle and an exact
|
||||
rollback, tested on both server and client, twice in one session. Assert on the
|
||||
removal count where the API returns one.
|
||||
|
||||
---
|
||||
|
||||
### NA-19 - Prediction without a correction path
|
||||
|
||||
**Mechanism.** A client-predicted value has no defined behaviour when the server
|
||||
disagrees. The prediction key reconciles ordering; nothing reconciles the value.
|
||||
|
||||
**Why it is silent.** The server rarely disagrees on a local network with an
|
||||
honest client. Disagreement requires latency, packet loss or a rejected action,
|
||||
and the development environment supplies none of them.
|
||||
|
||||
**Why the obvious check misses it.** The prediction machinery is present and
|
||||
correct, and prediction keys are a real feature that really works. Reviewing
|
||||
"is this predicted properly?" finds a correct implementation of prediction. The
|
||||
missing piece is a design question — what does the player see when we were
|
||||
wrong? — that has no code location to be missing from.
|
||||
|
||||
**Symptom.** The client keeps a wrong value indefinitely. UI and gameplay diverge
|
||||
with no error anywhere, and the player's experience is that the game lied to them
|
||||
once and never corrected.
|
||||
|
||||
**Detect.** Force disagreement and watch:
|
||||
|
||||
```bash
|
||||
rg -n "ScopedPredictionWindow|PredictionKey|GetPredictionKeyForNewAction" Source/
|
||||
```
|
||||
|
||||
For each predicted value, run the paid test: introduce 150 ms latency, make the
|
||||
server reject the action, and observe the client for ten seconds. A value that
|
||||
does not converge has no correction path. This one cannot be settled statically.
|
||||
|
||||
**Guardrail.** Define, per predicted value, what the correction looks like to the
|
||||
player. A prediction key reconciles ordering, not entitlement.
|
||||
|
||||
---
|
||||
|
||||
### NA-20 - PIE treated as a network test
|
||||
|
||||
**Mechanism.** The listen-server host shares memory with its client, so the
|
||||
serialisation, ordering and relevancy boundaries that produce networked defects do
|
||||
not exist during the test.
|
||||
|
||||
**Why it is silent.** PIE multiplayer looks like multiplayer. Two windows, two
|
||||
players, two views of one world. Everything that works there works convincingly,
|
||||
including the things that only work because the boundary is absent.
|
||||
|
||||
**Why the obvious check misses it.** The check *is* the test, and the test passes.
|
||||
There is no artefact to inspect and no code to review; the defect is in the
|
||||
choice of configuration, which is made once, early, by whoever set up the
|
||||
workflow, and never revisited.
|
||||
|
||||
**Symptom.** Features ship having never run in the configuration they will run in.
|
||||
Entries NA-01, NA-04, NA-11, NA-13, NA-17 and NA-19 above cannot occur in PIE at
|
||||
all.
|
||||
|
||||
**Detect.** Audit the test configuration rather than the code:
|
||||
|
||||
```bash
|
||||
rg -n "PlayNetMode|RunUnderOneProcess|bLaunchSeparateServer" Config/ Saved/Config/
|
||||
```
|
||||
|
||||
If the editor is configured for a listen server in one process, no networked
|
||||
feature in the project has been tested. Confirm by running one known-networked
|
||||
feature under a dedicated server with a remote client and comparing behaviour.
|
||||
|
||||
**Guardrail.** Dedicated server plus two remote clients, one with 150 ms latency
|
||||
and 2% packet loss, as the minimum configuration for accepting a networked
|
||||
feature — including a mid-match joiner.
|
||||
@@ -0,0 +1,178 @@
|
||||
# Patterns: authority in a measured reference product
|
||||
|
||||
A worked example of the authority model from `SKILL.md`, against one specific
|
||||
reference project read as source. It is here for one reason: the failure modes in
|
||||
[failure-modes.md](failure-modes.md) are easier to believe, and much easier to
|
||||
apply, when you can see how they cluster in a real codebase.
|
||||
|
||||
**Read this first.** The project audited below is an outstanding architectural
|
||||
reference and a poor network-security reference. Those are independent axes. None
|
||||
of what follows is a criticism of the sample — a learning project that ships no
|
||||
product has no reason to build lag compensation — but copying its structure
|
||||
without noticing its trust boundaries is a real and common accident, and it is the
|
||||
accident this skill exists to prevent.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not answerable without a dedicated-server build, and left
|
||||
open rather than guessed.
|
||||
|
||||
---
|
||||
|
||||
## 1. The headline: one absence explains almost everything
|
||||
|
||||
No custom saved-move type, no custom server-move path, no compressed-flags
|
||||
extension, and no historical position storage anywhere in the project
|
||||
**[measured]**.
|
||||
|
||||
Without stored per-tick positions the server has nothing to rewind to, and
|
||||
therefore **cannot** validate a client's hit claim even in principle
|
||||
**[derived]**. The client trust documented below is not carelessness at individual
|
||||
call sites; it is the necessary consequence of a subsystem that was never built.
|
||||
|
||||
This is the single most valuable thing to carry away from the audit, and it
|
||||
generalises: **before reading a weapon system, ask whether the machinery that
|
||||
would make validation possible exists at all.** If it does not, every "missing
|
||||
check" you find downstream is one symptom of one decision, not a list of
|
||||
independent bugs. Recipe: NA-08.
|
||||
|
||||
Practical implication for anyone adopting such code: "add server validation to the
|
||||
weapon system" is not a patch. It is building the rewind subsystem first.
|
||||
|
||||
---
|
||||
|
||||
## 2. Hit registration was client-authored end to end
|
||||
|
||||
Four findings, each of which reads as an isolated defect and is not:
|
||||
|
||||
| What was measured | Consequence |
|
||||
|---|---|
|
||||
| The targeting trace early-returns unless locally controlled — the server never traces **[measured]** | The client decides what was hit |
|
||||
| A target-data validity flag is a `const bool` initialised to `true`, consumed twice, with no computation between **[measured]** | The code path reads as validated in review; NA-06 |
|
||||
| Distance falloff is computed from a client-supplied trace origin; the damage multiplier is taken from a client-supplied surface material **[measured]** | A modified client controls falloff and multiplier directly **[derived]**; NA-05 |
|
||||
| A `bIsSimulated` parameter appears at ten sites, is passed a literal at its single call site, and is branched on nowhere; a hit-replacement helper is never called **[measured]** | Attachment points for the rewind subsystem that was never built **[derived]**; NA-07 |
|
||||
|
||||
One check in that path was genuinely server-side: the team comparison guarding
|
||||
whether damage may be caused at all **[measured]**. That detail matters more than
|
||||
it looks. **The model was incomplete, not absent** — which is the harder case,
|
||||
because the presence of one real server-side check makes the surrounding code read
|
||||
as server-side too.
|
||||
|
||||
---
|
||||
|
||||
## 3. Three mechanisms that look equivalent and are not
|
||||
|
||||
One equipment component demonstrated all three in a single file **[measured]**:
|
||||
|
||||
| Declaration | Actually enforces |
|
||||
|---|---|
|
||||
| `UFUNCTION(Server, Reliable, BlueprintCallable)` | routing only — no `WithValidation` |
|
||||
| `UFUNCTION(BlueprintAuthorityOnly, ...)` | blocks Blueprint callers; no effect on C++ |
|
||||
| plain C++ writes elsewhere in the same class | nothing |
|
||||
|
||||
`BlueprintAuthorityOnly` is an editor affordance. It appears in review as an
|
||||
authority guard and is not one **[derived]**. Across the whole project the
|
||||
specifier appeared at five sites in three subsystems, none of which carried an
|
||||
authority check in the guarded body **[measured]**. Recipe: NA-03.
|
||||
|
||||
Unguarded state transitions followed the same shape: the death-state writers were
|
||||
public, virtual, and unguarded, with the only barrier being a server-initiated
|
||||
activation policy on one ability **[measured]**. The guard belongs in the mutator,
|
||||
not in one of its callers **[derived]** — the call-site list is not a stable
|
||||
property, and virtual functions mean a subclass in a feature plugin inherits the
|
||||
hole. Recipe: NA-04.
|
||||
|
||||
---
|
||||
|
||||
## 4. Cheat surface and error handling
|
||||
|
||||
Two arbitrary-string cheat RPCs were declared outside any build-configuration
|
||||
guard **[measured]**. Their validators return `true` unconditionally
|
||||
**[measured]** — which is defensible for a cheat RPC and is also exactly the shape
|
||||
NA-02 looks for, so the recipe finds them first and you must read what you found.
|
||||
Whether the implementations are neutralised in a shipping configuration is
|
||||
**[open]** without a packaged build; the RPC surface itself is unconditional.
|
||||
|
||||
A team-comparison helper returned a permissive result on an invalid-argument
|
||||
outcome, carrying a comment marking it temporary **[measured]**. Since that
|
||||
function gates friendly fire, the failure mode is damage applied where team rules
|
||||
should have prevented it **[derived]**.
|
||||
|
||||
Rule extracted, and it is worth more than the finding: **failure to determine must
|
||||
deny.** Recipe: NA-09.
|
||||
|
||||
---
|
||||
|
||||
## 5. Replication configuration
|
||||
|
||||
| Setting | Measured | Reading |
|
||||
|---|---|---|
|
||||
| Ability-system replication mode | `Mixed` on all three ability-system components; `Minimal` and `Full` absent from the module **[measured]** | Correct for player-controlled actors; wasteful when inherited by AI **[derived]**; NA-16 |
|
||||
| Net cull distance | A squared value corresponding to a 300 m radius, set in the character base class **[measured]** | Every character relevant to every client at all times **[derived]** — correct for a small arena, fatal if inherited into a large map; NA-14 |
|
||||
| Replication graph | Ships disabled, with class routing configured **[measured]** | The routing is untested configuration, not working configuration **[derived]**; NA-15 |
|
||||
| A replicated fast-array | Empty removal callback, no callers in the project **[measured]** | Harmless while unused; a silent client-side desync the moment it is adopted **[derived]**; NA-13 |
|
||||
|
||||
The last row is the one worth pausing on. An unused structure with an incomplete
|
||||
callback set is invisible to every kind of testing, because nothing exercises it.
|
||||
It becomes a defect at adoption time — the moment when the person adopting it has
|
||||
the least context to recognise what they inherited.
|
||||
|
||||
---
|
||||
|
||||
## 6. What the reference got right
|
||||
|
||||
For balance, and because these are the parts worth copying directly:
|
||||
|
||||
1. **An explicit init-state chain** instead of timing assumptions. Its gate
|
||||
permits progression for actors without a controller, which is what allows
|
||||
simulated proxies to initialise at all **[measured]**. The naive version of the
|
||||
same gate deadlocks every simulated proxy on remote clients while working
|
||||
perfectly in PIE **[derived]** — study the correct one before writing your own.
|
||||
Context: NA-17.
|
||||
2. **A replication mode chosen deliberately** for player ability-system components
|
||||
rather than left at the engine default **[measured]**.
|
||||
3. **Replicate intent, realise presentation locally** — the cosmetic system sends
|
||||
which parts to wear, not the assembled result **[measured]**.
|
||||
4. **Correct teardown where it exists**: one pawn-extension path checks avatar
|
||||
identity before cancelling abilities, clearing input and removing cues
|
||||
**[measured]**. Note the shape — identity check first, then reverse-order
|
||||
removal.
|
||||
5. **The team check in the damage execution is genuinely server-side**
|
||||
**[measured]**. The model is not absent, it is incomplete.
|
||||
|
||||
---
|
||||
|
||||
## 7. Questions the audit could not answer
|
||||
|
||||
Left open deliberately. Each requires a dedicated-server or packaged build, which
|
||||
was out of scope:
|
||||
|
||||
1. Whether the cheat RPC implementations are neutralised in a shipping
|
||||
configuration.
|
||||
2. Actual bandwidth per client, and whether the relevancy radius holds at full
|
||||
player count.
|
||||
3. Behaviour of the replication graph when enabled — routing completeness,
|
||||
dormancy usage.
|
||||
4. Whether client and server tag registries match when feature plugins load
|
||||
asymmetrically.
|
||||
5. Real reconciliation behaviour of ability prediction under packet loss.
|
||||
|
||||
An audit that answers questions it cannot answer is worth less than one that
|
||||
enumerates them. If you adopt this material, these five are yours to close.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source in
|
||||
a single workspace. Source addresses stay in the research archive that produced
|
||||
this skill; each `NA-` identifier above resolves back to the audited location
|
||||
there. Numbers without an author are worth less than no numbers, so: this is where
|
||||
these came from, and they can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Every claim above refers to one project on one engine version. These are examples
|
||||
and failure evidence, not guarantees about other engine versions or other samples.
|
||||
Verify version-sensitive claims against your target engine before acting on them.
|
||||
In particular: PIE cannot reproduce most of the failure modes described here,
|
||||
because a listen-server host shares memory with its client.
|
||||
@@ -0,0 +1,343 @@
|
||||
---
|
||||
name: ue-reference-project-adoption
|
||||
description: >-
|
||||
Adopt architecture from an Unreal Engine reference or sample project (a
|
||||
first-party starter game, an engine template, an internal legacy title)
|
||||
without inheriting its shortcuts: classify each borrowed element as defect,
|
||||
showcase stub, deliberate sample-scope narrowing, architecture tax or
|
||||
foreign-product legacy, then decide copy/close/skip per element. Use when
|
||||
starting a project from a template, lifting a subsystem out of a sample,
|
||||
reviewing code that was copied from one, or explaining why a setting copied
|
||||
from the reference does nothing.
|
||||
---
|
||||
|
||||
# UE reference project adoption
|
||||
|
||||
The invariant:
|
||||
|
||||
> A reference project shows you a **shape**, not a **contract**. Anything you copy
|
||||
> without finding its consumer, its writer and its owner is a shape you now have to
|
||||
> finish yourself.
|
||||
|
||||
Worked classification, tax analysis and vocabulary hygiene:
|
||||
[patterns](references/patterns.md).
|
||||
Nineteen detection recipes for silent failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting or reviewing borrowed architecture. Each
|
||||
entry names the mechanism, why it stays silent, why the obvious check misses it,
|
||||
and a command you can run against your own tree.
|
||||
|
||||
Related skills: `ue-architecture-guardrails`, `ue-modular-gameplay`,
|
||||
`ue-multiplayer-authority`, `ue-gameplay-tag-governance`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. The one thing that makes this hard
|
||||
|
||||
Broken code announces itself. It crashes, it logs, it fails a test.
|
||||
|
||||
An unfinished mechanism announces nothing. It returns a neutral value forever, and
|
||||
that neutral value is indistinguishable from "the feature is off right now".
|
||||
|
||||
So the ranking that matters during adoption is **not severity**. It is *probability
|
||||
of copying it without noticing*. Those are close to inverted: the most dangerous
|
||||
elements in a mature sample project are the ones that never misbehave.
|
||||
|
||||
### Severity is not copyability
|
||||
|
||||
| | Pattern is safe to copy | Pattern is unsafe to copy |
|
||||
|---|---|---|
|
||||
| **Fails loudly** | plain defect — fix the line, keep the design | — |
|
||||
| **Never fails** | foreign legacy — rename and keep | **stubs and scope narrowings** |
|
||||
| **Works as intended** | tax you can afford | tax sized for someone else |
|
||||
|
||||
Only the bottom-right cell costs a subsystem rewrite instead of a line edit.
|
||||
|
||||
---
|
||||
|
||||
## 2. Five categories, and the decision each one implies
|
||||
|
||||
Classify every element you intend to borrow. The category, not the symptom,
|
||||
determines what you do.
|
||||
|
||||
### C1 — Defect
|
||||
|
||||
A mistake the reference's authors would fix if they noticed.
|
||||
|
||||
**Decision:** fix the line; keep the surrounding design. A defect is not evidence
|
||||
against the pattern that contains it.
|
||||
|
||||
### C2 — Showcase stub
|
||||
|
||||
Declared, editable, sometimes even read — but with no writer, no consumer, or a
|
||||
hardcoded neutral result. Usually marked with a TODO, often not.
|
||||
|
||||
**Decision:** either implement it before exposing it, or delete the field. Never
|
||||
ship an editable property whose consumer is a literal constant. Designers will
|
||||
spend days tuning it.
|
||||
|
||||
### C3 — Deliberate sample-scope narrowing
|
||||
|
||||
The reference intentionally did not build a subsystem, because a sample does not
|
||||
need it. Nothing is broken. The reference is internally consistent.
|
||||
|
||||
**Decision:** this is the expensive one. Decide explicitly, at adoption time,
|
||||
whether your product needs the missing subsystem — and budget it. The failure mode
|
||||
is not "we copied a bug", it is "we shipped without noticing there was a hole".
|
||||
|
||||
### C4 — Architecture tax
|
||||
|
||||
Correct, working, and sized for the reference author's team, release cadence and
|
||||
platform matrix.
|
||||
|
||||
**Decision:** adopt by *need*, not by *completeness*. Copying an entire modular
|
||||
stack because it is there means paying for an answer to a question you do not have.
|
||||
|
||||
### C5 — Foreign-product legacy
|
||||
|
||||
Constants, names and comments carried over from a different title.
|
||||
|
||||
**Decision:** rename on entry. Harmless as code, corrosive as vocabulary — in a
|
||||
year, half the team will believe the foreign term is an engine concept.
|
||||
|
||||
---
|
||||
|
||||
## 3. Four tests, run before you copy anything
|
||||
|
||||
Cheapest first. Most stubs die at test 3.1.
|
||||
|
||||
### 3.1 Find the writer, not the reader
|
||||
|
||||
For any tunable field, search for **assignment**, not usage.
|
||||
|
||||
```bash
|
||||
rg -n "\bMyField\b" Source/ # declaration plus two reads: looks alive
|
||||
rg -n "\bMyField\b\s*(=|\+=|-=)" Source/ # one hit, and it is the declaration
|
||||
```
|
||||
|
||||
Watch the second line. A member declared as `float MyField = 0.f;` matches every
|
||||
assignment pattern you can write, so the naive search answers "there is a writer"
|
||||
in exactly the case where there is none. Subtract the declaration before you
|
||||
conclude anything:
|
||||
|
||||
```bash
|
||||
rg -n "\bMyField\b\s*(=|\+=|-=)" Source/ \
|
||||
| rg -v "\b(bool|u?int\d+|float|double|F[A-Z]\w+|T\w+<)\s+MyField\b"
|
||||
```
|
||||
|
||||
A field that is read and compared but never written is worse than an unused field:
|
||||
reading the code convinces you it is wired up. This is the single highest-yield
|
||||
check in this skill — and, as the two commands above show, the one whose first
|
||||
draft most often lies to you.
|
||||
|
||||
### 3.2 Find the comparison, not the plumbing
|
||||
|
||||
For any field named `Priority`, `Order`, `Weight`, `Rank`, `SortKey`: search for its
|
||||
use inside a sort, predicate or comparison. Being passed through six functions and
|
||||
stored in a struct means nothing.
|
||||
|
||||
If there is no comparison, the effective ordering is *registration order*, which in
|
||||
a plugin-based project means *feature activation order* — non-deterministic from
|
||||
the designer's point of view.
|
||||
|
||||
### 3.3 Ask who fills this in
|
||||
|
||||
Any of these is a stub or a narrowing until proven otherwise:
|
||||
|
||||
- an empty container next to a TODO;
|
||||
- `const bool bSomething = false;` with the real expression in a comment;
|
||||
- `const bool bIsValid = true;` consumed by a validation branch;
|
||||
- a parameter threaded through several call layers that is never branched on;
|
||||
- a function that exists, is public, and has zero callers in the codebase.
|
||||
|
||||
The distinction between C2 and C3 is only whether the author ever intended to fill
|
||||
it in. Both cost you the same.
|
||||
|
||||
### 3.4 Ask what the adversary does
|
||||
|
||||
For anything on a trust boundary, ask: *what stops the other machine from lying
|
||||
about this fact?* If the answer is "nothing", you are looking at C3, and the cost
|
||||
of closing it is yours to price.
|
||||
|
||||
### 3.5 The paid test
|
||||
|
||||
When static reading is ambiguous, change the setting and measure whether anything
|
||||
moves. Expensive, and the only test that cannot be fooled. Use a written protocol
|
||||
with a restore point, one knob at a time.
|
||||
|
||||
---
|
||||
|
||||
## 4. Sample-scope narrowings a product must close
|
||||
|
||||
These are the holes that reference projects most commonly leave, in rough order of
|
||||
how much they cost to close later:
|
||||
|
||||
| Narrowing | Why a sample omits it | What it costs a product |
|
||||
|---|---|---|
|
||||
| No lag compensation / server rewind | needs a real server, real latency, real test rig | Cannot be retrofitted cheaply — it dictates how abilities, tracing and damage are structured |
|
||||
| Replication graph present but disabled | tuning it requires production-scale player counts | Traffic and CPU limits discovered in beta |
|
||||
| No tag/asset redirect discipline | a sample never renames anything after release | Every rename becomes a breaking change with no migration path |
|
||||
| Authority checks missing on state writers | single-process PIE never exercises the boundary | Silent state divergence, cheat surface |
|
||||
| Fail-open defaults on unknown input | sample never produces unknown input | Security and gameplay decisions default to "allowed" |
|
||||
| Preload/streaming machinery bypassed by config | sample content fits in memory | Hitching that looks like an engine bug |
|
||||
| No teardown for dynamic additions | sample never unloads a feature | Leaks and duplicate registrations on the second load |
|
||||
|
||||
None of these is a bug report against the reference. Each is a line item in your
|
||||
adoption budget.
|
||||
|
||||
### The rule
|
||||
|
||||
> Every C3 narrowing you accept must be written down as an accepted risk with an
|
||||
> owner, or scheduled. "We did not notice" is the only unacceptable outcome.
|
||||
|
||||
---
|
||||
|
||||
## 5. Architecture tax: adopt need, not completeness
|
||||
|
||||
A reference project's modularity is an answer to a specific question — usually
|
||||
"how do several teams ship seasonal content into one binary without merge
|
||||
conflicts". If you do not have that question, you pay the cost and receive nothing.
|
||||
|
||||
Before adopting a modular stack, state the question it answers for *you*:
|
||||
|
||||
- more than one team shipping into the same binary?
|
||||
- content added after launch without a client patch?
|
||||
- game modes swapped at runtime by data?
|
||||
- platform-specific feature sets?
|
||||
|
||||
Zero yes answers means the stack is tax. One or two means adopt the layers that
|
||||
answer them and stop there.
|
||||
|
||||
### The C++/Blueprint boundary is also tax
|
||||
|
||||
A reference project's C++/BP split reflects its author's staffing. A codebase where
|
||||
C++ exposes overwhelmingly *calls* rather than *override points* assumes a strong
|
||||
C++ team is always available to extend it. Copy that boundary without that team and
|
||||
designers hit a wall on their first non-trivial request.
|
||||
|
||||
Measure both sides before copying the ratio: how many `BlueprintCallable` entry
|
||||
points versus how many `BlueprintImplementableEvent` and `BlueprintNativeEvent`
|
||||
extension points, and who on your team can add to the second column. Measured
|
||||
counts from one reference project are in [patterns](references/patterns.md).
|
||||
|
||||
### Measure the whole mount surface, not the primary content root
|
||||
|
||||
In a plugin-based project, plugins mount their own content roots. Any census,
|
||||
budget or memory metric computed over the primary game root alone is wrong by
|
||||
however much the plugins hold — in one measured reference, by roughly a factor of
|
||||
six. Enumerate all mount roots, not just the obvious one.
|
||||
|
||||
---
|
||||
|
||||
## 6. Foreign legacy hygiene
|
||||
|
||||
On entry, rename:
|
||||
|
||||
- constants and type names carrying another product's brand;
|
||||
- asset type identifiers whose names describe another game's content;
|
||||
- comments copied from a different class;
|
||||
- tags whose comment admits they are in the wrong namespace.
|
||||
|
||||
Do this at import time. After the first sprint it becomes archaeology, and the
|
||||
foreign vocabulary starts appearing in design documents.
|
||||
|
||||
---
|
||||
|
||||
## 7. Fork and drift discipline
|
||||
|
||||
If you copy code rather than depend on a plugin:
|
||||
|
||||
1. Record the reference version and commit at import time, in a file next to the
|
||||
copy.
|
||||
2. Keep imported code in a directory that is obviously imported.
|
||||
3. Record every intentional divergence with a one-line reason.
|
||||
4. Re-diff against the upstream on each engine upgrade; an unrecorded divergence is
|
||||
indistinguishable from an upstream fix you are about to lose.
|
||||
|
||||
Without this, engine upgrades turn every borrowed subsystem into a manual merge
|
||||
with no ground truth.
|
||||
|
||||
---
|
||||
|
||||
## 8. Adoption procedure
|
||||
|
||||
For each subsystem you intend to borrow:
|
||||
|
||||
1. **Name the question** it answers for your product. If you cannot, stop.
|
||||
2. **Read the consumer, not the declaration.** Confirm each field has a writer and
|
||||
each ordering field has a comparison.
|
||||
3. **List the trust boundaries** it crosses and what validates each one.
|
||||
4. **Classify every surprising element** as C1–C5.
|
||||
5. **Decide per category:** fix / implement or delete / budget and schedule / trim
|
||||
to need / rename.
|
||||
6. **Write down accepted C3 risks** with an owner.
|
||||
7. **Record the import version** and directory.
|
||||
8. **Add a test that would fail** if the narrowing you accepted became a bug.
|
||||
|
||||
---
|
||||
|
||||
## 9. Review checklist
|
||||
|
||||
- [ ] Every borrowed tunable field has a verified writer.
|
||||
- [ ] Every ordering field participates in an actual comparison.
|
||||
- [ ] No editable property is consumed by a hardcoded literal.
|
||||
- [ ] Every parameter threaded through layers is branched on somewhere.
|
||||
- [ ] Trust boundaries are enumerated; each has a named validator or an accepted
|
||||
risk.
|
||||
- [ ] Missing subsystems (C3) are written down, not discovered later.
|
||||
- [ ] Modular layers adopted map to a stated product need.
|
||||
- [ ] C++/BP boundary matches the team that has to live with it.
|
||||
- [ ] Metrics computed over all mount roots.
|
||||
- [ ] Foreign-product vocabulary renamed at import.
|
||||
- [ ] Import version and divergences recorded.
|
||||
- [ ] Teardown exists for every dynamic addition adopted.
|
||||
|
||||
---
|
||||
|
||||
## 10. When copying wholesale is right
|
||||
|
||||
Adopt as-is, without this ceremony, when **all** of the following hold:
|
||||
|
||||
- the element is self-contained and has no trust boundary;
|
||||
- you can name its consumer within one file;
|
||||
- it has no editable tuning surface;
|
||||
- replacing it later is a local edit, not a migration.
|
||||
|
||||
Small utilities, math helpers, container types and single-purpose components
|
||||
qualify. Subsystems, lifecycle machinery and anything that touches replication,
|
||||
persistence or the tag registry do not.
|
||||
|
||||
---
|
||||
|
||||
## 11. What this skill does not say
|
||||
|
||||
It does not say reference projects are bad guidance. A mature sample is the best
|
||||
available evidence for *composition* — how definition and instance separate, how
|
||||
readiness becomes an explicit chain, how presentation decouples from gameplay.
|
||||
|
||||
It says: the sample is evidence for the shape and silent about the finish. The
|
||||
references exist so the silence is enumerated instead of discovered in production.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The classification, the four tests and every entry under `references/` come from a
|
||||
line-by-line audit of Epic's Lyra Starter Game on Unreal Engine 5.6, read as source
|
||||
rather than run. Source addresses stay in the research archive that produced this
|
||||
skill; what ships is the detection recipe, because an address in someone else's
|
||||
tree is not something you can act on and a recipe is.
|
||||
|
||||
Each entry carries a stable identifier (`RA-01`, `RA-02`, …) that resolves back to
|
||||
the audited location in that archive. If you need the original address to settle a
|
||||
dispute, it exists and can be produced.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one specific reference project on one engine
|
||||
version in one workspace. It is a worked example of the classification, not a
|
||||
defect list to carry into other engine versions or other samples. Re-run the four
|
||||
tests against your own reference before trusting any specific claim.
|
||||
+739
@@ -0,0 +1,739 @@
|
||||
# Failure modes when adopting a reference project
|
||||
|
||||
Nineteen ways a borrowed subsystem stays quiet while doing nothing.
|
||||
|
||||
They share one property: **each is discovered late by construction.** Every entry
|
||||
here compiles, passes review, produces no warning and no log line. What is absent
|
||||
is a writer, a comparison, a validator or an owner — and absence does not raise.
|
||||
|
||||
Each entry gives you a `Detect` you can run against your own tree today. The
|
||||
recipes are deliberately crude: a grep that returns nothing where reads exist is
|
||||
worth more than a subtle static analysis nobody runs.
|
||||
|
||||
`RA-nn` identifiers are stable and resolve to the audited location in the research
|
||||
archive behind this skill.
|
||||
|
||||
Categories: **C1** defect · **C2** showcase stub · **C3** deliberate sample-scope
|
||||
narrowing. Architecture tax (C4) and foreign vocabulary (C5) are not failure modes
|
||||
and live in [patterns](patterns.md).
|
||||
|
||||
---
|
||||
|
||||
## C1 — Defects
|
||||
|
||||
Fix the line, keep the design. A defect is not evidence against the pattern that
|
||||
contains it.
|
||||
|
||||
### RA-01 - Tag picker filtered to a namespace that does not exist
|
||||
|
||||
**Mechanism.** An editable gameplay-tag property carries a
|
||||
`meta = (Categories = "Some.Root")` filter, and no tag in the project is declared
|
||||
under that root. The real tags live under a different prefix.
|
||||
|
||||
**Why it is silent.** The filter's only job is to narrow an editor dropdown. A
|
||||
filter that matches nothing produces an empty dropdown, which is visually identical
|
||||
to "no tags of this kind have been authored yet". Nothing compiles differently,
|
||||
nothing logs, and the property simply stays unset forever.
|
||||
|
||||
**Why the obvious check misses it.** The filter string is syntactically valid and
|
||||
the property is genuinely declared, read and used — every "is this wired up?" search
|
||||
answers yes. Nobody searches for the *filter value* as a tag prefix, because it
|
||||
reads like a category label rather than a key that has to resolve. In the measured
|
||||
reference this was one of only fourteen places in the entire project where tag input
|
||||
was constrained at all, so the one guardrail present was also the one misconfigured.
|
||||
|
||||
**Symptom.** A designer opens the dropdown, sees nothing, concludes the feature is
|
||||
unimplemented, and leaves the field empty. The system then behaves as if the
|
||||
designer chose "none".
|
||||
|
||||
**Detect.** Extract every filter value and confirm each resolves to declared tags:
|
||||
|
||||
```bash
|
||||
rg -o 'Categories\s*=\s*"([^"]*)"' -r '$1' Source/ Plugins/ | sort -u
|
||||
# then, per root R:
|
||||
rg -c "\"R\." Config/*.ini Source/
|
||||
```
|
||||
|
||||
A root with zero declarations is a dead filter.
|
||||
|
||||
**Guardrail.** A `Categories` filter is a guardrail only if something proves it
|
||||
resolves. Add a startup or CI assertion that every filter value matches at least one
|
||||
declared tag; an empty picker must be an error, not a shrug.
|
||||
|
||||
---
|
||||
|
||||
### RA-02 - A misspelling that works because every side repeats it
|
||||
|
||||
**Mechanism.** A tag or string key is misspelled in its declaration and in every
|
||||
consumer. The string is the join between config, code and platform overrides, so
|
||||
identical errors on both sides join correctly.
|
||||
|
||||
**Why it is silent.** Correctness for a string join means *all sides agree*, not
|
||||
*the string is a word*. Six identical misspellings across four files function
|
||||
exactly as six correct ones would. There is no side that could disagree.
|
||||
|
||||
**Why the obvious check misses it.** Searching for the correct spelling returns
|
||||
nothing, and an empty result reads as "this feature is not implemented here" rather
|
||||
than "it is spelled differently". The compiler never sees the literal as a name, and
|
||||
spell-checkers do not run over identifiers. The bug only becomes visible at the
|
||||
moment someone *fixes* it at one site.
|
||||
|
||||
**Symptom.** Nothing at all — until a cleanup pass corrects the spelling in one
|
||||
file. The join then breaks with no error, and the regression is attributed to
|
||||
whatever else shipped that week.
|
||||
|
||||
**Detect.** Find tag literals repeated as raw strings instead of referenced through
|
||||
one constant:
|
||||
|
||||
```bash
|
||||
rg -oN '"[A-Za-z][A-Za-z0-9]*(\.[A-Za-z][A-Za-z0-9]*){2,}"' Source/ Config/ \
|
||||
| sort | uniq -c | sort -rn | head -40
|
||||
```
|
||||
|
||||
Any literal appearing more than once is a join with no single source of truth.
|
||||
|
||||
**Guardrail.** One declaration site per tag; every other site references the
|
||||
constant. Then a rename is a compile error instead of a silent disconnection.
|
||||
|
||||
---
|
||||
|
||||
### RA-03 - A loop over a container that has no producer
|
||||
|
||||
**Mechanism.** A local container is declared next to a TODO explaining how it will
|
||||
one day be filled, then iterated immediately. Nothing ever fills it.
|
||||
|
||||
**Why it is silent.** Zero iterations is a legal, successful outcome. The function
|
||||
returns normally and reports success, because doing nothing to an empty set is
|
||||
indistinguishable from doing the work correctly.
|
||||
|
||||
**Why the obvious check misses it.** The code reads as a complete mechanism:
|
||||
declaration, loop, and a body that does real work. Review attention goes to whether
|
||||
the body is correct. The missing piece is one level up — the container has a
|
||||
consumer and no producer, and reviewers check consumers.
|
||||
|
||||
**Symptom.** A documented behaviour ("these entries are always loaded") reports zero
|
||||
at runtime, and is diagnosed as a content or configuration problem for days.
|
||||
|
||||
**Detect.** For every container that is iterated, look for a writer:
|
||||
|
||||
```bash
|
||||
rg -n "for\s*\(.*:\s*(\w+)\)" -r '$1' Source/ | sort -u > /tmp/iterated
|
||||
# per name N:
|
||||
rg -n "\bN\b\s*\.\s*(Add|Append|Emplace|Push|Insert|Reserve)" Source/
|
||||
```
|
||||
|
||||
Reads without writes mean the loop is decoration.
|
||||
|
||||
**Guardrail.** Do not ship a loop over a container that has no producer in the same
|
||||
translation unit or an obvious injection point. If the producer is future work, the
|
||||
loop is future work too.
|
||||
|
||||
---
|
||||
|
||||
### RA-04 - Progress arithmetic with inverted operands
|
||||
|
||||
**Mechanism.** A startup progress fraction computes each sub-step's contribution
|
||||
with the operands the wrong way round, so the reported value does not advance while
|
||||
a single long step runs.
|
||||
|
||||
**Why it is silent.** The result stays inside the valid range and still increases
|
||||
across step boundaries. It is a wrong number, not an invalid one, so no assertion
|
||||
and no clamp fires.
|
||||
|
||||
**Why the obvious check misses it.** On a developer machine with a warm cache the
|
||||
whole sequence finishes in under a second, and nobody watches one step long enough
|
||||
to see it stall. The code is exercised on every launch and observed on none.
|
||||
|
||||
**Symptom.** On a cold shader cache or a slow disk the bar freezes for a long time
|
||||
and players report a hang. The engineering response is to look for a deadlock,
|
||||
because the bar is trusted.
|
||||
|
||||
**Detect.** Two options. Force the slow path with whatever artificial-delay cvar the
|
||||
loading system provides and watch whether the value moves within a step. Or unit-test
|
||||
the progress function directly: feed it a fixed step count and assert the value is
|
||||
strictly increasing at every sub-step, not just at step boundaries.
|
||||
|
||||
**Guardrail.** Progress arithmetic gets a test. It is the canonical example of code
|
||||
that runs constantly and is observed only under conditions developers do not have.
|
||||
|
||||
---
|
||||
|
||||
### RA-05 - Listener removed using the wrong channel key
|
||||
|
||||
**Mechanism.** A tag-addressed pub/sub subsystem walks parent tags when
|
||||
broadcasting. On finding an expired listener it removes the handle using the
|
||||
*original broadcast channel* rather than the ancestor tag currently being visited.
|
||||
Handle identifiers are unique only within a channel.
|
||||
|
||||
**Why it is silent.** Both outcomes are legal operations on a map. Either the key is
|
||||
absent and removal is a no-op — the stale listener survives in the parent bucket — or
|
||||
a different listener happens to hold the same channel-local id and is removed from
|
||||
the wrong bucket. Neither path raises.
|
||||
|
||||
**Why the obvious check misses it.** The removal call is present, correctly typed,
|
||||
and looks right. The defect is in *which key* is passed, and both candidate keys are
|
||||
in scope, same-typed, and similarly named. Type checking cannot separate them, and
|
||||
review reads the statement as "remove the expired listener".
|
||||
|
||||
**Symptom.** A listener that was unregistered keeps receiving messages, or an
|
||||
unrelated listener silently stops receiving them. The failure surfaces in whichever
|
||||
system owned the collateral listener, arbitrarily far from the subsystem at fault.
|
||||
|
||||
**Detect.** Inside any parent-tag or hierarchy traversal, confirm every mutation
|
||||
keys on the loop cursor rather than the function parameter:
|
||||
|
||||
```bash
|
||||
rg -n -B12 "(Unregister|Remove\w*Listener|RemoveAt)" Source/ \
|
||||
| rg -n "for\s*\(|ParentTag|Ancestor"
|
||||
```
|
||||
|
||||
Then read each hit and check the key.
|
||||
|
||||
**Guardrail.** Make listener identity an explicit pair type — channel plus id — so
|
||||
passing the wrong channel fails to compile instead of silently mis-keying.
|
||||
|
||||
---
|
||||
|
||||
## C2 — Showcase stubs
|
||||
|
||||
The defining property: **the declaration and the consumer both exist; the writer
|
||||
does not** — or the consumer is a literal. Neither a compiler warning nor an
|
||||
"unused symbol" search finds these.
|
||||
|
||||
### RA-06 - Editable tunable with a live reader and no writer
|
||||
|
||||
**Mechanism.** A designer-facing cooldown is compared against a "last event"
|
||||
timestamp field. The timestamp is declared and read, and is never assigned anywhere
|
||||
in the codebase.
|
||||
|
||||
**Why it is silent.** Time elapsed since an unset timestamp is time since world
|
||||
start, which always exceeds any sane threshold. The comparison is permanently true,
|
||||
so the gate the setting was meant to impose never rejects anything. Permanently
|
||||
allowing is a valid runtime state, so nothing errors.
|
||||
|
||||
**Why the obvious check misses it.** The field is read twice and participates in a
|
||||
comparison, so every "is this used?" heuristic — grep, IDE find-usages, unused-symbol
|
||||
warnings — answers yes. What is missing is the **writer**, not the reader, and no
|
||||
default tool asks that question. This is the highest-yield check in the whole skill
|
||||
precisely because it inverts the usual direction of the search.
|
||||
|
||||
**Symptom.** A designer tunes the value for days, reports it has no effect, and is
|
||||
told to check their data. The setting is then either abandoned or "fixed" by
|
||||
changing something else that happens to correlate.
|
||||
|
||||
**Detect.** For every `EditAnywhere` or `EditDefaultsOnly` numeric property and for
|
||||
every timestamp it is compared against, search for an assignment rather than a
|
||||
mention:
|
||||
|
||||
```bash
|
||||
rg -n "\bLastFireTime\b" Source/ # reads: several
|
||||
rg -n "\bLastFireTime\b\s*(=|\+=|-=)" Source/ # one hit - and it is the declaration
|
||||
```
|
||||
|
||||
The second search does **not** come back empty, and that is the trap. A member
|
||||
declared with an initializer (`double LastFireTime = 0.0;`) matches every
|
||||
assignment pattern you can write, so the naive search reports a writer that does
|
||||
not exist. Subtract the declaration before concluding:
|
||||
|
||||
```bash
|
||||
rg -n "\bLastFireTime\b\s*(=|\+=|-=)" Source/ \
|
||||
| rg -v "\b(bool|u?int\d+|float|double|F[A-Z]\w+|T\w+<)\s+LastFireTime\b"
|
||||
```
|
||||
|
||||
Empty after that subtraction, with live reads elsewhere, is a stub. Verified on
|
||||
the reference: the unsubtracted search returns one line, the subtracted search
|
||||
returns none.
|
||||
|
||||
**Guardrail.** Never ship an editor-exposed property whose value has no writer.
|
||||
Every tunable needs one owning assignment site and one test that moves the value and
|
||||
asserts an observable delta.
|
||||
|
||||
---
|
||||
|
||||
### RA-07 - Ordering field that is stored, copied and never compared
|
||||
|
||||
**Mechanism.** A `Priority` integer is accepted by eight registration overloads,
|
||||
stored on the entry, copied into the request struct, and never appears in a sort, a
|
||||
predicate or a comparison anywhere.
|
||||
|
||||
**Why it is silent.** The field has a defined value at all times and travels
|
||||
correctly through the entire API. Ordering still happens — it is just registration
|
||||
order. A plausible order is produced on every run, so nothing looks wrong.
|
||||
|
||||
**Why the obvious check misses it.** The value is written, read, copied and passed
|
||||
across a large public surface. Usage counts are high. Any review that asks "is
|
||||
`Priority` used?" finds a dozen sites and stops. The right question is narrower: does
|
||||
it appear inside a `Sort`, `<`, `>` or predicate? That question is not one anyone
|
||||
thinks to ask about a field that is obviously plumbed.
|
||||
|
||||
**Symptom.** Widget or handler order changes between runs and between machines. In a
|
||||
plugin-based project, registration order is feature activation order, so the symptom
|
||||
reads as a race condition and gets chased in the wrong subsystem.
|
||||
|
||||
**Detect.** Search for the comparison, not the plumbing:
|
||||
|
||||
```bash
|
||||
F='(Priority|Order|Weight|Rank|SortKey)'
|
||||
rg -n "\b$F\b" Source/ Plugins/ | wc -l # plumbing: many
|
||||
rg -n "(Sort|StableSort|Algo::)" Source/ Plugins/ | rg "\b$F\b" # ordering: ?
|
||||
rg -n "\b$F\s*[<>]=?\s*\w|\w\s*[<>]=?\s*\w*$F\b" Source/ Plugins/ \
|
||||
| rg -v -- "->" # comparisons: ?
|
||||
```
|
||||
|
||||
The `-v -- "->"` matters. Without it, every `Entry->Priority = X` is counted as a
|
||||
comparison because the arrow contains `>`, and a field that is only ever assigned
|
||||
looks like a field that is ordered. On the measured reference the naive form
|
||||
reported two "comparisons"; both were assignments.
|
||||
|
||||
A large first number with empty second and third lines is an ordering field that
|
||||
does not order.
|
||||
|
||||
**Guardrail.** An ordering field ships with the comparator that consumes it, in the
|
||||
same change. If ordering is not implemented yet, do not expose the field.
|
||||
|
||||
---
|
||||
|
||||
### RA-08 - Tag-driven visibility with the check hardcoded off
|
||||
|
||||
**Mechanism.** A widget base class exposes an editable tag container for
|
||||
"hide when these tags are present". Both consumer sites read
|
||||
`const bool bHasHiddenTags = false;` with the real expression parked in a comment
|
||||
next to a TODO. The listener registration that would supply the tags is also a
|
||||
comment.
|
||||
|
||||
**Why it is silent.** `false` is a legitimate value: it means "no hiding tags are
|
||||
active right now". The widget is visible, which is a normal state, so no test and no
|
||||
reviewer can tell the difference between "correctly not hidden" and "never able to
|
||||
hide".
|
||||
|
||||
**Why the obvious check misses it.** The TODO admits *one* missing piece and thereby
|
||||
misdirects. The obvious fix — implement `bHasHiddenTags` — does not restore the
|
||||
intended behaviour, because a nearby visibility setter overwrites the
|
||||
designer-authored shown and hidden visibility values on its first call, which arrives
|
||||
during construction. A stub can be deeper than its own TODO says it is, and a reviewer
|
||||
who trusts the TODO stops one layer too early.
|
||||
|
||||
**Symptom.** A designer fills the tag container, observes nothing, and the class is
|
||||
recorded in team lore as "the tag hiding does not work" without anyone establishing
|
||||
why.
|
||||
|
||||
**Detect.** Find neutral literals feeding decision branches:
|
||||
|
||||
```bash
|
||||
rg -n "const bool b\w+ = (false|true);\s*//" Source/ Plugins/
|
||||
```
|
||||
|
||||
For each hit, check whether the property that should feed it is `EditAnywhere`. Then
|
||||
check that nothing else overwrites the same state before the branch is reached — the
|
||||
second step is the one people skip.
|
||||
|
||||
**Guardrail.** An editable property whose consumer is a literal constant must not
|
||||
ship. When you do implement one, re-derive the whole path rather than the single line
|
||||
the TODO names.
|
||||
|
||||
---
|
||||
|
||||
### RA-09 - Parameters threaded through for a subsystem that was never built
|
||||
|
||||
**Mechanism.** A `bIsSimulated` flag appears in three method signatures and is
|
||||
passed through six call sites. Its sole caller passes a literal `false`. A companion
|
||||
"replace this hit" method exists and is never called; the boolean recording its result
|
||||
is only ever read.
|
||||
|
||||
**Why it is silent.** A flag that is always `false` produces one consistent code
|
||||
path, which is the path everything is tested on. A method with no callers has no
|
||||
behaviour to be wrong. Both are inert rather than broken.
|
||||
|
||||
**Why the obvious check misses it.** The threading through multiple layers is exactly
|
||||
what a real, wired-up feature looks like — that shape is itself the disguise. Usage
|
||||
searches return many hits. Only two narrower questions expose it: is the parameter
|
||||
ever branched on, and does the sole call site pass a variable or a literal?
|
||||
|
||||
**Symptom.** An engineer reads the signatures, concludes server-side hit rewind
|
||||
exists, and budgets zero for it. The gap is discovered when the game meets real
|
||||
latency, at which point the structure of abilities, tracing and damage is already
|
||||
committed.
|
||||
|
||||
**Detect.** Find parameters that are carried but never branched on:
|
||||
|
||||
```bash
|
||||
rg -n "\bbIsSimulated\b" Source/ # many hits: plumbed
|
||||
rg -n "if\s*\(\s*!?bIsSimulated" Source/ # none: never decides anything
|
||||
rg -n "\w+\(.*\bfalse\b.*\)" Source/ | rg "bIsSimulated|Simulated"
|
||||
```
|
||||
|
||||
Also list public functions with zero callers: `rg -n "FunctionName"` returning only
|
||||
the declaration and definition.
|
||||
|
||||
**Guardrail.** Do not ship attachment points for a subsystem that does not exist.
|
||||
An unused parameter is a claim about the architecture; if the claim is false, delete
|
||||
the parameter or write the subsystem.
|
||||
|
||||
---
|
||||
|
||||
### RA-10 - Field silently dropped in both directions of a conversion
|
||||
|
||||
**Mechanism.** Helper functions convert between a gameplay message and cue
|
||||
parameters. A context tag container present on both sides is not copied in either
|
||||
direction. Both sites carry a TODO.
|
||||
|
||||
**Why it is silent.** The conversion succeeds and produces a fully valid object. The
|
||||
dropped field simply arrives empty, and empty is the same value a caller who did not
|
||||
set it would produce. There is no partial-failure signal because nothing failed.
|
||||
|
||||
**Why the obvious check misses it.** Round-tripping is the natural test, and a
|
||||
round trip through both directions loses the same field consistently — so a
|
||||
comparison of "before" against "after" on the fields anyone thought to compare
|
||||
passes. Structural equality is the test that would catch it, and structural equality
|
||||
is what nobody writes for conversion helpers.
|
||||
|
||||
**Symptom.** Effects lose their contextual tags when routed through the conversion,
|
||||
so downstream selection by context silently picks the default variant.
|
||||
|
||||
**Detect.** For every conversion helper, compare field counts on both sides:
|
||||
|
||||
```bash
|
||||
rg -n "^\s*(FGameplayTagContainer|F\w+|TArray<\w+>)\s+\w+;" Source/**/Struct.h
|
||||
rg -n -A30 "ToOtherForm|FromOtherForm|Convert\w+" Source/ | rg "\bField\b"
|
||||
```
|
||||
|
||||
Any field declared on both types and mentioned in neither direction is dropped.
|
||||
|
||||
**Guardrail.** Conversion helpers get a structural round-trip test that enumerates
|
||||
fields by reflection rather than by hand. A hand-written comparison tests the fields
|
||||
the author remembered, which is the same set they remembered to copy.
|
||||
|
||||
---
|
||||
|
||||
### RA-11 - Replicated structure with no callers and an empty removal hook
|
||||
|
||||
**Mechanism.** A replicated fast-array serializer type is fully declared, with the
|
||||
removal callback implemented as an empty body. Nothing in the codebase constructs or
|
||||
uses it.
|
||||
|
||||
**Why it is silent.** Unused replicated types cost nothing at runtime and generate no
|
||||
warning. The empty callback is a valid override — many are legitimately empty.
|
||||
|
||||
**Why the obvious check misses it.** The type is complete and idiomatic, so it reads
|
||||
as infrastructure that some subsystem depends on. Deleting it feels risky, so it
|
||||
survives every cleanup. Meanwhile a reader treats its existence as evidence that
|
||||
replicated messaging is solved.
|
||||
|
||||
**Symptom.** A team builds on top of it, assuming it works, and discovers on the
|
||||
first multi-client test that no path ever populated it.
|
||||
|
||||
**Detect.** Find declared-but-unconstructed types:
|
||||
|
||||
```bash
|
||||
rg -ln "struct F\w+ : public FFastArraySerializer" Source/ Plugins/
|
||||
# per type T:
|
||||
rg -n "\bT\b" Source/ Plugins/ | rg -v "\.h:|struct|USTRUCT|template"
|
||||
```
|
||||
|
||||
Hits only in the header mean the type has no user.
|
||||
|
||||
**Guardrail.** Replication infrastructure ships with at least one caller and one
|
||||
test that observes a replicated change. Untested replication is not infrastructure,
|
||||
it is a plan.
|
||||
|
||||
---
|
||||
|
||||
## C3 — Deliberate sample-scope narrowings
|
||||
|
||||
Nothing here is broken. The reference is internally consistent with every one of
|
||||
these. They become expensive only when the sample becomes the base of a product.
|
||||
|
||||
### RA-12 - No lag compensation, and the whole hit chain follows from it
|
||||
|
||||
**Mechanism.** The project contains no custom saved-move, server-move or compressed
|
||||
flags implementation, and therefore no historical positions on the server. Local
|
||||
targeting early-returns when the pawn is not locally controlled, so the server never
|
||||
traces. A `bIsTargetDataValid` local is initialised to `true` and consumed as if it
|
||||
were a validation result. Damage falloff is computed from the client's supplied trace
|
||||
start, and the material multiplier from the client's supplied physical material.
|
||||
|
||||
**Why it is silent.** Every step is the only available behaviour given the missing
|
||||
subsystem. The server cannot validate positions it never recorded, so trusting the
|
||||
client is not a slip — it is the sole option. Correct-looking code all the way down,
|
||||
with the gap one level below the code.
|
||||
|
||||
**Why the obvious check misses it.** Reviewing the damage path finds an authority
|
||||
check and a validity boolean, which together look like a validation layer. The
|
||||
boolean is a literal and the authority check answers a different question ("may this
|
||||
instigator damage this target?") than the one that matters ("did this shot happen?").
|
||||
Reading for the presence of checks finds checks; reading for what the adversary can
|
||||
assert finds nothing stopping them.
|
||||
|
||||
**Symptom.** In production, clients report impossible hits and the diagnosis becomes
|
||||
"someone forgot a validation call" — which sends engineers to patch individual sites
|
||||
instead of pricing the missing subsystem.
|
||||
|
||||
**Detect.** Ask what the adversary supplies, then trace each such value:
|
||||
|
||||
```bash
|
||||
rg -n "class \w+ : public FSavedMove_Character|ServerMove|CompressedFlags" Source/
|
||||
rg -n "const bool b\w*Valid\w* = true" Source/
|
||||
rg -n "TargetData|HitResult" Source/ | rg -i "damage|falloff|physmat"
|
||||
```
|
||||
|
||||
No saved-move type plus client-supplied geometry in the damage math means full
|
||||
client trust.
|
||||
|
||||
**Guardrail.** Choose a hit-registration model *before* borrowing a weapon stack:
|
||||
server-authoritative with rewind, client-claim with plausibility bounds, or explicit
|
||||
full trust. Write the choice down. Each has a different cost and a different cheat
|
||||
surface, and the choice dictates structure you cannot cheaply change later.
|
||||
|
||||
---
|
||||
|
||||
### RA-13 - Replication graph implemented, configured, and shipped disabled
|
||||
|
||||
**Mechanism.** A replication graph implementation and around a dozen tuning cvars
|
||||
exist. Project config disables it and the class routing table is essentially empty,
|
||||
with the few entries present marked not-routed.
|
||||
|
||||
**Why it is silent.** The default replication path works fine at sample scale.
|
||||
Disabled infrastructure produces no error and no measurable difference until player
|
||||
counts and actor counts grow.
|
||||
|
||||
**Why the obvious check misses it.** The presence of the implementation, the cvars
|
||||
and the config section reads as "this is configured". Nobody diffs the routing table
|
||||
against the actor classes that actually exist, because a populated-looking config
|
||||
section satisfies the eye.
|
||||
|
||||
**Symptom.** Bandwidth and server CPU limits are discovered in beta, at the point
|
||||
where the actor class layout is fixed and re-routing is a large change.
|
||||
|
||||
**Detect.** Compare routing entries against replicated classes:
|
||||
|
||||
```bash
|
||||
rg -n "bDisableReplicationGraph|ClassNodeMapping" Config/
|
||||
rg -c "GetLifetimeReplicatedProps" Source/ Plugins/
|
||||
```
|
||||
|
||||
A handful of routing entries against dozens of replicated classes means the graph
|
||||
has never carried traffic.
|
||||
|
||||
**Guardrail.** Copying a disabled subsystem's configuration copies untested
|
||||
configuration. Either enable it and measure under load, or delete the config and
|
||||
record that you owe the work.
|
||||
|
||||
---
|
||||
|
||||
### RA-14 - Preload machinery made unreachable by a load-mode constant
|
||||
|
||||
**Mechanism.** A cue manager supports delayed loading with garbage-collect and
|
||||
map-load hooks. A load-mode constant is set to "load upfront", which short-circuits
|
||||
every switch site — including the one that installs the three delegates.
|
||||
|
||||
**Why it is silent.** Loading everything upfront is correct behaviour. The
|
||||
machinery is not broken; it is bypassed. Memory is higher and nothing reports it,
|
||||
because nothing is measuring against a budget.
|
||||
|
||||
**Why the obvious check misses it.** Searching for the delegates finds them bound in
|
||||
source, so the wiring appears present. The early return that makes the binding
|
||||
unreachable is in a different function, guarded by a constant that reads like a
|
||||
development convenience. The project's own diagnostic command reports zeros, which is
|
||||
then read as "no cues are preloaded yet" rather than "this counter is dead".
|
||||
|
||||
**Symptom.** A team diagnosing memory or hitching concludes the engine's preload
|
||||
system is broken and descends into engine code, when the sample simply opted out.
|
||||
|
||||
**Detect.** Find switch sites gated by a constant, and confirm delegate binding is
|
||||
reachable:
|
||||
|
||||
```bash
|
||||
rg -n "LoadMode|ELoadMode|EEditorLoadMode" Source/ Plugins/
|
||||
rg -n -B6 "AddUObject|AddRaw|BindUObject" Source/ | rg "return;|LoadUpfront"
|
||||
```
|
||||
|
||||
**Guardrail.** Distinguish "the mechanism is broken" from "the sample configured it
|
||||
off" before you spend a day in engine code. Record which subsystems the reference
|
||||
bypasses; that list is part of what you are adopting.
|
||||
|
||||
---
|
||||
|
||||
### RA-15 - Public virtual state writers with no authority check
|
||||
|
||||
**Mechanism.** Death-state transitions are written by two public virtual methods on
|
||||
a health component, neither of which checks for network authority.
|
||||
|
||||
**Why it is silent.** Single-process play-in-editor never separates authority from
|
||||
simulation, so the code path is exercised constantly and always on the authority. The
|
||||
missing check has no observable consequence in the environment where the sample runs.
|
||||
|
||||
**Why the obvious check misses it.** Authority checks are usually reviewed at the
|
||||
RPC boundary, and these methods are not RPCs — they are ordinary virtual functions
|
||||
that a caller in the right context invokes correctly. Being public and virtual is
|
||||
what makes it dangerous: the class invites overriding and calling from anywhere, and
|
||||
the invitation carries no precondition.
|
||||
|
||||
**Symptom.** A client-side subclass or a Blueprint calls the method during
|
||||
prediction, the client's death state diverges from the server's, and the resulting
|
||||
desync is chased as a replication bug.
|
||||
|
||||
**Detect.** List state writers and check each for an authority guard:
|
||||
|
||||
```bash
|
||||
rg -n "void (Start|Finish|Set)\w*(Death|State|Team)\w*\(" Source/
|
||||
rg -n -A6 "void (Start|Finish|Set)\w*(Death|State|Team)\w*\(" Source/ | rg -c "HasAuthority"
|
||||
```
|
||||
|
||||
A count below the number of writers names the unguarded ones.
|
||||
|
||||
**Guardrail.** Every writer of replicated state either checks authority or is
|
||||
private with a single guarded caller. Public plus virtual plus unguarded is an
|
||||
invitation.
|
||||
|
||||
---
|
||||
|
||||
### RA-16 - Faction query fails open on invalid input
|
||||
|
||||
**Mechanism.** A team comparison returns an "invalid argument" result for unknown
|
||||
membership, and the damage path treats anything that is not an explicit "same team"
|
||||
as damageable. The site carries a TODO calling itself temporary.
|
||||
|
||||
**Why it is silent.** Unknown membership does not occur in a sample where every pawn
|
||||
is assigned a team at spawn. The fail-open branch is never taken, so it is never
|
||||
observed to be wrong.
|
||||
|
||||
**Why the obvious check misses it.** The function returns a three-valued result,
|
||||
which reads as careful design. The defect is in the *consumer's* collapse of three
|
||||
values into two, in a different file. Reviewing the query finds correct code;
|
||||
reviewing the damage path finds a plausible boolean.
|
||||
|
||||
**Symptom.** Neutral, spectating or mid-transition actors take or deal damage. The
|
||||
bug appears only in modes the sample does not have, which is to say, in yours.
|
||||
|
||||
**Detect.** Find three-valued results collapsed to booleans:
|
||||
|
||||
```bash
|
||||
rg -n "Invalid|Unknown|Indeterminate" Source/ | rg -i "team|faction|relation"
|
||||
rg -n -B4 -A4 "CanCauseDamage|CanDamage|IsHostile" Source/
|
||||
```
|
||||
|
||||
Read the branch: does the invalid case land with "allowed" or "denied"?
|
||||
|
||||
**Guardrail.** Unknown must be denied, never allowed, on any decision that costs
|
||||
something. If denial is not acceptable, the unknown state itself is the bug.
|
||||
|
||||
---
|
||||
|
||||
### RA-17 - No tag or asset redirects declared at all
|
||||
|
||||
**Mechanism.** Config declares zero gameplay tag redirects. Every tag name is a hard
|
||||
string with no migration path.
|
||||
|
||||
**Why it is silent.** A project that has never renamed a tag needs no redirects. The
|
||||
absence is invisible right up to the first rename, and samples do not rename after
|
||||
release.
|
||||
|
||||
**Why the obvious check misses it.** This is an absence with no consumer to inspect
|
||||
— there is no line of code to review. It is only visible if you go looking for a
|
||||
capability you have not needed yet, which is exactly the thing nobody does during
|
||||
adoption.
|
||||
|
||||
**Symptom.** The first tag rename after content exists breaks every asset that
|
||||
referenced it. Since references live in binary assets, the breakage is silent data
|
||||
loss rather than a compile error, discovered per-asset over weeks.
|
||||
|
||||
**Detect.**
|
||||
|
||||
```bash
|
||||
rg -c "GameplayTagRedirects|\+ActiveGameRedirects|ClassRedirects" Config/
|
||||
```
|
||||
|
||||
Zero, in a project that has shipped content, means renames have not yet happened —
|
||||
not that they are safe.
|
||||
|
||||
**Guardrail.** Establish redirect discipline before the first rename, not after.
|
||||
Add a CI check that a removed tag name has a corresponding redirect entry.
|
||||
|
||||
---
|
||||
|
||||
### RA-18 - Server RPC declared without validation
|
||||
|
||||
**Mechanism.** A quick-bar style component exposes a `Server, Reliable,
|
||||
BlueprintCallable` function with no `WithValidation`, accepting an index from the
|
||||
client.
|
||||
|
||||
**Why it is silent.** Well-behaved clients send valid indices, and the sample only
|
||||
ever has well-behaved clients. Out-of-range access either clamps harmlessly or hits a
|
||||
path no test covers.
|
||||
|
||||
**Why the obvious check misses it.** Nearby functions are marked
|
||||
`BlueprintAuthorityOnly`, which reads as an access restriction and satisfies a
|
||||
reviewer scanning for guards. That specifier constrains Blueprint execution context
|
||||
only — it does not validate parameters, and it does not constrain C++ callers at all.
|
||||
A guard that answers a different question is worse than no guard, because it stops
|
||||
the search.
|
||||
|
||||
**Symptom.** A modified client sends an arbitrary index. Depending on the container,
|
||||
this is a crash, a read of unrelated memory, or equipping something the player never
|
||||
earned.
|
||||
|
||||
**Detect.**
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*\bServer\b[^)]*\)" Source/ Plugins/ | rg -v "WithValidation"
|
||||
```
|
||||
|
||||
Every hit takes client-controlled input on trust.
|
||||
|
||||
**Guardrail.** Every server RPC validates its parameters, including range and
|
||||
ownership, and `BlueprintAuthorityOnly` is never counted as validation.
|
||||
|
||||
---
|
||||
|
||||
### RA-19 - Readiness gate admits simulated proxies with no controller
|
||||
|
||||
**Mechanism.** An initialisation-state transition treats "has a controller" as a
|
||||
precondition, but admits simulated proxies that have none, so readiness is reached
|
||||
on different grounds depending on net role.
|
||||
|
||||
**Why it is silent.** Both branches produce "ready", and ready is the state everything
|
||||
downstream waits for. No consumer asks *why* readiness was granted.
|
||||
|
||||
**Why the obvious check misses it.** The gate is short and reads as a careful
|
||||
role-aware special case — which is what it is. The problem is that "ready" then means
|
||||
two different things, and the divergence lives in every consumer rather than in the
|
||||
gate. Reviewing the gate finds nothing wrong.
|
||||
|
||||
**Symptom.** Systems that assume a controller exists after readiness crash or
|
||||
no-op on simulated proxies. The failure appears only with a remote client, so it
|
||||
survives all local testing.
|
||||
|
||||
**Detect.** Find role-dependent readiness and enumerate its consumers:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A10 "CanChangeInitState|HasReachedInitState" Source/ | rg "IsLocallyControlled|GetController|ROLE_SimulatedProxy"
|
||||
rg -n "HasReachedInitState|GameplayReady" Source/ | wc -l
|
||||
```
|
||||
|
||||
**Guardrail.** One readiness state means one set of guarantees. If proxies reach it
|
||||
by a different route, they need a different state name, so consumers must choose.
|
||||
|
||||
---
|
||||
|
||||
## How to use this list
|
||||
|
||||
Do not read it as a defect report about someone else's project. Read it as a list of
|
||||
**question shapes**:
|
||||
|
||||
1. Which of my editable properties has no writer?
|
||||
2. Which of my ordering fields never reaches a comparison?
|
||||
3. Which of my decision branches reads a literal?
|
||||
4. Which of my parameters is threaded but never branched on?
|
||||
5. Which of my unknown-input paths fails open?
|
||||
6. Which of my public state writers has no authority check?
|
||||
7. Which of my capabilities exists only as configuration that is switched off?
|
||||
|
||||
Each has a one-line `Detect` above. Run all seven against your own tree before you
|
||||
run any of them against the reference.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Every entry was read in source in one reference project on one engine version. They
|
||||
are worked examples of the classification, not a defect list to carry into other
|
||||
engine versions or other samples. An entry that does not reproduce under its own
|
||||
`Detect` in your tree does not apply to you.
|
||||
@@ -0,0 +1,294 @@
|
||||
# Adoption patterns: a worked classification
|
||||
|
||||
This is the C1-C5 classification from `SKILL.md` applied end to end against one
|
||||
real reference project, plus the tax analysis and vocabulary hygiene that follow
|
||||
from it.
|
||||
|
||||
Individual silent failures are not repeated here. They live in
|
||||
[failure modes](failure-modes.md), one entry each, with a detection recipe you
|
||||
can run against your own tree. This file is about *how the categories behave* and
|
||||
what each one costs, so that when you meet a new element you can place it.
|
||||
|
||||
---
|
||||
|
||||
## 1. What each category feels like from the inside
|
||||
|
||||
The categories are not severity levels. They are answers to a different question:
|
||||
**what does the author of the reference believe about this element?**
|
||||
|
||||
| | Author's belief | Your evidence | Your move |
|
||||
|---|---|---|---|
|
||||
| **C1 Defect** | "this works" | it demonstrably does not | fix the line, keep the design |
|
||||
| **C2 Showcase stub** | "this is not finished" | reads as finished | implement or delete the field |
|
||||
| **C3 Scope narrowing** | "a sample does not need this" | nothing is wrong at all | budget it or accept the risk in writing |
|
||||
| **C4 Architecture tax** | "this is worth it for us" | correct, working, expensive | adopt by need, not completeness |
|
||||
| **C5 Foreign legacy** | nobody believes anything | a name from another product | rename on entry |
|
||||
|
||||
The dangerous boundary is C2 against C3, because from the outside they look
|
||||
identical: a mechanism that produces a neutral result forever. The distinction is
|
||||
whether the author intended to come back. It changes nothing about your cost, and
|
||||
everything about how you talk to your team: a C2 is a hole in the reference, a C3
|
||||
is a hole in your understanding of what you bought.
|
||||
|
||||
---
|
||||
|
||||
## 2. C1 - Defects
|
||||
|
||||
A defect in a reference project is the cheapest category and the least
|
||||
interesting. Fix the line, keep the surrounding design, and do not let it
|
||||
discredit the pattern that contains it.
|
||||
|
||||
Three shapes recur often enough to be worth naming:
|
||||
|
||||
**The filter that filters nothing.** A field constrains designer input to a
|
||||
namespace, and the namespace does not exist in this project. The picker comes up
|
||||
empty, which reads as "no matching entries yet" rather than "this constraint is
|
||||
misconfigured". Where such filters are rare in a codebase, each one is load-bearing
|
||||
and worth checking individually.
|
||||
|
||||
**The symmetric typo.** A misspelled identifier used as a string join across
|
||||
config, code and platform overrides works perfectly, because every side misspells
|
||||
it identically. It is only discovered by someone fixing the spelling in one place.
|
||||
The lesson is not "check spelling"; it is that when a string literal is the join,
|
||||
correctness means *all sides agree*, and the way to get that is one declaration
|
||||
site referenced everywhere else.
|
||||
|
||||
**The loop over an empty container.** A local collection is declared next to a
|
||||
TODO, never populated, and then iterated. The loop reads as a mechanism. It is a
|
||||
placeholder for one.
|
||||
|
||||
None of these survives a careful read of the consumer. They survive a read of the
|
||||
declaration, which is why they get copied.
|
||||
|
||||
---
|
||||
|
||||
## 3. C2 - Showcase stubs
|
||||
|
||||
The defining property is precise, and worth memorizing:
|
||||
|
||||
> The declaration exists. The consumer exists. The **writer** does not - or the
|
||||
> consumer is a literal.
|
||||
|
||||
Neither a compiler warning nor an "unused symbol" search finds these, because the
|
||||
symbol is used. Reading the code convinces you it is wired up. This is why the
|
||||
highest-yield check in the whole skill is searching for assignment rather than
|
||||
mention.
|
||||
|
||||
Sub-shapes, each with an entry in the failure-mode index:
|
||||
|
||||
- a tunable compared against a timestamp that is never assigned (RA-06);
|
||||
- an ordering field stored, copied through the entire API surface, and never
|
||||
compared (RA-07);
|
||||
- an editable tag container whose consumer is a hardcoded `false` (RA-08);
|
||||
- a parameter threaded through several call layers, whose only call site passes a
|
||||
literal, waiting for a subsystem that was never written (RA-09);
|
||||
- a field dropped silently in both directions of a conversion helper (RA-10);
|
||||
- a replicated structure with no callers at all (RA-11).
|
||||
|
||||
### The trap inside the trap
|
||||
|
||||
A stub can be deeper than its TODO admits. In one measured case, implementing the
|
||||
obvious missing condition would still not restore the intended behaviour, because
|
||||
a second method on the same class overwrites the designer-authored values on its
|
||||
first call. The TODO points at one line; the repair is two.
|
||||
|
||||
Treat a TODO as evidence that the author knew about *something*, not as a
|
||||
specification of what is missing. Read the whole class before pricing the fix.
|
||||
|
||||
### What to do
|
||||
|
||||
Either implement it before exposing it, or delete the field. There is no third
|
||||
option that is honest. An editable property whose consumer is a literal constant
|
||||
is a trap for designers, who will tune it for days and then report that the build
|
||||
is broken.
|
||||
|
||||
---
|
||||
|
||||
## 4. C3 - Deliberate sample-scope narrowings
|
||||
|
||||
Nothing here is broken. The reference is internally consistent with every one of
|
||||
these. They become expensive only when the sample becomes the base of a product.
|
||||
|
||||
This is the category that costs a subsystem rather than a line, and the one that
|
||||
adoption discussions habitually skip, because there is nothing to point at.
|
||||
|
||||
### The largest one: no lag compensation
|
||||
|
||||
When a reference project has no server-side rewind - no historical position
|
||||
storage, no custom saved-move pipeline - everything downstream follows
|
||||
mechanically:
|
||||
|
||||
- the server does not trace, because tracing is gated on being locally
|
||||
controlled, and the server controls nobody;
|
||||
- the client's target data is consumed as authoritative, because there is nothing
|
||||
to compare it against;
|
||||
- damage falloff is computed from the client's own reported trace origin;
|
||||
- surface-material multipliers are taken from the client's reported hit.
|
||||
|
||||
Read as a defect list, this is four bugs. Read correctly, it is one design
|
||||
decision with four consequences. "The reference forgot to validate hits" is the
|
||||
wrong diagnosis, and it leads to four local patches that do not add up to server
|
||||
authority.
|
||||
|
||||
**Adoption decision:** choose a hit-registration model *before* borrowing a weapon
|
||||
stack. Server-authoritative with rewind, client-claim with plausibility checks, or
|
||||
full trust. A sample usually implements the third. Each has a different cost and a
|
||||
different cheat surface, and the choice dictates how abilities, tracing and damage
|
||||
are structured - which is why it cannot be retrofitted cheaply.
|
||||
|
||||
### Configuration that ships disabled
|
||||
|
||||
A subsystem can be present, implemented, documented by its own console variables,
|
||||
and switched off in shipped configuration with an essentially empty routing table.
|
||||
The implementation being there tells you nothing about whether it was ever
|
||||
exercised. Copying that configuration copies untested settings, and the settings
|
||||
are the part that needs production-scale traffic to tune.
|
||||
|
||||
### Machinery made unreachable by a mode switch
|
||||
|
||||
A load-mode enum set to "load everything upfront" can short-circuit every branch
|
||||
that would install a delay-load subsystem's delegates. The inherited machinery is
|
||||
functional; the sample bypasses it. The diagnostic tools built for that subsystem
|
||||
then report zeros forever, which reads as "nothing to report" rather than "this
|
||||
code never runs".
|
||||
|
||||
Important qualification: this is a *configuration choice of the sample*, not an
|
||||
engine defect. Diagnosing it as a broken engine feature sends you into engine
|
||||
source for no reason. Check the mode before you check the mechanism.
|
||||
|
||||
### The rest, in one table
|
||||
|
||||
| Narrowing | What it costs a product |
|
||||
|---|---|
|
||||
| State transitions on a health or death component written without an authority check, on public virtual methods | Death becomes a state with no owner; every caller is an authority |
|
||||
| Team or faction lookup that fails open on an invalid argument | Unknown membership resolves to "damage allowed" |
|
||||
| No tag or asset redirects configured anywhere | Every rename is a breaking change with no migration path |
|
||||
| A reliable server RPC declared without validation | Unvalidated input from the network; a Blueprint-authority marker does not constrain C++ callers |
|
||||
| An init-state gate that admits simulated proxies with no controller | "Ready" means something different per net role |
|
||||
|
||||
Each row is a line item in your adoption budget, not a bug report against the
|
||||
reference.
|
||||
|
||||
### The rule
|
||||
|
||||
> Every C3 narrowing you accept must be written down as an accepted risk with an
|
||||
> owner, or scheduled. "We did not notice" is the only unacceptable outcome.
|
||||
|
||||
---
|
||||
|
||||
## 5. C4 - Architecture tax
|
||||
|
||||
Correct, working, and sized for the reference author's team, release cadence and
|
||||
platform matrix. The question is never "is this good architecture". It is "does my
|
||||
product have the problem this solves".
|
||||
|
||||
### Measured shape of one reference project
|
||||
|
||||
| Element | Measured | When it stops paying |
|
||||
|---|---|---|
|
||||
| Definition-to-instance stack: experience selects feature plugins, which contribute action sets, which grant abilities and data | 5 gameplay-feature plugins out of 81 plugins total | One game mode and no post-launch content: you buy distributed control flow and receive no modularity in return |
|
||||
| C++ / Blueprint boundary | 493 Blueprints scanned, median 5 nodes each; 1256 `BlueprintCallable` entry points against 45 `BlueprintImplementableEvent` plus `BlueprintNativeEvent` override points | The 28:1 ratio means C++ exposes *calls*, not *extension points*. That assumes a permanently available C++ team. Without one, designers stall on their first non-trivial request |
|
||||
| Plugin-mounted content roots | primary game root holds 2 838 assets; all 105 mount roots hold 17 256; project-owned content is 3 876 | Any census, budget or memory number computed over the primary root alone understates by roughly six times |
|
||||
|
||||
The second row is the one people copy without noticing they copied it. A C++ to
|
||||
Blueprint ratio is a staffing decision wearing an architecture costume. Measure
|
||||
both columns before adopting the boundary, and name the person who will add to the
|
||||
second one.
|
||||
|
||||
The third row is a measurement discipline, not an architecture choice, and it
|
||||
generalizes: in any plugin-based project, enumerate all mount roots. Tooling that
|
||||
defaults to the primary game root will quietly report a fraction of reality.
|
||||
|
||||
### How to decide
|
||||
|
||||
State the question the modular stack answers *for you*:
|
||||
|
||||
- more than one team shipping into the same binary?
|
||||
- content added after launch without a client patch?
|
||||
- game modes swapped at runtime by data?
|
||||
- platform-specific feature sets?
|
||||
|
||||
Zero yes answers means the whole stack is tax. One or two means adopt the layers
|
||||
that answer them and stop there. Adopting for completeness is paying for an answer
|
||||
to a question you do not have.
|
||||
|
||||
---
|
||||
|
||||
## 6. C5 - Foreign-product legacy
|
||||
|
||||
Three recurring shapes:
|
||||
|
||||
- global constants and type identifiers carrying another title's brand, usually
|
||||
because the subsystem was lifted from that title;
|
||||
- a tag whose own descriptive comment admits it is in the wrong namespace and
|
||||
needs to move;
|
||||
- a class comment copy-pasted from a different class, describing something the
|
||||
class is not.
|
||||
|
||||
Harmless as code. Corrosive as vocabulary. Within a year nobody remembers why a
|
||||
primary asset type is named after another game's content, and part of the team
|
||||
assumes it is an engine term. The comment case is worse than the constant case: a
|
||||
wrong name is eventually questioned, a wrong explanation is believed.
|
||||
|
||||
Rename at import time. After the first sprint it becomes archaeology.
|
||||
|
||||
---
|
||||
|
||||
## 7. The counterexample worth copying
|
||||
|
||||
A register-and-unregister pair, in which the unregister path removes exactly the
|
||||
paths that registration added and then asserts on the removal count:
|
||||
|
||||
```cpp
|
||||
// on unregister
|
||||
const int32 NumRemoved = Registry.RemovePathsAddedBy(Feature);
|
||||
ensure(NumRemoved == PathsAddedBy(Feature).Num());
|
||||
```
|
||||
|
||||
In an audit whose dominant finding was "add works, remove is incomplete", this was
|
||||
the one place where teardown was both implemented and asserted. Copy the shape:
|
||||
paired add/remove, plus an assertion that the counts match. The assertion is what
|
||||
turns a silent leak into a test failure.
|
||||
|
||||
Note the asymmetry that remained even there: registration went through the
|
||||
project's own manager subclass while unregistration reached for the engine base
|
||||
class, and a refresh call made on the way in had no counterpart on the way out.
|
||||
Even the good example is worth reading twice - copy the shape, fix the asymmetry.
|
||||
|
||||
---
|
||||
|
||||
## 8. What source reading cannot settle
|
||||
|
||||
Three limits are worth stating, because they bound every claim in this skill:
|
||||
|
||||
1. **Blueprint call sites are invisible to source search.** A function with zero
|
||||
C++ callers may be called from a Blueprint graph. Absence of C++ callers is not
|
||||
absence of callers; settling it requires opening the graphs.
|
||||
2. **Data tables and data assets are binary.** How many entries a table
|
||||
contributes to a registry cannot be read from text. It requires the editor.
|
||||
3. **An empty search result proves the pattern, not the absence.** Widen the
|
||||
pattern, check the path, and check what your search tool excludes by default
|
||||
before concluding that something does not exist.
|
||||
|
||||
The third has caught real errors in this material more than once. Treat it as a
|
||||
standing rule rather than a caveat.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The classification and every measured number above come from a line-by-line audit
|
||||
of Epic's Lyra Starter Game on Unreal Engine 5.6, read as source rather than run,
|
||||
during a research project whose product was this skill set.
|
||||
|
||||
Source addresses stay in that archive. What ships is the recipe, because an
|
||||
address in a tree you do not have is not actionable and a recipe is. Entry
|
||||
identifiers (`RA-01` and up) in [failure modes](failure-modes.md) resolve back to
|
||||
the audited locations, so any specific claim can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One reference project, one engine version, one workspace. This is a worked example
|
||||
of the classification, not a defect list to carry into other engine versions or
|
||||
other samples. Re-run the tests from `SKILL.md` against your own reference before
|
||||
trusting any specific claim here.
|
||||
@@ -0,0 +1,265 @@
|
||||
---
|
||||
name: ue-runtime-allocation-and-caching
|
||||
description: >-
|
||||
Design or audit runtime allocation, pooling, caching and replication cost in
|
||||
Unreal Engine: per-frame and per-event churn, object pools, bounded caches and
|
||||
ring buffers, pointers into container storage, strong versus weak retention of
|
||||
spawned components, fast-array and subobject replication, dormancy and
|
||||
relevancy, and dedicated-server exclusions. Use when profiling hitches,
|
||||
investigating memory that grows during a session, effect or UI churn, or
|
||||
deciding how to reduce replicated object count.
|
||||
---
|
||||
|
||||
# UE runtime allocation and caching
|
||||
|
||||
The invariant:
|
||||
|
||||
> Every cache needs an owner, a bound and an eviction rule. Every pooled object
|
||||
> needs a stable handle. Every replication optimization must name which cost it
|
||||
> reduces — memory, CPU or bandwidth — because they are not the same.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes: [failure modes](references/failure-modes.md).
|
||||
|
||||
Related skills: `ue-asset-loading-and-memory`, `ue-cosmetics-and-teams`,
|
||||
`ue-gameplay-messaging`, `ue-streaming-and-platform-budgets`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Never store a raw pointer into a container's storage
|
||||
|
||||
This is the highest-severity defect class in this skill, and it is the one that
|
||||
looks most harmless:
|
||||
|
||||
```cpp
|
||||
FPooledList& Pool = PooledMap.FindOrAdd(Key); // address of a map value
|
||||
LiveEntries.Emplace(Component, &Pool, ReleaseTime); // held across frames
|
||||
```
|
||||
|
||||
Maps and arrays reallocate on insertion. A later insert invalidates every stored
|
||||
address. The failure is **data-dependent**: it needs a second key inserted while
|
||||
an earlier entry is still live, so it survives casual testing and appears under
|
||||
load.
|
||||
|
||||
A null check does not help. The pointer is not null; it is non-null garbage.
|
||||
|
||||
### Rules
|
||||
|
||||
- store a **key or an index**, and resolve at the point of use;
|
||||
- if a stable address is genuinely required, use a container that guarantees
|
||||
stable element addresses, and say so in a comment at the declaration;
|
||||
- treat "pointer obtained from a container, stored beyond the current scope" as a
|
||||
review blocker, not a style note;
|
||||
- when you suspect one, the allocator's stomp mode converts a probabilistic
|
||||
corruption into a deterministic crash at the right line.
|
||||
|
||||
The same shape appears with lazily-growing statistics maps whose entries are
|
||||
cached by a widget and dereferenced every frame — the map grows a key when a new
|
||||
data source appears, and the widget's pointer dies mid-session. Recipes: RC-01,
|
||||
RC-02.
|
||||
|
||||
---
|
||||
|
||||
## 2. Bound every cache and name its eviction
|
||||
|
||||
A cache without a bound is a leak with good intentions. For each one, write down:
|
||||
|
||||
```text
|
||||
owner who invalidates it
|
||||
key space bounded by what?
|
||||
capacity a number, or "unbounded" with a justification
|
||||
eviction time, count, or explicit removal
|
||||
retention strong or weak
|
||||
teardown what empties it, and when
|
||||
```
|
||||
|
||||
Three shapes recur and each is silent:
|
||||
|
||||
**Read-modify-write growth.** Fetch a collection, append one element, write the
|
||||
whole thing back. Per-event cost grows linearly and total work is quadratic over
|
||||
a session — and if nothing ever clears it, the collection is also a leak. Recipe:
|
||||
RC-03.
|
||||
|
||||
**Re-append without pruning.** A component that re-adds every previously spawned
|
||||
effect on each trigger and never removes finished ones. When the arrays are
|
||||
strong references, every one-shot effect is pinned for the whole match. Recipe:
|
||||
RC-04.
|
||||
|
||||
**Prefilled buffers reported as data.** A fixed-capacity sample cache filled with
|
||||
zeroes reports a misleading minimum and average until it wraps. Track the fill
|
||||
count, or exclude the prefix. Recipe: RC-08.
|
||||
|
||||
---
|
||||
|
||||
## 3. Pools need a lifecycle, not just reuse
|
||||
|
||||
A pool is correct when acquire and release are symmetric and idempotent; released
|
||||
objects are reset to a known state; there is a high-water policy — grow, cap, or
|
||||
trim on idle; pooled objects remain reachable by the collector while pooled; the
|
||||
pool empties on teardown; and a leaked acquire is detectable rather than silent.
|
||||
|
||||
Fixed-size pools must define what happens on exhaustion: block, drop, or grow.
|
||||
Silent drop is acceptable for cosmetic effects only, and it should be counted.
|
||||
|
||||
**A pool that only grows is not a leak in the technical sense and behaves like one
|
||||
in practice.** One burst leaves its peak allocated for the rest of the session.
|
||||
|
||||
---
|
||||
|
||||
## 4. Churn: diff, do not rebuild
|
||||
|
||||
A periodic scan that destroys and recreates its descriptor objects every interval
|
||||
allocates, invalidates observers, and hides real changes in noise.
|
||||
|
||||
```text
|
||||
compute the desired set
|
||||
diff against the current set
|
||||
add new, remove missing, leave the rest alone
|
||||
```
|
||||
|
||||
Reserve capacity once when the size is predictable. In hot paths prefer the
|
||||
reset-without-freeing operation over the one that releases the buffer — the
|
||||
latter guarantees a reallocation on the next frame. Do not allocate in paint or
|
||||
per-frame render-thread paths at all. Recipes: RC-05, RC-06, RC-07.
|
||||
|
||||
---
|
||||
|
||||
## 5. Protocol identifiers need width and uniqueness discipline
|
||||
|
||||
A batch identifier assigned as *the current in-flight count* rather than a
|
||||
monotonic sequence produces reuse: A gets 0, B gets 1, A confirms and is removed,
|
||||
C gets 1 again — and a handler that matches the first plausible entry applies the
|
||||
confirmation to the wrong batch.
|
||||
|
||||
Compounding it, the same value can be one width at the source, another in the
|
||||
message and a third in storage — which turns a wraparound into a permanent
|
||||
mismatch, and an unbounded pending list into an unbounded leak.
|
||||
|
||||
### Rules
|
||||
|
||||
- identifiers are sequences, never counts;
|
||||
- one width for the whole path, asserted at every boundary;
|
||||
- define wrap behaviour explicitly;
|
||||
- handlers match on identity, not on "the first entry that could be it";
|
||||
- an unbounded pending list needs a timeout and a drain policy.
|
||||
|
||||
Recipe: RC-09.
|
||||
|
||||
---
|
||||
|
||||
## 6. Replication: name the axis you are optimizing
|
||||
|
||||
Three different costs, routinely conflated:
|
||||
|
||||
| Mechanism | Memory | CPU | Bandwidth |
|
||||
|---|---|---|---|
|
||||
| Dedicated-server presentation exclusion | ↓ | ↓ | **unchanged** |
|
||||
| Fast-array delta serialization | **↑** per-item key and per-connection state | **↑** delta walk | ↓ |
|
||||
| Subobject replication | ↓ | ↓ | ↓ |
|
||||
| Dormancy | — | ↓↓ | ↓ |
|
||||
| Relevancy and replication graph | — | ↓↓ | ↓ |
|
||||
|
||||
Two counterintuitive facts worth internalizing:
|
||||
|
||||
- excluding cosmetics on a dedicated server saves memory and CPU and changes
|
||||
bandwidth by **zero**, because the intent list still replicates;
|
||||
- fast arrays are a **bandwidth** optimization that costs memory and CPU.
|
||||
|
||||
The genuine object-count reducer is subobject replication: N owned items add zero
|
||||
actors and zero channels.
|
||||
|
||||
### Before claiming a replication optimization
|
||||
|
||||
- [ ] state which axis improves and which regresses;
|
||||
- [ ] measure that axis specifically, not "performance";
|
||||
- [ ] confirm the mechanism is actually **enabled** rather than merely present —
|
||||
see `ue-multiplayer-authority` for a measured case where the replication
|
||||
graph ships disabled, dormancy is unused, and a force-update list is
|
||||
declared, read, and never populated;
|
||||
- [ ] verify net-mode gating is consistent across presentation systems;
|
||||
- [ ] check replicated history structures for a removal path.
|
||||
|
||||
---
|
||||
|
||||
## 7. Dedicated-server exclusions must be systematic
|
||||
|
||||
List every client-only construction and guard it in one reviewable place:
|
||||
cosmetic actors and meshes, widgets and layouts, effect and audio components,
|
||||
preview and capture systems, input and settings objects.
|
||||
|
||||
**An inconsistent set is worse than none**, because server memory profiles become
|
||||
unpredictable — one system guarded while a sibling spawning real actors is not.
|
||||
Audit by searching for the net-mode guard and diffing against the list of
|
||||
presentation systems.
|
||||
|
||||
---
|
||||
|
||||
## 8. Review checklist
|
||||
|
||||
### Allocation
|
||||
|
||||
- [ ] no stored pointers into container storage;
|
||||
- [ ] hot paths reset rather than free, with capacity reserved;
|
||||
- [ ] no allocation in paint or render-thread paths;
|
||||
- [ ] periodic work diffs rather than rebuilds.
|
||||
|
||||
### Caches and pools
|
||||
|
||||
- [ ] owner, bound, eviction and teardown written down for each;
|
||||
- [ ] key space provably bounded;
|
||||
- [ ] retention strength intentional;
|
||||
- [ ] acquire and release symmetric; exhaustion policy defined;
|
||||
- [ ] prefilled buffers do not distort statistics.
|
||||
|
||||
### Protocol
|
||||
|
||||
- [ ] identifiers are sequences with one consistent width;
|
||||
- [ ] pending lists have a timeout and a drain;
|
||||
- [ ] handlers match by identity.
|
||||
|
||||
### Replication
|
||||
|
||||
- [ ] the optimization names its axis;
|
||||
- [ ] the mechanism is verified enabled;
|
||||
- [ ] net-mode gating consistent across presentation systems;
|
||||
- [ ] replicated history bounded.
|
||||
|
||||
---
|
||||
|
||||
## 9. Measurement plan
|
||||
|
||||
- **long-session soak**: fifty to a hundred spawn, despawn and transition cycles,
|
||||
looking for monotonic growth rather than an absolute number;
|
||||
- per-frame allocation counters in the hot scenes;
|
||||
- the allocator's stomp mode **first** when a dangling-pointer class is
|
||||
suspected — it is a binary answer, and it costs one run;
|
||||
- profiler counters for boundedness curves over time, which is the cheapest way
|
||||
to turn "is this cache bounded?" into a graph;
|
||||
- network statistics and a network profiler for the replication axes, measured
|
||||
separately from each other;
|
||||
- server and client measured independently; never generalize one to the other.
|
||||
|
||||
**Report growth curves, not single samples.** A cache that is 4 MB after one
|
||||
minute and 40 MB after ten is a defect regardless of where it eventually stops.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The patterns and failure modes in `references/` come from a source audit of the
|
||||
runtime, feedback, UI and weapon systems of Epic's Lyra Starter Game on Unreal
|
||||
Engine 5.6, read as source rather than profiled. No profiler was run, so there
|
||||
are no timing numbers here by construction — the findings are structural, and the
|
||||
recipes tell you how to measure them yourself.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill. Each
|
||||
entry carries a stable identifier (`RC-01` and up) that resolves back to the
|
||||
audited location, so any specific claim can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The defect classes are general;
|
||||
the specific instances are evidence, not guarantees about other versions. Re-run
|
||||
the detection recipes against your own tree before acting on any specific claim.
|
||||
+384
@@ -0,0 +1,384 @@
|
||||
# Failure modes: runtime allocation and caching
|
||||
|
||||
Ten ways a running game allocates more each minute, keeps what it should drop, or
|
||||
dereferences memory that moved.
|
||||
|
||||
The shared property splits this list in two, and the split is worth stating
|
||||
because it changes how you look for each half.
|
||||
|
||||
**The dangling-pointer half is data-dependent.** It needs a second key inserted
|
||||
while a first entry is still live, or a new statistic to appear mid-session.
|
||||
Neither happens in a short test, both happen in a match, and the crash lands far
|
||||
from the cause. A null check never catches it, because the pointer is not null.
|
||||
|
||||
**The growth half is time-dependent.** Nothing is wrong at any instant; the
|
||||
defect is a slope. A single memory sample cannot show it and a single frame
|
||||
cannot contain it, which is why every recipe below asks for two measurements
|
||||
rather than one.
|
||||
|
||||
Recipes use `rg` from a project source root, and were executed against the
|
||||
audited project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## Pointers into moving storage
|
||||
|
||||
### RC-01 - A raw pointer into a map value, held across frames
|
||||
|
||||
**Mechanism.** A reference to a map value is taken, its address is stored in a
|
||||
long-lived entry, and the map later gains a key. The map reallocates; the stored
|
||||
address points into freed memory.
|
||||
|
||||
**Why it is silent.** Until a second key arrives, the address stays valid and
|
||||
everything works. The window is narrow and specific — a live entry plus an insert
|
||||
— so ordinary play produces it and ordinary testing does not.
|
||||
|
||||
**Why the obvious check misses it.** The code often *has* a check, and the check
|
||||
passes: the pointer is not null, it is non-null garbage. Reading the storing line
|
||||
shows a reference to a pool, which is exactly what the entry needs. The defect is
|
||||
a property of the container's reallocation behaviour, which is nowhere in view.
|
||||
|
||||
**Symptom.** A crash inside the pooled object's own method, under load, with a
|
||||
stack that looks like the pool is corrupt. In the audited project the window is a
|
||||
one-second component lifespan plus a second visual style — trivially produced by
|
||||
mixed normal and critical damage in the same second.
|
||||
|
||||
**Detect.** Find addresses taken from containers and stored:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A4 "FindOrAdd\(|\.Find\(" --glob "*.cpp" . | rg "&\w+|Emplace\(|= &"
|
||||
rg -n "\w+\*\s+\w+ = nullptr;" --glob "*.h" . | rg -i "pool|cache|list|entry"
|
||||
```
|
||||
|
||||
The second search finds the *declaration* side — a bare pointer member whose name
|
||||
suggests it points at a collection — which is often easier to spot than the
|
||||
assignment.
|
||||
|
||||
**Guardrail.** Store a key or an index and resolve at use. If a stable address is
|
||||
genuinely required, use a container that guarantees stable element addresses and
|
||||
say so at the declaration. When you suspect one, run under the allocator's stomp
|
||||
mode: it converts a probabilistic corruption into a deterministic crash on the
|
||||
right line, and costs one run.
|
||||
|
||||
---
|
||||
|
||||
### RC-02 - A cache pointer handed out and stored by a widget
|
||||
|
||||
**Mechanism.** A subsystem returns a pointer into its own map so a widget can
|
||||
read samples cheaply. The widget stores it. The map grows a key when a new data
|
||||
source appears — a network connection, an optional module — and reallocates.
|
||||
|
||||
**Why it is silent.** The optimization is real and the pointer is valid at hand-
|
||||
out. Growth is lazy and event-driven: the map gains keys minutes into a session,
|
||||
when a connection is established or a feature is enabled, long after the widget
|
||||
cached its pointer.
|
||||
|
||||
**Why the obvious check misses it.** Both sides are idiomatic. Returning a
|
||||
pointer to avoid copying a sample buffer is good practice; caching it to avoid a
|
||||
lookup per frame is also good practice. Neither side knows the map is still
|
||||
growing, and nothing in either signature says the pointer has a lifetime.
|
||||
|
||||
**Symptom.** A crash in the widget's paint path, appearing only after the session
|
||||
has been running long enough to acquire a new statistic — and therefore never in
|
||||
a short repro. Compounded when the subsystem's teardown resets the tracker while
|
||||
the widget still holds the pointer.
|
||||
|
||||
**Detect.** Find getters returning interior pointers, then find who stores them:
|
||||
|
||||
```bash
|
||||
rg -n "const \w+\* \w+::Get\w*Data\(" --glob "*.cpp" . -A4 | rg "\.Find\("
|
||||
rg -n "const \w+\* \w+ = nullptr;" --glob "*.h" . | rg -i "widget|graph|view"
|
||||
```
|
||||
|
||||
A getter returning `Find` on a member map, plus a member pointer of that type in
|
||||
a widget, is the pair.
|
||||
|
||||
**Guardrail.** Return a copy of the small value, or a handle the subsystem can
|
||||
invalidate, or require the caller to re-fetch each frame. If a pointer must
|
||||
escape, give the subsystem an invalidation broadcast and make subscribing
|
||||
mandatory.
|
||||
|
||||
---
|
||||
|
||||
## Unbounded growth
|
||||
|
||||
### RC-03 - Read, append, write back — with no clear
|
||||
|
||||
**Mechanism.** Each event fetches an entire collection from another system,
|
||||
appends one element, and writes the whole thing back. Nothing ever clears it.
|
||||
|
||||
**Why it is silent.** Every individual call is correct and fast. The array is a
|
||||
legitimate data channel to the effects system, and its contents are consumed
|
||||
visually, so nobody looks for a removal path.
|
||||
|
||||
**Why the obvious check misses it.** The two lines read as an idiomatic
|
||||
accessor pair. The cost — two full copies per event — is invisible at small sizes,
|
||||
and the growth is invisible without a second measurement. Reviewing the function
|
||||
shows three lines of correct code.
|
||||
|
||||
**Symptom.** Cost per event rising through a match, total work quadratic in event
|
||||
count, and memory that never returns. In a damage-number path this means the
|
||||
hundredth hit costs measurably more than the first.
|
||||
|
||||
**Detect.** Find get-append-set triples on the same key:
|
||||
|
||||
```bash
|
||||
rg -n -A3 "= \w*::Get\w*Array\w*\(" --glob "*.cpp" . | rg "Add\(|SetNiagara|Set\w*Array"
|
||||
```
|
||||
|
||||
Then, for each hit, search the file for any clear:
|
||||
|
||||
```bash
|
||||
rg -n "Empty\(\)|Reset\(\)|SetNum\(0\)|Clear" <that-file>
|
||||
```
|
||||
|
||||
An append path with no clear anywhere in the file is the finding.
|
||||
|
||||
**Guardrail.** Give the collection an owner and an eviction rule — a cap, a
|
||||
time-out, or a clear on the event that logically ends its contents. If the
|
||||
consuming system owns the lifetime, ask it for a removal API rather than
|
||||
assuming one.
|
||||
|
||||
---
|
||||
|
||||
### RC-04 - Re-append every previous element on every trigger
|
||||
|
||||
**Mechanism.** A component that spawns audio and effect components copies all
|
||||
previously active ones into temporary arrays, appends the new ones, and writes
|
||||
the union back — with no removal of finished components.
|
||||
|
||||
**Why it is silent.** Effects play and stop correctly, because stopping is the
|
||||
component's own business. What accumulates is the *references*, and those are
|
||||
invisible until something counts them.
|
||||
|
||||
**Why the obvious check misses it.** The function reads as careful state
|
||||
management: gather, combine, store. The absent operation is a filter for finished
|
||||
components, and absence in a function that is otherwise thorough is the hardest
|
||||
kind to see.
|
||||
|
||||
**Symptom.** Thousands of retained effect components after a match. Because the
|
||||
arrays are strong references, none of them can be collected, so this is
|
||||
simultaneously a memory leak and a quadratic allocation profile in the hottest
|
||||
cosmetic path — footsteps, at a few per second.
|
||||
|
||||
**Detect.** Find union-rebuild patterns and check for a prune:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A6 "\.Append\(" --glob "*.cpp" . | rg "Empty\(\)|Reset\(\)"
|
||||
rg -n "UPROPERTY\(Transient\)" -A2 --glob "*.h" . | rg "TArray<TObjectPtr<U\w*(Audio|Niagara|Particle)"
|
||||
```
|
||||
|
||||
The second search alone is worth running in any project: a strong transient array
|
||||
of effect components is a retention decision, and it needs a removal path.
|
||||
|
||||
**Guardrail.** Filter finished components on each pass, or use weak references
|
||||
and let the components own their own lifetime. Prefer the reset-without-freeing
|
||||
operation to the one that releases the buffer — see RC-06.
|
||||
|
||||
---
|
||||
|
||||
### RC-05 - Rebuilding descriptors on a timer instead of diffing
|
||||
|
||||
**Mechanism.** A periodic scan removes all its descriptor objects and creates new
|
||||
ones for the same targets, several times a second.
|
||||
|
||||
**Why it is silent.** The result is correct: the right descriptors exist
|
||||
afterwards. Object creation is fast enough that no single frame stands out.
|
||||
|
||||
**Why the obvious check misses it.** The function reads as a clean rebuild —
|
||||
clear, then populate — which is a legitimate and common shape. The cost is in the
|
||||
frequency, which is set elsewhere as a scan rate, and in the downstream churn:
|
||||
each create and destroy propagates through observers into widget pooling and
|
||||
canvas slot management.
|
||||
|
||||
**Symptom.** Steady allocation and observer churn while the player stands still
|
||||
in a crowd of interactive objects. Profiles show constant low-level cost with no
|
||||
obvious owner.
|
||||
|
||||
**Detect.** Find object creation inside timer-driven update functions:
|
||||
|
||||
```bash
|
||||
rg -n -B10 "NewObject<" --glob "*.cpp" . | rg "Update\w+\(|OnTimer|ScanRate"
|
||||
rg -n "ScanRate|Interval|SetTimer" --glob "*.h" --glob "*.cpp" . | rg -i "0\.[0-9]"
|
||||
```
|
||||
|
||||
Also check the change-detection that gates the rebuild: comparison logic that
|
||||
sorts only when sizes match will report a change on any reorder, making the
|
||||
"only when changed" guard ineffective.
|
||||
|
||||
**Guardrail.** Diff against a stable key and reuse descriptors. If a rebuild is
|
||||
genuinely simpler, measure it at the worst-case target count before accepting the
|
||||
scan rate.
|
||||
|
||||
---
|
||||
|
||||
### RC-06 - Freeing the buffer in a hot path
|
||||
|
||||
**Mechanism.** The container operation that releases memory is used where the one
|
||||
that keeps capacity was intended, in code that runs many times per second.
|
||||
|
||||
**Why it is silent.** Both operations leave the container empty and both are
|
||||
correct. The difference is whether the next fill reallocates.
|
||||
|
||||
**Why the obvious check misses it.** The two calls are one word apart and read
|
||||
identically at a glance. Nothing distinguishes a hot path from a cold one in the
|
||||
call itself, and in teardown code the freeing version is the right choice — so the
|
||||
same line is correct in one file and wrong in another.
|
||||
|
||||
**Symptom.** Allocator churn proportional to event frequency, showing up as
|
||||
generic allocation cost rather than attributable to the system causing it.
|
||||
|
||||
**Detect.** Find the freeing operation and classify each site by frequency:
|
||||
|
||||
```bash
|
||||
rg -n "\.Empty\(\)" --glob "*.cpp" . -B6 | rg -i "tick|paint|onhit|effect|update"
|
||||
```
|
||||
|
||||
Teardown sites are fine. Anything reached more than once a second is the finding.
|
||||
In the audited project the two hottest cosmetic arrays are freed and refilled on
|
||||
every footstep.
|
||||
|
||||
**Guardrail.** Reset in hot paths, free at teardown, and write the reason in a
|
||||
comment when it is not obvious which one a site is.
|
||||
|
||||
---
|
||||
|
||||
### RC-07 - Allocating in a paint or arrange path
|
||||
|
||||
**Mechanism.** A per-frame layout or paint function builds a local array, without
|
||||
reserving, and lets it reallocate as it fills.
|
||||
|
||||
**Why it is silent.** It is correct, local and easy to read. The cost is a
|
||||
handful of small allocations per frame, which no single profile sample
|
||||
attributes.
|
||||
|
||||
**Why the obvious check misses it.** A local array in a function is the most
|
||||
ordinary thing in the file. Knowing it matters requires knowing the function is
|
||||
called once or more per frame — which is a property of the caller.
|
||||
|
||||
**Symptom.** A frame-time floor that nothing accounts for, and allocator pressure
|
||||
that scales with the number of on-screen elements.
|
||||
|
||||
**Detect.** Find containers declared in per-frame widget paths:
|
||||
|
||||
```bash
|
||||
rg -n -A12 "::OnPaint\(|::OnArrangeChildren\(|::Tick\(" --glob "*.cpp" . \
|
||||
| rg "TArray<|TMap<"
|
||||
rg -n "static TArray|static TMap" --glob "*.cpp" . | rg -i "widget|slate|paint"
|
||||
```
|
||||
|
||||
The second search finds the wrong fix for the first: a static container avoids the
|
||||
allocation and introduces sharing between every instance of the widget, plus a
|
||||
thread-safety assumption nobody wrote down. In the audited project both appear —
|
||||
one path allocates per frame, another uses a static.
|
||||
|
||||
**Guardrail.** A mutable member with a reset at the top of the function. Not a
|
||||
local, not a static.
|
||||
|
||||
---
|
||||
|
||||
## Semantics
|
||||
|
||||
### RC-08 - A prefilled ring buffer reporting statistics
|
||||
|
||||
**Mechanism.** A fixed-capacity sample cache is filled with zeroes at
|
||||
construction. Minimum, average and related statistics include the prefix until
|
||||
the buffer wraps.
|
||||
|
||||
**Why it is silent.** Every value returned is a real number computed correctly
|
||||
from the buffer's contents. There is no error state for "not enough samples yet",
|
||||
because the buffer is always full.
|
||||
|
||||
**Why the obvious check misses it.** The implementation is correct for a full
|
||||
buffer, and a reviewer checks the arithmetic rather than the initial condition.
|
||||
The distortion disappears after a few seconds, so it is gone before anyone
|
||||
investigates.
|
||||
|
||||
**Symptom.** A minimum of zero and a depressed average on any graph for the first
|
||||
seconds after it appears — which is exactly when someone is looking at it.
|
||||
|
||||
**Detect.** Find prefilled buffers and check whether the statistics account for
|
||||
fill level:
|
||||
|
||||
```bash
|
||||
rg -n "AddZeroed\(|SetNumZeroed\(" --glob "*.h" --glob "*.cpp" .
|
||||
rg -n -A6 "GetMin\(\)|GetAverage\(\)" --glob "*.h" . | rg "Num\(\)|SampleSize|NumValid"
|
||||
```
|
||||
|
||||
Statistics dividing by capacity rather than by a fill count is the finding.
|
||||
|
||||
**Guardrail.** Track the number of valid samples and compute over that window.
|
||||
Also check the accessor names: an index-based "current" that returns the *next
|
||||
write position* is the oldest sample, not the newest — a naming trap worth
|
||||
checking in any ring buffer.
|
||||
|
||||
---
|
||||
|
||||
### RC-09 - An identifier that is a count, at three different widths
|
||||
|
||||
**Mechanism.** A batch identifier is assigned from the number of in-flight
|
||||
batches rather than a monotonic sequence, so identifiers are reused as batches
|
||||
complete. Separately, the same value is one width at the source, another in the
|
||||
message, and a third in storage.
|
||||
|
||||
**Why it is silent.** With one or two batches in flight, counts and sequences are
|
||||
indistinguishable. The width mismatch is invisible until the value exceeds the
|
||||
smallest type — which needs sustained fire and a degraded connection.
|
||||
|
||||
**Why the obvious check misses it.** Each half is defensible alone: a count is a
|
||||
plausible identifier when everything completes promptly, and each width is
|
||||
reasonable in its own context. Confirming means reading a declaration, a
|
||||
constructor and a message signature that are not adjacent — in the audited project
|
||||
the storage width and the message width are sixteen lines apart in the same
|
||||
header, which is exactly far enough.
|
||||
|
||||
**Symptom.** Confirmations applied to the wrong batch when identifiers are
|
||||
reused, and — once the counter passes the narrow type's range — confirmations that
|
||||
match nothing at all, so the pending list grows without bound.
|
||||
|
||||
**Detect.** Trace one identifier through the whole path and compare widths:
|
||||
|
||||
```bash
|
||||
I='UniqueId'
|
||||
rg -n "\b$I\b" --glob "*.h" --glob "*.cpp" .
|
||||
```
|
||||
|
||||
Read the output as a table of types. Then check the assignment: an identifier
|
||||
taken from a `Num()` or a count getter is a reused identifier.
|
||||
|
||||
**Guardrail.** One monotonic sequence, one width, asserted at every boundary.
|
||||
Define wrap behaviour. Give the pending list a timeout and a drain, so a matching
|
||||
failure degrades instead of accumulating.
|
||||
|
||||
---
|
||||
|
||||
### RC-10 - Broadcasting before the state is consistent
|
||||
|
||||
**Mechanism.** A registry broadcasts an "added" event before inserting the item
|
||||
into its own collection. A listener that responds by enumerating the collection
|
||||
sees a state that does not include the item it was just told about.
|
||||
|
||||
**Why it is silent.** For a listener that only uses the event payload, the order
|
||||
is irrelevant and everything works. The defect needs a listener that also reads
|
||||
the collection — typically one doing late binding, catching up on existing items.
|
||||
|
||||
**Why the obvious check misses it.** Two adjacent lines, both correct, in an
|
||||
order that reads naturally: announce, then record. Reversing them looks like a
|
||||
stylistic preference until you know a listener enumerates.
|
||||
|
||||
**Symptom.** An item lost or added twice, depending on which listener ran when.
|
||||
In the audited project a canvas both handles the event and enumerates existing
|
||||
items on attach, and it adds to its own list without a duplicate check.
|
||||
|
||||
**Detect.** Find broadcasts that precede the mutation they announce:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A2 "\.Broadcast\(" --glob "*.cpp" . | rg "Add\(|Remove\(|Emplace\("
|
||||
```
|
||||
|
||||
A broadcast on the line *before* the insert is the finding. Then check the
|
||||
listener for a duplicate guard.
|
||||
|
||||
**Guardrail.** Mutate, then broadcast. Listeners must observe a consistent
|
||||
container, and any listener that also enumerates on attach needs an idempotent
|
||||
add.
|
||||
+260
@@ -0,0 +1,260 @@
|
||||
# Patterns: runtime allocation and caching in a measured product
|
||||
|
||||
A worked reading of the runtime allocation, pooling, caching and replication-cost
|
||||
surface of one shipped project.
|
||||
|
||||
This file differs from the others in one respect worth flagging: **two of its
|
||||
findings are memory-safety defects, not design smells.** Everywhere else in this
|
||||
bundle the failures are silent-but-benign until adopted. Here, two of them are
|
||||
dangling pointers in shipping paths, and they are silent for the ordinary reason —
|
||||
freed memory usually still contains the right bytes.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — requires a runtime experiment.
|
||||
|
||||
**Nothing here was profiled.** Every claim is structural. That is a limitation
|
||||
and also the point: both critical findings were found by reading, and both are
|
||||
confirmable in one run under an allocator that poisons freed pages.
|
||||
|
||||
---
|
||||
|
||||
## 1. The same defect twice, in two unrelated subsystems
|
||||
|
||||
A raw pointer taken from a container's storage and held across frames. It appears
|
||||
independently in a cosmetic feedback pool and in a statistics cache
|
||||
**[measured]**.
|
||||
|
||||
**Case one — the pool.** A map from mesh to a pooled component list. Code takes a
|
||||
reference to a map value, stores its address in a live-entry struct, and
|
||||
dereferences it up to a second later when the entry expires **[measured]**.
|
||||
|
||||
**Case two — the statistics cache.** A map from statistic type to a fixed-size
|
||||
sample cache. A getter returns the address of a map value; a Slate widget stores
|
||||
that pointer as a member and dereferences it **every frame during paint**
|
||||
**[measured]**.
|
||||
|
||||
**[derived]** Both are correct until the map rehashes, and both maps grow after
|
||||
the pointer is taken:
|
||||
|
||||
- the pool map gains a key the first time a second visual style appears — mixed
|
||||
normal and critical damage within one second is enough;
|
||||
- the statistics map is **populated lazily**, gaining keys as a game state
|
||||
appears, a player state appears, a network connection appears, and a latency
|
||||
module is enabled **[measured]**. A widget that cached a pointer before any of
|
||||
those events is holding a stale address afterwards.
|
||||
|
||||
The second case is worse in a specific way: the dereference happens in a paint
|
||||
path that the widget declares as always volatile, so it runs every frame with no
|
||||
invalidation mechanism **[measured]**.
|
||||
|
||||
**[derived]** The general rule this yields is worth more than the two instances:
|
||||
**a pointer obtained from a container and stored beyond the current scope is a
|
||||
review blocker, regardless of how the container is used.** Store a key or an
|
||||
index and resolve at use. The audited code even guards one of them with a
|
||||
non-null check, which cannot detect this failure — the pointer is non-null
|
||||
garbage. Recipes: RC-01, RC-02.
|
||||
|
||||
---
|
||||
|
||||
## 2. Unbounded growth, three shapes
|
||||
|
||||
| Shape **[measured]** | Where | What it costs |
|
||||
|---|---|---|
|
||||
| Read whole array, append one, write whole array back — never cleared | damage number effect data | Linear per event, quadratic per session, two full copies per hit |
|
||||
| Re-append every previously active component on each trigger, never prune finished ones | footstep and impact effects | Strong references pin every one-shot sound and effect for the whole match |
|
||||
| Pending protocol entries removed only on a matching confirmation | hit-marker batches | Grows without bound once matching stops working (§4) |
|
||||
|
||||
**[derived]** The first two share a structure that is easy to miss in review: the
|
||||
code that grows the collection is also the code that appears to manage it. Reading
|
||||
`Empty()` followed by `Append()` looks like a refresh. It is a refresh that
|
||||
re-adds everything, and the array is a strong property, so nothing is ever
|
||||
released.
|
||||
|
||||
The second one also uses `Empty()` rather than `Reset()` in a path that runs at
|
||||
footstep frequency **[measured]**, which frees the buffer and guarantees a
|
||||
reallocation on the next call — a smaller problem sitting inside a larger one.
|
||||
Recipes: RC-03, RC-04, RC-06.
|
||||
|
||||
---
|
||||
|
||||
## 3. Pools without lifecycle
|
||||
|
||||
The component pool grows to the peak concurrent count and **never shrinks**
|
||||
**[measured]**. One burst of two hundred simultaneous hits leaves two hundred
|
||||
components allocated for the rest of the match.
|
||||
|
||||
A second pool in the indicator system is fixed at ten entries with surplus
|
||||
requests simply hidden **[measured]** — which is a legitimate policy, stated
|
||||
nowhere.
|
||||
|
||||
**[derived]** The contrast between the two is the lesson: one pool has no
|
||||
high-water policy and the other has an undocumented one. Both are the same
|
||||
omission — **a pool is not correct until its exhaustion and trim behaviour are
|
||||
written down.** Neither project reader can answer "what happens at peak?" without
|
||||
reading the implementation.
|
||||
|
||||
Related and cheaper to fix: the live queue removes from the front of an array,
|
||||
shifting the tail on every timer tick **[measured]**, where a ring buffer would
|
||||
not move anything.
|
||||
|
||||
---
|
||||
|
||||
## 4. An identifier that is a count
|
||||
|
||||
The clearest single defect in the audit, and a good example of why protocol
|
||||
identifiers deserve their own review rule.
|
||||
|
||||
A hit-marker batch is identified by **the current number of unconfirmed batches**
|
||||
rather than by a monotonically increasing sequence **[measured]**.
|
||||
|
||||
**[derived]** The failure needs no packet loss to reason about. Batch A takes
|
||||
identifier 0. Batch B takes 1. A is confirmed and removed. C now takes 1 — the
|
||||
same identifier B is still holding. The confirmation handler matches on the first
|
||||
entry with that identifier, so a confirmation for one batch can be applied to
|
||||
another.
|
||||
|
||||
Compounding it, the same value is carried at **three different widths** along one
|
||||
path: signed 32-bit at the source, unsigned 16-bit in the network call, and
|
||||
unsigned 8-bit in storage **[measured]** — the last two declared sixteen lines
|
||||
apart in the same header. Past 255 unconfirmed entries the storage wraps and the
|
||||
network value does not, so matching stops entirely and the pending list grows
|
||||
without bound.
|
||||
|
||||
Recipe: RC-09. The rules that fall out: identifiers are sequences, one width for
|
||||
the whole path, and pending lists get a timeout and a drain policy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Allocation in paint and arrange paths
|
||||
|
||||
Three instances, all in per-frame Slate code **[measured]**:
|
||||
|
||||
- a local array built and sorted on every arrange, without reserving;
|
||||
- a full array copy returned by value from a getter called during paint;
|
||||
- a **static** array used as scratch space in a draw helper.
|
||||
|
||||
**[derived]** The third is the interesting one, because it is an optimisation. It
|
||||
avoids per-call allocation and introduces two new problems: the buffer is shared
|
||||
by every instance of the widget in every window, and it is not thread-safe in a
|
||||
context where paint is not guaranteed to be on one thread. The correct form —
|
||||
a mutable member reset per call — is the same cost and none of the risk.
|
||||
|
||||
Recipe: RC-07. Also visible in the same code: a component lookup by class
|
||||
performed both in paint and in tick, twice per frame, uncached **[measured]**.
|
||||
|
||||
---
|
||||
|
||||
## 6. Rebuild instead of diff
|
||||
|
||||
An interaction system destroys every descriptor object and creates new ones on a
|
||||
tenth-of-a-second timer, for the same targets **[measured]**. Each cycle
|
||||
allocates objects, invalidates observers, and drives a full add/remove cycle
|
||||
through the indicator canvas.
|
||||
|
||||
**[derived]** The change-detection that gates this is itself unreliable: the
|
||||
comparison sorts the candidate list only when the two lists have equal length
|
||||
**[measured]**, so an equal-length reordering reports a change that did not
|
||||
happen. In a crowd of interactable objects the rebuild runs ten times a second.
|
||||
|
||||
The fix is the general one: compute the desired set, diff against the current
|
||||
set, and add and remove the difference. Recipe: RC-05.
|
||||
|
||||
A related ordering defect sits in the same subsystem: the add notification is
|
||||
broadcast **before** the element is inserted **[measured]**, so a listener that
|
||||
enumerates existing elements during that window either misses it or receives it
|
||||
twice — and the receiving side adds without a duplicate check. Recipe: RC-10.
|
||||
|
||||
---
|
||||
|
||||
## 7. Statistics that are wrong before the window fills
|
||||
|
||||
The sample cache is pre-filled with zeroes at construction **[measured]**. Until
|
||||
the ring wraps — which takes as many frames as its capacity — the minimum reads
|
||||
zero and the average is divided by the full capacity rather than by the number of
|
||||
real samples **[measured]**.
|
||||
|
||||
**[derived]** So the first seconds of every session report a minimum of zero and
|
||||
an artificially low average, on a performance overlay whose entire purpose is to
|
||||
be read during those seconds.
|
||||
|
||||
There is a second, quieter trap in the same class: a method named for the current
|
||||
sample returns the **oldest** one, because the index points at the next write
|
||||
position **[measured]**. The correctly-named sibling exists and is what the
|
||||
project uses, which is why nobody noticed. Recipe: RC-08.
|
||||
|
||||
The cache itself is well bounded — capacity times sample size times statistic
|
||||
count is on the order of tens of kilobytes **[derived]** — so this is a
|
||||
correctness finding, not a memory one. Both matter; conflating them is how
|
||||
"performance work" produces no performance.
|
||||
|
||||
---
|
||||
|
||||
## 8. Replication: name the axis
|
||||
|
||||
The audit's replication half produced no new defects and one clarification worth
|
||||
carrying, because it is routinely got backwards:
|
||||
|
||||
| Mechanism | Memory | CPU | Bandwidth |
|
||||
|---|---|---|---|
|
||||
| Excluding presentation on a dedicated server | down | down | **unchanged** |
|
||||
| Delta-serialised arrays | **up** (per-item key, per-connection state) | **up** (delta walk) | down |
|
||||
| Subobject replication | down | down | down |
|
||||
| Dormancy, relevancy, replication graph | — | down | down |
|
||||
|
||||
**[derived]** Two of these surprise people. Excluding cosmetics on a server saves
|
||||
memory and CPU and changes bandwidth by exactly zero, because the intent still
|
||||
replicates. And delta serialisation is a **bandwidth** optimisation that costs
|
||||
memory and CPU — adopting it to "reduce replication cost" without naming the axis
|
||||
can make the measured problem worse.
|
||||
|
||||
The genuine object-count reducer is subobject replication: many owned items, zero
|
||||
additional actors and channels.
|
||||
|
||||
Before claiming any replication optimisation: state which axis improves and which
|
||||
regresses, measure that axis specifically, and **confirm the mechanism is
|
||||
actually enabled** — see `ue-multiplayer-authority`, NA-15, for a replication
|
||||
graph that ships disabled with configured routing.
|
||||
|
||||
---
|
||||
|
||||
## 9. What to do with this list
|
||||
|
||||
In order of ratio of certainty to effort:
|
||||
|
||||
1. **Run once under an allocator that poisons freed memory**, with a mixed-style
|
||||
damage burst and with a latency module toggled on while the performance
|
||||
overlay is open. Those two scenarios turn RC-01 and RC-02 from arguments into
|
||||
crashes at the right line. **[open]** — this is the one experiment that
|
||||
settles the two critical findings.
|
||||
2. **Add five counters to the existing profiling category** — pool size, live
|
||||
entries, pending protocol entries, active effect components, indicator count.
|
||||
The category and an example already exist in the audited project
|
||||
**[measured]**; extending it is three lines each and produces boundedness
|
||||
curves over a match, which is what actually distinguishes a cache from a leak.
|
||||
3. **Grep for the two structural patterns** — stored container pointers, and
|
||||
`Empty()` in hot paths. Both are one command and both found real instances
|
||||
here.
|
||||
4. Everything else is ordinary optimisation and can wait for a profile.
|
||||
|
||||
**[derived]** Note the shape of that list: the cheapest actions are the ones that
|
||||
convert reading into evidence. A structural audit without step 1 produces a
|
||||
plausible argument; with it, a defect report.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source in
|
||||
a single workspace: feedback and effect components, the indicator system, the
|
||||
performance statistics subsystem, weapon state and targeting, equipment and
|
||||
inventory, and the replication surface of each. Source addresses stay in the
|
||||
research archive that produced this skill; each `RC-` identifier resolves back to
|
||||
the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
No profiler was run and no allocator instrumentation was used. Every finding is
|
||||
structural, derived from reading; the growth characterisations are arguments
|
||||
about code shape, not measurements. Quantities such as pool capacities and buffer
|
||||
sizes are declared values read from source. Re-run the recipes and the two
|
||||
scenarios above against your own tree before acting on any of it.
|
||||
@@ -0,0 +1,291 @@
|
||||
---
|
||||
name: ue-streaming-and-platform-budgets
|
||||
description: >-
|
||||
Design or audit GPU and streaming budgets in Unreal Engine: texture, mesh and
|
||||
virtual texture streaming, streaming pool sizing, texture LOD groups,
|
||||
scalability buckets, device profiles and console-variable priority, world
|
||||
partition and level instance lifetime, and the boundary between disk, CPU
|
||||
memory and video memory. Use when targeting a platform memory ceiling, tuning
|
||||
quality tiers, adding a background or preview level, or investigating why a
|
||||
quality setting has no effect.
|
||||
---
|
||||
|
||||
# UE streaming and platform budgets
|
||||
|
||||
The invariant:
|
||||
|
||||
> A streaming pool is a **permission**, not a measurement. A quality setting is a
|
||||
> **request**, not an outcome. Both are arbitrated by device profiles and
|
||||
> console-variable priority before anything reaches the GPU.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes: [failure modes](references/failure-modes.md).
|
||||
|
||||
Related skills: `ue-asset-loading-and-memory`, `ue-game-settings-architecture`,
|
||||
`ue-runtime-allocation-and-caching`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Four budgets, four instruments
|
||||
|
||||
| Budget | What it physically is | Instrument |
|
||||
|---|---|---|
|
||||
| **A. Disk / cook** | bytes in the shipped package | build and chunk sizes |
|
||||
| **B. CPU memory** | object graph plus bulk data in system RAM | memory report, object listing |
|
||||
| **C. Video memory** | live render resources: textures, buffers, render-target pool, virtual-texture pool, shadow pages, ray-tracing structures | platform GPU profiler |
|
||||
| **D. Streaming pool** | an accounting **ceiling** the streamer tries to stay under | streaming stat |
|
||||
|
||||
Four rules that prevent most wrong conclusions:
|
||||
|
||||
- **A is not D.** Clamping a texture group's maximum size shrinks what is cooked.
|
||||
The pool is unchanged.
|
||||
- **D is not C.** The pool covers streamable mips only. Render targets, virtual
|
||||
texture pools, shadow pages and acceleration structures are in C and invisible
|
||||
to the streaming stat.
|
||||
- **B is not C.** A loaded texture object need not have any resident mips.
|
||||
- **Never report a pool size as memory used.** It is the single most common
|
||||
error in this area, and it makes every downstream number meaningless.
|
||||
|
||||
---
|
||||
|
||||
## 2. Console-variable priority is the real arbiter
|
||||
|
||||
Effective value is not "last write wins"; sources are ranked. In practice:
|
||||
|
||||
```text
|
||||
command line > console > device profile > system ini > project ini
|
||||
> game setting > scalability > constructor
|
||||
```
|
||||
|
||||
with one exception that matters enormously: a device profile applies variables
|
||||
whose names begin with the scalability prefix at **scalability** priority, and
|
||||
everything else at **device-profile** priority.
|
||||
|
||||
### The consequence
|
||||
|
||||
A value set by a device profile **cannot be changed** by anything the
|
||||
scalability system does afterwards — which includes every in-game quality
|
||||
setting. The same setting is therefore mutable on one tier and frozen on
|
||||
another, with nothing in the UI explaining the difference. Recipe: ST-01.
|
||||
|
||||
### Gates
|
||||
|
||||
- [ ] every quality-affecting variable has a documented intended priority source;
|
||||
- [ ] prefixed and unprefixed names are used deliberately, not by habit;
|
||||
- [ ] per-tier overrides are complete across tiers, or partial with a comment
|
||||
saying why;
|
||||
- [ ] the settings UI can express "clamped by platform" rather than silently
|
||||
failing;
|
||||
- [ ] requested and effective values are both observable at runtime.
|
||||
|
||||
**The one command that settles it:** dump variables with their source. Anything
|
||||
else is inference.
|
||||
|
||||
---
|
||||
|
||||
## 3. Pool sizing, and its detachment from real memory
|
||||
|
||||
Two knobs interact: the pool size, and whether the pool is clamped by actual
|
||||
video memory.
|
||||
|
||||
The trap is structural rather than accidental: **the highest quality buckets
|
||||
commonly disable the video-memory clamp**, and automatic detection promotes
|
||||
hardware to the top bucket on a compute benchmark that does not consider memory
|
||||
size at all. A fast card with modest memory therefore lands in the bucket where
|
||||
the pool budget is largest and least constrained. Recipe: ST-02.
|
||||
|
||||
Rules:
|
||||
|
||||
- decide per platform whether the pool is derived from memory or fixed;
|
||||
- if fixed, write down the assumed memory floor;
|
||||
- benchmark thresholds must consider memory, not only throughput;
|
||||
- test the top bucket on a low-memory card, not only on the reference machine.
|
||||
|
||||
### Sharing the pool
|
||||
|
||||
If mesh streaming is enabled and no separate mesh pool is configured, geometry
|
||||
and textures compete inside one budget. That is a legitimate choice; it is not a
|
||||
legitimate surprise. Check both values on every platform. Recipe: ST-03.
|
||||
|
||||
---
|
||||
|
||||
## 4. Texture groups are the per-content lever
|
||||
|
||||
Scalability moves everything at once; groups shape individual content classes.
|
||||
Use them to encode intent: hero content larger, world content platform-scaled,
|
||||
UI unstreamed and sized exactly, debug content excluded from cook.
|
||||
|
||||
Gates: every group has a stated purpose; mobile tiers override group caps rather
|
||||
than only global quality; no content relies on a group it does not belong to;
|
||||
streaming-exempt content is deliberate and small.
|
||||
|
||||
Watch for group settings that describe a delivery mechanism the project does not
|
||||
have — optional mip levels configured with no chunk installation, for instance.
|
||||
Configuration that cannot take effect is worse than absent configuration,
|
||||
because it reads as a decision. Recipe: ST-08.
|
||||
|
||||
---
|
||||
|
||||
## 5. Build-level decisions that look like runtime knobs
|
||||
|
||||
Some streaming variables are read-only: their value is fixed at startup and
|
||||
influences what is cooked. Changing them at runtime, or from a device profile,
|
||||
does nothing and says nothing.
|
||||
|
||||
Before treating any streaming variable as tunable, check its flags. And never
|
||||
A/B a read-only variable mid-session and compare the numbers — you are comparing
|
||||
one configuration with itself. Recipe: ST-04.
|
||||
|
||||
Virtual textures belong in the same category. They change residency
|
||||
characteristics fundamentally, their pool is separate from the streaming budget,
|
||||
and a pool that does not grow on over-subscription produces thrashing rather
|
||||
than an allocation. Enabling them for some content and not other content makes
|
||||
budget comparisons invalid.
|
||||
|
||||
---
|
||||
|
||||
## 6. Device profiles: reachability before content
|
||||
|
||||
A profile hierarchy usually encodes platform base → tier → device. Three
|
||||
failures recur, and all three are silent:
|
||||
|
||||
**A section that is never selected.** Defining a profile section does not make it
|
||||
reachable; something must map to it. A tier section absent from the registry is
|
||||
configuration that reviewers read as live. Recipe: ST-05.
|
||||
|
||||
**Headers that contradict assignment.** Comment banners grouping devices under a
|
||||
tier label, while the devices beneath resolve to a different base profile. The
|
||||
comment is documentation; the base profile is behaviour. Recipe: ST-06.
|
||||
|
||||
**An override list that no configuration populates.** A user-facing profile
|
||||
variant list declared in code and set in no file means the selection logic
|
||||
composes an empty name, the override is never applied, and the preset UI never
|
||||
appears. Recipe: ST-07.
|
||||
|
||||
Gates: every profile section is reachable from the registry; tier assignment is
|
||||
asserted by a dump, not read from comments; unknown hardware has a documented
|
||||
fallback; every override hook has a real caller and real data.
|
||||
|
||||
---
|
||||
|
||||
## 7. Requested versus effective quality
|
||||
|
||||
A settings screen showing the user's request while a device profile clamps the
|
||||
result is lying by omission. Model it explicitly:
|
||||
|
||||
```text
|
||||
requested quality → platform clamp → effective quality → applied variables
|
||||
```
|
||||
|
||||
Expose both, and give a reason when they differ.
|
||||
|
||||
### Unit-test every clamp at its range boundaries
|
||||
|
||||
The measured defect worth guarding against: a compatibility check comparing a
|
||||
quality **level** (0–3) against a resolution **percentage** (default 100), so the
|
||||
constraint can never fire. Its sibling function used the correct accessor, which
|
||||
is what makes this findable — and what makes it invisible to review, since both
|
||||
lines read identically. Recipe: ST-09.
|
||||
|
||||
Also check every normalisation for a zero denominator. A remap between two
|
||||
ranges divides by their difference; a configuration that makes those equal is
|
||||
reachable from data. Recipe: ST-10.
|
||||
|
||||
---
|
||||
|
||||
## 8. World content lifetime
|
||||
|
||||
### Level instances, not second worlds
|
||||
|
||||
Background and preview content is usually a level instanced into the **current**
|
||||
world, not a second world. The cost is duplicated level content in one scene:
|
||||
actors, components, primitives and their render resources.
|
||||
|
||||
Measure by scene composition — actor and primitive counts, visible material slots
|
||||
— not by counting world contexts. "There is only one world" is true and
|
||||
irrelevant.
|
||||
|
||||
### Ownership rules
|
||||
|
||||
Every streamed-in instance has one owner and an unload path; async handles are
|
||||
retained and cancellable; failure surfaces a reason rather than a permanent wait;
|
||||
teardown runs on travel, mode change and error.
|
||||
|
||||
### Hidden render targets
|
||||
|
||||
Scene-capture preview systems allocate render targets per capture object and may
|
||||
keep rendering state resident between captures. Budget them explicitly: count ×
|
||||
resolution × format, and verify release. A subsystem with no callers that still
|
||||
instantiates per world and installs a ticker is paying for a feature nobody uses
|
||||
— unused systems must decline creation.
|
||||
|
||||
### If runtime layer state is never set, it is not runtime streaming
|
||||
|
||||
Shipping editor-authored initial state is a valid choice. Describing it as
|
||||
dynamic streaming is not, and a team that believes it has a reference
|
||||
implementation will look for one that does not exist.
|
||||
|
||||
---
|
||||
|
||||
## 9. Review checklist
|
||||
|
||||
- [ ] Four budgets separated in every claim and every report.
|
||||
- [ ] Pool settings never presented as consumption.
|
||||
- [ ] Variable priority documented per quality knob, verified by a source dump.
|
||||
- [ ] Tier overrides complete, or partial with a stated reason.
|
||||
- [ ] Memory-clamp policy explicit per platform.
|
||||
- [ ] Mesh and texture pool split decided per platform.
|
||||
- [ ] Read-only variables identified before anyone tries to tune them.
|
||||
- [ ] Every profile section reachable; assignment asserted, not commented.
|
||||
- [ ] Requested and effective quality both visible, with a reason on divergence.
|
||||
- [ ] Clamp logic unit-tested at both ends of its real range.
|
||||
- [ ] Level instances owned, released and counted.
|
||||
- [ ] Preview and capture render targets budgeted and released.
|
||||
- [ ] Unused subsystems decline creation.
|
||||
|
||||
---
|
||||
|
||||
## 10. Measurement plan
|
||||
|
||||
Static first — cheap, and finds the structural problems:
|
||||
|
||||
1. Dump effective variables **with their source** per device profile, and diff
|
||||
against intent.
|
||||
2. Verify every profile section is reachable from the registry.
|
||||
3. List texture groups and their per-tier caps.
|
||||
4. Search for runtime streaming and layer-state calls to confirm what is
|
||||
actually dynamic.
|
||||
|
||||
Runtime, on target hardware, in a packaged build:
|
||||
|
||||
- streaming stat for pool budget versus used and over-budget;
|
||||
- platform GPU profiler for real video memory — the streaming stat cannot see
|
||||
most of it;
|
||||
- **fixed camera poses** for any comparison; view-dependent metrics are otherwise
|
||||
noise;
|
||||
- one run per tier on real devices, not emulated profiles alone;
|
||||
- scene composition counts for level-instance cost.
|
||||
|
||||
**Do not compare numbers taken at different camera poses, different tiers, or
|
||||
between editor and packaged builds.** Those comparisons look quantitative and
|
||||
mean nothing.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings come from a configuration and source audit of Epic's Lyra Starter
|
||||
Game on Unreal Engine 5.6, cross-read against the engine's own scalability and
|
||||
device-profile configuration. Source addresses stay in the research archive that
|
||||
produced this skill; each entry carries a stable identifier (`ST-01`, `ST-02`, …)
|
||||
resolving back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Static configuration analysis only: no profiler was run, no device was measured,
|
||||
and no number in the failure modes is a memory measurement. Engine defaults
|
||||
quoted describe one engine version. The recipes are written to be re-run against
|
||||
your own project and engine, which is the only way any of these values become
|
||||
true for you.
|
||||
+388
@@ -0,0 +1,388 @@
|
||||
# Failure modes: streaming and platform budgets
|
||||
|
||||
Ten ways a platform budget is configured, reviewed, approved and inert.
|
||||
|
||||
The shared property here is different from the rest of this bundle, and it is the
|
||||
reason this file exists: **configuration cannot fail.** A line in an ini file
|
||||
that names an unreachable section, a variable set at a priority nothing can
|
||||
override, a clamp that compares two different units — none of these produce an
|
||||
error, a warning, or a log line. They produce a game that runs, on hardware that
|
||||
was budgeted for, using numbers nobody set.
|
||||
|
||||
A second property: **the instrument lies more often than the code.** Most wrong
|
||||
conclusions in this area come from reading a pool setting as consumption, or a
|
||||
scalability level as an outcome. Several recipes below check the instrument
|
||||
rather than the project.
|
||||
|
||||
Recipes use `rg` from a project root and were executed against the audited
|
||||
project while this file was written. Runtime recipes name the console command,
|
||||
because static analysis genuinely cannot answer some of these.
|
||||
|
||||
---
|
||||
|
||||
## Priority and arbitration
|
||||
|
||||
### ST-01 - A value set at a priority nothing can override
|
||||
|
||||
**Mechanism.** A device profile sets a streaming variable. Device-profile
|
||||
priority outranks scalability priority, so every later quality change — including
|
||||
every in-game setting — is silently discarded for that variable.
|
||||
|
||||
**Why it is silent.** Both halves work exactly as designed. The profile applies
|
||||
its value; the settings system applies its value; the higher priority wins and
|
||||
nobody is told. The setting's UI shows the user's choice, which is stored
|
||||
correctly and simply has no effect.
|
||||
|
||||
**Why the obvious check misses it.** Reading either file shows a sensible value.
|
||||
Reading both shows two sensible values with no visible conflict, because priority
|
||||
is not expressed in the files at all — it is a property of the mechanism that
|
||||
applied them. And the same setting **does** work on a tier where the profile is
|
||||
silent, so "it works on my device" is true and useless.
|
||||
|
||||
**Symptom.** A quality setting that works on some hardware and not on other
|
||||
hardware, with no pattern the player or the tester can see.
|
||||
|
||||
**Detect.** Ask the running game who set the value — this cannot be done
|
||||
statically:
|
||||
|
||||
```text
|
||||
DumpCVars # every variable with its SetBy source
|
||||
r.Streaming.PoolSize # prints the value and its source
|
||||
```
|
||||
|
||||
Then find the static side and compare coverage:
|
||||
|
||||
```bash
|
||||
rg -n "r\.Streaming\.PoolSize" Config/
|
||||
rg -n "^\[TextureQuality@" Config/ | wc -l
|
||||
```
|
||||
|
||||
In the audited project the pool is set in exactly two places — the mid tier of
|
||||
each mobile platform — and the project declares **no** per-tier texture-quality
|
||||
sections at all, so every other tier inherits the engine's values. One setting,
|
||||
opposite behaviour across tiers.
|
||||
|
||||
**Guardrail.** Document the intended source per quality variable, and assert it
|
||||
in a startup dump on each platform. If a value must be immutable, say so in the
|
||||
settings UI rather than accepting the write and dropping it.
|
||||
|
||||
---
|
||||
|
||||
### ST-02 - The top quality bucket detaches the pool from real memory
|
||||
|
||||
**Mechanism.** The highest quality buckets disable the clamp that limits the
|
||||
streaming pool to available video memory, while automatic detection promotes
|
||||
hardware into those buckets on a compute benchmark that never considers memory
|
||||
size.
|
||||
|
||||
**Why it is silent.** Both decisions are defensible alone. Disabling the clamp on
|
||||
the top bucket assumes a machine with plenty of memory; a compute benchmark is a
|
||||
reasonable proxy for GPU class. The combination targets exactly the hardware the
|
||||
assumption is wrong about: a fast chip with modest memory.
|
||||
|
||||
**Why the obvious check misses it.** The clamp lives in engine configuration and
|
||||
the thresholds live in project configuration. Each is reviewed by a different
|
||||
person, in a different file, at a different time. Neither file mentions the
|
||||
other.
|
||||
|
||||
**Symptom.** Stutter and driver-level eviction on mid-range cards, at the quality
|
||||
tier automatic detection chose for them. Reproduces on hardware the team does not
|
||||
own, because the team's machines have enough memory for the assumption to hold.
|
||||
|
||||
**Detect.** Read the thresholds, then read what the top bucket does to the clamp:
|
||||
|
||||
```bash
|
||||
rg -n "PerfIndexThresholds_TextureQuality" Config/
|
||||
rg -n "LimitPoolSizeToVRAM|r\.Streaming\.PoolSize" \
|
||||
"<engine>/Engine/Config/BaseScalability.ini"
|
||||
```
|
||||
|
||||
Confirm on hardware: set the quality level manually and watch the pool.
|
||||
|
||||
```text
|
||||
sg.TextureQuality 3
|
||||
stat streaming # pool size, required pool, over budget
|
||||
```
|
||||
|
||||
**Guardrail.** Include a memory term in the promotion thresholds, or keep the
|
||||
clamp enabled on the top bucket. Test the top bucket on a low-memory card before
|
||||
shipping the thresholds.
|
||||
|
||||
---
|
||||
|
||||
### ST-03 - Meshes and textures share a budget nobody split
|
||||
|
||||
**Mechanism.** Mesh streaming is enabled, and the separate mesh pool is left at
|
||||
its default, which means one shared pool.
|
||||
|
||||
**Why it is silent.** Sharing is a supported configuration and works. Textures and
|
||||
geometry simply compete inside one ceiling, and the streamer resolves the
|
||||
competition by lowering mips — which looks like a texture problem.
|
||||
|
||||
**Why the obvious check misses it.** Two separate settings, one enabled in a
|
||||
renderer section and one absent from every platform except mobile. Nothing
|
||||
connects them, and the absent one has a default whose meaning ("share") is only
|
||||
documented in the engine's variable description.
|
||||
|
||||
**Symptom.** Texture quality degrading on scenes with heavy geometry, on
|
||||
platforms where nobody configured a mesh budget.
|
||||
|
||||
**Detect.** Check both values on every platform, and check where each was set:
|
||||
|
||||
```bash
|
||||
rg -n "r\.MeshStreaming" Config/
|
||||
rg -n "r\.Streaming\.PoolSizeForMeshes" Config/
|
||||
```
|
||||
|
||||
In the audited project mesh streaming is enabled globally in the renderer
|
||||
settings while the mesh pool is configured only for mobile — so desktop and
|
||||
console share one pool by default.
|
||||
|
||||
**Guardrail.** Set both explicitly per platform, even when the value equals the
|
||||
default. An explicit default is a decision; an absent one is an assumption.
|
||||
|
||||
---
|
||||
|
||||
### ST-04 - A read-only variable treated as a runtime knob
|
||||
|
||||
**Mechanism.** Several streaming variables are marked read-only: their value is
|
||||
captured at startup and influences what is cooked. Setting them later — from a
|
||||
console, a device profile, or a settings screen — is accepted and ignored.
|
||||
|
||||
**Why it is silent.** The write succeeds. Nothing rejects it, nothing warns, and
|
||||
the variable reads back as unchanged, which is easy to miss when you are looking
|
||||
at frame time rather than at the variable.
|
||||
|
||||
**Why the obvious check misses it.** The variable has a name, a value and a
|
||||
description like any other. Its flags are in engine source, and nothing in the
|
||||
project's configuration indicates the difference.
|
||||
|
||||
**Symptom.** An A/B comparison of a "setting" that produces two identical
|
||||
configurations, and therefore a confident conclusion drawn from noise.
|
||||
|
||||
**Detect.** Check the flags before tuning anything:
|
||||
|
||||
```bash
|
||||
rg -n -B2 -A6 "r\.MeshStreaming|CoarseMeshStreaming" \
|
||||
"<engine>/Engine/Source/Runtime/Engine/Private/ContentStreaming.cpp" \
|
||||
| rg "ECVF_ReadOnly|ECVF_Scalability"
|
||||
```
|
||||
|
||||
At runtime, the same question is one command: print the variable and confirm the
|
||||
value changed after you set it.
|
||||
|
||||
**Guardrail.** Maintain a short list of the build-level streaming decisions in
|
||||
your project, next to the platform matrix. Anything on that list is changed in a
|
||||
build, never in a session, and never compared mid-session.
|
||||
|
||||
---
|
||||
|
||||
## Reachability
|
||||
|
||||
### ST-05 - A profile section that nothing can select
|
||||
|
||||
**Mechanism.** A device profile section is defined with a full set of values, and
|
||||
no registry entry or matching rule leads to it.
|
||||
|
||||
**Why it is silent.** The file parses, the section is well formed, and the values
|
||||
inside it are correct. It is simply never chosen, so it produces no behaviour to
|
||||
be wrong.
|
||||
|
||||
**Why the obvious check misses it.** Reviewers read sections, not the registry.
|
||||
A tier section that sits beside three reachable siblings looks exactly like them
|
||||
— same shape, same variables, same care — and the difference is a missing line in
|
||||
a list at the top of the file, or an absent mapping in engine configuration.
|
||||
|
||||
**Symptom.** A platform that never reaches its top tier, with per-tier values
|
||||
that were reviewed, approved and never applied.
|
||||
|
||||
**Detect.** Diff declared profiles against defined sections:
|
||||
|
||||
```bash
|
||||
rg -n "DeviceProfileNameAndTypes" Config/DefaultDeviceProfiles.ini
|
||||
rg -n "^\[\w+ DeviceProfile\]" Config/DefaultDeviceProfiles.ini
|
||||
```
|
||||
|
||||
Every defined section must be either declared here or present in the engine's
|
||||
base profiles. In the audited project one platform's top tier is defined in full
|
||||
and appears in neither list — every value in it is unreachable.
|
||||
|
||||
Confirm on hardware, which is the only authoritative answer:
|
||||
|
||||
```text
|
||||
dumpdeviceprofile # the active profile and its variables
|
||||
```
|
||||
|
||||
**Guardrail.** Assert reachability in a test that parses both lists. A profile
|
||||
that cannot be selected should fail the build, not be discovered on a device.
|
||||
|
||||
---
|
||||
|
||||
### ST-06 - Section headers that contradict the assignment
|
||||
|
||||
**Mechanism.** Comment banners group device sections under tier labels. The
|
||||
sections beneath name a different base profile.
|
||||
|
||||
**Why it is silent.** Comments do not execute. The devices get the tier their
|
||||
base profile names, which is a real and consistent assignment — just not the one
|
||||
the banner announces.
|
||||
|
||||
**Why the obvious check misses it.** The banner is the fastest way to read a
|
||||
600-line configuration file, and it is the wrong instrument. Checking means
|
||||
reading the base profile line of every device section, which is exactly the work
|
||||
the banner appears to save.
|
||||
|
||||
**Symptom.** Device tier assumptions in planning documents that do not match
|
||||
runtime. Performance reports grouped by a tier the devices are not in.
|
||||
|
||||
**Detect.** Ignore banners; extract the pairs:
|
||||
|
||||
```bash
|
||||
rg -n -A2 "^\[\w+ DeviceProfile\]" Config/DefaultDeviceProfiles.ini \
|
||||
| rg "DeviceProfile\]|BaseProfileName"
|
||||
```
|
||||
|
||||
Read the output as a table. In the audited project several devices sitting under
|
||||
a high-tier banner resolve to the low tier.
|
||||
|
||||
**Guardrail.** Generate the device-to-tier table from configuration rather than
|
||||
maintaining banners, and put the generated table in the performance
|
||||
documentation.
|
||||
|
||||
---
|
||||
|
||||
### ST-07 - An override list that no configuration populates
|
||||
|
||||
**Mechanism.** A platform settings object declares a list of user-facing profile
|
||||
variants. Selection logic composes a profile name from that list. No
|
||||
configuration file ever sets it.
|
||||
|
||||
**Why it is silent.** The composition handles an empty list gracefully: the
|
||||
resulting name is empty, no matching profile is found, and the override is not
|
||||
applied. The engine's startup selection stays in effect, which is correct
|
||||
behaviour and looks like intent.
|
||||
|
||||
**Why the obvious check misses it.** The code is complete and reads well —
|
||||
fallback chains, refresh-rate filtering, logging. Confirming the absence means
|
||||
searching every platform's configuration for a property that is never mentioned
|
||||
in the source that consumes it.
|
||||
|
||||
**Symptom.** A console-style quality preset UI that never appears, because it is
|
||||
only added when the variant list has more than one entry. The feature is
|
||||
"implemented" and has never run.
|
||||
|
||||
**Detect.** Separate reads from writes:
|
||||
|
||||
```bash
|
||||
P='UserFacingDeviceProfileOptions'
|
||||
rg -n "\b$P\b" Source/ Config/ Platforms/
|
||||
```
|
||||
|
||||
Hits only in a declaration and in read sites — with nothing in any configuration
|
||||
file — is the finding. In the audited project the property is read in three
|
||||
places and set in none.
|
||||
|
||||
**Guardrail.** For every data-driven selection list, add a startup log naming its
|
||||
size, and treat zero as a warning. The logging line already exists in most such
|
||||
systems; make someone read it.
|
||||
|
||||
---
|
||||
|
||||
## Configuration that cannot take effect
|
||||
|
||||
### ST-08 - Settings for a delivery mechanism the project does not have
|
||||
|
||||
**Mechanism.** Texture groups configure optional mip levels, which are cooked
|
||||
into a separate chunk. The project has no chunk installation and does not acquire
|
||||
missing chunks on load.
|
||||
|
||||
**Why it is silent.** The settings are valid, the cook honours them, and nothing
|
||||
at runtime asks for the chunk that would contain the extra mips. Everything
|
||||
behaves as if the settings were absent.
|
||||
|
||||
**Why the obvious check misses it.** The group configuration is reviewed as a
|
||||
block of texture settings, where the value is plausible. The dependency on a
|
||||
delivery mechanism is two systems away, in packaging configuration nobody reads
|
||||
during a texture pass.
|
||||
|
||||
**Symptom.** A team believes it has a high-resolution optional tier. It does not,
|
||||
and the belief survives because nothing contradicts it.
|
||||
|
||||
**Detect.** Check the consumer of the mechanism, not the producer:
|
||||
|
||||
```bash
|
||||
rg -n "OptionalMaxLODSize|OptionalLODBias" Config/
|
||||
rg -n "bShouldAcquireMissingChunksOnLoad|bBuildHttpChunkInstallData" Config/
|
||||
```
|
||||
|
||||
Optional mips configured while chunk acquisition is disabled is the finding.
|
||||
|
||||
**Guardrail.** Configuration that depends on a delivery mechanism gets a comment
|
||||
naming that mechanism, and a check in the packaging review. Better: remove it
|
||||
until the mechanism exists.
|
||||
|
||||
---
|
||||
|
||||
## Clamps
|
||||
|
||||
### ST-09 - A clamp comparing two different units
|
||||
|
||||
**Mechanism.** A compatibility check obtains a **resolution percentage** and
|
||||
compares it against a **quality level**. The percentage defaults to one hundred
|
||||
and the level ranges to three, so the comparison is always satisfied.
|
||||
|
||||
**Why it is silent.** Both values are integers, both accessors exist, and the
|
||||
comparison is well typed. The constraint simply never binds, which is
|
||||
indistinguishable from a constraint that is never violated.
|
||||
|
||||
**Why the obvious check misses it.** The line reads correctly. The two accessor
|
||||
names differ by one word, the sibling function two hundred lines away uses the
|
||||
correct one, and nothing about either line suggests a unit mismatch. This is a
|
||||
type error that the type system cannot express.
|
||||
|
||||
**Symptom.** A frame-rate and quality compatibility rule that never fires, so a
|
||||
platform can select a combination the rule exists to prevent.
|
||||
|
||||
**Detect.** Find clamps and check the accessor against the value being compared:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A4 "GetApplicable\w*Limit\(" Source/ | rg "Quality|Resolution|Limit"
|
||||
```
|
||||
|
||||
Read each hit for unit agreement. Then unit-test every clamp at both ends of its
|
||||
real range — the test that would have caught this takes three lines.
|
||||
|
||||
**Guardrail.** Give the two ranges distinct types, or name the accessors so a
|
||||
mismatch is visible in one line. Test each clamp with a value that must be
|
||||
rejected; a clamp with no failing test case has never been shown to work.
|
||||
|
||||
---
|
||||
|
||||
### ST-10 - A remap with a reachable zero denominator
|
||||
|
||||
**Mechanism.** A normalisation maps a value between two ranges by dividing by the
|
||||
difference between a maximum and a minimum. Nothing guarantees they differ.
|
||||
|
||||
**Why it is silent.** For every configuration anyone has tried, they differ. The
|
||||
division is correct and the result is correct.
|
||||
|
||||
**Why the obvious check misses it.** The formula is standard and reads as
|
||||
obviously right. Reachability of the degenerate case depends on data — a device
|
||||
profile setting a limit equal to the scale minimum — and data is not reviewed with
|
||||
the arithmetic.
|
||||
|
||||
**Symptom.** An infinity or a not-a-number propagating into a quality value on
|
||||
one device configuration, producing behaviour that is hard to attribute and
|
||||
impossible to reproduce elsewhere.
|
||||
|
||||
**Detect.** Find range normalisations and check their guards:
|
||||
|
||||
```bash
|
||||
rg -n -B3 -A1 "\) / \(\w+ - \w+\)|/ \(Max\w+ - Min\w+\)" Source/
|
||||
```
|
||||
|
||||
Any division by a difference with no preceding equality check is the finding.
|
||||
|
||||
**Guardrail.** Guard the denominator and decide the degenerate result
|
||||
deliberately. Then check whether data can actually produce it — and if it can,
|
||||
validate the data too.
|
||||
+202
@@ -0,0 +1,202 @@
|
||||
# Patterns: platform budgets in a measured reference product
|
||||
|
||||
A worked reading of one shipped platform-budget configuration: scalability
|
||||
buckets, device profile tiers, streaming pools and the code that applies them.
|
||||
|
||||
This file is almost entirely about **configuration**, which makes it unlike the
|
||||
others in this bundle in one important way: none of the findings below are bugs
|
||||
in the usual sense. Every value is valid, every file parses, and the game runs.
|
||||
What the audit produced is a list of places where the configuration says
|
||||
something the runtime does not do — and the only reason any of it is findable is
|
||||
that someone read the engine's own configuration alongside the project's.
|
||||
|
||||
Markers: **[measured]** — read in project or engine configuration and source;
|
||||
**[derived]** — conclusion from measured facts; **[open]** — requires hardware.
|
||||
|
||||
**No number here is a memory measurement.** Nothing was profiled. Everything is
|
||||
a declared budget, which is exactly the distinction the skill is about.
|
||||
|
||||
---
|
||||
|
||||
## 1. The single most useful table in the audit
|
||||
|
||||
What the project declares about the texture streaming pool, per mobile tier
|
||||
**[measured]**:
|
||||
|
||||
| Tier | Pool declared by the project | Effective source |
|
||||
|---|---:|---|
|
||||
| Low | — | engine scalability bucket |
|
||||
| **Mid** | **85 MB** | **device profile** |
|
||||
| High | — | engine scalability bucket |
|
||||
| Epic | — | engine scalability bucket |
|
||||
|
||||
The project sets the pool in exactly two lines — the mid tier of each mobile
|
||||
platform — and declares **no** per-tier texture-quality sections of its own
|
||||
**[measured]**. Every other tier inherits engine values, which on this engine
|
||||
version run from several hundred megabytes to a gigabyte, with the
|
||||
video-memory clamp disabled at the top **[measured]**.
|
||||
|
||||
**[derived]** Two consequences, and the second is the one that matters:
|
||||
|
||||
1. The tier with the smallest declared budget is the only tier the project
|
||||
actually configured. The rest are engine defaults that nobody chose.
|
||||
2. Because a device profile applies at a higher priority than scalability, the
|
||||
mid tier's pool **cannot be changed by the in-game texture setting**, while
|
||||
on every other tier the same setting works. One control, opposite behaviour,
|
||||
no explanation anywhere in the UI.
|
||||
|
||||
That asymmetry is the reason ST-01 leads the failure modes. It is not a mistake
|
||||
in either file; it is a mistake that only exists between them.
|
||||
|
||||
---
|
||||
|
||||
## 2. Priority, in one paragraph you can act on
|
||||
|
||||
Device profiles apply variables whose names carry the scalability prefix at
|
||||
**scalability** priority, and every other variable at **device-profile**
|
||||
priority — which outranks everything the settings system does **[measured, from
|
||||
engine source]**.
|
||||
|
||||
**[derived]** So the same file produces two completely different mutability
|
||||
outcomes depending on a naming convention:
|
||||
|
||||
- a prefixed quality bucket set in a profile is a **default** the player can
|
||||
override;
|
||||
- an unprefixed variable set in a profile is a **lock**.
|
||||
|
||||
Nothing in the ini syntax distinguishes them. The audited project uses both
|
||||
forms, mostly correctly and once consequentially (§1).
|
||||
|
||||
**The practical rule:** you cannot determine effective configuration by reading
|
||||
files. Dump the variables with their source on the target device. That command is
|
||||
the only authoritative answer, and it takes ten seconds.
|
||||
|
||||
---
|
||||
|
||||
## 3. Automatic detection promotes on the wrong axis
|
||||
|
||||
Promotion thresholds are expressed against a GPU performance index, for every
|
||||
quality group, with **no memory term anywhere** **[measured]**. The texture group
|
||||
reaches its top level at a low index value, and that top level is the one where
|
||||
the engine disables the video-memory clamp on the pool **[measured]**.
|
||||
|
||||
**[derived]** A fast chip with modest memory therefore lands in the bucket that
|
||||
grants the largest pool and removes the guard against it. The two decisions were
|
||||
made in different files by different reasoning, and each is defensible alone.
|
||||
|
||||
The project's own comment on the thresholds is honest about their intent — tuning
|
||||
so a specific generation of hardware lands in a specific mix **[measured]** — and
|
||||
that intent is about throughput, which is exactly the axis that misses this.
|
||||
Recipe: ST-02.
|
||||
|
||||
---
|
||||
|
||||
## 4. Dead configuration, and how much of it there is
|
||||
|
||||
Four independent instances, all silent, all reviewed at some point:
|
||||
|
||||
| Finding **[measured]** | What it means |
|
||||
|---|---|
|
||||
| A top-tier device profile section defined in full, absent from the profile registry and from the engine's base profiles | Every value in it is unreachable |
|
||||
| Comment banners grouping devices under tier labels that contradict the base profile named two lines below | The documentation and the behaviour disagree; several devices sit under the wrong banner |
|
||||
| A user-facing profile variant list read in three places and set in no configuration file | The profile override never applies; the console preset UI never appears |
|
||||
| Optional mip levels configured in every mobile texture group, with chunk acquisition disabled | A high-resolution optional tier that cannot be delivered |
|
||||
|
||||
**[derived]** The common structure is worth naming, because it generalises far
|
||||
beyond streaming: **configuration that names a mechanism the project does not
|
||||
have.** It reads as a decision, survives review indefinitely, and produces
|
||||
nothing. It is the configuration-file form of the stub described in
|
||||
`ue-reference-project-adoption`.
|
||||
|
||||
Recipes: ST-05, ST-06, ST-07, ST-08.
|
||||
|
||||
---
|
||||
|
||||
## 5. Two arithmetic defects in the clamp layer
|
||||
|
||||
**A comparison between a percentage and a level.** A frame-rate compatibility
|
||||
check obtains a resolution-quality limit — a percentage defaulting to one hundred
|
||||
— and compares it against an overall quality level in the range zero to three
|
||||
**[measured]**. The comparison is always satisfied, so the constraint never
|
||||
binds.
|
||||
|
||||
**[derived]** What makes this findable rather than invisible is that the sibling
|
||||
function two hundred lines away uses the correct accessor for the same purpose
|
||||
**[measured]**. One of the two is wrong, and the codebase contains its own
|
||||
counterexample.
|
||||
|
||||
**A remap with no guard on its denominator.** A normalisation between two
|
||||
resolution ranges divides by the difference between a maximum and a minimum, with
|
||||
no check that they differ, and the degenerate case is reachable from device
|
||||
profile data **[measured]**.
|
||||
|
||||
Recipes: ST-09, ST-10. Both are three-line fixes and both would be caught by one
|
||||
unit test each — which is the argument for testing clamps at their range
|
||||
boundaries rather than at their typical values.
|
||||
|
||||
---
|
||||
|
||||
## 6. What the project got right
|
||||
|
||||
Worth stating, because the list above is long and the configuration is not bad:
|
||||
|
||||
1. **Cook-time budget levers are used deliberately.** Mobile texture groups clamp
|
||||
the maximum size sixteen-fold on a side, animation key stripping is enabled,
|
||||
and unused detail-mode components and emitters are pruned at cook
|
||||
**[measured]**. These are budget-A decisions, correctly separated from
|
||||
runtime budgets.
|
||||
2. **Foliage density was moved between scalability groups deliberately**, with the
|
||||
old group's keys removed and a different curve applied in the new one
|
||||
**[measured]**. That is a real understanding of how the groups compose.
|
||||
3. **Per-device resolution overrides on specific hardware** rather than a blanket
|
||||
platform value **[measured]**.
|
||||
4. **Dynamic resolution configured for one mobile platform** with an explicit
|
||||
comment stating the target frame rate and output resolution **[measured]** —
|
||||
the kind of comment that makes a number reviewable.
|
||||
5. **A rich set of runtime diagnostics** already wired: the profile selection
|
||||
logs its inputs and its chosen result, frame pacing logs its target, quality
|
||||
clamping logs both the before and after values **[measured]**. Someone
|
||||
arriving on a device has the answers available.
|
||||
|
||||
**[derived]** Point 5 deserves emphasis. Most of the findings in this file are
|
||||
detectable in ten seconds on a device by reading log lines the project already
|
||||
prints. The failure is not instrumentation; it is that nobody read the output.
|
||||
|
||||
---
|
||||
|
||||
## 7. What to copy, and what to check first
|
||||
|
||||
Copy: the four-budget separation as a documentation convention; cook-time levers
|
||||
used per platform; the diagnostic logging around profile selection and clamping;
|
||||
per-device overrides for outlier hardware.
|
||||
|
||||
Check first, in this order, on your own project:
|
||||
|
||||
1. **Dump variables with their source** on each target. Everything else is
|
||||
inference. (ST-01)
|
||||
2. **Diff declared profiles against defined sections.** One command. (ST-05)
|
||||
3. **Separate reads from writes** for every data-driven selection list. (ST-07)
|
||||
4. **Unit-test every clamp** at both ends of its range. (ST-09, ST-10)
|
||||
5. **Check both pool values** — texture and mesh — on every platform. (ST-03)
|
||||
|
||||
**[derived]** Steps 1 and 2 take minutes and would have found the two largest
|
||||
findings in this file. That ratio is the argument for doing configuration audits
|
||||
at all.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as project
|
||||
configuration and source, cross-referenced against the engine's own scalability
|
||||
and device-profile configuration and streaming source in the same installation.
|
||||
Source addresses and engine paths stay in the research archive that produced this
|
||||
skill; each `ST-` identifier resolves back to the audited location there.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
**Nothing here was profiled.** No device was measured, no memory figure is an
|
||||
observation, and every quantity is a declared budget or a configured threshold.
|
||||
Engine values describe one engine version and will drift. The recipes exist
|
||||
precisely because these numbers must be re-derived on your project and your
|
||||
hardware before they mean anything.
|
||||
@@ -0,0 +1,380 @@
|
||||
---
|
||||
name: ue-ui-architecture
|
||||
description: >-
|
||||
Design or review scalable Unreal Engine UI: a root layout per local player,
|
||||
tagged layer stacks, extension points for slot-based widget injection from
|
||||
feature plugins, input and focus policy per screen, and centrally arbitrated
|
||||
loading screens. Use when building HUD, menu or modal architecture, adding
|
||||
mode or plugin UI without base-class dependencies, supporting gamepad or local
|
||||
multiplayer, or debugging focus, ordering, duplicate-widget and teardown
|
||||
issues.
|
||||
---
|
||||
|
||||
# UE UI architecture
|
||||
|
||||
The invariant:
|
||||
|
||||
> UI ownership, placement and content are separate. The shell owns slots and
|
||||
> layers; features contribute content through handles and contracts.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Detection recipes for silent UI failures: [failure modes](references/failure-modes.md).
|
||||
|
||||
Related skills: `ue-modular-gameplay`, `ue-input-architecture`,
|
||||
`ue-gameplay-messaging`, `ue-gameplay-tag-governance`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. One root layout per local player
|
||||
|
||||
```text
|
||||
UI manager subsystem
|
||||
└─ UI policy
|
||||
└─ primary layout, one per local player
|
||||
├─ UI.Layer.Game
|
||||
├─ UI.Layer.GameMenu
|
||||
├─ UI.Layer.Menu
|
||||
└─ UI.Layer.Modal
|
||||
```
|
||||
|
||||
Responsibilities: the **subsystem** is the process-level entry point and owns the
|
||||
active policy; the **policy** decides which layout class to create, where to
|
||||
attach it, and how local players behave; the **layout** is a named layer registry
|
||||
with a push/pop API; **widgets** are content only.
|
||||
|
||||
Do not let gameplay systems add widgets to the viewport with a guessed draw
|
||||
order. A widget that bypasses the layout is invisible to every mechanism the
|
||||
layout provides — visibility toggles, input suspension, focus restoration.
|
||||
|
||||
### Let the base subsystem stand aside
|
||||
|
||||
A base UI subsystem should decline to be created when a project subclass exists,
|
||||
so the two never coexist. The idiom — *do not create me if I have derived
|
||||
classes* — is worth copying anywhere a plugin offers a replaceable subsystem.
|
||||
|
||||
### Local players own their UI
|
||||
|
||||
Each local player needs an independent layout, focus state and input routing. One
|
||||
global HUD root breaks split-screen and player-specific modal stacks, and the
|
||||
breakage appears only when a second player joins.
|
||||
|
||||
---
|
||||
|
||||
## 2. Tagged layer stacks
|
||||
|
||||
Layers express **modality**, not individual screens:
|
||||
|
||||
| Layer | Use |
|
||||
|---|---|
|
||||
| Game | HUD and persistent gameplay presentation |
|
||||
| GameMenu | pause and in-game menus |
|
||||
| Menu | front-end navigation |
|
||||
| Modal | confirmations, errors, blockers |
|
||||
|
||||
Each layer is an activatable widget stack: screens are pushed by soft class,
|
||||
activated and deactivated, and popped by lifecycle.
|
||||
|
||||
### Rules
|
||||
|
||||
- register each layer tag exactly once per layout;
|
||||
- no screen hard-codes a sibling widget path;
|
||||
- the modal layer carries the strongest input and focus policy;
|
||||
- popping restores the previous screen and its focus;
|
||||
- async load failure reports a reason and clears any loading state;
|
||||
- **pushing to an unregistered layer tag must not fail silently.** A typo in a
|
||||
layer tag that returns null with no log is a screen that never appears and
|
||||
nothing to search for. Recipe: UI-04.
|
||||
|
||||
### Layers are not draw-order buckets
|
||||
|
||||
If two widgets belong to one HUD composition, use extension slots inside the HUD
|
||||
layout. Inventing `UI.Layer.HUDLeft` and `UI.Layer.HUDRight` turns a modality
|
||||
concept into a layout concept and loses both.
|
||||
|
||||
---
|
||||
|
||||
## 3. Extension points decouple slot owners from contributors
|
||||
|
||||
A subsystem holds two maps keyed by tag:
|
||||
|
||||
```text
|
||||
extension point tag → registered points
|
||||
extension point tag → registered extensions
|
||||
```
|
||||
|
||||
**A point** is owned by a shell widget: its tag, an optional context object, the
|
||||
data or widget classes it accepts, its tag match mode, and a callback.
|
||||
|
||||
**An extension** is owned by a feature: the target tag, an optional context, a
|
||||
data object or widget class, and a registration handle.
|
||||
|
||||
### Registration must be order-independent
|
||||
|
||||
Both sides catch up on registration — a new point receives already-registered
|
||||
matching extensions, and a new extension notifies already-registered matching
|
||||
points. This is what removes initialization-order coupling between the HUD and
|
||||
feature activation, and it is the single most valuable property of the pattern.
|
||||
|
||||
### Both sides copy before iterating
|
||||
|
||||
Callbacks can register or unregister anything, including themselves. Copy the
|
||||
list before notifying, and define what happens to an entry removed mid-notify.
|
||||
|
||||
---
|
||||
|
||||
## 4. Contract matching
|
||||
|
||||
An extension matches a point when all applicable rules pass:
|
||||
|
||||
1. **tag** — exact, or ancestor-ward for partial matching;
|
||||
2. **context** — same object, or both explicitly global;
|
||||
3. **class contract** — the data derives from, or implements, an allowed class;
|
||||
4. **ownership** — local player and world are compatible.
|
||||
|
||||
### Context tiers
|
||||
|
||||
One slot widget can register points for the global context, for its local player,
|
||||
and for that player's player state. The same slot tag then serves global HUD
|
||||
elements and player-specific elements with no cross-player leakage — and
|
||||
split-screen works without a special case.
|
||||
|
||||
### Know which direction partial matching runs
|
||||
|
||||
Ascending from the extension's tag to its ancestors means a point tagged `A.B`
|
||||
with partial matching receives an extension tagged `A.B.C`. **The reverse does
|
||||
not happen.** Prefer exact tags for production HUD placement, and give every
|
||||
parent tag used for partial matching a coherent data contract.
|
||||
|
||||
### A dead context object is not a global context
|
||||
|
||||
If context matching compares pointers and the context is held weakly, an
|
||||
extension whose context died matches nothing at all — it does not fall back to
|
||||
global, and it does not clean itself up. It sits in the map until someone
|
||||
unregisters it. Recipe: UI-06.
|
||||
|
||||
---
|
||||
|
||||
## 5. The slot widget owns rendering
|
||||
|
||||
The slot widget registers and unregisters points, converts extension data into
|
||||
entry widgets, discovers its context, enforces the class whitelist, and defines
|
||||
empty-state behaviour.
|
||||
|
||||
The feature does not find the slot and does not add children. It registers an
|
||||
extension and retains the handle. That is dependency inversion:
|
||||
|
||||
```text
|
||||
feature knows: tag + widget class
|
||||
shell knows: tag + layout container
|
||||
neither knows the other's concrete type
|
||||
```
|
||||
|
||||
Register in the widget's build step rather than its constructor, and unregister
|
||||
when Slate resources are released — otherwise a rebuilt widget leaves a
|
||||
registration behind.
|
||||
|
||||
---
|
||||
|
||||
## 6. Feature-driven UI has two distinct paths
|
||||
|
||||
**Layout requests** push an activatable screen into a layer.
|
||||
**Element extensions** register a widget into a named slot.
|
||||
|
||||
They are different mechanisms with different removal semantics, and putting both
|
||||
in one feature action is convenient and dangerous. Retain separately: handles to
|
||||
pushed layouts, an extension handle per slot registration, actor extension
|
||||
request handles, and per-actor ownership records.
|
||||
|
||||
### The two teardowns are not equivalent
|
||||
|
||||
Unregistering an extension handle removes the entry widget **synchronously**.
|
||||
Deactivating a layout **asks** the stack to release it. Do not assume the second
|
||||
is removal — verify the stack no longer retains or reactivates it. Recipe: UI-05.
|
||||
|
||||
Removal must also be keyed to the specific actor that was registered: more than
|
||||
one HUD actor can be alive on a client at once.
|
||||
|
||||
---
|
||||
|
||||
## 7. Ordering must be a real contract
|
||||
|
||||
If the API exposes a priority, the implementation must sort or insert by it.
|
||||
**Copying a priority into a request does not create ordering.**
|
||||
|
||||
Define: which direction wins; the tie-break for equal values; whether priority is
|
||||
per point, per tag, or per context; what happens when it changes; and who owns
|
||||
the sort — the subsystem or the slot widget.
|
||||
|
||||
Where ordering is unimplemented, the effective order is registration order, which
|
||||
in a plugin-based project is feature activation order — non-deterministic from the
|
||||
designer's point of view, and made worse by any removal that swaps elements.
|
||||
Recipe: UI-01.
|
||||
|
||||
---
|
||||
|
||||
## 8. The loading screen is arbitration, not a layer
|
||||
|
||||
A loading screen sits above the entire UI, including the root layout. A manager
|
||||
aggregates reasons from objects implementing a loading-process interface:
|
||||
|
||||
```text
|
||||
ShouldShowLoadingScreen(out Reason) → bool
|
||||
```
|
||||
|
||||
Blockers typically include: experience or mode still loading, map travel, pending
|
||||
network game, world not begun, seamless travel, missing player controllers, and
|
||||
explicitly registered async tasks.
|
||||
|
||||
Manager behaviour: inspect all providers, show while any blocker exists, retain a
|
||||
human-readable reason, enforce an optional minimum display time, hide when clear,
|
||||
and survive world teardown.
|
||||
|
||||
### Rules
|
||||
|
||||
- providers report need and reason; the manager owns presentation;
|
||||
- no subsystem toggles the loading widget directly;
|
||||
- **require a non-empty reason.** The reason string is the only diagnostic that
|
||||
answers "why is this screen still up";
|
||||
- stale tasks have a timeout, a cancel path and an owner;
|
||||
- editor and packaged behaviour are both tested — an input blocker that is
|
||||
disabled in the editor means the editor is not testing the shipped behaviour.
|
||||
Recipe: UI-08;
|
||||
- **measure before copying performance side effects.** Forcing garbage collection
|
||||
when the screen hides puts a hitch exactly at the moment control returns to the
|
||||
player. Recipe: UI-07.
|
||||
|
||||
---
|
||||
|
||||
## 9. Input and focus policy per screen
|
||||
|
||||
Each activatable screen declares its input configuration: game-only, menu-only or
|
||||
both; mouse capture and cursor behaviour; a desired focus target; back-action
|
||||
handling; whether lower layers receive input.
|
||||
|
||||
The framework owns the activation stack, focus restoration, gamepad navigation,
|
||||
the bound-action bar, and input suspension during transitions.
|
||||
|
||||
### Validate what the compiler cannot
|
||||
|
||||
- **every activatable screen has a desired focus target.** A screen without one
|
||||
opens and leaves the gamepad with nowhere to go. If your framework only warns
|
||||
about this, treat the warning as an error — a screen with no focus is not
|
||||
shippable. Recipe: UI-09;
|
||||
- the input configuration matches the layer's role;
|
||||
- a required back action exists;
|
||||
- no raw key handling bypasses the UI input system;
|
||||
- extension point contracts are non-empty where required.
|
||||
|
||||
### Transitions and focus interact
|
||||
|
||||
A non-zero stack transition duration can break focus hand-off to the incoming
|
||||
screen on a gamepad. If you disable transitions to fix focus, write down that
|
||||
this is why — otherwise someone will re-enable them.
|
||||
|
||||
---
|
||||
|
||||
## 10. The communication boundary
|
||||
|
||||
UI may listen to local messages for **transient facts**: an elimination for the
|
||||
feed, an accolade for a toast, an inventory delta for a notification.
|
||||
|
||||
Persistent display reads **observable state**: health, score, inventory,
|
||||
settings. Do not rebuild state from a stream of messages — a widget created
|
||||
mid-match will have missed them. See `ue-gameplay-messaging`.
|
||||
|
||||
UI sends commands through accountable controllers and services, never as a
|
||||
broadcast with no owner.
|
||||
|
||||
---
|
||||
|
||||
## 11. Test matrix
|
||||
|
||||
### Extension ordering
|
||||
|
||||
Point first then extension; extension first then point; several points with
|
||||
distinct contexts; global plus local-player plus player-state points;
|
||||
unregistering an extension while its point lives; destroying a point before the
|
||||
extension's owner.
|
||||
|
||||
### Modular lifecycle
|
||||
|
||||
```text
|
||||
baseline HUD
|
||||
→ activate feature → exactly one layout, one entry per extension
|
||||
→ deactivate → back to baseline, verified not assumed
|
||||
→ reactivate → no duplicates
|
||||
```
|
||||
|
||||
### Local players
|
||||
|
||||
One player; split-screen; adding and removing a player at runtime; confirming
|
||||
player-specific context does not leak; independent modal and focus stacks.
|
||||
|
||||
### Input and focus
|
||||
|
||||
Keyboard and mouse; gamepad navigation; device hot-swap; back through nested
|
||||
stacks; modal blocking lower layers; the loading screen capturing input **in a
|
||||
packaged build**; focus restored after a pop.
|
||||
|
||||
---
|
||||
|
||||
## 12. Review checklist
|
||||
|
||||
- [ ] One root layout per local player.
|
||||
- [ ] Every screen enters through a named layer.
|
||||
- [ ] HUD contributions use extension slots, not direct widget references.
|
||||
- [ ] Registration order is irrelevant in both directions.
|
||||
- [ ] Context and class contract are explicit at every point.
|
||||
- [ ] Features retain extension, layout and request handles.
|
||||
- [ ] Deactivation removes layouts and entries completely, verified.
|
||||
- [ ] Ordering is deterministic, or documented as unspecified.
|
||||
- [ ] Pushing to an unknown layer tag is loud.
|
||||
- [ ] The loading screen is centrally arbitrated and every reason is non-empty.
|
||||
- [ ] Every activatable widget has a desired focus target.
|
||||
- [ ] Device presentation stays in the input/UI layer.
|
||||
- [ ] Persistent display reads state; only transient facts come from messages.
|
||||
- [ ] Split-screen behaviour is tested, not assumed.
|
||||
|
||||
---
|
||||
|
||||
## 13. When to simplify
|
||||
|
||||
For a small single-player game with one fixed HUD and no modular modes, a root
|
||||
widget with a few direct child references is enough and this indirection is
|
||||
overhead.
|
||||
|
||||
Introduce layers and extension points when at least one applies: plugins or modes
|
||||
add UI independently; local multiplayer is required; modal, menu and game input
|
||||
policies conflict; screens need async loading and a stack; contributors must not
|
||||
depend on the shell widget class; gamepad focus restoration must be systematic.
|
||||
|
||||
The goal is not maximum indirection. It is explicit ownership and replaceable
|
||||
placement contracts — and the cost of that is a debugging path that crosses
|
||||
several assets, which you pay every time something does not appear.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
the UI layer of Epic's Lyra Starter Game on Unreal Engine 5.6: the extension
|
||||
subsystem, the layer and policy plugins, the loading-screen plugin, and the
|
||||
project's own widgets and feature action.
|
||||
|
||||
The framework the project sits on — the activatable widget stack, focus
|
||||
management, the input subsystem — was **not** in the audited tree, so every
|
||||
statement about it is inferred from call sites and is marked as such rather than
|
||||
presented as measured. Widget blueprints are binary and were not read.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill; what
|
||||
ships is the detection recipe. Each entry carries a stable identifier (`UI-01`
|
||||
and up) that resolves back to the audited location.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. Layer-to-container binding lives
|
||||
in binary assets, so which layers actually exist in that project is out of scope
|
||||
rather than concluded. Re-run the recipes against your own tree before acting on
|
||||
any specific claim.
|
||||
@@ -0,0 +1,518 @@
|
||||
# Failure modes: UI architecture
|
||||
|
||||
Fourteen ways a layered, slot-based UI puts the wrong thing in the wrong place,
|
||||
in the wrong order, or nowhere at all — without reporting anything.
|
||||
|
||||
The shared property: **this architecture is built out of joins that nothing
|
||||
checks.** A layer is a tag looked up in a map. A slot is a tag matched against
|
||||
another tag. A contribution is a class pointer that may or may not be loaded.
|
||||
Every one of those joins fails by producing nothing, and producing nothing is
|
||||
also what a correctly configured, currently-empty UI looks like.
|
||||
|
||||
A second property worth stating: **UI failures are reported by people, not by
|
||||
tools.** Nobody writes a test that asserts a widget appeared in a slot. So the
|
||||
detection recipes below matter more here than in most areas, because they are
|
||||
frequently the only instrument that exists.
|
||||
|
||||
Recipes use `rg` from a project root and were executed against the audited
|
||||
project while this file was written.
|
||||
|
||||
---
|
||||
|
||||
## Ordering and placement
|
||||
|
||||
### UI-01 - A priority field that never sorts
|
||||
|
||||
**Mechanism.** The extension API accepts a priority on every registration
|
||||
overload, stores it on the entry, and copies it into the request struct handed to
|
||||
the slot widget. Nothing ever compares two priorities.
|
||||
|
||||
**Why it is silent.** An order is produced on every run — registration order —
|
||||
and it is stable for as long as the set of contributors is stable. The UI looks
|
||||
deliberate. Nothing is missing; things are merely arranged by accident.
|
||||
|
||||
**Why the obvious check misses it.** The value travels through the entire public
|
||||
surface: eight registration entry points, the entry struct, the request struct.
|
||||
Any review asking "is priority used?" finds a dozen sites and stops. The right
|
||||
question is much narrower — *does it appear inside a sort or a comparison?* — and
|
||||
that is not a question anyone thinks to ask about a field that is so obviously
|
||||
plumbed.
|
||||
|
||||
**Symptom.** Widget order inside a slot changes between runs and between
|
||||
machines. In a plugin-based project registration order is feature activation
|
||||
order, so the symptom reads as a race condition and gets investigated in the
|
||||
loading system.
|
||||
|
||||
**Detect.** Count the plumbing, then look for the comparison:
|
||||
|
||||
```bash
|
||||
rg -n "\bPriority\b" --glob "*.cpp" --glob "*.h" <ui-extension-module>/ | wc -l
|
||||
rg -n "Sort|StableSort|Algo::" <ui-extension-module>/ | wc -l
|
||||
```
|
||||
|
||||
In the audited project the first number was substantial and **the second was
|
||||
exactly zero** — no sort of any kind exists in the module. A removal that uses
|
||||
swap-with-last compounds it by reordering the array on every unregistration.
|
||||
|
||||
**Guardrail.** An ordering field ships with its comparator in the same change, or
|
||||
it does not ship. If order genuinely does not matter, do not expose a field that
|
||||
promises it — the field is a more expensive lie than the missing feature.
|
||||
|
||||
---
|
||||
|
||||
### UI-02 - A duplicated key in the plugin descriptor silently drops a dependency
|
||||
|
||||
**Mechanism.** The descriptor is JSON. It contains the dependency array key twice.
|
||||
By JSON semantics the second occurrence replaces the first, so half the declared
|
||||
dependencies do not exist as far as any parser is concerned.
|
||||
|
||||
**Why it is silent.** The project builds and runs, because the lost dependency
|
||||
arrives transitively through the surviving one and is also enabled at the project
|
||||
level. The declaration was redundant in this configuration — which is exactly why
|
||||
nobody notices it is gone.
|
||||
|
||||
**Why the obvious check misses it.** Reading the file shows both blocks. A human
|
||||
reader sees two lists and mentally unions them; a parser sees one key assigned
|
||||
twice and keeps the last. The file *looks* more correct than it is, and no editor
|
||||
or build step warns about duplicate keys.
|
||||
|
||||
**Symptom.** Nothing, until the transitive path disappears — someone disables the
|
||||
intermediate plugin, or extracts this plugin into another project. Then a
|
||||
dependency that the descriptor appears to declare is simply not there.
|
||||
|
||||
**Detect.** Do not read it. Parse it, and compare against the raw text:
|
||||
|
||||
```bash
|
||||
rg -c '"Plugins"' <plugin>/<Plugin>.uplugin # occurrences in the text
|
||||
python -c "import json,sys; d=json.load(open(sys.argv[1],encoding='utf-8-sig')); \
|
||||
print([p.get('Name') for p in d.get('Plugins',[])])" <plugin>/<Plugin>.uplugin
|
||||
```
|
||||
|
||||
A count above one from the first command means the second command is
|
||||
authoritative. In the audited project the text contained two dependency blocks
|
||||
and the parser reported exactly one surviving dependency — the other was
|
||||
discarded. This is the single cheapest recipe in this file and it generalises to
|
||||
every hand-edited descriptor in a project.
|
||||
|
||||
**Guardrail.** Validate every descriptor by parsing it in CI, not by reading it.
|
||||
Assert that the parsed dependency set equals the intended set.
|
||||
|
||||
---
|
||||
|
||||
### UI-03 - Unknown layer tag resolves to a silent null
|
||||
|
||||
**Mechanism.** Pushing a screen looks the layer tag up in a map. A miss returns
|
||||
null and the push does not happen.
|
||||
|
||||
**Why it is silent.** Returning null is the correct behaviour for a container
|
||||
that does not exist. The caller usually ignores the return value, because in the
|
||||
working case it is a widget nobody needs a reference to.
|
||||
|
||||
**Why the obvious check misses it.** The call site reads perfectly: a tag, a
|
||||
class, a push. The tag is a constant that exists in the vocabulary, so it is
|
||||
spelled correctly — it simply names a layer that this particular root layout did
|
||||
not register. Layers are registered from a Blueprint layout asset, which no code
|
||||
search can inspect, so the set of valid tags is not knowable from source.
|
||||
|
||||
**Symptom.** A screen never appears. No log, no assertion. Investigation starts
|
||||
from the screen's own class, which is fine, and from the push call, which is also
|
||||
fine.
|
||||
|
||||
**Detect.** Compare the tags that are pushed against the tags that are
|
||||
registered, accepting that the second half may require the editor:
|
||||
|
||||
```bash
|
||||
rg -o -N 'PushWidgetToLayerStack\w*\(\s*([A-Za-z_:]+)' -r '$1' --glob "*.cpp" .
|
||||
rg -n "RegisterLayer" --glob "*.cpp" --glob "*.h" .
|
||||
rg -o -N 'Tag="(UI\.Layer\.[^"]+)"' -r '$1' Config/ | sort -u
|
||||
```
|
||||
|
||||
The third command gives the declared vocabulary; if a pushed tag is not in it,
|
||||
the defect is certain. If it is in the vocabulary, the layout asset still has to
|
||||
be checked in the editor — which is the finding in itself.
|
||||
|
||||
**Guardrail.** Make the push assert or log an error on an unknown layer, naming
|
||||
the tag and listing the registered ones. A lookup miss on a hardcoded key is
|
||||
never a legitimate runtime state.
|
||||
|
||||
---
|
||||
|
||||
## Contribution lifetime
|
||||
|
||||
### UI-04 - Contribution class resolved with a bare accessor and no fallback
|
||||
|
||||
**Mechanism.** The feature action resolves its widget classes with an accessor
|
||||
that returns whatever is already loaded, relying on asset bundles to have loaded
|
||||
them, and passes the result on without checking.
|
||||
|
||||
**Why it is silent.** When bundles work — which is always, inside the pipeline
|
||||
that was designed for them — the classes are loaded and everything is correct. A
|
||||
null is rejected downstream with at most a verbose log.
|
||||
|
||||
**Why the obvious check misses it.** The bundle annotation is right there on the
|
||||
property, so the loading question looks answered. The failure requires the action
|
||||
to run outside its intended pipeline, which is not a state the code was reviewed
|
||||
against.
|
||||
|
||||
**Symptom.** Widgets silently missing for one feature, in one configuration
|
||||
— typically a packaged build, or a manual invocation of the action — while
|
||||
everything works in the editor.
|
||||
|
||||
**Detect.** Find bare accessors on soft references in contribution paths:
|
||||
|
||||
```bash
|
||||
rg -n "\.Get\(\)" --glob "*.cpp" . | rg -i "widget|layout|class"
|
||||
rg -n -B4 "\.Get\(\)" --glob "*.cpp" . | rg "AssetBundles"
|
||||
```
|
||||
|
||||
Then check each hit for a null branch. In the audited project both contribution
|
||||
loops used the bare accessor; the layout loop tested for null and the slot loop
|
||||
did not — the null travelled one function further and was rejected there with a
|
||||
verbose-level log.
|
||||
|
||||
**Guardrail.** Either load explicitly with a failure path, or check the result
|
||||
and log at warning level with the asset name. "The bundle guarantees it" is a
|
||||
statement about one pipeline, not about the code.
|
||||
|
||||
---
|
||||
|
||||
### UI-05 - Two teardown paths with different strength
|
||||
|
||||
**Mechanism.** The feature action removes its two kinds of contribution
|
||||
differently: slot extensions are unregistered, which synchronously removes the
|
||||
entry widget; layouts are asked to deactivate, which is a request the stack may
|
||||
honour later or differently.
|
||||
|
||||
**Why it is silent.** Both calls succeed. Deactivation is a real operation with a
|
||||
visible effect, so the layout does disappear from view in the common case.
|
||||
Whether the stack still retains it, and whether it will reactivate when the layer
|
||||
is next popped to, is not observable from the action.
|
||||
|
||||
**Why the obvious check misses it.** Both halves of teardown exist and are
|
||||
symmetric in shape — a loop over layouts, a loop over handles. Reading the
|
||||
function gives no reason to suspect that one loop is weaker than the other. The
|
||||
difference is in the semantics of two APIs in a different plugin.
|
||||
|
||||
**Symptom.** A layout that reappears, or that is still in the stack's history
|
||||
after its owning feature is gone. On reactivation the layer may hold two.
|
||||
|
||||
**Detect.** Compare the verbs used in the two removal loops:
|
||||
|
||||
```bash
|
||||
rg -n -A20 "::Remove\w*\(" --glob "*.cpp" . | rg "Deactivate|Unregister|RemoveWidget"
|
||||
```
|
||||
|
||||
A `Deactivate` in one loop and an `Unregister` in the other is the finding.
|
||||
Confirm at runtime: activate, deactivate, reactivate, and count entries in both
|
||||
the slot and the layer.
|
||||
|
||||
**Guardrail.** Removal means removal. If the stack API distinguishes deactivation
|
||||
from removal, call the removing one and verify the stack no longer retains the
|
||||
widget. Assert the count returns to baseline.
|
||||
|
||||
---
|
||||
|
||||
### UI-06 - Manual garbage-collection reference tracking
|
||||
|
||||
**Mechanism.** The object held by a registration is not a reflected property, so
|
||||
the subsystem walks its own maps in the reference-collection hook to keep those
|
||||
objects alive.
|
||||
|
||||
**Why it is silent.** It works, exactly and completely, for as long as the hook
|
||||
matches the data structure. There is no degraded mode — either the object is
|
||||
collected or it is not.
|
||||
|
||||
**Why the obvious check misses it.** The hook is correct when written and is far
|
||||
from the data structure it mirrors. Adding a third map, or a new object-holding
|
||||
field, is a local change that does not visibly relate to a function elsewhere in
|
||||
the file. Nothing links them.
|
||||
|
||||
**Symptom.** After a refactor, an object is collected while still registered.
|
||||
Symptoms are arbitrary: a null contribution, a crash on a stale pointer, a
|
||||
contribution that silently stops matching.
|
||||
|
||||
**Detect.** Find manual reference collection and check it covers every container:
|
||||
|
||||
```bash
|
||||
rg -n -A15 "AddReferencedObjects" --glob "*.cpp" .
|
||||
rg -n "TMap<.*TSharedPtr<|TObjectPtr<\w+> \w+;" --glob "*.h" <ui-extension-module>/
|
||||
```
|
||||
|
||||
Compare the containers walked by the first against the containers declared by the
|
||||
second. Any container holding objects and not walked is a leak of correctness.
|
||||
|
||||
**Guardrail.** Prefer reflected properties. Where manual collection is
|
||||
unavoidable, put the hook immediately adjacent to the declarations it mirrors and
|
||||
add a comment at each declaration pointing at it.
|
||||
|
||||
---
|
||||
|
||||
### UI-07 - Context matched by pointer, held weakly
|
||||
|
||||
**Mechanism.** A contribution is scoped to a context object — usually a local
|
||||
player — and matching compares pointers. The stored reference is weak.
|
||||
|
||||
**Why it is silent.** When the context dies, the contribution stops matching
|
||||
anything, which looks exactly like a contribution that was correctly removed. No
|
||||
stale widget appears, so nothing draws attention.
|
||||
|
||||
**Why the obvious check misses it.** The weak pointer is the *right* choice and
|
||||
prevents the dangerous failure. What it does not do is remove the record, and
|
||||
"does this leak memory?" is answered by the weak pointer while "does this leak
|
||||
records?" is not asked at all.
|
||||
|
||||
**Symptom.** Registration maps that grow across sessions with entries that can
|
||||
never match again. Only visible if something counts them.
|
||||
|
||||
**Detect.** Find weak context fields and check for a cleanup path keyed on their
|
||||
expiry:
|
||||
|
||||
```bash
|
||||
rg -n "TWeakObjectPtr<UObject>\s+ContextObject|ContextObject ==" --glob "*.h" --glob "*.cpp" .
|
||||
rg -n "IsExplicitlyNull|IsValid\(\)" --glob "*.cpp" <ui-extension-module>/
|
||||
```
|
||||
|
||||
The finding is a comparison path with no eviction path. Confirm by logging map
|
||||
sizes across a local-player add/remove cycle.
|
||||
|
||||
**Guardrail.** Evict on context expiry, or key the map by context so an entire
|
||||
context can be dropped in one operation.
|
||||
|
||||
---
|
||||
|
||||
## Diagnosis surface
|
||||
|
||||
### UI-08 - Swapped logging branches
|
||||
|
||||
**Mechanism.** Two log statements in the two arms of a conditional have their
|
||||
format strings exchanged: the branch that has a context logs the format without
|
||||
it, and the branch that has none logs the format with it — printing a safe-name
|
||||
of null.
|
||||
|
||||
**Why it is silent.** It is a log line. It costs nothing at runtime, breaks no
|
||||
behaviour, and is emitted at verbose level, so it is invisible unless someone
|
||||
turns the category up.
|
||||
|
||||
**Why the obvious check misses it.** Reviewers read log statements for the
|
||||
message, not for which branch they are in. Both lines are individually plausible.
|
||||
The defect only exists in the pairing, and the pairing is exactly what the eye
|
||||
skips.
|
||||
|
||||
**Symptom.** When someone finally enables verbose logging — necessarily during a
|
||||
difficult investigation — the log lies about which registrations have a context.
|
||||
The cost is paid at the worst possible moment.
|
||||
|
||||
**Detect.** Read both arms of every logging conditional together:
|
||||
|
||||
```bash
|
||||
rg -n -B3 -A8 "if\s*\(\s*ContextObject\s*\)" --glob "*.cpp" . | rg "UE_LOG|GetNameSafe"
|
||||
```
|
||||
|
||||
Look for a format string mentioning the variable inside the branch where it is
|
||||
null. In the audited project this is present and unmistakable once the two lines
|
||||
are read side by side.
|
||||
|
||||
**Guardrail.** Log through one statement with a conditional argument rather than
|
||||
two statements with divergent formats. One format string cannot disagree with
|
||||
itself.
|
||||
|
||||
---
|
||||
|
||||
### UI-09 - Missing focus target caught by a warning, not an error
|
||||
|
||||
**Mechanism.** A screen that does not declare a focus target compiles with a
|
||||
Blueprint warning.
|
||||
|
||||
**Why it is silent.** Warnings scroll. A project with any warning debt has
|
||||
normalised them, and this one appears only when the specific widget is compiled.
|
||||
|
||||
**Why the obvious check misses it.** The check exists — someone wrote it
|
||||
deliberately and worded it well. It is diligence that has been graded as
|
||||
optional. The failure is not the absence of a check but its severity, which no
|
||||
review of the checking code would flag.
|
||||
|
||||
**Symptom.** On a gamepad, opening the screen leaves focus nowhere: no
|
||||
navigation, no visible selection, and often no way back. It is one of the most
|
||||
expensive UI bugs to receive from a player and one of the cheapest to prevent.
|
||||
|
||||
**Detect.** Find the validation and check its severity, then find the screens
|
||||
that would trip it:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A8 "ValidateCompiledWidgetTree|ValidateCompiledDefaults" --glob "*.cpp" . \
|
||||
| rg "Warning|Error|Note"
|
||||
rg -n "GetDesiredFocusTarget|BP_GetDesiredFocusTarget" --glob "*.cpp" --glob "*.h" .
|
||||
```
|
||||
|
||||
A validation that emits a warning for a gamepad-blocking condition is the
|
||||
finding.
|
||||
|
||||
**Guardrail.** Promote it to an error, or gate packaging on it. A check that
|
||||
cannot fail a build is documentation with extra steps.
|
||||
|
||||
---
|
||||
|
||||
### UI-10 - The loading screen's only diagnostic is a string nobody must break
|
||||
|
||||
**Mechanism.** Arbitration polls a set of providers, each of which returns
|
||||
whether it needs the screen and why. The accumulated reason is the sole
|
||||
explanation for a stuck loading screen.
|
||||
|
||||
**Why it is silent.** A permanent loading screen *is* the failure, and it is
|
||||
loud. The silence is in the cause: the screen is a correct response to a provider
|
||||
that never releases.
|
||||
|
||||
**Why the obvious check misses it.** There is no artefact to inspect. The screen
|
||||
is showing because at least one of a dozen conditions is true, and only one of
|
||||
them names itself. A team that bypasses the interface with its own flag removes
|
||||
even that.
|
||||
|
||||
**Symptom.** A build that hangs on the loading screen with no information, in a
|
||||
configuration that cannot be attached to a debugger.
|
||||
|
||||
**Detect.** Enumerate the providers and check that nothing shows the screen
|
||||
outside them:
|
||||
|
||||
```bash
|
||||
rg -n "ShouldShowLoadingScreen" --glob "*.cpp" --glob "*.h" .
|
||||
rg -n "AddViewportWidgetContent|bShowLoadingScreen|LoadingWidget" --glob "*.cpp" . \
|
||||
| rg -v "LoadingScreenManager"
|
||||
```
|
||||
|
||||
The second command should be empty. Any hit is a path that shows or hides the
|
||||
screen without contributing a reason, and it is the one you will be unable to
|
||||
diagnose. In the audited project the arbitration enforced a non-empty reason with
|
||||
an assertion — a design worth copying exactly.
|
||||
|
||||
**Guardrail.** One owner presents; everyone else reports need plus reason.
|
||||
Require the reason with an assertion, and expose the current reason to a console
|
||||
command.
|
||||
|
||||
---
|
||||
|
||||
### UI-11 - Forced collection at the moment control returns to the player
|
||||
|
||||
**Mechanism.** Hiding the loading screen forces a full garbage collection.
|
||||
|
||||
**Why it is silent.** It is deliberate, it is defensible, and it is not a bug. The
|
||||
cost lands as a hitch precisely when the screen disappears, which is a moment the
|
||||
player already expects to be uneven.
|
||||
|
||||
**Why the obvious check misses it.** Nobody profiles the loading screen. It is
|
||||
infrastructure that ran before the thing you are measuring. The hitch is
|
||||
attributed to the level, to shader compilation, or to streaming.
|
||||
|
||||
**Symptom.** A consistent stall at the first frame of gameplay that resists
|
||||
explanation from gameplay profiling.
|
||||
|
||||
**Detect.** Find global side effects in the hide path:
|
||||
|
||||
```bash
|
||||
rg -n -B6 -A6 "ForceGarbageCollection|SetBatchMode|SuspendHeartBeat|bDisableWorldRendering" \
|
||||
--glob "*.cpp" .
|
||||
```
|
||||
|
||||
Every hit is a process-wide effect owned by the loading screen. Each may be
|
||||
correct; all of them are worth knowing about before copying the subsystem.
|
||||
|
||||
**Guardrail.** Measure before adopting. If a forced collection is wanted, do it
|
||||
while the screen is still up, not as it comes down.
|
||||
|
||||
---
|
||||
|
||||
### UI-12 - Behaviour that differs in the editor by design
|
||||
|
||||
**Mechanism.** The loading screen's input blocker declines to consume input when
|
||||
running under the editor.
|
||||
|
||||
**Why it is silent.** It makes editor iteration bearable, which is why it was
|
||||
written. The divergence is invisible because the editor is where everyone looks.
|
||||
|
||||
**Why the obvious check misses it.** Testing happens in the configuration that
|
||||
has the exception. Confirming the difference requires running a packaged build
|
||||
specifically to compare input behaviour during loading, which is not on anyone's
|
||||
list.
|
||||
|
||||
**Symptom.** Input handled during loading behaves differently in a packaged
|
||||
build. Bugs that only reproduce for QA, in a build the developer cannot iterate
|
||||
on.
|
||||
|
||||
**Detect.** Find editor-conditional behaviour in shared runtime paths:
|
||||
|
||||
```bash
|
||||
rg -n "GIsEditor|WITH_EDITOR" --glob "*.cpp" <loading-and-ui-modules>/ \
|
||||
| rg -v "^.*Editor.*\.cpp"
|
||||
```
|
||||
|
||||
Each hit is a place where the editor is not representative. List them and decide
|
||||
which matter; the point is to hold the list, not to remove it.
|
||||
|
||||
**Guardrail.** Keep a written list of deliberate editor divergences in the UI and
|
||||
loading paths, and test each one in a packaged build at least once per milestone.
|
||||
|
||||
---
|
||||
|
||||
## Structure
|
||||
|
||||
### UI-13 - Layers used as z-order buckets
|
||||
|
||||
**Mechanism.** New layers are added to place widgets relative to each other,
|
||||
rather than to express modality.
|
||||
|
||||
**Why it is silent.** It works. A new layer is one registration and one tag, and
|
||||
the widget appears where intended.
|
||||
|
||||
**Why the obvious check misses it.** Each addition is individually reasonable and
|
||||
locally minimal. The cost is structural and accrues over releases: input and
|
||||
focus policy is per layer, so every added layer multiplies the modality matrix
|
||||
that must be reasoned about and tested.
|
||||
|
||||
**Symptom.** Focus and input behaviour that nobody can predict, because there are
|
||||
nine layers and the policy interaction between them was never designed.
|
||||
|
||||
**Detect.** Count layers against modalities:
|
||||
|
||||
```bash
|
||||
rg -o -N 'Tag="(UI\.Layer\.[^"]+)"' -r '$1' Config/ Plugins/*/Config/ | sort -u
|
||||
```
|
||||
|
||||
More than a handful, or names that describe position rather than modality, is the
|
||||
finding. In the audited project there were exactly four, named for modality.
|
||||
|
||||
**Guardrail.** Layers express modality: gameplay, game menu, menu, modal. Relative
|
||||
placement within one modality belongs to slots inside a layout, not to new layers.
|
||||
|
||||
---
|
||||
|
||||
### UI-14 - The composition is knowable only from binary assets
|
||||
|
||||
**Mechanism.** Layers are registered, and slots are placed, from widget assets.
|
||||
Source contains the vocabulary and the API but not the arrangement.
|
||||
|
||||
**Why it is silent.** It is the intended design and its benefit is real: the
|
||||
shell's layout is authored by designers without engineering involvement.
|
||||
|
||||
**Why the obvious check misses it.** Reading all the source and finding no
|
||||
arrangement feels like an incomplete search rather than a property of the system.
|
||||
Time is lost widening greps that cannot succeed.
|
||||
|
||||
**Symptom.** "Which layers exist?" and "which slots does the HUD have?" cannot be
|
||||
answered from a repository checkout. Onboarding, audits and automated validation
|
||||
all stop at this boundary.
|
||||
|
||||
**Detect.** Get what source can give, and mark the rest as open:
|
||||
|
||||
```bash
|
||||
rg -o -N 'Tag="(UI\.(Layer|Slot)\.[^"]+)"' -r '$1' Config/ Plugins/*/Config/ | sort -u
|
||||
rg -n "RegisterLayer|ExtensionPointTag" --glob "*.cpp" --glob "*.h" .
|
||||
```
|
||||
|
||||
The first command yields the declared vocabulary — the set of names that *may*
|
||||
exist. The mapping from those names to actual containers requires the editor.
|
||||
**Record that as an open question rather than concluding from the empty result**;
|
||||
this is the same rule as any other empty search, and it is the honest form of the
|
||||
answer.
|
||||
|
||||
**Guardrail.** Emit the registered layer and slot inventory to the log at
|
||||
startup, or to a console command. One dump converts a permanently unanswerable
|
||||
question into a cheap one.
|
||||
@@ -0,0 +1,331 @@
|
||||
# UI architecture patterns: a worked analysis
|
||||
|
||||
This is the shell-and-contributor model applied end to end in one shipped
|
||||
product, with the parts worth copying separated from the parts worth knowing
|
||||
about.
|
||||
|
||||
Individual defects live in [failure modes](failure-modes.md), one entry each with
|
||||
a detection recipe. This file is about the shape: what each layer owns, why the
|
||||
seams are where they are, and what it costs.
|
||||
|
||||
---
|
||||
|
||||
## 1. Three things that are usually conflated
|
||||
|
||||
Most UI trouble in a modular project comes from treating these as one concern:
|
||||
|
||||
| Concern | Owner | Changes when |
|
||||
|---|---|---|
|
||||
| **Modality** — does this block input, capture focus, dismiss on back | the layer stack | a screen opens or closes |
|
||||
| **Placement** — where on screen this element sits | the slot inside a layout | the layout is redesigned |
|
||||
| **Content** — what the element is | the contributor | a feature ships |
|
||||
|
||||
A single root widget with direct child references collapses all three into one
|
||||
asset. That asset then belongs to everyone, which means it belongs to nobody, and
|
||||
every feature branch touches it.
|
||||
|
||||
The split that works: **the shell owns modality and placement; features own
|
||||
content and reach the shell only through tags.**
|
||||
|
||||
---
|
||||
|
||||
## 2. Layers: modality, not z-order
|
||||
|
||||
Four layers cover almost every product:
|
||||
|
||||
```text
|
||||
UI.Layer.Game HUD and persistent gameplay presentation
|
||||
UI.Layer.GameMenu pause and in-game menus
|
||||
UI.Layer.Menu front-end navigation
|
||||
UI.Layer.Modal confirmations, errors, blockers
|
||||
```
|
||||
|
||||
Each is an activatable widget stack, not a canvas slot. Pushing a screen
|
||||
activates it and deactivates the one below; popping restores the previous
|
||||
activation and its focus.
|
||||
|
||||
**The rule that keeps this honest:** if you are tempted to add
|
||||
`UI.Layer.HUDLeft`, you are using layers as z-order buckets. Two widgets in one
|
||||
HUD composition belong in extension slots inside the HUD layout, not in separate
|
||||
layers. Layers are for things that change what input means.
|
||||
|
||||
### The root layout is per local player
|
||||
|
||||
One root layout per local player, each with its own layer map, its own focus
|
||||
state and its own modal stack. A single global HUD root cannot support
|
||||
split-screen and cannot give two players independent confirmation dialogs. This
|
||||
is cheap to build in from the start and expensive to retrofit.
|
||||
|
||||
### Registration comes from data, not from code
|
||||
|
||||
Layers register themselves by tag from the layout asset rather than being
|
||||
hardcoded. That keeps the shell designer-editable — and it means no C++ search
|
||||
can tell you which layers actually exist. Write them down somewhere a human
|
||||
reads, because the tag list in configuration is the closest thing to a registry
|
||||
you will have.
|
||||
|
||||
---
|
||||
|
||||
## 3. Extension points: dependency inversion by tag
|
||||
|
||||
The subsystem holds two maps:
|
||||
|
||||
```text
|
||||
ExtensionPointTag → registered points (slots that exist)
|
||||
ExtensionPointTag → registered extensions (content wanting a slot)
|
||||
```
|
||||
|
||||
A **point** is owned by a shell widget: a tag, an optional context object, a
|
||||
whitelist of accepted classes, a match mode, and a callback. An **extension** is
|
||||
owned by a feature: a target tag, an optional context, a widget class or data
|
||||
object, and a handle it must keep.
|
||||
|
||||
Neither side names the other. The only shared vocabulary is the tag.
|
||||
|
||||
### Registration must be order-independent
|
||||
|
||||
This is the property that makes the pattern usable, and it costs two lines:
|
||||
|
||||
- a newly registered point immediately receives all matching extensions that
|
||||
already exist;
|
||||
- a newly registered extension immediately notifies all matching points.
|
||||
|
||||
Without both directions, whether a feature's widget appears depends on whether
|
||||
the feature activated before or after the HUD was constructed — which in a
|
||||
plugin-based project is not deterministic.
|
||||
|
||||
### The contract has three clauses
|
||||
|
||||
An extension matches a point when all of these hold:
|
||||
|
||||
1. **tag** — exact, or the point matched a parent of the extension's tag;
|
||||
2. **context** — both explicitly global, or the same object;
|
||||
3. **class** — the data derives from, or implements, something on the point's
|
||||
whitelist.
|
||||
|
||||
The class check has one subtlety worth copying: if the registered data *is* a
|
||||
class object, use it directly; otherwise use the class of the instance. That one
|
||||
branch is what lets "a widget class" and "a data object" travel through the same
|
||||
registry.
|
||||
|
||||
### Context tiers are what make split-screen work
|
||||
|
||||
A slot widget registers the same tag three times: once globally, once for its
|
||||
local player, once for that player's player state (subscribing for the player
|
||||
state if it has not arrived yet, and firing immediately if it has).
|
||||
|
||||
The result is that one slot tag serves global HUD elements and player-specific
|
||||
elements with no cross-player leakage, and no feature has to know which kind it
|
||||
is contributing.
|
||||
|
||||
---
|
||||
|
||||
## 4. What the slot widget owns
|
||||
|
||||
The container that renders a slot owns registration and unregistration, the
|
||||
conversion of extension data into entry widgets, entry creation and removal,
|
||||
context discovery, the class whitelist, and empty-state layout.
|
||||
|
||||
The feature does not find the slot and does not add children to it. It registers
|
||||
and keeps a handle.
|
||||
|
||||
Two modes fall out of this naturally:
|
||||
|
||||
- **class mode** — the extension is a widget class; the slot instantiates it;
|
||||
- **data mode** — the extension is a data object; the slot asks a factory which
|
||||
widget class fits, creates it, and then configures it with the data.
|
||||
|
||||
Data mode is what makes a HUD list data-driven — the same slot renders whatever
|
||||
the feature describes, and the widget choice is a lookup rather than a branch.
|
||||
|
||||
### Register in the build step, not the constructor
|
||||
|
||||
Registration belongs where the widget's underlying Slate widget is built, and
|
||||
unregistration belongs where those resources are released. Registering in a
|
||||
constructor registers a design-time object; registering in construction-time
|
||||
callbacks that do not pair with a release path leaks.
|
||||
|
||||
---
|
||||
|
||||
## 5. The feature action has two different jobs
|
||||
|
||||
A feature that contributes UI does two unrelated things, and conflating them is a
|
||||
common source of teardown bugs:
|
||||
|
||||
```text
|
||||
LayoutClass → UI.Layer.Game pushed onto a stack
|
||||
WidgetClass → HUD.Slot.Reticle registered into a slot
|
||||
```
|
||||
|
||||
They need separate bookkeeping — handles to pushed layouts, extension handles for
|
||||
slots, actor extension request handles — and they come off by **different
|
||||
mechanisms**. In the measured reference, slot extensions unregister synchronously
|
||||
and completely, while pushed layouts are asked to deactivate and left for the
|
||||
stack to dispose of. Both are defensible; treating them as the same operation is
|
||||
not.
|
||||
|
||||
### Attach to the shell actor through the extension protocol
|
||||
|
||||
Do not search the world for the HUD. Register a handler for the HUD actor class
|
||||
with the component manager and react to both the generic extension event and the
|
||||
shell's own ready event. That covers both orderings — feature before shell, shell
|
||||
before feature — without a retry loop.
|
||||
|
||||
### Assets must already be loaded
|
||||
|
||||
Feature actions typically resolve soft class references with a non-loading
|
||||
accessor, relying on the client asset bundle to have loaded them. That is the
|
||||
right performance choice and it has a failure mode: if the bundle did not load,
|
||||
the widget silently does not appear. Mark the bundle at the declaration, and
|
||||
prefer a code path that logs when the accessor returns null over one that passes
|
||||
null onward.
|
||||
|
||||
---
|
||||
|
||||
## 6. Ordering is a contract or it is nothing
|
||||
|
||||
If the API exposes a priority field, the implementation sorts by it. If it does
|
||||
not sort, the field is a lie with a plausible name — and in the measured
|
||||
reference it is exactly that: stored, copied into the request, never compared
|
||||
(recipe UI-01).
|
||||
|
||||
Before shipping an ordering field, define:
|
||||
|
||||
- whether higher or lower wins;
|
||||
- what happens for equal values;
|
||||
- whether ordering is per point, per tag, or per context;
|
||||
- what happens when an extension changes priority after registration;
|
||||
- who owns the sort — the subsystem or the container.
|
||||
|
||||
If ordering does not matter, do not expose the field. An unimplemented ordering
|
||||
knob costs more than no knob, because designers will spend a day tuning it.
|
||||
|
||||
---
|
||||
|
||||
## 7. Loading screens are arbitration, not a layer
|
||||
|
||||
A loading screen sits above everything, including the root layout, and it is not
|
||||
pushed onto a layer. The manager polls every registered participant:
|
||||
|
||||
```text
|
||||
ShouldShowLoadingScreen(out Reason) → bool
|
||||
```
|
||||
|
||||
Participants: the experience or mode loader, the front-end flow, travel and
|
||||
session state, and explicitly registered tasks. The manager owns presentation;
|
||||
participants only report need and reason.
|
||||
|
||||
**Two properties are worth copying verbatim.**
|
||||
|
||||
*Every reason is a human-readable string, and the helper asserts it is non-empty.*
|
||||
This is the single best diagnostic in the whole UI stack: when a loading screen
|
||||
hangs, one query answers "why", and it answers in words.
|
||||
|
||||
*The screen holds for a configurable extra period after the last blocker clears,
|
||||
with world rendering re-enabled during the hold* — so streaming has time to catch
|
||||
up and the player does not arrive to a texture pop. Note the documented gap: if a
|
||||
new blocker appears during the hold, the rendering flag is not restored.
|
||||
|
||||
Costs to measure rather than inherit: the manager polls every frame through a
|
||||
long chain of checks, and hiding the screen may force a full garbage collection —
|
||||
a guaranteed hitch at exactly the moment control returns to the player
|
||||
(recipe UI-07).
|
||||
|
||||
---
|
||||
|
||||
## 8. Input and focus belong to the screen
|
||||
|
||||
Each activatable screen declares its own input configuration: game-only,
|
||||
menu-only or both; mouse capture behaviour; desired focus target; back action.
|
||||
The stack applies the configuration of whatever is on top.
|
||||
|
||||
This is what makes "HUD means game input, menu means UI input" a declarative
|
||||
property rather than imperative code sprinkled through open and close handlers.
|
||||
|
||||
### The gamepad failure nobody catches in review
|
||||
|
||||
A screen with no desired focus target opens, and the gamepad has nothing to move
|
||||
from. The player is stuck with a visible screen and no way to interact.
|
||||
|
||||
The mitigation worth copying is an editor-time validation that warns when a
|
||||
screen subclass does not implement the focus-target hook. Note what it is: a
|
||||
**warning**, not an error. A screen with no focus target still compiles and still
|
||||
ships (recipe UI-09). If gamepad support matters, promote it.
|
||||
|
||||
### Transitions and focus interact
|
||||
|
||||
Stack transition animations can break focus hand-off to the incoming screen. The
|
||||
measured reference sets transition duration to zero at layer registration, with a
|
||||
comment explaining exactly that. The consequence — the product has no screen
|
||||
transition animations at all — is a real design cost, accepted deliberately.
|
||||
Know that you are making the same trade.
|
||||
|
||||
---
|
||||
|
||||
## 9. What UI may read, and what it may not
|
||||
|
||||
UI may subscribe to local messages for **transient facts**: an elimination for
|
||||
the feed, an accolade for a toast, an inventory delta for a notification.
|
||||
|
||||
UI reads **state** from an observable model: health, score, inventory contents,
|
||||
settings.
|
||||
|
||||
The line: do not rebuild state from a stream of events. A widget created after
|
||||
the third elimination must show the correct score, and a message-derived counter
|
||||
cannot promise that.
|
||||
|
||||
In the other direction, UI sends commands through an accountable service or
|
||||
controller, not as a broadcast. "Please equip weapon" with no named handler has
|
||||
no failure result and no owner.
|
||||
|
||||
---
|
||||
|
||||
## 10. Costs of this architecture
|
||||
|
||||
Stated plainly, so the decision is informed:
|
||||
|
||||
- **Indirection.** Finding "what draws this HUD element" means following a tag
|
||||
through a registry rather than a reference through an asset. There is no
|
||||
"find references" for a tag.
|
||||
- **No compile-time check on the join.** A slot with a typo in its tag renders
|
||||
nothing and reports nothing. Same for a contributor.
|
||||
- **Runtime debugging is push-based.** If the subsystem exposes no query of its
|
||||
current registrations, the only tool is verbose logging.
|
||||
- **World-scoped registries do not survive travel.** Correct for
|
||||
feature-scoped UI, wrong for anything that must persist across maps.
|
||||
- **Order is not guaranteed** unless you implement it (§6).
|
||||
|
||||
### When to simplify
|
||||
|
||||
A single-player game with one fixed HUD, no local multiplayer and no modular
|
||||
modes does not need this. A root widget with direct children is honest and
|
||||
smaller.
|
||||
|
||||
Introduce layers and extension points when at least one is true: plugins or modes
|
||||
add UI independently; local multiplayer is required; screens need distinct input
|
||||
policies; screens load asynchronously and stack; contributors must not depend on
|
||||
the shell widget class; gamepad focus restoration must be systematic.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The analysis above comes from a line-by-line audit of the UI stack of Epic's Lyra
|
||||
Starter Game on Unreal Engine 5.6 — its extension subsystem, its layer and policy
|
||||
plugins, its loading screen plugin, and the game module that consumes them — read
|
||||
as source rather than run.
|
||||
|
||||
Two boundaries applied to that reading and apply to these claims: the engine's
|
||||
own CommonUI implementation is not part of that project, so statements about
|
||||
focus and activation internals are inferred from usage rather than read; and
|
||||
widget assets are binary, so the concrete layout of any screen was not inspected.
|
||||
|
||||
Source addresses stay in the research archive that produced this skill. Entry
|
||||
identifiers in [failure modes](failure-modes.md) resolve back to the audited
|
||||
locations, so any specific claim can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
One project, one engine version, one workspace. The architecture is transferable;
|
||||
the specific defects are evidence, not guarantees about other versions. Re-run
|
||||
the detection recipes against your own tree before acting on any specific claim.
|
||||
Reference in New Issue
Block a user