Files
ue-toolchain/plugins/ue-design-skills/skills/ue-runtime-allocation-and-caching/references/failure-modes.md
T
ue-toolchain dab3f35079 feat(skills): ship ue-design-skills bundle, licensing and delivery gate
Phase 0 of the handoff plan, as a marketplace rather than a flat skills/
directory. Content moved out of the LyraResearch archive and depersonalised:
addresses stay in the archive, recipes ship.

- plugins/ue-design-skills: 17 skills, 232 failure-mode entries, each with the
  six required fields; catalog.json as the harness-neutral source of truth and
  .claude-plugin/ as one adapter over it.
- _gate: 16 rules, one poisoned fixture per rule, plus surface coverage so a
  declared file cannot silently miss the line rules.
- ADR-0002 (harness-neutral bundle behind a marketplace) and ADR-0003 (split
  licensing: CC BY-ND 4.0 prose, Apache-2.0 code and metadata).
- LICENSE files at both levels, CONTRIBUTING.md, docs/licensing-options.md as
  the material the licence decision grew from.

Verified: gate.py 0 violations; test_gate.py 16/16 rules redden on their
fixtures with a clean baseline and 2 root files reaching the line rules.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 23:48:55 +07:00

16 KiB

Failure modes: runtime allocation and caching

Ten ways a running game allocates more each minute, keeps what it should drop, or dereferences memory that moved.

The shared property splits this list in two, and the split is worth stating because it changes how you look for each half.

The dangling-pointer half is data-dependent. It needs a second key inserted while a first entry is still live, or a new statistic to appear mid-session. Neither happens in a short test, both happen in a match, and the crash lands far from the cause. A null check never catches it, because the pointer is not null.

The growth half is time-dependent. Nothing is wrong at any instant; the defect is a slope. A single memory sample cannot show it and a single frame cannot contain it, which is why every recipe below asks for two measurements rather than one.

Recipes use rg from a project source root, and were executed against the audited project while this file was written.


Pointers into moving storage

RC-01 - A raw pointer into a map value, held across frames

Mechanism. A reference to a map value is taken, its address is stored in a long-lived entry, and the map later gains a key. The map reallocates; the stored address points into freed memory.

Why it is silent. Until a second key arrives, the address stays valid and everything works. The window is narrow and specific — a live entry plus an insert — so ordinary play produces it and ordinary testing does not.

Why the obvious check misses it. The code often has a check, and the check passes: the pointer is not null, it is non-null garbage. Reading the storing line shows a reference to a pool, which is exactly what the entry needs. The defect is a property of the container's reallocation behaviour, which is nowhere in view.

Symptom. A crash inside the pooled object's own method, under load, with a stack that looks like the pool is corrupt. In the audited project the window is a one-second component lifespan plus a second visual style — trivially produced by mixed normal and critical damage in the same second.

Detect. Find addresses taken from containers and stored:

rg -n -B2 -A4 "FindOrAdd\(|\.Find\(" --glob "*.cpp" . | rg "&\w+|Emplace\(|= &"
rg -n "\w+\*\s+\w+ = nullptr;" --glob "*.h" . | rg -i "pool|cache|list|entry"

The second search finds the declaration side — a bare pointer member whose name suggests it points at a collection — which is often easier to spot than the assignment.

Guardrail. Store a key or an index and resolve at use. If a stable address is genuinely required, use a container that guarantees stable element addresses and say so at the declaration. When you suspect one, run under the allocator's stomp mode: it converts a probabilistic corruption into a deterministic crash on the right line, and costs one run.


RC-02 - A cache pointer handed out and stored by a widget

Mechanism. A subsystem returns a pointer into its own map so a widget can read samples cheaply. The widget stores it. The map grows a key when a new data source appears — a network connection, an optional module — and reallocates.

Why it is silent. The optimization is real and the pointer is valid at hand- out. Growth is lazy and event-driven: the map gains keys minutes into a session, when a connection is established or a feature is enabled, long after the widget cached its pointer.

Why the obvious check misses it. Both sides are idiomatic. Returning a pointer to avoid copying a sample buffer is good practice; caching it to avoid a lookup per frame is also good practice. Neither side knows the map is still growing, and nothing in either signature says the pointer has a lifetime.

Symptom. A crash in the widget's paint path, appearing only after the session has been running long enough to acquire a new statistic — and therefore never in a short repro. Compounded when the subsystem's teardown resets the tracker while the widget still holds the pointer.

Detect. Find getters returning interior pointers, then find who stores them:

rg -n "const \w+\* \w+::Get\w*Data\(" --glob "*.cpp" . -A4 | rg "\.Find\("
rg -n "const \w+\* \w+ = nullptr;" --glob "*.h" . | rg -i "widget|graph|view"

A getter returning Find on a member map, plus a member pointer of that type in a widget, is the pair.

Guardrail. Return a copy of the small value, or a handle the subsystem can invalidate, or require the caller to re-fetch each frame. If a pointer must escape, give the subsystem an invalidation broadcast and make subscribing mandatory.


Unbounded growth

RC-03 - Read, append, write back — with no clear

Mechanism. Each event fetches an entire collection from another system, appends one element, and writes the whole thing back. Nothing ever clears it.

Why it is silent. Every individual call is correct and fast. The array is a legitimate data channel to the effects system, and its contents are consumed visually, so nobody looks for a removal path.

Why the obvious check misses it. The two lines read as an idiomatic accessor pair. The cost — two full copies per event — is invisible at small sizes, and the growth is invisible without a second measurement. Reviewing the function shows three lines of correct code.

Symptom. Cost per event rising through a match, total work quadratic in event count, and memory that never returns. In a damage-number path this means the hundredth hit costs measurably more than the first.

Detect. Find get-append-set triples on the same key:

rg -n -A3 "= \w*::Get\w*Array\w*\(" --glob "*.cpp" . | rg "Add\(|SetNiagara|Set\w*Array"

Then, for each hit, search the file for any clear:

rg -n "Empty\(\)|Reset\(\)|SetNum\(0\)|Clear" <that-file>

An append path with no clear anywhere in the file is the finding.

Guardrail. Give the collection an owner and an eviction rule — a cap, a time-out, or a clear on the event that logically ends its contents. If the consuming system owns the lifetime, ask it for a removal API rather than assuming one.


RC-04 - Re-append every previous element on every trigger

Mechanism. A component that spawns audio and effect components copies all previously active ones into temporary arrays, appends the new ones, and writes the union back — with no removal of finished components.

Why it is silent. Effects play and stop correctly, because stopping is the component's own business. What accumulates is the references, and those are invisible until something counts them.

Why the obvious check misses it. The function reads as careful state management: gather, combine, store. The absent operation is a filter for finished components, and absence in a function that is otherwise thorough is the hardest kind to see.

Symptom. Thousands of retained effect components after a match. Because the arrays are strong references, none of them can be collected, so this is simultaneously a memory leak and a quadratic allocation profile in the hottest cosmetic path — footsteps, at a few per second.

Detect. Find union-rebuild patterns and check for a prune:

rg -n -B4 -A6 "\.Append\(" --glob "*.cpp" . | rg "Empty\(\)|Reset\(\)"
rg -n "UPROPERTY\(Transient\)" -A2 --glob "*.h" . | rg "TArray<TObjectPtr<U\w*(Audio|Niagara|Particle)"

The second search alone is worth running in any project: a strong transient array of effect components is a retention decision, and it needs a removal path.

Guardrail. Filter finished components on each pass, or use weak references and let the components own their own lifetime. Prefer the reset-without-freeing operation to the one that releases the buffer — see RC-06.


RC-05 - Rebuilding descriptors on a timer instead of diffing

Mechanism. A periodic scan removes all its descriptor objects and creates new ones for the same targets, several times a second.

Why it is silent. The result is correct: the right descriptors exist afterwards. Object creation is fast enough that no single frame stands out.

Why the obvious check misses it. The function reads as a clean rebuild — clear, then populate — which is a legitimate and common shape. The cost is in the frequency, which is set elsewhere as a scan rate, and in the downstream churn: each create and destroy propagates through observers into widget pooling and canvas slot management.

Symptom. Steady allocation and observer churn while the player stands still in a crowd of interactive objects. Profiles show constant low-level cost with no obvious owner.

Detect. Find object creation inside timer-driven update functions:

rg -n -B10 "NewObject<" --glob "*.cpp" . | rg "Update\w+\(|OnTimer|ScanRate"
rg -n "ScanRate|Interval|SetTimer" --glob "*.h" --glob "*.cpp" . | rg -i "0\.[0-9]"

Also check the change-detection that gates the rebuild: comparison logic that sorts only when sizes match will report a change on any reorder, making the "only when changed" guard ineffective.

Guardrail. Diff against a stable key and reuse descriptors. If a rebuild is genuinely simpler, measure it at the worst-case target count before accepting the scan rate.


RC-06 - Freeing the buffer in a hot path

Mechanism. The container operation that releases memory is used where the one that keeps capacity was intended, in code that runs many times per second.

Why it is silent. Both operations leave the container empty and both are correct. The difference is whether the next fill reallocates.

Why the obvious check misses it. The two calls are one word apart and read identically at a glance. Nothing distinguishes a hot path from a cold one in the call itself, and in teardown code the freeing version is the right choice — so the same line is correct in one file and wrong in another.

Symptom. Allocator churn proportional to event frequency, showing up as generic allocation cost rather than attributable to the system causing it.

Detect. Find the freeing operation and classify each site by frequency:

rg -n "\.Empty\(\)" --glob "*.cpp" . -B6 | rg -i "tick|paint|onhit|effect|update"

Teardown sites are fine. Anything reached more than once a second is the finding. In the audited project the two hottest cosmetic arrays are freed and refilled on every footstep.

Guardrail. Reset in hot paths, free at teardown, and write the reason in a comment when it is not obvious which one a site is.


RC-07 - Allocating in a paint or arrange path

Mechanism. A per-frame layout or paint function builds a local array, without reserving, and lets it reallocate as it fills.

Why it is silent. It is correct, local and easy to read. The cost is a handful of small allocations per frame, which no single profile sample attributes.

Why the obvious check misses it. A local array in a function is the most ordinary thing in the file. Knowing it matters requires knowing the function is called once or more per frame — which is a property of the caller.

Symptom. A frame-time floor that nothing accounts for, and allocator pressure that scales with the number of on-screen elements.

Detect. Find containers declared in per-frame widget paths:

rg -n -A12 "::OnPaint\(|::OnArrangeChildren\(|::Tick\(" --glob "*.cpp" . \
  | rg "TArray<|TMap<"
rg -n "static TArray|static TMap" --glob "*.cpp" . | rg -i "widget|slate|paint"

The second search finds the wrong fix for the first: a static container avoids the allocation and introduces sharing between every instance of the widget, plus a thread-safety assumption nobody wrote down. In the audited project both appear — one path allocates per frame, another uses a static.

Guardrail. A mutable member with a reset at the top of the function. Not a local, not a static.


Semantics

RC-08 - A prefilled ring buffer reporting statistics

Mechanism. A fixed-capacity sample cache is filled with zeroes at construction. Minimum, average and related statistics include the prefix until the buffer wraps.

Why it is silent. Every value returned is a real number computed correctly from the buffer's contents. There is no error state for "not enough samples yet", because the buffer is always full.

Why the obvious check misses it. The implementation is correct for a full buffer, and a reviewer checks the arithmetic rather than the initial condition. The distortion disappears after a few seconds, so it is gone before anyone investigates.

Symptom. A minimum of zero and a depressed average on any graph for the first seconds after it appears — which is exactly when someone is looking at it.

Detect. Find prefilled buffers and check whether the statistics account for fill level:

rg -n "AddZeroed\(|SetNumZeroed\(" --glob "*.h" --glob "*.cpp" .
rg -n -A6 "GetMin\(\)|GetAverage\(\)" --glob "*.h" . | rg "Num\(\)|SampleSize|NumValid"

Statistics dividing by capacity rather than by a fill count is the finding.

Guardrail. Track the number of valid samples and compute over that window. Also check the accessor names: an index-based "current" that returns the next write position is the oldest sample, not the newest — a naming trap worth checking in any ring buffer.


RC-09 - An identifier that is a count, at three different widths

Mechanism. A batch identifier is assigned from the number of in-flight batches rather than a monotonic sequence, so identifiers are reused as batches complete. Separately, the same value is one width at the source, another in the message, and a third in storage.

Why it is silent. With one or two batches in flight, counts and sequences are indistinguishable. The width mismatch is invisible until the value exceeds the smallest type — which needs sustained fire and a degraded connection.

Why the obvious check misses it. Each half is defensible alone: a count is a plausible identifier when everything completes promptly, and each width is reasonable in its own context. Confirming means reading a declaration, a constructor and a message signature that are not adjacent — in the audited project the storage width and the message width are sixteen lines apart in the same header, which is exactly far enough.

Symptom. Confirmations applied to the wrong batch when identifiers are reused, and — once the counter passes the narrow type's range — confirmations that match nothing at all, so the pending list grows without bound.

Detect. Trace one identifier through the whole path and compare widths:

I='UniqueId'
rg -n "\b$I\b" --glob "*.h" --glob "*.cpp" .

Read the output as a table of types. Then check the assignment: an identifier taken from a Num() or a count getter is a reused identifier.

Guardrail. One monotonic sequence, one width, asserted at every boundary. Define wrap behaviour. Give the pending list a timeout and a drain, so a matching failure degrades instead of accumulating.


RC-10 - Broadcasting before the state is consistent

Mechanism. A registry broadcasts an "added" event before inserting the item into its own collection. A listener that responds by enumerating the collection sees a state that does not include the item it was just told about.

Why it is silent. For a listener that only uses the event payload, the order is irrelevant and everything works. The defect needs a listener that also reads the collection — typically one doing late binding, catching up on existing items.

Why the obvious check misses it. Two adjacent lines, both correct, in an order that reads naturally: announce, then record. Reversing them looks like a stylistic preference until you know a listener enumerates.

Symptom. An item lost or added twice, depending on which listener ran when. In the audited project a canvas both handles the event and enumerates existing items on attach, and it adds to its own list without a duplicate check.

Detect. Find broadcasts that precede the mutation they announce:

rg -n -B2 -A2 "\.Broadcast\(" --glob "*.cpp" . | rg "Add\(|Remove\(|Emplace\("

A broadcast on the line before the insert is the finding. Then check the listener for a duplicate guard.

Guardrail. Mutate, then broadcast. Listeners must observe a consistent container, and any listener that also enumerates on attach needs an idempotent add.