feat(skills): ship ue-design-skills bundle, licensing and delivery gate
Phase 0 of the handoff plan, as a marketplace rather than a flat skills/ directory. Content moved out of the LyraResearch archive and depersonalised: addresses stay in the archive, recipes ship. - plugins/ue-design-skills: 17 skills, 232 failure-mode entries, each with the six required fields; catalog.json as the harness-neutral source of truth and .claude-plugin/ as one adapter over it. - _gate: 16 rules, one poisoned fixture per rule, plus surface coverage so a declared file cannot silently miss the line rules. - ADR-0002 (harness-neutral bundle behind a marketplace) and ADR-0003 (split licensing: CC BY-ND 4.0 prose, Apache-2.0 code and metadata). - LICENSE files at both levels, CONTRIBUTING.md, docs/licensing-options.md as the material the licence decision grew from. Verified: gate.py 0 violations; test_gate.py 16/16 rules redden on their fixtures with a clean baseline and 2 root files reaching the line rules. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,390 @@
|
||||
---
|
||||
name: ue-multiplayer-authority
|
||||
description: >-
|
||||
Design and review the authority model of a networked Unreal Engine product:
|
||||
which side owns each fact, what is predicted versus replicated versus
|
||||
validated, where responsibility boundaries run between GameState, PlayerState,
|
||||
Controller, Pawn and components, RPC and validation discipline, hit
|
||||
registration and lag compensation, replication modes, relevancy and cost, and
|
||||
the cheat surface. Use when starting a multiplayer feature, porting
|
||||
single-player code to network, reviewing trust boundaries, or investigating
|
||||
desync, rubber-banding, "works in PIE but not on dedicated server", or
|
||||
suspected cheating.
|
||||
---
|
||||
|
||||
# UE multiplayer authority
|
||||
|
||||
The invariant:
|
||||
|
||||
> Every fact in a networked game has exactly one owner. Authority is not a
|
||||
> property of code that says `HasAuthority` — it is the guarantee that no other
|
||||
> machine can produce that fact.
|
||||
|
||||
Measured patterns from a reference product: [patterns](references/patterns.md).
|
||||
Twenty detection recipes for silent trust failures:
|
||||
[failure modes](references/failure-modes.md).
|
||||
|
||||
Read the failure modes before adopting or reviewing a networked design. Each
|
||||
entry names the mechanism, why it stays silent, why the obvious check misses it,
|
||||
and a command you can run against your own tree.
|
||||
|
||||
Related skills: `ue-gas-architecture`, `ue-gameplay-messaging`,
|
||||
`ue-modular-gameplay`, `ue-cosmetics-and-teams`, `ue-architecture-guardrails`.
|
||||
|
||||
Method, not architecture — how to check any claim in this bundle before repeating it: `ue-evidence-discipline`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Start from one question
|
||||
|
||||
Before any code: **for each fact in the game, who is allowed to be wrong about
|
||||
it?**
|
||||
|
||||
- If a client being wrong is a cosmetic glitch → the client may own it.
|
||||
- If a client being wrong is an advantage → the server must own it, and must be
|
||||
able to produce the value independently.
|
||||
- If a client being wrong is unrecoverable → the server must own it *and*
|
||||
persist it.
|
||||
|
||||
This question, answered per fact and written down, is the authority model.
|
||||
Everything else in this document is enforcement of that table.
|
||||
|
||||
### The three states of a networked fact
|
||||
|
||||
| State | Meaning | Who can be wrong |
|
||||
|---|---|---|
|
||||
| **Predicted** | Client computes it immediately for responsiveness; server may overrule | Client, temporarily, visibly |
|
||||
| **Replicated** | Server computes it, clients receive it | Nobody — clients only observe |
|
||||
| **Validated** | Client proposes it, server independently verifies before accepting | Client may lie; server must detect |
|
||||
|
||||
**The third is the expensive one, and it is the one that is usually missing.**
|
||||
"Replicated" is often mistaken for "validated": a value that travelled from the
|
||||
client to the server and was then replicated back out to everyone is *client
|
||||
authored*, no matter how authoritative it looks on arrival.
|
||||
|
||||
---
|
||||
|
||||
## 2. What actually enforces authority
|
||||
|
||||
Unreal offers several mechanisms that look similar and enforce very different
|
||||
amounts. Know which is which before relying on one.
|
||||
|
||||
| Mechanism | Enforces | Does not enforce |
|
||||
|---|---|---|
|
||||
| `HasAuthority()` guard | that this code path runs on the server | anything about callers that skip the guard |
|
||||
| `UFUNCTION(Server)` | routing to the server | that the arguments are truthful |
|
||||
| `WithValidation` | that a `_Validate` function runs and can disconnect | anything, if `_Validate` returns `true` unconditionally |
|
||||
| `BlueprintAuthorityOnly` | that **Blueprint** cannot call this | nothing at all for C++ callers |
|
||||
| replicated property | one-way flow server → client | that the server computed the value itself |
|
||||
| `GetLifetimeReplicatedProps` conditions | who receives it | who may change it |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **A public mutator on a replicated fact must guard itself**, not rely on its
|
||||
callers. `if (!HasAuthority()) return;` at the top of the function, not at the
|
||||
call site — call sites multiply.
|
||||
2. **`WithValidation` with an empty `_Validate` is worse than none**: it passes
|
||||
review as validation while enforcing nothing.
|
||||
3. **`BlueprintAuthorityOnly` is a Blueprint editor affordance.** Never cite it as
|
||||
a security property.
|
||||
4. **Write the guard even when the only current caller is the server.** The guard
|
||||
is documentation the compiler enforces; the caller list is not stable.
|
||||
|
||||
---
|
||||
|
||||
## 3. Responsibility boundaries by layer
|
||||
|
||||
Put each fact in exactly one place. The most common source of networked bugs is a
|
||||
fact that lives in two.
|
||||
|
||||
| Layer | Exists on | Owns | Must not own |
|
||||
|---|---|---|---|
|
||||
| **GameMode** | server only | match rules, spawning, scoring decisions | anything a client must read |
|
||||
| **GameState** | server + all clients | match-wide replicated facts: phase, timer, team scores | per-player secrets |
|
||||
| **PlayerState** | server + all clients | per-player facts everyone may see: name, team, score, abilities | input, camera, UI state |
|
||||
| **PlayerController** | server + owning client | input, camera, client-directed commands, UI ownership | facts other players need |
|
||||
| **Pawn / Character** | server + all clients | embodiment: transform, movement, physical state | player identity, persistent progression |
|
||||
| **Components** | follow their owner | one cohesive slice of the owner's facts | facts belonging to a different layer |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **`GameMode` exists only on the server.** Any client code path referencing it
|
||||
is a bug that PIE will hide.
|
||||
2. **Persistent player facts belong to `PlayerState`, not the Pawn.** Pawns die.
|
||||
3. **Input and camera belong to the Controller.** A Pawn reading input directly
|
||||
cannot be possessed by AI or spectated.
|
||||
4. **A component's authority is its owner's authority.** A component attached to a
|
||||
client-owned actor cannot be authoritative regardless of its internal guards.
|
||||
5. **Simulated proxies have no controller.** Any logic gated on "has a controller"
|
||||
silently excludes them; any logic that assumes one will null-deref on them.
|
||||
|
||||
---
|
||||
|
||||
## 4. RPC discipline
|
||||
|
||||
### Choosing the direction
|
||||
|
||||
```text
|
||||
Client -> Server UFUNCTION(Server, Reliable/Unreliable, WithValidation)
|
||||
Server -> owner UFUNCTION(Client, ...)
|
||||
Server -> everyone UFUNCTION(NetMulticast, ...)
|
||||
```
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Every `Server` RPC carries `WithValidation` and a `_Validate` that actually
|
||||
checks.** Range, ownership, state legality, rate. An RPC without validation is
|
||||
a client-side API into your server.
|
||||
2. **Reliable is a budget, not a default.** Reliable RPCs are ordered and queued;
|
||||
overflowing the queue disconnects the client. Anything cosmetic, frequent or
|
||||
idempotent should be unreliable.
|
||||
3. **Never multicast what a replicated property already carries.** The property
|
||||
handles joiners and relevancy; a multicast does not.
|
||||
4. **Multicasts do not reach late joiners or newly relevant clients.** Anything a
|
||||
late joiner must know is state, not an event.
|
||||
5. **Validate on the server that the sender owns the thing it is acting on.**
|
||||
Routing to the server proves nothing about the *subject* of the call.
|
||||
6. **Rate-limit anything a client can call in a loop.**
|
||||
|
||||
### Debug and cheat commands
|
||||
|
||||
Cheat entry points reachable from a client are the highest-value target in the
|
||||
codebase. They must be compiled out of shipping builds — not merely hidden,
|
||||
disabled by a flag, or guarded by UI. A `Server` RPC that executes an arbitrary
|
||||
cheat string, present in a shipping binary, is a remote console.
|
||||
|
||||
---
|
||||
|
||||
## 5. Prediction
|
||||
|
||||
Prediction exists to hide latency, not to save server work. The server must still
|
||||
compute everything it accepts.
|
||||
|
||||
### Safe to predict
|
||||
|
||||
- movement through the engine's own prediction (correction path already exists);
|
||||
- animation, audio, VFX, camera;
|
||||
- UI affordances (cooldown wheels, ammo counters) that a correction can restore;
|
||||
- ability activation under a prediction key, where the server can reject.
|
||||
|
||||
### Never predict
|
||||
|
||||
- damage application;
|
||||
- score, currency, progression;
|
||||
- inventory grants;
|
||||
- death and elimination;
|
||||
- anything persisted.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Every predicted value needs a defined correction path** — what happens
|
||||
visually when the server disagrees. Undefined correction means the client
|
||||
keeps a wrong value forever.
|
||||
2. **Prediction is a client-side optimisation of a server-side truth.** If the
|
||||
server cannot independently produce the value, the feature is not predicted,
|
||||
it is client-authoritative.
|
||||
3. **A prediction key is not validation.** It reconciles ordering; it does not
|
||||
check whether the client was entitled to act.
|
||||
|
||||
---
|
||||
|
||||
## 6. Hit registration — the decision that shapes the whole product
|
||||
|
||||
This is where authority is usually lost, and it is a design decision that must be
|
||||
made explicitly and early, because retrofitting is expensive.
|
||||
|
||||
There are three positions:
|
||||
|
||||
| Model | Server does | Cost | Cheat exposure |
|
||||
|---|---|---|---|
|
||||
| **Server-authoritative trace** | traces from its own state at the current tick | cheap | none, but shots feel like they miss under latency |
|
||||
| **Lag-compensated rewind** | stores per-tick historical positions, rewinds to the client's timestamp, re-traces | expensive: memory, CPU, complexity | low; the standard for competitive shooters |
|
||||
| **Client-reported hit** | accepts the client's hit result | free | total |
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Choose consciously and write the choice down.** Most projects arrive at the
|
||||
third by default, because the first feels bad and the second is work.
|
||||
2. **If you accept client-reported hits, know exactly what a malicious client
|
||||
gains** — typically arbitrary damage, arbitrary range, arbitrary target,
|
||||
arbitrary surface multipliers.
|
||||
3. **The server cannot validate what it cannot reproduce.** Without stored
|
||||
historical positions there is nothing to validate a client's hit *against*,
|
||||
so bolting on "validation" later is not a patch, it is building the rewind
|
||||
subsystem.
|
||||
4. **Sanity checks are not validation, but they are not worthless.** Maximum
|
||||
range, line-of-sight from the shooter's replicated position, fire-rate limits,
|
||||
team check, and "was the target ever near there recently" raise the cost of
|
||||
trivial cheats even without full rewind.
|
||||
5. **Never let the client choose damage-modifying inputs.** The surface/material
|
||||
that multiplies damage, the distance used for falloff, and the trace origin
|
||||
must all come from server-known state, even in a client-reported-hit model.
|
||||
These are the cheapest things to move server-side and the most valuable.
|
||||
|
||||
### Recognising a stub
|
||||
|
||||
Watch for parameters threaded through many function signatures but never branched
|
||||
on, and for booleans initialised to a constant with a `//@TODO`. These are the
|
||||
attachment points of a rewind subsystem that was designed and never built. They
|
||||
make the code *look* like it validates. Recipes: NA-06, NA-07 in the failure
|
||||
modes.
|
||||
|
||||
---
|
||||
|
||||
## 7. Replication cost and relevancy
|
||||
|
||||
Authority decisions determine correctness; relevancy decisions determine whether
|
||||
the product ships.
|
||||
|
||||
### Rules
|
||||
|
||||
1. **Pick an ability-system replication mode deliberately.** `Full` replicates all
|
||||
gameplay effects to all clients; `Mixed` replicates effects to the owner and
|
||||
tags/cues to everyone; `Minimal` replicates no effects to anyone.
|
||||
Player-controlled actors want `Mixed`; AI want `Minimal`.
|
||||
2. **Net cull distance is a gameplay decision, not a performance knob.** Setting
|
||||
it beyond the map size disables relevancy culling entirely — every actor
|
||||
replicates to every client, which is correct for a small arena and fatal for a
|
||||
large map.
|
||||
3. **Prefer replicated state over event spam.** State is idempotent, survives
|
||||
packet loss, and reaches late joiners.
|
||||
4. **FastArraySerializer for collections that change incrementally**, with the
|
||||
`PreReplicatedRemove` / `PostReplicatedAdd` / `PostReplicatedChange` callbacks
|
||||
all implemented — an empty callback is a client-side desync with no symptom on
|
||||
the server.
|
||||
5. **Replication graph before you need it, or never.** Enabling it late changes
|
||||
which actors clients see, and every relevancy assumption in gameplay code has
|
||||
to be re-verified. If it ships disabled, treat routing configuration as
|
||||
untested.
|
||||
6. **Cosmetic state should not replicate.** Replicate intent (which skin, which
|
||||
team), realise presentation locally.
|
||||
|
||||
---
|
||||
|
||||
## 8. Lifecycle across the network
|
||||
|
||||
1. **Readiness is a state chain, not a timing assumption.** Actor spawned,
|
||||
controller assigned, PlayerState replicated, data initialised, gameplay ready —
|
||||
each an explicit state with explicit transitions.
|
||||
2. **Client and server reach each state in a different order.** Any code that
|
||||
assumes an order will work for the listen-server host and fail for remote
|
||||
clients.
|
||||
3. **A simulated proxy passes through the chain without ever acquiring a
|
||||
controller.** Gates must handle that case explicitly rather than by omission.
|
||||
4. **Every dynamic addition needs its exact removal**: granted abilities, applied
|
||||
effects, registered listeners, spawned cues, added tag paths. Test the removal
|
||||
on both server and client, in the same session, more than once.
|
||||
5. **Death is not destruction.** Define the full sequence — death started, effects
|
||||
removed, input released, cues cleaned, actor destroyed or respawned — and give
|
||||
each step one owner.
|
||||
|
||||
---
|
||||
|
||||
## 9. The cheat surface review
|
||||
|
||||
Enumerate, for the shipping build, every path a malicious client can reach:
|
||||
|
||||
```text
|
||||
1. Every Server RPC -> does _Validate independently verify?
|
||||
2. Every RPC argument -> is any of it trusted without recomputation?
|
||||
3. Every cheat/debug entry point -> is it compiled out, not just disabled?
|
||||
4. Every damage input -> which come from the client?
|
||||
5. Every fail-open error path -> what does the code do when a check errors out?
|
||||
6. Every client-side gate -> is there a matching server-side one?
|
||||
```
|
||||
|
||||
Item 5 deserves attention: a permission check that returns "allowed" when it
|
||||
cannot determine the answer converts an unknown state into an exploit. Failure to
|
||||
determine must deny.
|
||||
|
||||
---
|
||||
|
||||
## 10. Test matrix
|
||||
|
||||
**PIE is not a network test.** A listen-server host shares memory with its client;
|
||||
half of the failure modes here cannot occur in it.
|
||||
|
||||
### Mandatory configurations
|
||||
|
||||
- dedicated server + two remote clients;
|
||||
- one client with 150 ms latency and 2% packet loss;
|
||||
- a client joining mid-match;
|
||||
- a client disconnecting mid-action;
|
||||
- an actor becoming relevant and then irrelevant.
|
||||
|
||||
### Per feature
|
||||
|
||||
- does it work for a remote client, not just the host?
|
||||
- does a late joiner reach the correct state?
|
||||
- what does a simulated proxy see?
|
||||
- what happens if the client sends the RPC twice, out of order, or with absurd
|
||||
arguments?
|
||||
- is the state correct after death and respawn, twice?
|
||||
|
||||
### Authority regression
|
||||
|
||||
- for each fact in the authority table, assert on the server that the client
|
||||
cannot change it — with a test client that tries.
|
||||
|
||||
---
|
||||
|
||||
## 11. Review checklist
|
||||
|
||||
- [ ] Every fact has a named owner in a written authority table.
|
||||
- [ ] Each fact is classified predicted / replicated / validated.
|
||||
- [ ] Public mutators of replicated state guard themselves.
|
||||
- [ ] Every `Server` RPC has `WithValidation` with a real `_Validate`.
|
||||
- [ ] No trust in client-supplied damage inputs: origin, distance, surface, target.
|
||||
- [ ] Hit registration model is chosen and documented.
|
||||
- [ ] Cheat entry points are compiled out of shipping.
|
||||
- [ ] Error paths deny rather than allow.
|
||||
- [ ] Replication mode and net cull distance chosen per class, not inherited.
|
||||
- [ ] FastArray callbacks all implemented.
|
||||
- [ ] Every addition has a verified removal.
|
||||
- [ ] Feature tested on a dedicated server with a remote client under latency.
|
||||
- [ ] `BlueprintAuthorityOnly` is not cited as enforcement anywhere.
|
||||
|
||||
---
|
||||
|
||||
## 12. When client authority is the right answer
|
||||
|
||||
Deliberate client authority is legitimate when:
|
||||
|
||||
- the fact is purely presentational;
|
||||
- the game is cooperative or single-player-with-friends and the threat model is
|
||||
"no adversary";
|
||||
- the cost of server validation exceeds the value of the protected fact;
|
||||
- an anti-cheat layer outside the game handles the threat.
|
||||
|
||||
In every one of these cases, the decision is written down with its threat model.
|
||||
The failure worth avoiding is not "we trusted the client" — it is **"we did not
|
||||
know we trusted the client."**
|
||||
|
||||
That distinction is what this skill exists to preserve. A codebase can be an
|
||||
excellent architectural reference and simultaneously an unsafe network reference;
|
||||
those are separate axes, and copying the first without noticing the second is the
|
||||
common accident.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
The findings behind the failure modes come from a line-by-line source audit of
|
||||
Epic's Lyra Starter Game on Unreal Engine 5.6, read rather than run. That project
|
||||
is an outstanding architectural reference and a poor network-security reference;
|
||||
those are independent axes, and the audit exists to keep them apart. Source
|
||||
addresses stay in the research archive that produced this skill; what ships is the
|
||||
detection recipe, because an address in someone else's tree is not something you
|
||||
can act on and a recipe is.
|
||||
|
||||
Each entry carries a stable identifier (`NA-01`, `NA-02`, …) that resolves back to
|
||||
the audited location in that archive. If you need the original address to settle a
|
||||
dispute, it exists and can be produced.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
The measured material refers to one specific reference project on one engine
|
||||
version in one workspace. It is evidence of failure modes, not a guarantee about
|
||||
other engine versions or other samples. Several claims in it are explicitly marked
|
||||
as unanswerable without a dedicated-server build, and they stay unanswered. Re-run
|
||||
the recipes against your own tree before trusting any specific claim.
|
||||
@@ -0,0 +1,764 @@
|
||||
# Failure modes: multiplayer authority
|
||||
|
||||
Twenty ways a networked Unreal project loses authority without anything going
|
||||
wrong on screen. Each entry names the mechanism, why it stays silent, why the
|
||||
obvious check misses it, the symptom, a recipe you can run against your own tree,
|
||||
and the guardrail.
|
||||
|
||||
Two properties recur across this list and are worth stating once, because they
|
||||
explain why these survive review:
|
||||
|
||||
- **PIE cannot reproduce most of them.** A listen-server host shares memory with
|
||||
its client, so the client/server split that produces the bug does not exist
|
||||
during the test that would have caught it. Entries NA-01, NA-04, NA-11, NA-13,
|
||||
NA-17 and NA-19 are invisible in PIE by construction.
|
||||
- **Several of them look correct in review.** A `WithValidation` marker, a
|
||||
`BlueprintAuthorityOnly` specifier, a `bIsSimulated` parameter threaded through
|
||||
signatures — each reads as evidence of a server-side path that is not there.
|
||||
The specifier is real; the enforcement is not.
|
||||
|
||||
Recipes below use `rg`. They are written to be run from the root of a project's
|
||||
source tree and to be read as a *comparison of two numbers*, not as a single
|
||||
search: the whole class of defect here is "the thing is present, and the thing
|
||||
that would make it mean something is absent".
|
||||
|
||||
---
|
||||
|
||||
## Trust boundaries
|
||||
|
||||
### NA-01 - Replicated is mistaken for validated
|
||||
|
||||
**Mechanism.** A value travels client to server and is then replicated back out to
|
||||
everyone. On arrival it lives in an authoritative-looking replicated property, and
|
||||
every later reader treats it as a server fact.
|
||||
|
||||
**Why it is silent.** Replication works perfectly. The value is consistent on every
|
||||
machine, arrives in order, survives relevancy changes and late joins. Consistency
|
||||
is not authorship, but consistency is what everyone observes and tests.
|
||||
|
||||
**Why the obvious check misses it.** Review asks "is this replicated?" and the
|
||||
answer is yes. There is no compiler concept, specifier or naming convention that
|
||||
distinguishes a server-computed replicated value from a server-relayed one. Both
|
||||
compile to the same declaration.
|
||||
|
||||
**Symptom.** Nothing at all, until someone modifies a client. Then the value is
|
||||
whatever that client chose, and it is authoritative across the whole session.
|
||||
|
||||
**Detect.** This one is a table, not a grep. List every replicated property and
|
||||
name, for each, the server-side computation that produced it:
|
||||
|
||||
```bash
|
||||
rg -n "UPROPERTY\([^)]*Replicated" Source/ | wc -l # facts to account for
|
||||
rg -n "GetLifetimeReplicatedProps" Source/ # where they are declared
|
||||
```
|
||||
|
||||
For each result, answer in writing: *what did the server compute to get this?* If
|
||||
the answer is "the client sent it", the fact is client-authored. A property whose
|
||||
only writer is an RPC handler that assigns the parameter directly is the shape to
|
||||
look for.
|
||||
|
||||
**Guardrail.** Classify every fact as predicted / replicated / validated in a
|
||||
written table before the feature is built. "Validated" requires the server to
|
||||
independently reproduce the value, not merely to receive it.
|
||||
|
||||
---
|
||||
|
||||
### NA-02 - `WithValidation` with an empty `_Validate`
|
||||
|
||||
**Mechanism.** The RPC carries the `WithValidation` specifier and the paired
|
||||
`_Validate` function exists, so the engine calls it. Its body returns `true`
|
||||
unconditionally.
|
||||
|
||||
**Why it is silent.** The mechanism is fully wired and working. The engine really
|
||||
does call the validator and really would disconnect a client that failed it. There
|
||||
is simply no input it rejects, and a validator that accepts everything is
|
||||
indistinguishable at runtime from one that has nothing to reject yet.
|
||||
|
||||
**Why the obvious check misses it.** The specifier is what reviewers grep for, and
|
||||
it is present. Declarations and implementations live in different files, so a
|
||||
header review — the natural place to audit an RPC surface — never sees the body.
|
||||
|
||||
**Symptom.** The RPC accepts anything. Review signs off on the subsystem as
|
||||
validated, and the estimate for "add validation" is later scoped as a patch when
|
||||
it is a design.
|
||||
|
||||
**Detect.** Read every validator body, not the declarations:
|
||||
|
||||
```bash
|
||||
rg -n --stats "WithValidation" Source/ # declared validators
|
||||
rg -n -A3 "_Validate\(" Source/ | rg -B1 "^\s*return true;" # bodies that accept all
|
||||
```
|
||||
|
||||
Compare the counts. On the measured reference the second command returned two
|
||||
sites, both belonging to cheat RPCs. Any validator whose body is a bare
|
||||
`return true;` is an unimplemented validator regardless of the specifier.
|
||||
|
||||
**Guardrail.** Require every `_Validate` body to check range, ownership, state
|
||||
legality and rate. Treat a one-line `return true;` in a validator as a review
|
||||
blocker, the same way an empty `catch` block is.
|
||||
|
||||
---
|
||||
|
||||
### NA-03 - `BlueprintAuthorityOnly` cited as enforcement
|
||||
|
||||
**Mechanism.** A function is marked `BlueprintAuthorityOnly`. The specifier stops
|
||||
Blueprint graphs from calling it on a non-authoritative actor. C++ callers are
|
||||
entirely unaffected.
|
||||
|
||||
**Why it is silent.** In a project where the function is only ever called from
|
||||
Blueprint, the specifier does exactly what everyone believes it does. The gap
|
||||
opens the first time a C++ caller appears, and that caller works, so nothing draws
|
||||
attention to it.
|
||||
|
||||
**Why the obvious check misses it.** The word "authority" is in the specifier.
|
||||
Grepping for authority-related terms finds it and it looks like a guard. It is an
|
||||
editor affordance whose name describes intent rather than enforcement.
|
||||
|
||||
**Symptom.** A function documented and reviewed as authority-guarded is reachable
|
||||
from any C++ path, including client code, and mutates replicated state from there.
|
||||
|
||||
**Detect.** Enumerate the sites and check each for a real guard in the body:
|
||||
|
||||
```bash
|
||||
rg -n "BlueprintAuthorityOnly" Source/ # claimed guards
|
||||
rg -n -A6 "BlueprintAuthorityOnly" Source/ \
|
||||
| rg "HasAuthority|GetLocalRole" # actual guards
|
||||
```
|
||||
|
||||
A large gap between the two counts is the finding. On the measured reference the
|
||||
specifier appeared at five sites across three subsystems, with no `HasAuthority`
|
||||
in the guarded bodies.
|
||||
|
||||
**Guardrail.** Never cite `BlueprintAuthorityOnly` as a security property in a
|
||||
review, a comment or a design document. Add an explicit `HasAuthority()` guard
|
||||
inside the function body; keep the specifier only as an editor hint.
|
||||
|
||||
---
|
||||
|
||||
### NA-04 - Guard placed at the call site rather than in the mutator
|
||||
|
||||
**Mechanism.** A public mutator of replicated state relies on its callers to check
|
||||
authority. Today every caller does.
|
||||
|
||||
**Why it is silent.** The invariant holds for as long as the caller list is what it
|
||||
was when the code was written. Nothing degrades gradually; the code is correct
|
||||
until one day it is not, and the transition is a new caller, not an edit to the
|
||||
function.
|
||||
|
||||
**Why the obvious check misses it.** Reading the function shows nothing wrong —
|
||||
there is no missing line, because the check was never meant to be there. Reading
|
||||
the call sites shows a check at each one. Both halves pass review independently.
|
||||
The defect only exists in the relationship between them, which no single file
|
||||
review inspects.
|
||||
|
||||
**Symptom.** Works until a new caller appears — commonly a subclass in a feature
|
||||
plugin, a Blueprint node, or a second caller added by someone who reasonably
|
||||
assumed the mutator guards itself. Then a client writes replicated state directly.
|
||||
|
||||
**Detect.** Find public or virtual writers of replicated state whose own body has
|
||||
no guard:
|
||||
|
||||
```bash
|
||||
# 1. names of replicated properties
|
||||
rg -n -o "UPROPERTY\([^)]*Replicated[^)]*\)\s*\n\s*\w+\s+(\w+)" -r '$1' Source/
|
||||
# 2. for each name, find assignments and check the enclosing function
|
||||
rg -n -B12 "\bDeathState\s*=" Source/ | rg "HasAuthority|void |virtual"
|
||||
```
|
||||
|
||||
Substitute your own property name in step 2. The finding is an assignment whose
|
||||
preceding function signature is public or virtual and whose body contains no
|
||||
authority check.
|
||||
|
||||
**Guardrail.** `if (!HasAuthority()) return;` as the first line of the mutator.
|
||||
Write it even when the only current caller is the server: the guard is
|
||||
documentation the compiler enforces, and the caller list is not a stable property.
|
||||
|
||||
---
|
||||
|
||||
### NA-05 - Client-supplied damage inputs
|
||||
|
||||
**Mechanism.** Distance, trace origin, surface material or target identity arrive
|
||||
in the RPC payload and feed the damage calculation directly.
|
||||
|
||||
**Why it is silent.** Every value is in a plausible range, because a normal client
|
||||
sends plausible values. The damage numbers that result look exactly like correct
|
||||
damage numbers. There is no error state and no log line, because nothing failed.
|
||||
|
||||
**Why the obvious check misses it.** The damage code is usually reviewed for
|
||||
correctness of the formula, and the formula is correct. The question that matters
|
||||
is not "does this compute the right number?" but "who chose the inputs?", and the
|
||||
inputs arrive as ordinary function parameters several layers away from where they
|
||||
entered the process.
|
||||
|
||||
**Symptom.** Arbitrary damage multipliers from a modified client, invisible in
|
||||
logs. In a competitive game this surfaces as player reports that a specific
|
||||
weapon or surface "feels wrong", long before anyone suspects the trust boundary.
|
||||
|
||||
**Detect.** Trace each damage-modifying input back to its origin:
|
||||
|
||||
```bash
|
||||
rg -n "PhysicalMaterial|PhysMat|TraceStart|HitLocation|Distance" \
|
||||
Source/ --glob "*Damage*" --glob "*Execution*"
|
||||
```
|
||||
|
||||
For each hit, answer: did the server compute this, or did it arrive from the
|
||||
client? On the measured reference the falloff distance and the surface multiplier
|
||||
both came from client-supplied fields, while the team check beside them was
|
||||
genuinely server-side — the model was incomplete, not absent, which is the harder
|
||||
case to spot.
|
||||
|
||||
**Guardrail.** Recompute every damage-modifying input from server-known state. The
|
||||
shooter's replicated position and a server-side collision query cover most of it
|
||||
and are cheap. Client-supplied values may inform presentation, never arithmetic
|
||||
that changes an outcome.
|
||||
|
||||
---
|
||||
|
||||
### NA-06 - Validation flag initialised to a constant
|
||||
|
||||
**Mechanism.** A boolean local is initialised to a literal — `const bool bIsValid
|
||||
= true;` — and then consumed by branches that read as validation. There is no
|
||||
computation between the initialisation and the use.
|
||||
|
||||
**Why it is silent.** A permanently-valid result is a legal runtime state. Every
|
||||
branch takes the path it would take on a successful validation, which is also the
|
||||
path taken during all normal play, so the code behaves exactly as intended for
|
||||
every non-malicious input.
|
||||
|
||||
**Why the obvious check misses it.** The flag has a meaningful name, appears in
|
||||
conditionals, and is used more than once. Any search for "is validation present?"
|
||||
finds a variable called `bIsTargetDataValid` being consumed by a gate and stops
|
||||
there. The tell is the *initialiser*, which is three lines above and looks like
|
||||
boilerplate.
|
||||
|
||||
**Symptom.** The code path reads as validated in review and in estimates. Nothing
|
||||
is observable at runtime until a modified client sends implausible data and it is
|
||||
accepted.
|
||||
|
||||
**Detect.** Look for boolean locals initialised to a literal and later consumed as
|
||||
gates, restricted to network-relevant code:
|
||||
|
||||
```bash
|
||||
rg -n "const bool b\w+\s*=\s*(true|false)\s*;" Source/
|
||||
```
|
||||
|
||||
Then, for each hit, check whether the name promises a computation
|
||||
(`bIsValid`, `bCanApply`, `bIsAuthorized`) and whether the variable is read by a
|
||||
branch. On the measured reference this returned eight sites; most were legitimate
|
||||
call-parameter constants, and exactly one — a target-data validity flag inside a
|
||||
weapon ability — was a validation stub. The recipe's value is that it reduces the
|
||||
whole codebase to a list short enough to read.
|
||||
|
||||
**Guardrail.** Treat any boolean local in a network path whose initialiser is a
|
||||
literal and whose name asserts a property as unimplemented until proven otherwise.
|
||||
If validation is not written yet, name the variable for what it is.
|
||||
|
||||
---
|
||||
|
||||
### NA-07 - Dead parameters that imply a server path
|
||||
|
||||
**Mechanism.** A parameter such as `bIsSimulated` is threaded through many function
|
||||
signatures, passed a literal at the only call site, and never branched on
|
||||
anywhere.
|
||||
|
||||
**Why it is silent.** The parameter is passed and received correctly at every
|
||||
level. Nothing is unused in the compiler's view, so no warning fires. The code
|
||||
does what it does today with complete consistency.
|
||||
|
||||
**Why the obvious check misses it.** Signature review is exactly the wrong
|
||||
instrument here, and it is the instrument reviewers reach for when auditing a
|
||||
subsystem's shape. A parameter appearing in ten signatures is strong evidence to a
|
||||
human that a feature exists. Confirming it requires following the value to a
|
||||
branch, which no IDE affordance does for you.
|
||||
|
||||
**Symptom.** A reviewer concludes a simulated or authoritative path exists.
|
||||
Estimates for "just add server validation" come out an order of magnitude low,
|
||||
because the attachment points were mistaken for the subsystem.
|
||||
|
||||
**Detect.** Count mentions, then count branches:
|
||||
|
||||
```bash
|
||||
P='bIsSimulated'
|
||||
rg -n "\b$P\b" Source/ | wc -l # mentions: many
|
||||
rg -n "if\s*\(\s*!?\s*$P\b|\?\s*.*:\s*.*\b$P\b" Source/ # branches: ?
|
||||
```
|
||||
|
||||
On the measured reference the parameter appeared at ten sites and was branched on
|
||||
at zero. A high first number with an empty second is a mounting point, not a
|
||||
feature.
|
||||
|
||||
**Guardrail.** Trace every network-relevant parameter to a branch. No branch means
|
||||
no feature: either implement it or delete the parameter, because its presence is
|
||||
an active lie to the next reader.
|
||||
|
||||
---
|
||||
|
||||
### NA-08 - Retrofitting hit validation without rewind
|
||||
|
||||
**Mechanism.** The server holds no historical positions, so there is nothing to
|
||||
validate a timestamped client hit against. Validation is scheduled as a task on
|
||||
top of the existing weapon code.
|
||||
|
||||
**Why it is silent.** The absence is an absence. Nothing in the codebase points at
|
||||
the missing subsystem; there is no stub, no TODO, no failing test. The weapon
|
||||
system works, feels good, and passes playtests, because every client is honest
|
||||
during a playtest.
|
||||
|
||||
**Why the obvious check misses it.** The question "do we validate hits?" gets
|
||||
answered by looking at the hit path, which is full of plausible machinery. The
|
||||
right question is one level down and is about a subsystem nobody expects to find
|
||||
by reading weapon code: does the server store per-tick positions at all?
|
||||
|
||||
**Symptom.** "Add server validation" is scoped as a patch and turns out to be a
|
||||
subsystem. Partial attempts produce shots that visibly miss under latency, which
|
||||
is then treated as a regression and reverted.
|
||||
|
||||
**Detect.** Ask whether the machinery exists at all, before reading any weapon
|
||||
code:
|
||||
|
||||
```bash
|
||||
rg -n "FSavedMove|ServerMove|CompressedFlags|SavedMoves|PositionHistory" Source/
|
||||
```
|
||||
|
||||
Empty output means there is no lag compensation and no custom movement
|
||||
prediction, and therefore no basis for hit validation in principle. On the
|
||||
measured reference this returned nothing at all, which explains every other
|
||||
finding in this file about client trust: they are consequences of one absent
|
||||
subsystem, not independent oversights.
|
||||
|
||||
**Guardrail.** Decide the hit-registration model — server trace, lag-compensated
|
||||
rewind, or client-reported — before the weapon system exists, and record the
|
||||
decision with its threat model. Interim sanity checks (max range, line of sight
|
||||
from the replicated position, fire rate, team) raise cheat cost without full
|
||||
rewind and are worth having in every model.
|
||||
|
||||
---
|
||||
|
||||
### NA-09 - Fail-open permission check
|
||||
|
||||
**Mechanism.** A check returns "allowed" on an indeterminate or error result. The
|
||||
error branch was written to avoid blocking legitimate play during development.
|
||||
|
||||
**Why it is silent.** The indeterminate case does not arise in normal play. Both
|
||||
participants are always fully initialised, always have a PlayerState, always have
|
||||
a resolved team. The permissive branch is dead code for the entire duration of
|
||||
every test.
|
||||
|
||||
**Why the obvious check misses it.** The function has a check, the check is
|
||||
correct in the determinate cases, and unit tests cover those cases. Reviewing the
|
||||
error branch requires asking "what if this returns the error value?", and the
|
||||
answer during development is "it doesn't".
|
||||
|
||||
**Symptom.** Rules bypassed exactly in the edge cases nobody tested — early
|
||||
initialisation, missing PlayerState, a disconnecting client, an actor spawned
|
||||
before its owner. For a friendly-fire check this means damage applied where team
|
||||
rules should have prevented it.
|
||||
|
||||
**Detect.** Find the error branches of permission checks and read their return
|
||||
value:
|
||||
|
||||
```bash
|
||||
rg -n -B3 -A3 "InvalidArgument|Indeterminate|Unknown" Source/ \
|
||||
| rg "return true|return .*Allow|bCanDamage\s*=\s*true"
|
||||
rg -n "//\s*@?TODO.*temporar" -i Source/
|
||||
```
|
||||
|
||||
The second command is worth running on its own: a permissive branch marked
|
||||
temporary is the highest-confidence instance of this class.
|
||||
|
||||
**Guardrail.** Failure to determine must deny. Make the error branch explicit and
|
||||
restrictive, and if a permissive default is genuinely needed during development,
|
||||
gate it on a build configuration so it cannot ship.
|
||||
|
||||
---
|
||||
|
||||
### NA-10 - Cheat RPC surface in shipping
|
||||
|
||||
**Mechanism.** `Server` RPCs taking arbitrary command strings are declared outside
|
||||
any build-configuration guard. The implementation may be neutralised; the surface
|
||||
is not.
|
||||
|
||||
**Why it is silent.** The commands are useful, work as intended, and are reached
|
||||
through a UI that only appears in development. Hiding the entry point in the UI
|
||||
feels like removing it, and the difference never shows up in play.
|
||||
|
||||
**Why the obvious check misses it.** The guard that reviewers look for is usually
|
||||
present *somewhere* — around the cheat manager, around the console UI, around the
|
||||
implementation body. Confirming the RPC declaration itself is unguarded means
|
||||
checking a header for the absence of a preprocessor directive, which is not
|
||||
something a code search naturally expresses.
|
||||
|
||||
**Symptom.** A remote console in the shipped binary. Any client that can encode
|
||||
the RPC can execute arbitrary cheat strings on the server.
|
||||
|
||||
**Detect.** Check whether the declarations sit inside a build guard:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*Server[^)]*\)" -A2 Source/ | rg -i "cheat|exec|command"
|
||||
rg -n -B8 "void Server(Cheat|Exec)" Source/ | rg "#if|UE_BUILD_SHIPPING"
|
||||
```
|
||||
|
||||
An empty second result with a non-empty first is the finding. Note the
|
||||
distinction the recipe forces: whether the *implementation* is neutralised in a
|
||||
shipping configuration is a separate question that only a packaged build answers,
|
||||
and it should stay marked as open until someone answers it.
|
||||
|
||||
**Guardrail.** Wrap declaration *and* implementation in `#if !UE_BUILD_SHIPPING`
|
||||
or an equivalent cheat-manager guard. Hiding the UI is not removal.
|
||||
|
||||
---
|
||||
|
||||
## Replication shape
|
||||
|
||||
### NA-11 - Multicast used where state belongs
|
||||
|
||||
**Mechanism.** A one-shot multicast conveys a fact that a late joiner needs. It
|
||||
reaches everyone connected and relevant at the moment it fires.
|
||||
|
||||
**Why it is silent.** For everyone present, it works. The fact arrives, the visual
|
||||
plays, the state is right. The failure is scoped precisely to players who were not
|
||||
there, and those players see a plausible world — just the wrong one.
|
||||
|
||||
**Why the obvious check misses it.** Testing is done by players who join before
|
||||
the match starts, which is the configuration in which the bug cannot occur.
|
||||
Joining mid-match is a separate deliberate test that has to be designed.
|
||||
|
||||
**Symptom.** Players who join, or become relevant, after the event see the wrong
|
||||
state permanently. It never self-corrects, because nothing re-sends an event.
|
||||
|
||||
**Detect.** Enumerate multicasts and ask what each one carries:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*NetMulticast" -A3 Source/
|
||||
```
|
||||
|
||||
For each, answer: if a player joined one second after this fired, what would they
|
||||
see? Then run the test — join mid-round and inspect the state of every feature.
|
||||
This is a runtime check by nature; the grep only produces the list to test.
|
||||
|
||||
**Guardrail.** Anything a late joiner must know is replicated state. Multicasts
|
||||
carry transient presentation only — a sound, an impact, a one-frame effect whose
|
||||
absence is invisible a second later.
|
||||
|
||||
---
|
||||
|
||||
### NA-12 - Reliable RPC overflow
|
||||
|
||||
**Mechanism.** Frequent or per-frame reliable RPCs fill the reliable queue.
|
||||
|
||||
**Why it is silent.** The queue is generous. At two players on a local network it
|
||||
never fills. The cost scales with player count, latency and packet loss
|
||||
simultaneously, so it stays invisible across the entire range of configurations
|
||||
used during development.
|
||||
|
||||
**Why the obvious check misses it.** Each individual call site is defensible —
|
||||
one reliable RPC is nothing. The failure is a sum over call sites and time, and no
|
||||
review reads all of them together.
|
||||
|
||||
**Symptom.** Clients disconnect under load, typically first observed in a
|
||||
playtest at full player count, and usually attributed to the network rather than
|
||||
to a specific call site.
|
||||
|
||||
**Detect.** Count reliable RPCs and look for the ones a client can trigger
|
||||
repeatedly:
|
||||
|
||||
```bash
|
||||
rg -n "UFUNCTION\([^)]*\bReliable\b" Source/ | wc -l
|
||||
rg -n "UFUNCTION\([^)]*\bReliable\b" -A3 Source/ \
|
||||
| rg -i "tick|update|move|aim|look|input|cosmetic"
|
||||
```
|
||||
|
||||
The second command finds reliable calls on names that imply frequency. Confirm
|
||||
with a load test at maximum player count and 2% packet loss, watching for
|
||||
queue-overflow disconnects.
|
||||
|
||||
**Guardrail.** Reliable is a budget, not a default. Cosmetic, frequent or
|
||||
idempotent calls are unreliable; rate-limit anything a client can call in a loop.
|
||||
|
||||
---
|
||||
|
||||
### NA-13 - Incomplete FastArray callbacks
|
||||
|
||||
**Mechanism.** `PreReplicatedRemove`, `PostReplicatedAdd` and
|
||||
`PostReplicatedChange` are partially implemented. A derived cache or local
|
||||
listener list is updated in some of them and not others.
|
||||
|
||||
**Why it is silent.** The server is always right, because the server never runs
|
||||
the callbacks that are missing — it mutates the array directly. Only clients
|
||||
apply deltas, so only clients drift, and they drift into a state that is
|
||||
internally consistent.
|
||||
|
||||
**Why the obvious check misses it.** All three callbacks exist as overrides, so a
|
||||
structural review finds a complete-looking implementation. The empty one has a
|
||||
valid signature and compiles. Server-side tests pass because the defect cannot
|
||||
exist on the server.
|
||||
|
||||
**Symptom.** Client-side desync with no symptom on the server — the hardest class
|
||||
of bug to attribute, because every server-side log and dump shows correct state.
|
||||
|
||||
**Detect.** Compare the three callbacks per FastArray type:
|
||||
|
||||
```bash
|
||||
rg -n "FFastArraySerializer" -l Source/ | while read -r f; do
|
||||
echo "== $f"
|
||||
rg -n -A4 "Pre(ReplicatedRemove)|Post(ReplicatedAdd|ReplicatedChange)" "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Read the bodies. An empty body next to two implemented ones is the finding.
|
||||
On the measured reference one replicated structure had an empty removal hook and
|
||||
no callers at all — harmless while unused, a silent desync the moment it is
|
||||
adopted, which is the worst possible time to discover it.
|
||||
|
||||
**Guardrail.** Implement all three or none. Assert cache/array consistency in
|
||||
development builds so the drift becomes an error rather than a mystery.
|
||||
|
||||
---
|
||||
|
||||
### NA-14 - Relevancy disabled by an inherited constant
|
||||
|
||||
**Mechanism.** A net cull distance chosen for a small arena is inherited by a
|
||||
large map. Beyond the map's own size, the setting disables relevancy culling
|
||||
entirely.
|
||||
|
||||
**Why it is silent.** It is not a bug on the map it was chosen for — it is the
|
||||
correct setting there, and it makes gameplay better. The value carries forward
|
||||
because it is in a base class, and base-class defaults are not re-examined when a
|
||||
new map is added.
|
||||
|
||||
**Why the obvious check misses it.** The number is a squared distance, which
|
||||
makes it unreadable at a glance: nine hundred million does not announce itself as
|
||||
three hundred metres. Reviewers see a tuned-looking constant and move on.
|
||||
|
||||
**Symptom.** Bandwidth and server CPU scale with total actor count instead of
|
||||
local density. Discovered at player-count scale-up, late, when the fix requires
|
||||
re-verifying every gameplay assumption about who can see whom.
|
||||
|
||||
**Detect.** Find the value and convert it:
|
||||
|
||||
```bash
|
||||
rg -n "SetNetCullDistanceSquared|NetCullDistanceSquared\s*=" Source/ Config/
|
||||
```
|
||||
|
||||
Take the square root of each result and compare it to the largest map's diagonal.
|
||||
On the measured reference the value corresponded to a 300 m radius on arena maps —
|
||||
correct there, fatal if inherited into an open world.
|
||||
|
||||
**Guardrail.** Treat net cull distance as a per-class gameplay decision with a
|
||||
stated map-size assumption written next to it, not as an inherited performance
|
||||
knob.
|
||||
|
||||
---
|
||||
|
||||
### NA-15 - Replication graph enabled late
|
||||
|
||||
**Mechanism.** The graph implementation and its routing configuration exist, and
|
||||
the graph ships disabled. The routing is therefore never exercised.
|
||||
|
||||
**Why it is silent.** Disabled means the default replication path runs, and the
|
||||
default path is correct. The configuration file looks maintained. Nothing
|
||||
degrades, because nothing is using it.
|
||||
|
||||
**Why the obvious check misses it.** The presence of a configured routing table is
|
||||
read as evidence that the feature is in use. Confirming otherwise means finding
|
||||
the one boolean that disables the whole thing, which lives in a settings class
|
||||
rather than next to the routing.
|
||||
|
||||
**Symptom.** Enabling it changes which actors clients see, breaking gameplay code
|
||||
that had silently assumed universal relevancy. The breakage appears far from the
|
||||
change, in features nobody associated with replication.
|
||||
|
||||
**Detect.** Check the enable flag before trusting the routing:
|
||||
|
||||
```bash
|
||||
rg -n "bDisableReplicationGraph|bEnableReplicationGraph" Source/ Config/
|
||||
rg -n "ClassNodeMapping|RepGraph" Config/
|
||||
```
|
||||
|
||||
A disabled flag beside a populated routing table means the routing is untested
|
||||
configuration, not working configuration.
|
||||
|
||||
**Guardrail.** Enable it early or treat it as unavailable. Re-verify every
|
||||
relevancy assumption at the moment it is turned on, and treat that as a project
|
||||
milestone rather than a settings change.
|
||||
|
||||
---
|
||||
|
||||
### NA-16 - Wrong ability-system replication mode for AI
|
||||
|
||||
**Mechanism.** `Mixed` is inherited for AI-controlled actors. `Mixed` replicates
|
||||
gameplay effects to the owning client — and an AI has no owning client to
|
||||
receive them.
|
||||
|
||||
**Why it is silent.** It is correct for player-controlled actors, which is where
|
||||
it was chosen, and it is not wrong for AI either — merely wasteful. Waste has no
|
||||
symptom until bandwidth is the constraint.
|
||||
|
||||
**Why the obvious check misses it.** All three modes are one-line calls in
|
||||
constructors spread across different classes. There is no place where the choice
|
||||
is expressed as a policy, so nobody reviews the set of choices together.
|
||||
|
||||
**Symptom.** Bandwidth spent replicating gameplay effects nobody reads. Shows up
|
||||
as an unexplained baseline in profiling, attributed to player count.
|
||||
|
||||
**Detect.** List every mode choice in one place:
|
||||
|
||||
```bash
|
||||
rg -n "SetReplicationMode|EGameplayEffectReplicationMode::" Source/ \
|
||||
| sed 's/.*EGameplayEffectReplicationMode:://' | sort | uniq -c
|
||||
```
|
||||
|
||||
On the measured reference all three ability-system components used `Mixed`, with
|
||||
`Minimal` and `Full` absent from the module entirely — a default applied
|
||||
uniformly rather than a decision made three times.
|
||||
|
||||
**Guardrail.** `Mixed` for player-controlled, `Minimal` for AI, `Full` only with a
|
||||
written reason. Put the three choices in one comment block so they are reviewed
|
||||
as a policy.
|
||||
|
||||
---
|
||||
|
||||
## Lifecycle
|
||||
|
||||
### NA-17 - Readiness gate that assumes a controller
|
||||
|
||||
**Mechanism.** An initialisation chain waits for a controller. Simulated proxies
|
||||
on remote clients never receive one.
|
||||
|
||||
**Why it is silent.** The host is fine. The listen-server's own pawn has a
|
||||
controller, and so does every pawn in a single-process test. The stall happens
|
||||
only on machines that are not running the authority, which is the configuration
|
||||
least often exercised.
|
||||
|
||||
**Why the obvious check misses it.** The gate reads as obviously correct — of
|
||||
course initialisation waits for the controller. The missing knowledge is not in
|
||||
the code, it is that simulated proxies exist and never satisfy the condition. A
|
||||
reviewer who knows that spots it instantly; one who does not cannot find it by
|
||||
reading more carefully.
|
||||
|
||||
**Symptom.** Remote clients see uninitialised characters forever. Reproduces only
|
||||
with a dedicated server, so it is often first reported by QA and not reproducible
|
||||
by the developer.
|
||||
|
||||
**Detect.** Find readiness conditions that mention a controller:
|
||||
|
||||
```bash
|
||||
rg -n -B4 -A4 "GetController\(\)|Controller\s*!=\s*nullptr|IsPlayerControlled" \
|
||||
Source/ --glob "*Init*" --glob "*Extension*" --glob "*Ready*"
|
||||
```
|
||||
|
||||
For each, ask what a simulated proxy does at that line. The measured reference
|
||||
handles this correctly — its init-state gate permits progression without a
|
||||
controller — and is worth reading precisely because the naive version of the same
|
||||
gate deadlocks every simulated proxy while working perfectly in PIE.
|
||||
|
||||
**Guardrail.** An explicit named state chain, with the no-controller case handled
|
||||
deliberately rather than by omission. Write down which states a simulated proxy
|
||||
reaches.
|
||||
|
||||
---
|
||||
|
||||
### NA-18 - Addition without verified removal
|
||||
|
||||
**Mechanism.** Granted abilities, applied effects, registered listeners and
|
||||
spawned cues accumulate across death/respawn or feature activation cycles,
|
||||
because the addition path was built and the removal path was assumed.
|
||||
|
||||
**Why it is silent.** One cycle is fine. Two cycles produce a duplicate, and a
|
||||
duplicate of an idempotent-looking effect is usually invisible. The cost grows
|
||||
linearly with session length, and development sessions are short.
|
||||
|
||||
**Why the obvious check misses it.** Both halves exist in most cases — there is a
|
||||
removal function, it is called, and it removes *something*. Whether it removes
|
||||
exactly what was added requires holding the handle, and a review that sees
|
||||
`Remove` next to `Add` reasonably concludes the pair is complete.
|
||||
|
||||
**Symptom.** Duplicated effects after the second respawn; leaks that only appear
|
||||
in long sessions; a second feature activation that adds a second copy of every
|
||||
widget.
|
||||
|
||||
**Detect.** Compare addition and removal sites per resource type, then test:
|
||||
|
||||
```bash
|
||||
rg -n "GiveAbility|ApplyGameplayEffect|AddLoadedGameplayCuePath|RegisterListener" \
|
||||
Source/ | wc -l
|
||||
rg -n "ClearAbility|RemoveActiveGameplayEffect|RemoveLoadedGameplayCuePath|Unregister" \
|
||||
Source/ | wc -l
|
||||
```
|
||||
|
||||
Unequal counts are a starting point, not a verdict. The verdict is runtime:
|
||||
baseline, activate, deactivate, compare to baseline, reactivate — on both server
|
||||
and client, twice in one session.
|
||||
|
||||
**Guardrail.** Every dynamic addition retains an ownership handle and an exact
|
||||
rollback, tested on both server and client, twice in one session. Assert on the
|
||||
removal count where the API returns one.
|
||||
|
||||
---
|
||||
|
||||
### NA-19 - Prediction without a correction path
|
||||
|
||||
**Mechanism.** A client-predicted value has no defined behaviour when the server
|
||||
disagrees. The prediction key reconciles ordering; nothing reconciles the value.
|
||||
|
||||
**Why it is silent.** The server rarely disagrees on a local network with an
|
||||
honest client. Disagreement requires latency, packet loss or a rejected action,
|
||||
and the development environment supplies none of them.
|
||||
|
||||
**Why the obvious check misses it.** The prediction machinery is present and
|
||||
correct, and prediction keys are a real feature that really works. Reviewing
|
||||
"is this predicted properly?" finds a correct implementation of prediction. The
|
||||
missing piece is a design question — what does the player see when we were
|
||||
wrong? — that has no code location to be missing from.
|
||||
|
||||
**Symptom.** The client keeps a wrong value indefinitely. UI and gameplay diverge
|
||||
with no error anywhere, and the player's experience is that the game lied to them
|
||||
once and never corrected.
|
||||
|
||||
**Detect.** Force disagreement and watch:
|
||||
|
||||
```bash
|
||||
rg -n "ScopedPredictionWindow|PredictionKey|GetPredictionKeyForNewAction" Source/
|
||||
```
|
||||
|
||||
For each predicted value, run the paid test: introduce 150 ms latency, make the
|
||||
server reject the action, and observe the client for ten seconds. A value that
|
||||
does not converge has no correction path. This one cannot be settled statically.
|
||||
|
||||
**Guardrail.** Define, per predicted value, what the correction looks like to the
|
||||
player. A prediction key reconciles ordering, not entitlement.
|
||||
|
||||
---
|
||||
|
||||
### NA-20 - PIE treated as a network test
|
||||
|
||||
**Mechanism.** The listen-server host shares memory with its client, so the
|
||||
serialisation, ordering and relevancy boundaries that produce networked defects do
|
||||
not exist during the test.
|
||||
|
||||
**Why it is silent.** PIE multiplayer looks like multiplayer. Two windows, two
|
||||
players, two views of one world. Everything that works there works convincingly,
|
||||
including the things that only work because the boundary is absent.
|
||||
|
||||
**Why the obvious check misses it.** The check *is* the test, and the test passes.
|
||||
There is no artefact to inspect and no code to review; the defect is in the
|
||||
choice of configuration, which is made once, early, by whoever set up the
|
||||
workflow, and never revisited.
|
||||
|
||||
**Symptom.** Features ship having never run in the configuration they will run in.
|
||||
Entries NA-01, NA-04, NA-11, NA-13, NA-17 and NA-19 above cannot occur in PIE at
|
||||
all.
|
||||
|
||||
**Detect.** Audit the test configuration rather than the code:
|
||||
|
||||
```bash
|
||||
rg -n "PlayNetMode|RunUnderOneProcess|bLaunchSeparateServer" Config/ Saved/Config/
|
||||
```
|
||||
|
||||
If the editor is configured for a listen server in one process, no networked
|
||||
feature in the project has been tested. Confirm by running one known-networked
|
||||
feature under a dedicated server with a remote client and comparing behaviour.
|
||||
|
||||
**Guardrail.** Dedicated server plus two remote clients, one with 150 ms latency
|
||||
and 2% packet loss, as the minimum configuration for accepting a networked
|
||||
feature — including a mid-match joiner.
|
||||
@@ -0,0 +1,178 @@
|
||||
# Patterns: authority in a measured reference product
|
||||
|
||||
A worked example of the authority model from `SKILL.md`, against one specific
|
||||
reference project read as source. It is here for one reason: the failure modes in
|
||||
[failure-modes.md](failure-modes.md) are easier to believe, and much easier to
|
||||
apply, when you can see how they cluster in a real codebase.
|
||||
|
||||
**Read this first.** The project audited below is an outstanding architectural
|
||||
reference and a poor network-security reference. Those are independent axes. None
|
||||
of what follows is a criticism of the sample — a learning project that ships no
|
||||
product has no reason to build lag compensation — but copying its structure
|
||||
without noticing its trust boundaries is a real and common accident, and it is the
|
||||
accident this skill exists to prevent.
|
||||
|
||||
Markers: **[measured]** — read in source; **[derived]** — conclusion from measured
|
||||
facts; **[open]** — not answerable without a dedicated-server build, and left
|
||||
open rather than guessed.
|
||||
|
||||
---
|
||||
|
||||
## 1. The headline: one absence explains almost everything
|
||||
|
||||
No custom saved-move type, no custom server-move path, no compressed-flags
|
||||
extension, and no historical position storage anywhere in the project
|
||||
**[measured]**.
|
||||
|
||||
Without stored per-tick positions the server has nothing to rewind to, and
|
||||
therefore **cannot** validate a client's hit claim even in principle
|
||||
**[derived]**. The client trust documented below is not carelessness at individual
|
||||
call sites; it is the necessary consequence of a subsystem that was never built.
|
||||
|
||||
This is the single most valuable thing to carry away from the audit, and it
|
||||
generalises: **before reading a weapon system, ask whether the machinery that
|
||||
would make validation possible exists at all.** If it does not, every "missing
|
||||
check" you find downstream is one symptom of one decision, not a list of
|
||||
independent bugs. Recipe: NA-08.
|
||||
|
||||
Practical implication for anyone adopting such code: "add server validation to the
|
||||
weapon system" is not a patch. It is building the rewind subsystem first.
|
||||
|
||||
---
|
||||
|
||||
## 2. Hit registration was client-authored end to end
|
||||
|
||||
Four findings, each of which reads as an isolated defect and is not:
|
||||
|
||||
| What was measured | Consequence |
|
||||
|---|---|
|
||||
| The targeting trace early-returns unless locally controlled — the server never traces **[measured]** | The client decides what was hit |
|
||||
| A target-data validity flag is a `const bool` initialised to `true`, consumed twice, with no computation between **[measured]** | The code path reads as validated in review; NA-06 |
|
||||
| Distance falloff is computed from a client-supplied trace origin; the damage multiplier is taken from a client-supplied surface material **[measured]** | A modified client controls falloff and multiplier directly **[derived]**; NA-05 |
|
||||
| A `bIsSimulated` parameter appears at ten sites, is passed a literal at its single call site, and is branched on nowhere; a hit-replacement helper is never called **[measured]** | Attachment points for the rewind subsystem that was never built **[derived]**; NA-07 |
|
||||
|
||||
One check in that path was genuinely server-side: the team comparison guarding
|
||||
whether damage may be caused at all **[measured]**. That detail matters more than
|
||||
it looks. **The model was incomplete, not absent** — which is the harder case,
|
||||
because the presence of one real server-side check makes the surrounding code read
|
||||
as server-side too.
|
||||
|
||||
---
|
||||
|
||||
## 3. Three mechanisms that look equivalent and are not
|
||||
|
||||
One equipment component demonstrated all three in a single file **[measured]**:
|
||||
|
||||
| Declaration | Actually enforces |
|
||||
|---|---|
|
||||
| `UFUNCTION(Server, Reliable, BlueprintCallable)` | routing only — no `WithValidation` |
|
||||
| `UFUNCTION(BlueprintAuthorityOnly, ...)` | blocks Blueprint callers; no effect on C++ |
|
||||
| plain C++ writes elsewhere in the same class | nothing |
|
||||
|
||||
`BlueprintAuthorityOnly` is an editor affordance. It appears in review as an
|
||||
authority guard and is not one **[derived]**. Across the whole project the
|
||||
specifier appeared at five sites in three subsystems, none of which carried an
|
||||
authority check in the guarded body **[measured]**. Recipe: NA-03.
|
||||
|
||||
Unguarded state transitions followed the same shape: the death-state writers were
|
||||
public, virtual, and unguarded, with the only barrier being a server-initiated
|
||||
activation policy on one ability **[measured]**. The guard belongs in the mutator,
|
||||
not in one of its callers **[derived]** — the call-site list is not a stable
|
||||
property, and virtual functions mean a subclass in a feature plugin inherits the
|
||||
hole. Recipe: NA-04.
|
||||
|
||||
---
|
||||
|
||||
## 4. Cheat surface and error handling
|
||||
|
||||
Two arbitrary-string cheat RPCs were declared outside any build-configuration
|
||||
guard **[measured]**. Their validators return `true` unconditionally
|
||||
**[measured]** — which is defensible for a cheat RPC and is also exactly the shape
|
||||
NA-02 looks for, so the recipe finds them first and you must read what you found.
|
||||
Whether the implementations are neutralised in a shipping configuration is
|
||||
**[open]** without a packaged build; the RPC surface itself is unconditional.
|
||||
|
||||
A team-comparison helper returned a permissive result on an invalid-argument
|
||||
outcome, carrying a comment marking it temporary **[measured]**. Since that
|
||||
function gates friendly fire, the failure mode is damage applied where team rules
|
||||
should have prevented it **[derived]**.
|
||||
|
||||
Rule extracted, and it is worth more than the finding: **failure to determine must
|
||||
deny.** Recipe: NA-09.
|
||||
|
||||
---
|
||||
|
||||
## 5. Replication configuration
|
||||
|
||||
| Setting | Measured | Reading |
|
||||
|---|---|---|
|
||||
| Ability-system replication mode | `Mixed` on all three ability-system components; `Minimal` and `Full` absent from the module **[measured]** | Correct for player-controlled actors; wasteful when inherited by AI **[derived]**; NA-16 |
|
||||
| Net cull distance | A squared value corresponding to a 300 m radius, set in the character base class **[measured]** | Every character relevant to every client at all times **[derived]** — correct for a small arena, fatal if inherited into a large map; NA-14 |
|
||||
| Replication graph | Ships disabled, with class routing configured **[measured]** | The routing is untested configuration, not working configuration **[derived]**; NA-15 |
|
||||
| A replicated fast-array | Empty removal callback, no callers in the project **[measured]** | Harmless while unused; a silent client-side desync the moment it is adopted **[derived]**; NA-13 |
|
||||
|
||||
The last row is the one worth pausing on. An unused structure with an incomplete
|
||||
callback set is invisible to every kind of testing, because nothing exercises it.
|
||||
It becomes a defect at adoption time — the moment when the person adopting it has
|
||||
the least context to recognise what they inherited.
|
||||
|
||||
---
|
||||
|
||||
## 6. What the reference got right
|
||||
|
||||
For balance, and because these are the parts worth copying directly:
|
||||
|
||||
1. **An explicit init-state chain** instead of timing assumptions. Its gate
|
||||
permits progression for actors without a controller, which is what allows
|
||||
simulated proxies to initialise at all **[measured]**. The naive version of the
|
||||
same gate deadlocks every simulated proxy on remote clients while working
|
||||
perfectly in PIE **[derived]** — study the correct one before writing your own.
|
||||
Context: NA-17.
|
||||
2. **A replication mode chosen deliberately** for player ability-system components
|
||||
rather than left at the engine default **[measured]**.
|
||||
3. **Replicate intent, realise presentation locally** — the cosmetic system sends
|
||||
which parts to wear, not the assembled result **[measured]**.
|
||||
4. **Correct teardown where it exists**: one pawn-extension path checks avatar
|
||||
identity before cancelling abilities, clearing input and removing cues
|
||||
**[measured]**. Note the shape — identity check first, then reverse-order
|
||||
removal.
|
||||
5. **The team check in the damage execution is genuinely server-side**
|
||||
**[measured]**. The model is not absent, it is incomplete.
|
||||
|
||||
---
|
||||
|
||||
## 7. Questions the audit could not answer
|
||||
|
||||
Left open deliberately. Each requires a dedicated-server or packaged build, which
|
||||
was out of scope:
|
||||
|
||||
1. Whether the cheat RPC implementations are neutralised in a shipping
|
||||
configuration.
|
||||
2. Actual bandwidth per client, and whether the relevancy radius holds at full
|
||||
player count.
|
||||
3. Behaviour of the replication graph when enabled — routing completeness,
|
||||
dormancy usage.
|
||||
4. Whether client and server tag registries match when feature plugins load
|
||||
asymmetrically.
|
||||
5. Real reconciliation behaviour of ability prediction under packet loss.
|
||||
|
||||
An audit that answers questions it cannot answer is worth less than one that
|
||||
enumerates them. If you adopt this material, these five are yours to close.
|
||||
|
||||
---
|
||||
|
||||
## Provenance
|
||||
|
||||
Measured against Epic's Lyra Starter Game on Unreal Engine 5.6, read as source in
|
||||
a single workspace. Source addresses stay in the research archive that produced
|
||||
this skill; each `NA-` identifier above resolves back to the audited location
|
||||
there. Numbers without an author are worth less than no numbers, so: this is where
|
||||
these came from, and they can be produced on request.
|
||||
|
||||
## Evidence boundary
|
||||
|
||||
Every claim above refers to one project on one engine version. These are examples
|
||||
and failure evidence, not guarantees about other engine versions or other samples.
|
||||
Verify version-sensitive claims against your target engine before acting on them.
|
||||
In particular: PIE cannot reproduce most of the failure modes described here,
|
||||
because a listen-server host shares memory with its client.
|
||||
Reference in New Issue
Block a user