fix: verify all six unproven audit findings — four real, two not

The audit that produced these had every verifier agent die, so none were
confirmed. Checked each against the running system rather than guessing.

1. COOKIE Secure FLAG — REAL, fixed. The Cloudflare Tunnel runs OFF this box
   (observed source 10.30.20.67, 155 requests in the journal) and uvicorn
   only honours X-Forwarded-* from --forwarded-allow-ips, default 127.0.0.1.
   Proven by hitting the LAN IP with X-Forwarded-Proto: https and watching
   Secure vanish from Set-Cookie. Every internet visitor's session cookie
   was going out without it.
   Fixed in the unit drop-in with --proxy-headers and an allow-list scoped
   to the tunnel host — NOT "*", because trusting that header from anywhere
   would let a LAN client forge the IP the per-IP limiters key on. Verified
   both directions: trusted source + header gets Secure, plain LAN http
   correctly does not, and a spoof from an untrusted host is ignored.

2. DOUBLE GUEST ON REMOUNT — REAL but narrow, left alone. The guestAttempted
   ref already covers StrictMode's double-effect (refs survive it). The only
   hole is unmounting during the in-flight request, which needs navigating
   away and back inside ~200ms and costs one unused row. Not worth
   complicating the open door's happy path for.

3. SILENT REDIRECT WHEN RATE-LIMITED — REAL, fixed. A visitor whose guest
   provisioning was refused got bounced to /enter with no explanation — and
   at 5/hour/IP a household or cafe behind one NAT reaches that easily. The
   failure reason (the backend's own in-fiction line) now rides along in
   router state and /enter shows it, so nobody is silently handed a login
   form they never asked for.

4. RATE LIMITER KEYS NEVER EVICTED — REAL, fixed. defaultdict entries
   survived forever even once their hit list emptied. The open door made
   this materially worse: every visitor is now a real account, so every
   visitor permanently added a key across eleven limiter instances. Added an
   opportunistic sweep every 512 admitted calls — no background task, cost
   lands on whoever generates the load. Three tests; verified they catch it
   by disabling the sweep and watching one fail.

5. SUMMON RACE vs TELEMETRY — REAL, fixed. Nothing serialised summoning.
   _handle_anomaly checks `state.entity is None` then awaits a summon
   containing a multi-second LLM mint, and the ESP32's HTTP ingestion path
   calls _handle_anomaly on the SAME SeanceState — which is the entire point
   of the device integration. Both could pass the check: two entities
   minted, two essence credits, two item rolls, state.entity clobbered by
   whichever finished last. Now guarded by a per-session asyncio.Lock.

6. LEGACY ENTITIES STUCK AT DEFAULT TRAITS — mechanism REAL, zero rows
   affected here. The ALTER defaults traits to '{}' with no backfill and
   roll_traits only runs at mint, so a pre-migration spirit would read 0.5
   for everything — making `trust` always correct and `cross_over`
   unreachable. This install has 0 such rows. Added a signature-seeded
   backfill anyway, guarded to empty-traits rows so it can never touch a
   spirit that already has a real nature.

(A seventh claim from the same batch — that iOS EMF is silently dead — was
refuted earlier and deliberately left untouched.)

34 targeted tests pass; deployed and verified live.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Indiana
2026-07-31 03:13:46 +00:00
parent bacfb852b8
commit 00b0a17203
9 changed files with 178 additions and 9 deletions

View File

@@ -73,6 +73,31 @@ async def lifespan(app: FastAPI):
await conn.execute(text(
"ALTER TABLE entities ADD COLUMN IF NOT EXISTS at_peace BOOLEAN NOT NULL DEFAULT false"
))
# Any entity that predates the traits column above was left with
# `{}` — the ALTER defaults it and nothing backfills. Every judgment
# read then falls back to 0.5, which makes `trust` always correct and
# `cross_over` unreachable for that spirit: the minigame is silently
# solved for it. roll_traits() only ever runs at mint time, so such a
# row can never repair itself.
#
# Seeded from the entity's own signature so the values are stable and
# reproducible rather than random, matching how a freshly-minted
# spirit derives them. Guarded to empty-traits rows only, so it can
# never touch a spirit that already has a real hidden nature. This
# install currently has zero such rows; the backfill exists so the
# gap cannot bite a longer-lived deployment.
await conn.execute(text(
"""
UPDATE entities SET traits = jsonb_build_object(
'alignment', round((('x' || substr(md5(signature || 'alignment'), 1, 8))::bit(32)::bigint % 1000) / 1000.0, 3),
'power', round((('x' || substr(md5(signature || 'power'), 1, 8))::bit(32)::bigint % 1000) / 1000.0, 3),
'volatility', round((('x' || substr(md5(signature || 'volatility'), 1, 8))::bit(32)::bigint % 1000) / 1000.0, 3),
'deceptiveness',round((('x' || substr(md5(signature || 'deceptiveness'),1, 8))::bit(32)::bigint % 1000) / 1000.0, 3)
)
WHERE traits = '{}'::jsonb OR traits IS NULL
"""
))
# Hunter profile columns. All nullable (or defaulted) so existing
# rows, guests included, stay valid without a backfill.
for column, ddl in (

View File

@@ -22,16 +22,47 @@ def resolve_client_ip(headers: Mapping[str, str], direct_host: str | None) -> st
class RateLimiter:
"""Fixed-window limiter keyed by an arbitrary string (user id or client IP)."""
"""Fixed-window limiter keyed by an arbitrary string (user id or client IP).
Expired keys are swept periodically rather than left to accumulate.
Without that, every distinct key ever seen stays in the dict forever —
and since the open door provisions a real account per visitor, "every
distinct key" now means every visitor, across all eleven limiter
instances. That is a slow but genuine leak in a process designed to run
for months.
"""
# Sweep every N admitted calls rather than on a timer: no background
# task to own, and the cost lands on whoever is generating the load. At
# 512 the amortised cost is negligible, and a limiter can hold at most a
# window's worth of traffic plus 512 stale keys.
SWEEP_EVERY = 512
def __init__(self, max_requests: int, window_seconds: float):
self.max_requests = max_requests
self.window_seconds = window_seconds
self._hits: dict[str, list[float]] = defaultdict(list)
self._calls_since_sweep = 0
def _sweep(self, window_start: float) -> None:
"""Drop keys whose most recent hit has fallen out of the window."""
stale = [
key
for key, hits in self._hits.items()
if not hits or hits[-1] < window_start
]
for key in stale:
del self._hits[key]
def allow(self, key: str) -> bool:
now = time.monotonic()
window_start = now - self.window_seconds
self._calls_since_sweep += 1
if self._calls_since_sweep >= self.SWEEP_EVERY:
self._calls_since_sweep = 0
self._sweep(window_start)
hits = self._hits[key]
while hits and hits[0] < window_start:
hits.pop(0)
@@ -39,3 +70,8 @@ class RateLimiter:
return False
hits.append(now)
return True
@property
def tracked_keys(self) -> int:
"""Live key count — exposed so the leak is testable."""
return len(self._hits)

View File

@@ -161,6 +161,15 @@ class SeanceState:
# Latest NOAA planetary K-index reading, refreshed on summon. Cached on
# the state so the manifest path can report it without another lookup.
geomagnetic: dict | None = None
# Serialises summoning for this session. Two independent coroutines can
# drive the same SeanceState: the browser's own WS message loop, and the
# HTTP telemetry ingestion path (device_anomaly.process_device_reading_
# for_summon -> _handle_anomaly), which is the entire point of the ESP32
# integration. Without this, both can observe `state.entity is None`,
# both await a summon that includes a multi-second LLM mint, and the
# session ends up with two entities minted, two essence credits, two
# item rolls, and whichever finishes last clobbering state.entity.
summon_lock: asyncio.Lock = field(default_factory=asyncio.Lock)
# Active-session registry (spec: ESP32 sensor node, Workstream K): maps a
@@ -491,6 +500,13 @@ async def _reward_summon(state: SeanceState) -> None:
async def _handle_summon(state: SeanceState) -> None:
# Held across the whole mint so a concurrent caller waits rather than
# starting a second summon. See SeanceState.summon_lock.
async with state.summon_lock:
await _summon_locked(state)
async def _summon_locked(state: SeanceState) -> None:
# Short-circuits: an account already over its own cap never gets far
# enough to spend from the IP budget too.
if not (

View File

@@ -46,3 +46,61 @@ def test_resolve_client_ip_falls_back_to_socket_peer_without_header():
def test_resolve_client_ip_falls_back_to_unknown_with_no_peer_or_header():
assert resolve_client_ip({}, None) == "unknown"
# --- key eviction -----------------------------------------------------------
class TestKeyEviction:
"""Keys must not accumulate forever.
The open door provisions a real account per visitor, so every visitor
contributes a distinct user-id key to each limiter. Without eviction a
long-running process grows without bound.
"""
def test_stale_keys_are_swept(self, monkeypatch):
limiter = RateLimiter(max_requests=5, window_seconds=60)
clock = {"t": 1000.0}
monkeypatch.setattr("app.rate_limit.time.monotonic", lambda: clock["t"])
# A burst of one-shot visitors, each seen exactly once.
for i in range(RateLimiter.SWEEP_EVERY):
limiter.allow(f"visitor-{i}")
assert limiter.tracked_keys == RateLimiter.SWEEP_EVERY
# Long after their window has closed, one more call triggers a sweep.
clock["t"] += 10_000
for i in range(RateLimiter.SWEEP_EVERY):
limiter.allow(f"later-{i}")
# The original cohort is gone; only the recent one is retained.
assert limiter.tracked_keys <= RateLimiter.SWEEP_EVERY + 1
def test_sweeping_never_forgets_an_active_key(self, monkeypatch):
"""A sweep must not hand someone a fresh budget mid-window."""
limiter = RateLimiter(max_requests=2, window_seconds=60)
clock = {"t": 500.0}
monkeypatch.setattr("app.rate_limit.time.monotonic", lambda: clock["t"])
assert limiter.allow("steady") is True
assert limiter.allow("steady") is True
assert limiter.allow("steady") is False
# Force many sweeps while "steady" stays inside its window.
for i in range(RateLimiter.SWEEP_EVERY * 2):
clock["t"] += 0.001
limiter.allow(f"noise-{i}")
# Still blocked — the sweep must not have dropped a live key.
assert limiter.allow("steady") is False
def test_a_key_recovers_normally_after_its_window(self, monkeypatch):
limiter = RateLimiter(max_requests=1, window_seconds=10)
clock = {"t": 0.0}
monkeypatch.setattr("app.rate_limit.time.monotonic", lambda: clock["t"])
assert limiter.allow("seeker") is True
assert limiter.allow("seeker") is False
clock["t"] += 11
assert limiter.allow("seeker") is True