Files
qtalker---/backend/tests/test_ws_passage.py
Indiana c22b2f9c08 fix: reward guards now survive a reconnect
The three replay guards shipped earlier lived on the in-memory SeanceState,
which made them per-CONNECTION. I flagged that as an open residual at the
time: drop the socket and reconnect — or just open a second one — and the
client got a fresh empty guard and could be paid again for the same spirit.
summon_limiter bounded the rate of that, never the total.

`award_claims` is the durable form: one row per (seeker, presence,
milestone), with a UNIQUE constraint doing the actual enforcement. The claim
is a bare INSERT and losing the race raises IntegrityError, which is caught
and read as "already paid" — a check-then-insert would let two sockets both
read "unclaimed" and both pay. `crossing` is claimed by BOTH roads, so a
spirit crosses once whichever road arrives first.

Measured with the durable claim disabled: 5 reconnects paid 75 extra essence
on the ritual, 60 on a verdict, 140 on the passage, and two simultaneous
sockets paid 30 for one 15-essence ritual.

The four tests that were failing were the tests, not the guard. They compared
raw balances across reconnects, but re-opening a channel IS a summon, and
SUMMON_ESSENCE_TRICKLE is paid per summon by design (inventory.py:42, bounded
by summon_limiter rather than by any once-per-presence rule). The expected
trickle is now stated explicitly so the assertion speaks about the milestone
it is actually testing. Favor has no trickle, so it must not move at all —
asserted separately.

Anti-overshoot covered in both directions: a genuinely fresh presence still
pays in full across a reconnect, a corrected verdict still pays on a second
connection, and `test` stays freely repeatable since it never touches the
ledger.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 09:37:32 +00:00

328 lines
12 KiB
Python

"""WS integration tests for the Passage handlers in app.ws.
Why this file exists: `tests/test_passage.py` covers app/passage.py, which is
pure — it decides outcomes and returns them. It cannot see the handler that
banks the essence those outcomes describe, and that handler is where an
unbounded essence faucet lived undetected through a 400-test suite: nothing
exercised `passage_start` / `passage_layer` over a real connection.
The faucet: `passage_start` rewinds the rite to `listen` at any time, and the
UI offers exactly that button after a resisted release (PassagePanel.tsx's
"walk it again"). Every replayed layer re-credited its essence, so one summon
funded an endless loop. The rate limiter caps how FAST that runs, never how
much it totals.
These tests are written against the exploit path, not against the
implementation: they drive the same frames a browser sends and then read the
ledger straight out of the database.
"""
import uuid
import pytest
import app.ws
from app.entities import fallback_profile
from app.inventory import SUMMON_ESSENCE_TRICKLE
from app.models.user import User
from app.passage import LAYERS, LAYER_ESSENCE
from app.rate_limit import RateLimiter
class FakeSpiritService:
async def mint_profile(
self, signature, channel, anomalies, language="en", entropy=None, sky=None
):
return fallback_profile(signature)
async def fragment(self, source, anomaly, language="en"):
return "listen"
async def wire_whisper(self, telemetry, language="en"):
return "the wire hums"
def chat_stream(self, entity, question, history, language="en"):
async def gen():
yield "here."
return gen()
def ambient_ready(self):
return False
async def _fake_synth(text, voice, profile, instability=0.0):
return b"RIFFfake wav bytes"
@pytest.fixture(autouse=True)
def _fake_spirits(monkeypatch):
monkeypatch.setattr(app.ws, "spirit_service", FakeSpiritService())
monkeypatch.setattr(app.ws, "synthesize_spirit_voice", _fake_synth)
# Module-level limiters are shared singletons that accumulate real hit
# counts across the whole session (precedent: test_ws_ritual_judgment.py).
# The passage limiter matters most here — the exploit these tests replay
# is deliberately high-volume, and the real 20/60s cap would mask it
# behind a rate limit rather than letting us observe the actual total.
for name in (
"summon_limiter",
"summon_ip_limiter",
"question_limiter",
"question_ip_limiter",
"fragment_limiter",
"fragment_ip_limiter",
"ritual_limiter",
"ritual_ip_limiter",
"judgment_limiter",
"judgment_ip_limiter",
"passage_limiter",
"passage_ip_limiter",
):
monkeypatch.setattr(app.ws, name, RateLimiter(max_requests=10_000, window_seconds=60))
def _no_twists(monkeypatch):
"""Pin both Passage draws to "no twist", leaving every other draw real.
A PassageDraw of 1.0 never fires, since both twists trigger on
`draw < chance`. Only the `passage:` contexts are pinned — blanketing
`veil_float` would also steer the summon and manifest draws, and these
tests are about the ledger, not about the room's noise.
"""
real = app.ws.veil_float
def selective(entropy, context, *args, **kwargs):
if context.startswith("passage:"):
return 1.0
return real(entropy, context, *args, **kwargs)
monkeypatch.setattr(app.ws, "veil_float", selective)
def _force_new_presence(monkeypatch):
"""Pin the summon draw so the NEXT summon mints a genuinely new spirit.
`_summon` draws against RETURN_CHANCE to decide whether the presence
already on this channel answers again, and re-calling the same channel
usually returns the SAME entity row — measured on this suite, 8 of 11
consecutive re-summons came back with `is_new: false` and an identical id.
That matters now that the purse is keyed on the entity (award_claims),
because "summon again" and "the same spirit answers again" are then two
different things. This test is about what a genuinely FRESH presence is
worth, so the draw is pinned rather than left to a coin flip; a draw of
1.0 is >= any possible `return_chance`, so the familiar presence never
answers. Layered over `_no_twists` (call it after), leaving every other
draw real.
"""
real = app.ws.veil_float
def selective(entropy, context, *args, **kwargs):
if context == "answers":
return 1.0
return real(entropy, context, *args, **kwargs)
monkeypatch.setattr(app.ws, "veil_float", selective)
# The exploit never reaches `release`. Crossing latches `passage_crossed`,
# which correctly blocks replay — the faucet ran on the four beats BEFORE
# release, rewound by the "walk it again" button after a resisted release.
BEATS_BEFORE_RELEASE = len(LAYERS) - 1
def _read_until(ws, msg_type, max_frames=80, **match):
for _ in range(max_frames):
frame = ws.receive_json()
if frame.get("type") != msg_type:
continue
if all(frame.get(key) == value for key, value in match.items()):
return frame
raise AssertionError(f"never saw frame of type {msg_type!r} matching {match!r}")
def _login(sync_client, username):
sync_client.post("/auth/register", json={"username": username, "password": "spookyspooky"})
sync_client.post("/auth/login", json={"username": username, "password": "spookyspooky"})
return sync_client.cookies.get("qm_session")
def _ws_connect(sync_client, token):
return sync_client.websocket_connect(
"/ws/session", headers={"cookie": f"qm_session={token}"}
)
def _summon(ws):
ws.send_json({"type": "summon"})
frame = _read_until(ws, "entity")
_read_until(ws, "utterance", kind="greeting")
return frame
def _settle(ws):
"""Force a full round trip so any post-frame DB write has landed.
The essence credit runs before the frame is queued, but the sender task
drains that queue concurrently with the handler's remaining awaits. The
connection's message loop is sequential, so a pong is a hard guarantee
that the previous handler returned — not a poll-and-hope.
"""
ws.send_json({"type": "ping"})
_read_until(ws, "pong")
def _walk(ws, beats):
"""Send `beats` passage_layer frames, returning the result frames."""
frames = []
for _ in range(beats):
ws.send_json({"type": "passage_layer"})
frames.append(_read_until(ws, "passage_result"))
_settle(ws)
return frames
async def _essence(db_session, user_id):
db_session.expire_all()
user = await db_session.get(User, user_id)
return user.essence
def _user_id(sync_client, token):
return uuid.UUID(
sync_client.get("/auth/me", headers={"cookie": f"qm_session={token}"}).json()["id"]
)
# --- the faucet ------------------------------------------------------------
@pytest.mark.asyncio
async def test_replaying_the_rite_never_mints_essence_twice(
sync_client, db_session, monkeypatch
):
"""The exploit, executed literally.
Walk every layer, hit `passage_start` to rewind, walk them all again —
ten times over. The seeker must finish with exactly one rite's worth of
essence, because there was only ever one spirit.
Before the fix this credited the full ladder on every pass; with the
ladder summing to 53, ten passes paid 530 instead of 53.
"""
# No collapses and no crossing: this test is about the ledger, and a
# random rewind or an early `crossed` would make the totals depend on
# the draw rather than on the replay guard.
_no_twists(monkeypatch)
token = _login(sync_client, "faucet-walker")
user_id = _user_id(sync_client, token)
with _ws_connect(sync_client, token) as ws:
_read_until(ws, "session")
_summon(ws)
baseline = await _essence(db_session, user_id)
ws.send_json({"type": "passage_start"})
_settle(ws)
_walk(ws, BEATS_BEFORE_RELEASE)
after_one_rite = await _essence(db_session, user_id)
for _ in range(10):
ws.send_json({"type": "passage_start"})
_settle(ws)
_walk(ws, BEATS_BEFORE_RELEASE)
after_ten_replays = await _essence(db_session, user_id)
earned = after_one_rite - baseline
assert earned > 0, "the rite paid nothing at all — the test proves nothing"
assert earned <= sum(LAYER_ESSENCE.values())
assert after_ten_replays == after_one_rite, (
f"replaying the rite minted {after_ten_replays - after_one_rite} extra essence; "
"the passage is an unbounded faucet again"
)
@pytest.mark.asyncio
async def test_the_reported_essence_matches_what_was_actually_credited(
sync_client, db_session, monkeypatch
):
"""The UI sums the frame's `essence` field into its running total.
If a replayed layer reports the layer's face value while the ledger
credits nothing, the seeker watches a total climb that their account
never receives — the most corrosive kind of bug in a game about trust.
"""
_no_twists(monkeypatch)
token = _login(sync_client, "honest-ledger")
user_id = _user_id(sync_client, token)
with _ws_connect(sync_client, token) as ws:
_read_until(ws, "session")
_summon(ws)
baseline = await _essence(db_session, user_id)
ws.send_json({"type": "passage_start"})
_settle(ws)
first = _walk(ws, BEATS_BEFORE_RELEASE)
ws.send_json({"type": "passage_start"})
_settle(ws)
second = _walk(ws, BEATS_BEFORE_RELEASE)
final = await _essence(db_session, user_id)
assert sum(f["essence"] for f in first) == final - baseline
assert all(f["essence"] == 0 for f in second), (
"a replayed layer advertised essence it did not pay"
)
@pytest.mark.asyncio
async def test_a_genuinely_new_presence_reopens_the_purse(
sync_client, db_session, monkeypatch
):
"""The guard must not overshoot: each spirit is worth its own rite.
A durable claim that survived a summon would silently make every spirit
after the first worthless, which is a worse bug than the faucet.
"""
_no_twists(monkeypatch)
# The second summon must actually bring a DIFFERENT spirit — see
# `_force_new_presence`. Without this the test asserts a coin flip.
_force_new_presence(monkeypatch)
token = _login(sync_client, "second-spirit")
user_id = _user_id(sync_client, token)
with _ws_connect(sync_client, token) as ws:
_read_until(ws, "session")
first_entity = _summon(ws)
baseline = await _essence(db_session, user_id)
ws.send_json({"type": "passage_start"})
_settle(ws)
_walk(ws, BEATS_BEFORE_RELEASE)
after_first = await _essence(db_session, user_id)
second_entity = _summon(ws)
assert second_entity["entity"]["id"] != first_entity["entity"]["id"], (
"the second summon returned the same spirit — this test is about "
"a genuinely new presence"
)
ws.send_json({"type": "passage_start"})
_settle(ws)
_walk(ws, BEATS_BEFORE_RELEASE)
after_second = await _essence(db_session, user_id)
first_rite = after_first - baseline
# The second summon also pays its own trickle, so compare the rite only.
second_rite = after_second - after_first - SUMMON_ESSENCE_TRICKLE
assert first_rite > 0
assert second_rite == first_rite, (
"a fresh presence did not re-open the purse — every spirit after the "
"first is worth nothing"
)