feat: ritual + judgment + favor + cross-over (Workstream B)

Implements Workstream B of the character-depth-ghost-log spec:
ritual_start/ritual_step/judgment WS handlers, the pure judgment.py
logic module, and the User.favor / Entity.at_peace columns + migration.

- app/judgment.py: pure ritual success roll (base 65%, floored at 30%,
  driven by an entity's power+deceptiveness difficulty), the "stuck
  spirit" cross_over rule (alignment >= 0.5 and volatility > 0.6, ~20%
  of entities), judgment correctness/favor-delta/essence-delta/
  consequence resolution for all four verdicts, favor clamping, the
  favor-to-trait-roll bias applied at mint time, and tell-line
  generation (opaque behavioral flavor text, never a raw stat).
- app/ws.py: wires ritual_start/ritual_step/judgment frames, emits
  ritual_complete/tell/judgment_result/item_drop per the spec's
  Contract; traits are added to serialize_entity for internal
  server-side use but stripped from the outbound `entity` frame via a
  new _public_entity helper so hidden ground truth never reaches the
  client outside ritual_complete; _summon excludes at-peace entities
  from signature re-contact and mints a fresh (salted-signature) entity
  instead; new entities' traits are nudged by the discovering user's
  favor before being persisted.
- models/user.py, models/entity.py, main.py: User.favor and
  Entity.at_peace columns plus their idempotent ADD COLUMN IF NOT
  EXISTS migration lines in lifespan, alongside the existing ones.
- tests/test_judgment.py, tests/test_ws_ritual_judgment.py: 56 new
  tests covering the ritual/judgment correctness matrix, favor
  clamping/bias, essence crediting, at_peace persistence + re-contact,
  and the entity-frame trait leak guard.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Indiana
2026-07-24 21:27:30 +00:00
parent f023b591b5
commit 36c4a9e6e4
7 changed files with 1541 additions and 19 deletions

315
backend/app/judgment.py Normal file
View File

@@ -0,0 +1,315 @@
"""Ritual + judgment: pure logic for Workstream B of the Character Depth /
Ghost Log spec
(docs/superpowers/specs/2026-07-23-character-depth-ghost-log-design.md).
No I/O, no DB, no network — every function here takes plain values (and an
optional injectable `random.Random`) and returns plain values, so the whole
module is trivially unit-testable in isolation. `backend/app/ws.py`'s
`ritual_start`/`ritual_step`/`judgment` WS handlers are the only place these
outputs get persisted or sent over the wire.
"""
from __future__ import annotations
import random
from dataclasses import dataclass
from app.inventory import CORRECT_JUDGMENT_ESSENCE, CROSS_OVER_ESSENCE
# --- favor -------------------------------------------------------------
FAVOR_MIN = -1.0
FAVOR_MAX = 1.0
def clamp_favor(value: float) -> float:
"""Clamps a `User.favor` value to [-1.0, 1.0] — call this at every write
site, per the spec's Contract section."""
return max(FAVOR_MIN, min(FAVOR_MAX, value))
# --- ritual --------------------------------------------------------------
# Rituals should feel doable most of the time — this is a short rite in
# front of judgment, not a grindy gate. A highly resistant entity (high
# `power` *and* high `deceptiveness` — strong and cagey) drags the odds
# down, floored so even the worst-case entity stays beatable on a retry
# rather than a hard wall.
RITUAL_BASE_SUCCESS_CHANCE = 0.65
RITUAL_MIN_SUCCESS_CHANCE = 0.30
RITUAL_DIFFICULTY_SWING = RITUAL_BASE_SUCCESS_CHANCE - RITUAL_MIN_SUCCESS_CHANCE
def roll_ritual_success(traits: dict, rng: random.Random | None = None) -> bool:
"""One ritual attempt's pass/fail roll.
`traits` is the entity's hidden truth dict (`alignment`/`power`/
`volatility`/`deceptiveness`). Only `power` and `deceptiveness` affect
the odds — a strong, cagey spirit is the hardest to get an accurate
read on. `volatility` is deliberately left out: it's about how *noisy*
the entity's ambient tells are, a separate axis from whether a focused
ritual can pin it down. `alignment` is also left out — letting it
influence pass/fail would telegraph ground truth (easier ritual =
probably benevolent) before `ritual_complete` is supposed to reveal
anything, undermining the "judge blind or judge informed" choice the
contract is built around.
`rng` defaults to a fresh, unseeded `random.Random()` — deliberately
*not* signature-seeded like `entities.roll_traits`: a seeker can retry
the same entity's ritual repeatedly (`RitualPanel`'s "attempt again"),
and each attempt needs genuine variance, not a fixed pass/fail baked
into the entity's identity.
"""
rng = rng if rng is not None else random.Random()
power = float(traits.get("power", 0.5))
deceptiveness = float(traits.get("deceptiveness", 0.5))
difficulty = (power + deceptiveness) / 2.0 # 0..1
chance = RITUAL_BASE_SUCCESS_CHANCE - RITUAL_DIFFICULTY_SWING * difficulty
chance = max(RITUAL_MIN_SUCCESS_CHANCE, min(RITUAL_BASE_SUCCESS_CHANCE, chance))
return rng.random() < chance
# --- "stuck spirit" rule for cross_over -----------------------------------
# "Genuinely benevolent" mirrors the correct-trust threshold used below.
# "Reads as stuck" is modeled as above-average volatility: the existing
# fallback personas already lean on "died with something unfinished" themes
# (see entities.py's _PERSONA_TEMPLATES), and an unresolved spirit is
# exactly the kind whose behavioral tells would plausibly be noisy/
# inconsistent rather than settled, rather than tying "stuck" to a whole
# separate hidden dimension. 0.6 (vs. the 0.5 midpoint) keeps "stuck" a
# meaningfully above-average band rather than a coinflip: with both
# `alignment` and `volatility` rolled uniform(0, 1) independently,
# P(alignment >= 0.5) * P(volatility > 0.6) = 0.5 * 0.4 = 20% of all
# entities qualify — special enough to justify the single largest reward
# in the game (see CROSS_OVER_ESSENCE / FAVOR_CORRECT_CROSS_OVER) without
# being so rare it never comes up.
BENEVOLENT_ALIGNMENT_THRESHOLD = 0.5
STUCK_VOLATILITY_THRESHOLD = 0.6
def is_stuck_spirit(traits: dict) -> bool:
alignment = float(traits.get("alignment", 0.5))
volatility = float(traits.get("volatility", 0.5))
return (
alignment >= BENEVOLENT_ALIGNMENT_THRESHOLD
and volatility > STUCK_VOLATILITY_THRESHOLD
)
# --- judgment --------------------------------------------------------------
# Favor deltas: small, single-digit-percent nudges of the [-1, 1] range (a
# 2.0-wide band) — the caller always re-clamps via `clamp_favor`.
#
# Wrong-trust is punished harder than wrong-banish: trusting something that
# turns out malevolent is the reckless failure mode with real in-fiction
# consequences (the haunting escalates), while a wrong banish is merely
# over-cautious — you turned away something harmless, nothing lashes back.
# Correct cross_over pays the best favor *and* essence of any outcome,
# matching the contract's "largest essence reward of any outcome" language
# with an equivalently generous favor nudge: it's the hardest-to-spot,
# most compassionate correct call a seeker can make.
FAVOR_CORRECT_TRUST = 0.05
FAVOR_CORRECT_BANISH = 0.05
FAVOR_WRONG_TRUST = -0.10
FAVOR_WRONG_BANISH = -0.05
FAVOR_CORRECT_CROSS_OVER = 0.08
FAVOR_CROSS_OVER_FAIL = 0.0 # a naive read, not a reckless one — no penalty
FAVOR_TEST = 0.0 # a diagnostic pulse only — nothing risked, nothing gained
VERDICTS = ("trust", "banish", "test", "cross_over")
@dataclass(frozen=True)
class JudgmentOutcome:
correct: bool
favor_delta: float
essence_delta: int
at_peace: bool
consequence: str # "reward" | "escalation" | "withdrawal" | "crossed_over" | "resisted" | "neutral"
def judge_verdict(
verdict: str,
traits: dict,
*,
ritual_completed: bool = False,
ritual_success: bool = False,
) -> JudgmentOutcome:
"""Resolves one `judgment` frame against the entity's hidden truth.
`correct` per the contract: trust called on real alignment >= 0.5, or
banish called on alignment < 0.5, or cross_over called on a genuinely
"stuck" spirit (see `is_stuck_spirit`). `cross_over` on anything else
(a demon, or a benevolent-but-not-stuck spirit that simply isn't ready
to move on) resists — "resisted" covers both: neither is a reckless
misread the way a wrong trust/banish is, so neither costs favor.
`test` never touches favor/essence; its `correct` reflects whether a
completed-this-session ritual actually surfaced true information (no
completed ritual, or a failed one, means the diagnostic has nothing
real to go on).
"""
if verdict not in VERDICTS:
raise ValueError(f"unknown verdict: {verdict!r}")
alignment = float(traits.get("alignment", 0.5))
benevolent = alignment >= BENEVOLENT_ALIGNMENT_THRESHOLD
if verdict == "trust":
if benevolent:
return JudgmentOutcome(
True, FAVOR_CORRECT_TRUST, CORRECT_JUDGMENT_ESSENCE, False, "reward"
)
return JudgmentOutcome(False, FAVOR_WRONG_TRUST, 0, False, "escalation")
if verdict == "banish":
if not benevolent:
return JudgmentOutcome(
True, FAVOR_CORRECT_BANISH, CORRECT_JUDGMENT_ESSENCE, False, "reward"
)
return JudgmentOutcome(False, FAVOR_WRONG_BANISH, 0, False, "withdrawal")
if verdict == "cross_over":
if is_stuck_spirit(traits):
return JudgmentOutcome(
True, FAVOR_CORRECT_CROSS_OVER, CROSS_OVER_ESSENCE, True, "crossed_over"
)
return JudgmentOutcome(False, FAVOR_CROSS_OVER_FAIL, 0, False, "resisted")
# verdict == "test"
correct = bool(ritual_completed and ritual_success)
return JudgmentOutcome(correct, FAVOR_TEST, 0, False, "neutral")
# --- favor read-back bias on trait rolls ------------------------------------
# `entities.roll_traits` stays a pure, signature-seeded function with no
# session/user awareness (Workstream A's design — see the "roll_traits
# doesn't take a bias/rng parameter" check in ws.py's minting path). This is
# applied instead as a post-processing nudge, called from `app.ws._summon`
# right after a *new* entity's profile is minted, using the discovering
# user's current favor.
#
# It's applied once, at mint time, not per-viewing-session: the nudged
# traits become that entity's permanent, shared ground truth in the Codex
# (the same spirit reads the same way to every future seeker who contacts
# it). That matches the contract's "read back to mildly bias future summon
# trait rolls" framing — favor shapes what a seeker tends to *conjure*,
# not a private lens they view existing spirits through.
#
# Effect size is deliberately small and verifiable: at most a 0.12 shift at
# |favor| == 1.0, linear in between, applied only to volatility/
# deceptiveness (the two "how legible is this entity" axes) — alignment and
# power are left untouched so favor nudges readability, not who a spirit
# fundamentally is.
FAVOR_TRAIT_BIAS_MAX = 0.12
_BIASED_TRAIT_KEYS = ("volatility", "deceptiveness")
def apply_favor_bias(traits: dict, favor: float) -> dict:
"""Higher favor nudges volatility/deceptiveness down (more legible
entities); lower favor nudges them up. Returns a new dict; clamps each
nudged value back into [0.0, 1.0]."""
favor = clamp_favor(favor)
shift = -FAVOR_TRAIT_BIAS_MAX * favor
biased = dict(traits)
for key in _BIASED_TRAIT_KEYS:
if key in biased:
biased[key] = max(0.0, min(1.0, float(biased[key]) + shift))
return biased
# --- tells -----------------------------------------------------------------
# Flavor text describing *behavior*, never a stat number (per the contract
# and frontend/src/lib/evilMeter.ts's comments on how tells are consumed) —
# each line hints at one trait dimension without naming it. Picking one
# trait per call (rather than always the most extreme) keeps tells varied
# session to session even for the same entity; the high/low/neutral banding
# means a middling trait still produces a plausible, non-committal line
# instead of always screaming its most extreme dimension.
_TELL_LINES: dict[str, dict[str, tuple[str, ...]]] = {
"alignment": {
"high": (
"a warmth threads through the static, unmistakably kind.",
"it seems to want nothing more than to be heard.",
"the presence feels gentle, almost grateful for the company.",
),
"low": (
"something cold coils under the words.",
"you feel watched, not accompanied.",
"the air around the signal turns unfriendly.",
),
"neutral": (
"hard to say if it means well.",
"the intent behind it stays unreadable.",
),
},
"power": {
"high": (
"the signal surges, straining the line.",
"it pushes back against the questions, strong-willed.",
),
"low": (
"the presence feels thin, easily startled.",
"it flickers at the edge of hearing.",
),
"neutral": (
"steady, unremarkable strength.",
"neither weak nor overwhelming.",
),
},
"volatility": {
"high": (
"the tone lurches without warning.",
"it contradicts itself within the same breath.",
"the signal keeps slipping out from under itself.",
),
"low": (
"calm, consistent, almost rehearsed.",
"every answer lands the same measured way.",
),
"neutral": (
"mostly steady, with the odd hitch.",
"a little uneven, nothing alarming.",
),
},
"deceptiveness": {
"high": (
"the entity avoided a direct question.",
"an answer arrives that doesn't quite fit what was asked.",
"something about the reply feels rehearsed, not remembered.",
),
"low": (
"it answers plainly, almost bluntly.",
"nothing about the reply feels rehearsed.",
),
"neutral": (
"the reply seems straightforward enough.",
"no obvious dodge, no obvious tell.",
),
},
}
_TELL_TRAIT_KEYS = tuple(_TELL_LINES.keys())
TELL_HIGH_THRESHOLD = 0.6
TELL_LOW_THRESHOLD = 0.4
def generate_tell(traits: dict, rng: random.Random) -> str:
"""Picks one trait dimension at random and returns a short, opaque
flavor line hinting at it (never the raw number). `rng` should be a
per-session `random.Random` so a given session's tell sequence is
varied but the choice of *which* call sites emit a tell stays the
caller's (ws.py's) responsibility."""
trait_key = rng.choice(_TELL_TRAIT_KEYS)
value = float(traits.get(trait_key, 0.5))
bank = _TELL_LINES[trait_key]
if value >= TELL_HIGH_THRESHOLD:
pool = bank["high"]
elif value <= TELL_LOW_THRESHOLD:
pool = bank["low"]
else:
pool = bank["neutral"]
return rng.choice(pool)