feat: the lens — the camera as a channel the dead look through

A new séance mode. The seeker opens their camera, presses "let it look",
and the entity speaks about what is ACTUALLY in the room — the configured
chat model (minicpm-v4.5:8b) is vision-capable, so this is real perception,
not invented description. Same principle as every other channel here: real
measurement first, interpretation second.

Verified live end-to-end through the real WebSocket: given a synthetic room
(pale doorway, red flame on dark boards), "Bessie L. Carter" reported the
gray rectangle and red square on a dark surface with faint shadows, then
misread it as her pen feeling heavy the night before Mr. Edgerton's birdseed
arrived. Accuracy followed by wrongness, which is the whole effect.

Privacy is the load-bearing design constraint, not a footnote:
- "Camera open" and "the entity saw something" are deliberately separate
  states. Opening the lens transmits NOTHING; only an explicit press sends
  one still. There is no timer and no background capture path.
- Frames are downscaled to 768px and JPEG-compressed client-side, then
  passed to the model and dropped. Never written to disk, never logged,
  never attached to an event row — only the resulting utterance is stored,
  exactly like any other thing a spirit says.
- The prompt forbids describing faces or guessing anyone's identity, age or
  appearance; a person present is spoken of only as a presence.
- A closed lens is covered by an opaque veil in the UI, so there is never
  ambiguity about whether the camera is live.

Robustness:
- CameraEye carries the same generation guard the EVP listener needed:
  closing during the permission prompt releases the late-arriving stream
  instead of letting the camera go live after teardown.
- Failures are classified (denied / insecure / absent / busy / unknown)
  rather than always blaming the seeker for a refusal.
- Scrying is the heaviest request this app makes of a CPU-only Ollama box,
  so it gets the tightest limiter of any channel (4/min/user, 8/min/IP).
- Frames are size-capped BEFORE reaching the queue, and a vision failure
  emits an error frame instead of killing the socket — both covered by
  tests asserting the model was never called.

10 new frontend tests, 5 new backend tests. 385 frontend + backend suites
pass; i18n parity holds across both languages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Indiana
2026-07-30 02:08:02 +00:00
parent 16c8bae00d
commit d37bb71e5d
12 changed files with 759 additions and 11 deletions

View File

@@ -12,6 +12,10 @@ Protocol (client → server):
{"type": "ritual_step", "step": <int>} → (on the final step) ritual_complete
{"type": "judgment", "verdict": "trust" | "banish" | "test" | "cross_over"}
→ judgment_result
{"type": "scry", "image": "<base64 jpeg>"} → {"type": "utterance", kind: "scry"}
The seeker's camera, shown to the vision model so the entity can speak
about the real room. The frame is never stored or logged — only the
resulting utterance is, like any other spirit speech.
All server → client frames flow through a single sender task so concurrent
producers (ambient loop, reply streaming, TTS callbacks) never interleave on
@@ -65,7 +69,7 @@ router = APIRouter()
# Alias so tests can swap in the NullPool test session maker.
session_maker = _default_session_maker
MODES = {"wire", "evp", "radio", "ouija", "emf"}
MODES = {"wire", "evp", "radio", "ouija", "emf", "camera"}
# Per-user limiters for every LLM-triggering message type (spec §5).
fragment_limiter = RateLimiter(max_requests=30, window_seconds=60)
@@ -80,6 +84,10 @@ summon_limiter = RateLimiter(max_requests=4, window_seconds=60)
# same modest, human-plausible cadence as the other reward triggers.
ritual_limiter = RateLimiter(max_requests=6, window_seconds=60)
judgment_limiter = RateLimiter(max_requests=10, window_seconds=60)
# Scrying sends a real image to a vision model — by far the heaviest
# request this app makes of a CPU-only Ollama box, so it gets the
# tightest budget of any channel.
scry_limiter = RateLimiter(max_requests=4, window_seconds=60)
# Per-IP limiters for the same trigger points (spec §5). Ollama is a shared,
# single-instance, CPU-only resource — per-account limits alone don't stop
@@ -92,6 +100,7 @@ question_ip_limiter = RateLimiter(max_requests=12, window_seconds=60)
summon_ip_limiter = RateLimiter(max_requests=8, window_seconds=60)
ritual_ip_limiter = RateLimiter(max_requests=12, window_seconds=60)
judgment_ip_limiter = RateLimiter(max_requests=20, window_seconds=60)
scry_ip_limiter = RateLimiter(max_requests=8, window_seconds=60)
# Probability that a channel's familiar presence answers again rather than
# something new manifesting. High enough that the Codex stays collectable
@@ -920,6 +929,79 @@ async def _handle_judgment(state: SeanceState, message: dict) -> None:
await state.send_queue.put({"type": "item_drop", "item": item})
# A 768px JPEG at quality 0.72 is well under 200KB, so ~350KB of base64 is a
# generous ceiling that still refuses anything pathological before it reaches
# the model.
MAX_SCRY_B64_CHARS = 350_000
async def _handle_scry(state: SeanceState, message: dict) -> None:
"""The entity speaks about what the seeker's camera actually shows.
Unlike every other channel, the payload here is a photograph of a real
room, so this handler is deliberately strict: no entity means nothing to
look through, the image is size-capped before it touches the queue, and
the frame is never persisted or logged anywhere — it is passed to the
model and dropped. Only the resulting utterance is recorded, exactly like
any other thing a spirit says.
"""
if state.entity is None:
return
image = message.get("image")
if not isinstance(image, str) or not image.strip():
return
if len(image) > MAX_SCRY_B64_CHARS:
await state.send_queue.put(
{
"type": "error",
"code": "scry_too_large",
"message": "the lens showed too much at once — try again.",
}
)
return
if not (
scry_limiter.allow(str(state.user_id))
and scry_ip_limiter.allow(state.client_ip)
):
await state.send_queue.put(
{
"type": "error",
"code": "rate_limited",
"message": "the eye tires. let it rest a moment before looking again.",
}
)
return
await state.send_queue.put({"type": "status", "state": "gathering"})
try:
text = await spirit_service.scry(
state.entity, image, state.language, entropy=state.entropy
)
except SpiritBusyError:
await state.send_queue.put(
{
"type": "error",
"code": "veil_crowded",
"message": "too many eyes at once. try again shortly.",
}
)
return
except Exception:
# Deliberately broad, same reasoning as _maybe_manifest: a vision
# failure must leave the séance intact rather than surfacing a stack
# trace for something the seeker can't act on.
await state.send_queue.put(
{
"type": "error",
"code": "scry_failed",
"message": "the lens clouded over. nothing came through.",
}
)
return
if text:
await _speak(state, "scry", text)
@router.websocket("/ws/session")
async def session_socket(websocket: WebSocket) -> None:
user_id = await _authenticate(websocket)
@@ -995,6 +1077,8 @@ async def session_socket(websocket: WebSocket) -> None:
await _handle_ritual_step(state, message)
elif msg_type == "judgment":
await _handle_judgment(state, message)
elif msg_type == "scry":
await _handle_scry(state, message)
except WebSocketDisconnect:
pass
finally: