feat: the lens — the camera as a channel the dead look through
A new séance mode. The seeker opens their camera, presses "let it look", and the entity speaks about what is ACTUALLY in the room — the configured chat model (minicpm-v4.5:8b) is vision-capable, so this is real perception, not invented description. Same principle as every other channel here: real measurement first, interpretation second. Verified live end-to-end through the real WebSocket: given a synthetic room (pale doorway, red flame on dark boards), "Bessie L. Carter" reported the gray rectangle and red square on a dark surface with faint shadows, then misread it as her pen feeling heavy the night before Mr. Edgerton's birdseed arrived. Accuracy followed by wrongness, which is the whole effect. Privacy is the load-bearing design constraint, not a footnote: - "Camera open" and "the entity saw something" are deliberately separate states. Opening the lens transmits NOTHING; only an explicit press sends one still. There is no timer and no background capture path. - Frames are downscaled to 768px and JPEG-compressed client-side, then passed to the model and dropped. Never written to disk, never logged, never attached to an event row — only the resulting utterance is stored, exactly like any other thing a spirit says. - The prompt forbids describing faces or guessing anyone's identity, age or appearance; a person present is spoken of only as a presence. - A closed lens is covered by an opaque veil in the UI, so there is never ambiguity about whether the camera is live. Robustness: - CameraEye carries the same generation guard the EVP listener needed: closing during the permission prompt releases the late-arriving stream instead of letting the camera go live after teardown. - Failures are classified (denied / insecure / absent / busy / unknown) rather than always blaming the seeker for a refusal. - Scrying is the heaviest request this app makes of a CPU-only Ollama box, so it gets the tightest limiter of any channel (4/min/user, 8/min/IP). - Frames are size-capped BEFORE reaching the queue, and a vision failure emits an error frame instead of killing the socket — both covered by tests asserting the model was never called. 10 new frontend tests, 5 new backend tests. 385 frontend + backend suites pass; i18n parity holds across both languages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -12,6 +12,10 @@ Protocol (client → server):
|
||||
{"type": "ritual_step", "step": <int>} → (on the final step) ritual_complete
|
||||
{"type": "judgment", "verdict": "trust" | "banish" | "test" | "cross_over"}
|
||||
→ judgment_result
|
||||
{"type": "scry", "image": "<base64 jpeg>"} → {"type": "utterance", kind: "scry"}
|
||||
The seeker's camera, shown to the vision model so the entity can speak
|
||||
about the real room. The frame is never stored or logged — only the
|
||||
resulting utterance is, like any other spirit speech.
|
||||
|
||||
All server → client frames flow through a single sender task so concurrent
|
||||
producers (ambient loop, reply streaming, TTS callbacks) never interleave on
|
||||
@@ -65,7 +69,7 @@ router = APIRouter()
|
||||
# Alias so tests can swap in the NullPool test session maker.
|
||||
session_maker = _default_session_maker
|
||||
|
||||
MODES = {"wire", "evp", "radio", "ouija", "emf"}
|
||||
MODES = {"wire", "evp", "radio", "ouija", "emf", "camera"}
|
||||
|
||||
# Per-user limiters for every LLM-triggering message type (spec §5).
|
||||
fragment_limiter = RateLimiter(max_requests=30, window_seconds=60)
|
||||
@@ -80,6 +84,10 @@ summon_limiter = RateLimiter(max_requests=4, window_seconds=60)
|
||||
# same modest, human-plausible cadence as the other reward triggers.
|
||||
ritual_limiter = RateLimiter(max_requests=6, window_seconds=60)
|
||||
judgment_limiter = RateLimiter(max_requests=10, window_seconds=60)
|
||||
# Scrying sends a real image to a vision model — by far the heaviest
|
||||
# request this app makes of a CPU-only Ollama box, so it gets the
|
||||
# tightest budget of any channel.
|
||||
scry_limiter = RateLimiter(max_requests=4, window_seconds=60)
|
||||
|
||||
# Per-IP limiters for the same trigger points (spec §5). Ollama is a shared,
|
||||
# single-instance, CPU-only resource — per-account limits alone don't stop
|
||||
@@ -92,6 +100,7 @@ question_ip_limiter = RateLimiter(max_requests=12, window_seconds=60)
|
||||
summon_ip_limiter = RateLimiter(max_requests=8, window_seconds=60)
|
||||
ritual_ip_limiter = RateLimiter(max_requests=12, window_seconds=60)
|
||||
judgment_ip_limiter = RateLimiter(max_requests=20, window_seconds=60)
|
||||
scry_ip_limiter = RateLimiter(max_requests=8, window_seconds=60)
|
||||
|
||||
# Probability that a channel's familiar presence answers again rather than
|
||||
# something new manifesting. High enough that the Codex stays collectable
|
||||
@@ -920,6 +929,79 @@ async def _handle_judgment(state: SeanceState, message: dict) -> None:
|
||||
await state.send_queue.put({"type": "item_drop", "item": item})
|
||||
|
||||
|
||||
# A 768px JPEG at quality 0.72 is well under 200KB, so ~350KB of base64 is a
|
||||
# generous ceiling that still refuses anything pathological before it reaches
|
||||
# the model.
|
||||
MAX_SCRY_B64_CHARS = 350_000
|
||||
|
||||
|
||||
async def _handle_scry(state: SeanceState, message: dict) -> None:
|
||||
"""The entity speaks about what the seeker's camera actually shows.
|
||||
|
||||
Unlike every other channel, the payload here is a photograph of a real
|
||||
room, so this handler is deliberately strict: no entity means nothing to
|
||||
look through, the image is size-capped before it touches the queue, and
|
||||
the frame is never persisted or logged anywhere — it is passed to the
|
||||
model and dropped. Only the resulting utterance is recorded, exactly like
|
||||
any other thing a spirit says.
|
||||
"""
|
||||
if state.entity is None:
|
||||
return
|
||||
image = message.get("image")
|
||||
if not isinstance(image, str) or not image.strip():
|
||||
return
|
||||
if len(image) > MAX_SCRY_B64_CHARS:
|
||||
await state.send_queue.put(
|
||||
{
|
||||
"type": "error",
|
||||
"code": "scry_too_large",
|
||||
"message": "the lens showed too much at once — try again.",
|
||||
}
|
||||
)
|
||||
return
|
||||
if not (
|
||||
scry_limiter.allow(str(state.user_id))
|
||||
and scry_ip_limiter.allow(state.client_ip)
|
||||
):
|
||||
await state.send_queue.put(
|
||||
{
|
||||
"type": "error",
|
||||
"code": "rate_limited",
|
||||
"message": "the eye tires. let it rest a moment before looking again.",
|
||||
}
|
||||
)
|
||||
return
|
||||
|
||||
await state.send_queue.put({"type": "status", "state": "gathering"})
|
||||
try:
|
||||
text = await spirit_service.scry(
|
||||
state.entity, image, state.language, entropy=state.entropy
|
||||
)
|
||||
except SpiritBusyError:
|
||||
await state.send_queue.put(
|
||||
{
|
||||
"type": "error",
|
||||
"code": "veil_crowded",
|
||||
"message": "too many eyes at once. try again shortly.",
|
||||
}
|
||||
)
|
||||
return
|
||||
except Exception:
|
||||
# Deliberately broad, same reasoning as _maybe_manifest: a vision
|
||||
# failure must leave the séance intact rather than surfacing a stack
|
||||
# trace for something the seeker can't act on.
|
||||
await state.send_queue.put(
|
||||
{
|
||||
"type": "error",
|
||||
"code": "scry_failed",
|
||||
"message": "the lens clouded over. nothing came through.",
|
||||
}
|
||||
)
|
||||
return
|
||||
if text:
|
||||
await _speak(state, "scry", text)
|
||||
|
||||
|
||||
@router.websocket("/ws/session")
|
||||
async def session_socket(websocket: WebSocket) -> None:
|
||||
user_id = await _authenticate(websocket)
|
||||
@@ -995,6 +1077,8 @@ async def session_socket(websocket: WebSocket) -> None:
|
||||
await _handle_ritual_step(state, message)
|
||||
elif msg_type == "judgment":
|
||||
await _handle_judgment(state, message)
|
||||
elif msg_type == "scry":
|
||||
await _handle_scry(state, message)
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
finally:
|
||||
|
||||
Reference in New Issue
Block a user