fix: the test suite was destroying the production database

The worst bug of the session. tests/conftest.py built its engine from
settings.database_url — the live database — and an autouse fixture calls
drop_all() before EVERY test. So every backend run silently annihilated the
real install: accounts, discovered spirits, Ghost Logs, devices, all of it.
Found it because /sitemap.xml listed zero entities minutes after I had
watched live séances mint real ones.

Tests now use TEST_DATABASE_URL, or `<configured-db>_test` derived from it,
and refuse to start at all if that ever resolves back to the production URL
— this box both serves the app and holds the repo, so "don't run tests in
prod" is not a workable guard.

Proven: inserted a canary row into production, ran 50 tests, canary
survived. Before this it would have been dropped.

Also in this commit:

SEO (routes/seo.py, lib/pageMeta.ts)
- Live /sitemap.xml generated from real entity rows, and /robots.txt, both
  registered BEFORE the SPA catch-all or they'd be served index.html.
  Crawlers are disallowed from /seance specifically because the open door
  provisions a guest on arrival — a crawler would fill the users table with
  wanderers who never existed.
- Per-route <title>, description, canonical and JSON-LD. The Codex is the
  indexable asset here (every spirit is unique long-form prose) and all of
  it previously shared one static title, so entities competed with each
  other instead of ranking. Entities are marked up as fictional Persons so
  a rich result can never imply a record of a real dead human.
- public_base_url setting: absolute URLs for crawlers can't be derived from
  the request, since behind the tunnel the app only sees an internal host.

Camera channel, first half (lib/camera.ts, llm scry path)
- OllamaClient.generate() now accepts `images`; the configured chat model
  (minicpm-v4.5:8b) is vision-capable, so the entity can speak about what
  the seeker's camera actually shows. Verified against a synthetic room
  image: it named the pale column and the small red cube, then misread them
  as oak in a farmhouse parlor — real perception, in character.
- Frames are captured only on an explicit act, downscaled to 768px and
  JPEG-compressed, never stored, and the prompt forbids describing faces or
  guessing identity. CameraEye carries the same generation guard as the EVP
  listener so closing during the permission prompt can't leave the camera
  live after teardown.

338 backend tests pass; 375 frontend; i18n parity holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Indiana
2026-07-30 01:44:01 +00:00
parent 79401338d5
commit 16c8bae00d
13 changed files with 633 additions and 1 deletions

View File

@@ -21,5 +21,12 @@ class Settings(BaseSettings):
# Where synthesized utterance audio is written (served at /audio/).
data_dir: str = "data"
# Canonical public origin, used for absolute URLs that must be correct
# for machines rather than browsers: sitemap <loc> entries, the
# robots.txt Sitemap line, and rel=canonical. Cannot be derived from the
# request — behind the Cloudflare Tunnel the app sees an internal host,
# and advertising that to a crawler would publish unreachable URLs.
public_base_url: str = "https://spirit.thetempleofdoom.com"
settings = Settings()

View File

@@ -15,12 +15,19 @@ class OllamaClient:
prompt: str,
system: str | None = None,
options: dict | None = None,
images: list[str] | None = None,
) -> str:
"""`images` is a list of raw base64 JPEG/PNG strings (no data-URL
prefix) for vision-capable models — Ollama's /api/generate takes them
alongside the prompt. Ignored by text-only models, so passing them is
safe; the caller is responsible for choosing a model that can see."""
payload: dict = {"model": model, "prompt": prompt, "stream": False}
if system is not None:
payload["system"] = system
if options:
payload["options"] = options
if images:
payload["images"] = images
async with httpx.AsyncClient(base_url=self._base_url, timeout=120.0) as http_client:
response = await http_client.post("/api/generate", json=payload)

View File

@@ -63,6 +63,30 @@ MANIFEST_PROMPT = """The room right now:
Something shifted. Speak."""
SCRY_SYSTEM = (
"You are {name}, {epithet} — a spirit persona in an interactive horror "
"art installation.\n"
"Your nature: {persona}\n"
"The seeker has turned a lens toward the room and you can see through "
"it. Speak about what is ACTUALLY in the image — the real objects, the "
"real light, the real room. Name one or two specific things you see, "
"plainly enough that the seeker knows you are truly looking.\n"
"Then let one of them mean something to you: mistake an object for one "
"you owned, recognise a shape, notice what is missing, or refuse to look "
"at a particular corner. The unsettling part is accuracy followed by "
"wrongness — not vagueness.\n"
"Never describe a person's face or body, and never guess at anyone's "
"identity, age or appearance; if a person is present, speak only of "
"their presence. Under 40 words. Never break character, never mention "
"being an AI, never mention images, cameras or models."
"{language_clause}"
)
SCRY_PROMPT = """This is what the lens shows you right now.
Speak."""
MINT_SYSTEM = (
"You invent spirit personas for an interactive horror art installation. "
"The single biggest thing separating a convincing dead person from a "
@@ -197,6 +221,15 @@ def manifest_prompt(readings: dict) -> str:
return MANIFEST_PROMPT.format(readings="\n".join(lines))
def scry_system(entity: dict, language: str = "en") -> str:
return SCRY_SYSTEM.format(
name=entity.get("name", "an unnamed presence"),
epithet=entity.get("epithet", "a voice in the static"),
persona=entity.get("persona", "A drifting presence with no remembered past."),
language_clause=language_clause(language),
)
def mint_prompt(
signature: str,
channel: str,

View File

@@ -185,6 +185,43 @@ class SpiritService:
self._touch()
return raw.strip().strip('"')[:200]
async def scry(
self,
entity: dict,
image_b64: str,
language: str = "en",
entropy: object = None,
) -> str:
"""The entity speaks about what the seeker's camera actually shows.
The configured chat model (minicpm-v4.5) is vision-capable, so this
is a genuine look at the real room rather than an invented
description — the same principle as every other channel here: real
measurement first, interpretation second.
Seeded from physical entropy like manifest(), so two identical rooms
still produce different speech.
"""
seed = int.from_bytes(veil_seed(entropy, "scry")[:8], "big") % (2**63)
async def call() -> str:
return await self._client.generate(
settings.ollama_chat_model,
prompts.SCRY_PROMPT,
system=prompts.scry_system(entity, language),
options={
"num_predict": 90,
"temperature": 0.95,
"top_p": 0.95,
"seed": seed,
},
images=[image_b64],
)
raw = await self._queue.submit(call)
self._touch()
return raw.strip().strip('"')[:400]
async def mint_profile(
self,
signature: str,

View File

@@ -16,6 +16,7 @@ from app.routes.conditions import router as conditions_router
from app.routes.device import router as device_router
from app.routes.inventory import router as inventory_router
from app.routes.seances import router as seances_router
from app.routes.seo import router as seo_router
from app.routes.shop import router as shop_router
from app.session_cleanup import delete_expired_sessions
from app.ws import AUDIO_DIR
@@ -100,6 +101,9 @@ app.include_router(conditions_router)
app.include_router(device_router)
app.include_router(inventory_router)
app.include_router(seances_router)
# Registered before the SPA catch-all below, or /robots.txt and
# /sitemap.xml would be served index.html instead.
app.include_router(seo_router)
app.include_router(shop_router)
app.include_router(ws_router)

105
backend/app/routes/seo.py Normal file
View File

@@ -0,0 +1,105 @@
"""robots.txt and a live sitemap.
Served from the backend rather than dropped in `public/` because the
valuable, indexable surface of this site is the Codex, and the Codex grows
every time somebody summons. A static sitemap would be stale within an
hour; this one is generated from the actual entity rows.
What is deliberately NOT listed: /seance, /enter, /profile, /log,
/inventory, /devices — anything per-seeker or interactive. A crawler
hitting /seance would provision a guest account on arrival (the open
door), which would fill the users table with wanderers that never
existed as people. robots.txt disallows those paths for the same reason,
and the sitemap only advertises pages that are genuinely public,
stable, and worth a search result: the landing page, the shop, and every
discovered spirit.
"""
from datetime import datetime, timezone
from fastapi import APIRouter, Depends, Response
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.config import settings
from app.db import get_db
from app.models.entity import Entity
router = APIRouter(tags=["seo"])
# Cap the sitemap so a runaway Codex can't produce a multi-megabyte
# document. 5k is far inside the 50k/50MB sitemap limit and orders of
# magnitude beyond what this install will realistically hold.
MAX_SITEMAP_ENTITIES = 5000
# Paths that must never be crawled: they either mutate state (provisioning a
# guest) or are meaningless without a session.
DISALLOWED = (
"/seance",
"/enter",
"/profile",
"/hunters",
"/log",
"/inventory",
"/devices",
"/api/",
)
def _base_url() -> str:
return settings.public_base_url.rstrip("/")
def _iso(dt: datetime | None) -> str:
value = dt or datetime.now(timezone.utc)
if value.tzinfo is None:
value = value.replace(tzinfo=timezone.utc)
return value.date().isoformat()
def _xml_escape(text: str) -> str:
return (
text.replace("&", "&amp;")
.replace("<", "&lt;")
.replace(">", "&gt;")
.replace('"', "&quot;")
)
@router.get("/robots.txt", include_in_schema=False)
async def robots() -> Response:
lines = ["User-agent: *"]
lines += [f"Disallow: {path}" for path in DISALLOWED]
lines.append("Allow: /")
lines.append(f"Sitemap: {_base_url()}/sitemap.xml")
return Response("\n".join(lines) + "\n", media_type="text/plain")
@router.get("/sitemap.xml", include_in_schema=False)
async def sitemap(db: AsyncSession = Depends(get_db)) -> Response:
base = _base_url()
result = await db.execute(
select(Entity.id, Entity.discovered_at)
.order_by(Entity.discovered_at.desc())
.limit(MAX_SITEMAP_ENTITIES)
)
entities = result.all()
parts = [
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
f"<url><loc>{base}/</loc><changefreq>daily</changefreq>"
"<priority>1.0</priority></url>",
f"<url><loc>{base}/codex</loc><changefreq>hourly</changefreq>"
"<priority>0.9</priority></url>",
f"<url><loc>{base}/shop</loc><changefreq>weekly</changefreq>"
"<priority>0.7</priority></url>",
]
for entity_id, discovered_at in entities:
loc = _xml_escape(f"{base}/codex/{entity_id}")
parts.append(
f"<url><loc>{loc}</loc><lastmod>{_iso(discovered_at)}</lastmod>"
"<changefreq>weekly</changefreq><priority>0.6</priority></url>"
)
parts.append("</urlset>")
return Response("\n".join(parts), media_type="application/xml")