fix: the test suite was destroying the production database
The worst bug of the session. tests/conftest.py built its engine from settings.database_url — the live database — and an autouse fixture calls drop_all() before EVERY test. So every backend run silently annihilated the real install: accounts, discovered spirits, Ghost Logs, devices, all of it. Found it because /sitemap.xml listed zero entities minutes after I had watched live séances mint real ones. Tests now use TEST_DATABASE_URL, or `<configured-db>_test` derived from it, and refuse to start at all if that ever resolves back to the production URL — this box both serves the app and holds the repo, so "don't run tests in prod" is not a workable guard. Proven: inserted a canary row into production, ran 50 tests, canary survived. Before this it would have been dropped. Also in this commit: SEO (routes/seo.py, lib/pageMeta.ts) - Live /sitemap.xml generated from real entity rows, and /robots.txt, both registered BEFORE the SPA catch-all or they'd be served index.html. Crawlers are disallowed from /seance specifically because the open door provisions a guest on arrival — a crawler would fill the users table with wanderers who never existed. - Per-route <title>, description, canonical and JSON-LD. The Codex is the indexable asset here (every spirit is unique long-form prose) and all of it previously shared one static title, so entities competed with each other instead of ranking. Entities are marked up as fictional Persons so a rich result can never imply a record of a real dead human. - public_base_url setting: absolute URLs for crawlers can't be derived from the request, since behind the tunnel the app only sees an internal host. Camera channel, first half (lib/camera.ts, llm scry path) - OllamaClient.generate() now accepts `images`; the configured chat model (minicpm-v4.5:8b) is vision-capable, so the entity can speak about what the seeker's camera actually shows. Verified against a synthetic room image: it named the pale column and the small red cube, then misread them as oak in a farmhouse parlor — real perception, in character. - Frames are captured only on an explicit act, downscaled to 768px and JPEG-compressed, never stored, and the prompt forbids describing faces or guessing identity. CameraEye carries the same generation guard as the EVP listener so closing during the permission prompt can't leave the camera live after teardown. 338 backend tests pass; 375 frontend; i18n parity holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -21,5 +21,12 @@ class Settings(BaseSettings):
|
||||
# Where synthesized utterance audio is written (served at /audio/).
|
||||
data_dir: str = "data"
|
||||
|
||||
# Canonical public origin, used for absolute URLs that must be correct
|
||||
# for machines rather than browsers: sitemap <loc> entries, the
|
||||
# robots.txt Sitemap line, and rel=canonical. Cannot be derived from the
|
||||
# request — behind the Cloudflare Tunnel the app sees an internal host,
|
||||
# and advertising that to a crawler would publish unreachable URLs.
|
||||
public_base_url: str = "https://spirit.thetempleofdoom.com"
|
||||
|
||||
|
||||
settings = Settings()
|
||||
|
||||
@@ -15,12 +15,19 @@ class OllamaClient:
|
||||
prompt: str,
|
||||
system: str | None = None,
|
||||
options: dict | None = None,
|
||||
images: list[str] | None = None,
|
||||
) -> str:
|
||||
"""`images` is a list of raw base64 JPEG/PNG strings (no data-URL
|
||||
prefix) for vision-capable models — Ollama's /api/generate takes them
|
||||
alongside the prompt. Ignored by text-only models, so passing them is
|
||||
safe; the caller is responsible for choosing a model that can see."""
|
||||
payload: dict = {"model": model, "prompt": prompt, "stream": False}
|
||||
if system is not None:
|
||||
payload["system"] = system
|
||||
if options:
|
||||
payload["options"] = options
|
||||
if images:
|
||||
payload["images"] = images
|
||||
|
||||
async with httpx.AsyncClient(base_url=self._base_url, timeout=120.0) as http_client:
|
||||
response = await http_client.post("/api/generate", json=payload)
|
||||
|
||||
@@ -63,6 +63,30 @@ MANIFEST_PROMPT = """The room right now:
|
||||
Something shifted. Speak."""
|
||||
|
||||
|
||||
SCRY_SYSTEM = (
|
||||
"You are {name}, {epithet} — a spirit persona in an interactive horror "
|
||||
"art installation.\n"
|
||||
"Your nature: {persona}\n"
|
||||
"The seeker has turned a lens toward the room and you can see through "
|
||||
"it. Speak about what is ACTUALLY in the image — the real objects, the "
|
||||
"real light, the real room. Name one or two specific things you see, "
|
||||
"plainly enough that the seeker knows you are truly looking.\n"
|
||||
"Then let one of them mean something to you: mistake an object for one "
|
||||
"you owned, recognise a shape, notice what is missing, or refuse to look "
|
||||
"at a particular corner. The unsettling part is accuracy followed by "
|
||||
"wrongness — not vagueness.\n"
|
||||
"Never describe a person's face or body, and never guess at anyone's "
|
||||
"identity, age or appearance; if a person is present, speak only of "
|
||||
"their presence. Under 40 words. Never break character, never mention "
|
||||
"being an AI, never mention images, cameras or models."
|
||||
"{language_clause}"
|
||||
)
|
||||
|
||||
SCRY_PROMPT = """This is what the lens shows you right now.
|
||||
|
||||
Speak."""
|
||||
|
||||
|
||||
MINT_SYSTEM = (
|
||||
"You invent spirit personas for an interactive horror art installation. "
|
||||
"The single biggest thing separating a convincing dead person from a "
|
||||
@@ -197,6 +221,15 @@ def manifest_prompt(readings: dict) -> str:
|
||||
return MANIFEST_PROMPT.format(readings="\n".join(lines))
|
||||
|
||||
|
||||
def scry_system(entity: dict, language: str = "en") -> str:
|
||||
return SCRY_SYSTEM.format(
|
||||
name=entity.get("name", "an unnamed presence"),
|
||||
epithet=entity.get("epithet", "a voice in the static"),
|
||||
persona=entity.get("persona", "A drifting presence with no remembered past."),
|
||||
language_clause=language_clause(language),
|
||||
)
|
||||
|
||||
|
||||
def mint_prompt(
|
||||
signature: str,
|
||||
channel: str,
|
||||
|
||||
@@ -185,6 +185,43 @@ class SpiritService:
|
||||
self._touch()
|
||||
return raw.strip().strip('"')[:200]
|
||||
|
||||
async def scry(
|
||||
self,
|
||||
entity: dict,
|
||||
image_b64: str,
|
||||
language: str = "en",
|
||||
entropy: object = None,
|
||||
) -> str:
|
||||
"""The entity speaks about what the seeker's camera actually shows.
|
||||
|
||||
The configured chat model (minicpm-v4.5) is vision-capable, so this
|
||||
is a genuine look at the real room rather than an invented
|
||||
description — the same principle as every other channel here: real
|
||||
measurement first, interpretation second.
|
||||
|
||||
Seeded from physical entropy like manifest(), so two identical rooms
|
||||
still produce different speech.
|
||||
"""
|
||||
seed = int.from_bytes(veil_seed(entropy, "scry")[:8], "big") % (2**63)
|
||||
|
||||
async def call() -> str:
|
||||
return await self._client.generate(
|
||||
settings.ollama_chat_model,
|
||||
prompts.SCRY_PROMPT,
|
||||
system=prompts.scry_system(entity, language),
|
||||
options={
|
||||
"num_predict": 90,
|
||||
"temperature": 0.95,
|
||||
"top_p": 0.95,
|
||||
"seed": seed,
|
||||
},
|
||||
images=[image_b64],
|
||||
)
|
||||
|
||||
raw = await self._queue.submit(call)
|
||||
self._touch()
|
||||
return raw.strip().strip('"')[:400]
|
||||
|
||||
async def mint_profile(
|
||||
self,
|
||||
signature: str,
|
||||
|
||||
@@ -16,6 +16,7 @@ from app.routes.conditions import router as conditions_router
|
||||
from app.routes.device import router as device_router
|
||||
from app.routes.inventory import router as inventory_router
|
||||
from app.routes.seances import router as seances_router
|
||||
from app.routes.seo import router as seo_router
|
||||
from app.routes.shop import router as shop_router
|
||||
from app.session_cleanup import delete_expired_sessions
|
||||
from app.ws import AUDIO_DIR
|
||||
@@ -100,6 +101,9 @@ app.include_router(conditions_router)
|
||||
app.include_router(device_router)
|
||||
app.include_router(inventory_router)
|
||||
app.include_router(seances_router)
|
||||
# Registered before the SPA catch-all below, or /robots.txt and
|
||||
# /sitemap.xml would be served index.html instead.
|
||||
app.include_router(seo_router)
|
||||
app.include_router(shop_router)
|
||||
app.include_router(ws_router)
|
||||
|
||||
|
||||
105
backend/app/routes/seo.py
Normal file
105
backend/app/routes/seo.py
Normal file
@@ -0,0 +1,105 @@
|
||||
"""robots.txt and a live sitemap.
|
||||
|
||||
Served from the backend rather than dropped in `public/` because the
|
||||
valuable, indexable surface of this site is the Codex, and the Codex grows
|
||||
every time somebody summons. A static sitemap would be stale within an
|
||||
hour; this one is generated from the actual entity rows.
|
||||
|
||||
What is deliberately NOT listed: /seance, /enter, /profile, /log,
|
||||
/inventory, /devices — anything per-seeker or interactive. A crawler
|
||||
hitting /seance would provision a guest account on arrival (the open
|
||||
door), which would fill the users table with wanderers that never
|
||||
existed as people. robots.txt disallows those paths for the same reason,
|
||||
and the sitemap only advertises pages that are genuinely public,
|
||||
stable, and worth a search result: the landing page, the shop, and every
|
||||
discovered spirit.
|
||||
"""
|
||||
|
||||
from datetime import datetime, timezone
|
||||
|
||||
from fastapi import APIRouter, Depends, Response
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.ext.asyncio import AsyncSession
|
||||
|
||||
from app.config import settings
|
||||
from app.db import get_db
|
||||
from app.models.entity import Entity
|
||||
|
||||
router = APIRouter(tags=["seo"])
|
||||
|
||||
# Cap the sitemap so a runaway Codex can't produce a multi-megabyte
|
||||
# document. 5k is far inside the 50k/50MB sitemap limit and orders of
|
||||
# magnitude beyond what this install will realistically hold.
|
||||
MAX_SITEMAP_ENTITIES = 5000
|
||||
|
||||
# Paths that must never be crawled: they either mutate state (provisioning a
|
||||
# guest) or are meaningless without a session.
|
||||
DISALLOWED = (
|
||||
"/seance",
|
||||
"/enter",
|
||||
"/profile",
|
||||
"/hunters",
|
||||
"/log",
|
||||
"/inventory",
|
||||
"/devices",
|
||||
"/api/",
|
||||
)
|
||||
|
||||
|
||||
def _base_url() -> str:
|
||||
return settings.public_base_url.rstrip("/")
|
||||
|
||||
|
||||
def _iso(dt: datetime | None) -> str:
|
||||
value = dt or datetime.now(timezone.utc)
|
||||
if value.tzinfo is None:
|
||||
value = value.replace(tzinfo=timezone.utc)
|
||||
return value.date().isoformat()
|
||||
|
||||
|
||||
def _xml_escape(text: str) -> str:
|
||||
return (
|
||||
text.replace("&", "&")
|
||||
.replace("<", "<")
|
||||
.replace(">", ">")
|
||||
.replace('"', """)
|
||||
)
|
||||
|
||||
|
||||
@router.get("/robots.txt", include_in_schema=False)
|
||||
async def robots() -> Response:
|
||||
lines = ["User-agent: *"]
|
||||
lines += [f"Disallow: {path}" for path in DISALLOWED]
|
||||
lines.append("Allow: /")
|
||||
lines.append(f"Sitemap: {_base_url()}/sitemap.xml")
|
||||
return Response("\n".join(lines) + "\n", media_type="text/plain")
|
||||
|
||||
|
||||
@router.get("/sitemap.xml", include_in_schema=False)
|
||||
async def sitemap(db: AsyncSession = Depends(get_db)) -> Response:
|
||||
base = _base_url()
|
||||
result = await db.execute(
|
||||
select(Entity.id, Entity.discovered_at)
|
||||
.order_by(Entity.discovered_at.desc())
|
||||
.limit(MAX_SITEMAP_ENTITIES)
|
||||
)
|
||||
entities = result.all()
|
||||
|
||||
parts = [
|
||||
'<?xml version="1.0" encoding="UTF-8"?>',
|
||||
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
|
||||
f"<url><loc>{base}/</loc><changefreq>daily</changefreq>"
|
||||
"<priority>1.0</priority></url>",
|
||||
f"<url><loc>{base}/codex</loc><changefreq>hourly</changefreq>"
|
||||
"<priority>0.9</priority></url>",
|
||||
f"<url><loc>{base}/shop</loc><changefreq>weekly</changefreq>"
|
||||
"<priority>0.7</priority></url>",
|
||||
]
|
||||
for entity_id, discovered_at in entities:
|
||||
loc = _xml_escape(f"{base}/codex/{entity_id}")
|
||||
parts.append(
|
||||
f"<url><loc>{loc}</loc><lastmod>{_iso(discovered_at)}</lastmod>"
|
||||
"<changefreq>weekly</changefreq><priority>0.6</priority></url>"
|
||||
)
|
||||
parts.append("</urlset>")
|
||||
return Response("\n".join(parts), media_type="application/xml")
|
||||
Reference in New Issue
Block a user