The worst bug of the session. tests/conftest.py built its engine from settings.database_url — the live database — and an autouse fixture calls drop_all() before EVERY test. So every backend run silently annihilated the real install: accounts, discovered spirits, Ghost Logs, devices, all of it. Found it because /sitemap.xml listed zero entities minutes after I had watched live séances mint real ones. Tests now use TEST_DATABASE_URL, or `<configured-db>_test` derived from it, and refuse to start at all if that ever resolves back to the production URL — this box both serves the app and holds the repo, so "don't run tests in prod" is not a workable guard. Proven: inserted a canary row into production, ran 50 tests, canary survived. Before this it would have been dropped. Also in this commit: SEO (routes/seo.py, lib/pageMeta.ts) - Live /sitemap.xml generated from real entity rows, and /robots.txt, both registered BEFORE the SPA catch-all or they'd be served index.html. Crawlers are disallowed from /seance specifically because the open door provisions a guest on arrival — a crawler would fill the users table with wanderers who never existed. - Per-route <title>, description, canonical and JSON-LD. The Codex is the indexable asset here (every spirit is unique long-form prose) and all of it previously shared one static title, so entities competed with each other instead of ranking. Entities are marked up as fictional Persons so a rich result can never imply a record of a real dead human. - public_base_url setting: absolute URLs for crawlers can't be derived from the request, since behind the tunnel the app only sees an internal host. Camera channel, first half (lib/camera.ts, llm scry path) - OllamaClient.generate() now accepts `images`; the configured chat model (minicpm-v4.5:8b) is vision-capable, so the entity can speak about what the seeker's camera actually shows. Verified against a synthetic room image: it named the pale column and the small red cube, then misread them as oak in a farmhouse parlor — real perception, in character. - Frames are captured only on an explicit act, downscaled to 768px and JPEG-compressed, never stored, and the prompt forbids describing faces or guessing identity. CameraEye carries the same generation guard as the EVP listener so closing during the permission prompt can't leave the camera live after teardown. 338 backend tests pass; 375 frontend; i18n parity holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
106 lines
3.5 KiB
Python
106 lines
3.5 KiB
Python
"""robots.txt and a live sitemap.
|
|
|
|
Served from the backend rather than dropped in `public/` because the
|
|
valuable, indexable surface of this site is the Codex, and the Codex grows
|
|
every time somebody summons. A static sitemap would be stale within an
|
|
hour; this one is generated from the actual entity rows.
|
|
|
|
What is deliberately NOT listed: /seance, /enter, /profile, /log,
|
|
/inventory, /devices — anything per-seeker or interactive. A crawler
|
|
hitting /seance would provision a guest account on arrival (the open
|
|
door), which would fill the users table with wanderers that never
|
|
existed as people. robots.txt disallows those paths for the same reason,
|
|
and the sitemap only advertises pages that are genuinely public,
|
|
stable, and worth a search result: the landing page, the shop, and every
|
|
discovered spirit.
|
|
"""
|
|
|
|
from datetime import datetime, timezone
|
|
|
|
from fastapi import APIRouter, Depends, Response
|
|
from sqlalchemy import select
|
|
from sqlalchemy.ext.asyncio import AsyncSession
|
|
|
|
from app.config import settings
|
|
from app.db import get_db
|
|
from app.models.entity import Entity
|
|
|
|
router = APIRouter(tags=["seo"])
|
|
|
|
# Cap the sitemap so a runaway Codex can't produce a multi-megabyte
|
|
# document. 5k is far inside the 50k/50MB sitemap limit and orders of
|
|
# magnitude beyond what this install will realistically hold.
|
|
MAX_SITEMAP_ENTITIES = 5000
|
|
|
|
# Paths that must never be crawled: they either mutate state (provisioning a
|
|
# guest) or are meaningless without a session.
|
|
DISALLOWED = (
|
|
"/seance",
|
|
"/enter",
|
|
"/profile",
|
|
"/hunters",
|
|
"/log",
|
|
"/inventory",
|
|
"/devices",
|
|
"/api/",
|
|
)
|
|
|
|
|
|
def _base_url() -> str:
|
|
return settings.public_base_url.rstrip("/")
|
|
|
|
|
|
def _iso(dt: datetime | None) -> str:
|
|
value = dt or datetime.now(timezone.utc)
|
|
if value.tzinfo is None:
|
|
value = value.replace(tzinfo=timezone.utc)
|
|
return value.date().isoformat()
|
|
|
|
|
|
def _xml_escape(text: str) -> str:
|
|
return (
|
|
text.replace("&", "&")
|
|
.replace("<", "<")
|
|
.replace(">", ">")
|
|
.replace('"', """)
|
|
)
|
|
|
|
|
|
@router.get("/robots.txt", include_in_schema=False)
|
|
async def robots() -> Response:
|
|
lines = ["User-agent: *"]
|
|
lines += [f"Disallow: {path}" for path in DISALLOWED]
|
|
lines.append("Allow: /")
|
|
lines.append(f"Sitemap: {_base_url()}/sitemap.xml")
|
|
return Response("\n".join(lines) + "\n", media_type="text/plain")
|
|
|
|
|
|
@router.get("/sitemap.xml", include_in_schema=False)
|
|
async def sitemap(db: AsyncSession = Depends(get_db)) -> Response:
|
|
base = _base_url()
|
|
result = await db.execute(
|
|
select(Entity.id, Entity.discovered_at)
|
|
.order_by(Entity.discovered_at.desc())
|
|
.limit(MAX_SITEMAP_ENTITIES)
|
|
)
|
|
entities = result.all()
|
|
|
|
parts = [
|
|
'<?xml version="1.0" encoding="UTF-8"?>',
|
|
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
|
|
f"<url><loc>{base}/</loc><changefreq>daily</changefreq>"
|
|
"<priority>1.0</priority></url>",
|
|
f"<url><loc>{base}/codex</loc><changefreq>hourly</changefreq>"
|
|
"<priority>0.9</priority></url>",
|
|
f"<url><loc>{base}/shop</loc><changefreq>weekly</changefreq>"
|
|
"<priority>0.7</priority></url>",
|
|
]
|
|
for entity_id, discovered_at in entities:
|
|
loc = _xml_escape(f"{base}/codex/{entity_id}")
|
|
parts.append(
|
|
f"<url><loc>{loc}</loc><lastmod>{_iso(discovered_at)}</lastmod>"
|
|
"<changefreq>weekly</changefreq><priority>0.6</priority></url>"
|
|
)
|
|
parts.append("</urlset>")
|
|
return Response("\n".join(parts), media_type="application/xml")
|