Files
qtalker---/backend/tests/test_seo.py
Indiana 16c8bae00d fix: the test suite was destroying the production database
The worst bug of the session. tests/conftest.py built its engine from
settings.database_url — the live database — and an autouse fixture calls
drop_all() before EVERY test. So every backend run silently annihilated the
real install: accounts, discovered spirits, Ghost Logs, devices, all of it.
Found it because /sitemap.xml listed zero entities minutes after I had
watched live séances mint real ones.

Tests now use TEST_DATABASE_URL, or `<configured-db>_test` derived from it,
and refuse to start at all if that ever resolves back to the production URL
— this box both serves the app and holds the repo, so "don't run tests in
prod" is not a workable guard.

Proven: inserted a canary row into production, ran 50 tests, canary
survived. Before this it would have been dropped.

Also in this commit:

SEO (routes/seo.py, lib/pageMeta.ts)
- Live /sitemap.xml generated from real entity rows, and /robots.txt, both
  registered BEFORE the SPA catch-all or they'd be served index.html.
  Crawlers are disallowed from /seance specifically because the open door
  provisions a guest on arrival — a crawler would fill the users table with
  wanderers who never existed.
- Per-route <title>, description, canonical and JSON-LD. The Codex is the
  indexable asset here (every spirit is unique long-form prose) and all of
  it previously shared one static title, so entities competed with each
  other instead of ranking. Entities are marked up as fictional Persons so
  a rich result can never imply a record of a real dead human.
- public_base_url setting: absolute URLs for crawlers can't be derived from
  the request, since behind the tunnel the app only sees an internal host.

Camera channel, first half (lib/camera.ts, llm scry path)
- OllamaClient.generate() now accepts `images`; the configured chat model
  (minicpm-v4.5:8b) is vision-capable, so the entity can speak about what
  the seeker's camera actually shows. Verified against a synthetic room
  image: it named the pale column and the small red cube, then misread them
  as oak in a farmhouse parlor — real perception, in character.
- Frames are captured only on an explicit act, downscaled to 768px and
  JPEG-compressed, never stored, and the prompt forbids describing faces or
  guessing identity. CameraEye carries the same generation guard as the EVP
  listener so closing during the permission prompt can't leave the camera
  live after teardown.

338 backend tests pass; 375 frontend; i18n parity holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 01:44:01 +00:00

93 lines
3.3 KiB
Python

"""robots.txt and sitemap.xml.
The load-bearing property is that crawlers are kept OFF the interactive
routes. /seance provisions a guest account on arrival (the open door), so a
crawler wandering in would create real user rows for visitors who never
existed — these tests pin that it stays disallowed and unlisted.
"""
import xml.etree.ElementTree as ET
import pytest
from app.routes.seo import DISALLOWED
SITEMAP_NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
@pytest.mark.asyncio
async def test_robots_is_plain_text_and_not_the_spa(client):
response = await client.get("/robots.txt")
assert response.status_code == 200
assert response.headers["content-type"].startswith("text/plain")
# If the SPA catch-all had won the route we'd get HTML instead.
assert "<!doctype html" not in response.text.lower()
assert response.text.startswith("User-agent: *")
@pytest.mark.asyncio
async def test_robots_disallows_every_interactive_route(client):
body = await client.get("/robots.txt")
text = body.text
for path in DISALLOWED:
assert f"Disallow: {path}" in text
# The séance is the critical one: crawling it mints guest accounts.
assert "Disallow: /seance" in text
@pytest.mark.asyncio
async def test_robots_advertises_an_absolute_sitemap_url(client):
text = (await client.get("/robots.txt")).text
line = next(ln for ln in text.splitlines() if ln.startswith("Sitemap:"))
url = line.split(" ", 1)[1]
# Must be absolute and public — behind the tunnel the app sees an
# internal host, and a relative or internal URL is useless to a crawler.
assert url.startswith("https://")
assert url.endswith("/sitemap.xml")
@pytest.mark.asyncio
async def test_sitemap_is_valid_xml_with_the_public_pages(client):
response = await client.get("/sitemap.xml")
assert response.status_code == 200
assert "xml" in response.headers["content-type"]
root = ET.fromstring(response.text)
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
assert any(loc.endswith("/") for loc in locs)
assert any(loc.endswith("/codex") for loc in locs)
@pytest.mark.asyncio
async def test_sitemap_never_lists_an_interactive_route(client):
root = ET.fromstring((await client.get("/sitemap.xml")).text)
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
for loc in locs:
for path in DISALLOWED:
assert not loc.endswith(path.rstrip("/")), f"{loc} advertises {path}"
@pytest.mark.asyncio
async def test_sitemap_lists_discovered_spirits(client, db_session):
from app.models.entity import Entity
entity = Entity(
name="Sitemap Test Spirit",
epithet="the Indexed",
persona="A spirit that exists to be crawled.",
signature="sitemap-sig-1",
)
db_session.add(entity)
await db_session.commit()
await db_session.refresh(entity)
root = ET.fromstring((await client.get("/sitemap.xml")).text)
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
assert any(loc.endswith(f"/codex/{entity.id}") for loc in locs)
@pytest.mark.asyncio
async def test_sitemap_is_valid_with_an_empty_codex(client):
# A brand-new install must still serve a well-formed sitemap.
root = ET.fromstring((await client.get("/sitemap.xml")).text)
assert root.tag.endswith("urlset")