The worst bug of the session. tests/conftest.py built its engine from settings.database_url — the live database — and an autouse fixture calls drop_all() before EVERY test. So every backend run silently annihilated the real install: accounts, discovered spirits, Ghost Logs, devices, all of it. Found it because /sitemap.xml listed zero entities minutes after I had watched live séances mint real ones. Tests now use TEST_DATABASE_URL, or `<configured-db>_test` derived from it, and refuse to start at all if that ever resolves back to the production URL — this box both serves the app and holds the repo, so "don't run tests in prod" is not a workable guard. Proven: inserted a canary row into production, ran 50 tests, canary survived. Before this it would have been dropped. Also in this commit: SEO (routes/seo.py, lib/pageMeta.ts) - Live /sitemap.xml generated from real entity rows, and /robots.txt, both registered BEFORE the SPA catch-all or they'd be served index.html. Crawlers are disallowed from /seance specifically because the open door provisions a guest on arrival — a crawler would fill the users table with wanderers who never existed. - Per-route <title>, description, canonical and JSON-LD. The Codex is the indexable asset here (every spirit is unique long-form prose) and all of it previously shared one static title, so entities competed with each other instead of ranking. Entities are marked up as fictional Persons so a rich result can never imply a record of a real dead human. - public_base_url setting: absolute URLs for crawlers can't be derived from the request, since behind the tunnel the app only sees an internal host. Camera channel, first half (lib/camera.ts, llm scry path) - OllamaClient.generate() now accepts `images`; the configured chat model (minicpm-v4.5:8b) is vision-capable, so the entity can speak about what the seeker's camera actually shows. Verified against a synthetic room image: it named the pale column and the small red cube, then misread them as oak in a farmhouse parlor — real perception, in character. - Frames are captured only on an explicit act, downscaled to 768px and JPEG-compressed, never stored, and the prompt forbids describing faces or guessing identity. CameraEye carries the same generation guard as the EVP listener so closing during the permission prompt can't leave the camera live after teardown. 338 backend tests pass; 375 frontend; i18n parity holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
93 lines
3.3 KiB
Python
93 lines
3.3 KiB
Python
"""robots.txt and sitemap.xml.
|
|
|
|
The load-bearing property is that crawlers are kept OFF the interactive
|
|
routes. /seance provisions a guest account on arrival (the open door), so a
|
|
crawler wandering in would create real user rows for visitors who never
|
|
existed — these tests pin that it stays disallowed and unlisted.
|
|
"""
|
|
|
|
import xml.etree.ElementTree as ET
|
|
|
|
import pytest
|
|
|
|
from app.routes.seo import DISALLOWED
|
|
|
|
SITEMAP_NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_robots_is_plain_text_and_not_the_spa(client):
|
|
response = await client.get("/robots.txt")
|
|
assert response.status_code == 200
|
|
assert response.headers["content-type"].startswith("text/plain")
|
|
# If the SPA catch-all had won the route we'd get HTML instead.
|
|
assert "<!doctype html" not in response.text.lower()
|
|
assert response.text.startswith("User-agent: *")
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_robots_disallows_every_interactive_route(client):
|
|
body = await client.get("/robots.txt")
|
|
text = body.text
|
|
for path in DISALLOWED:
|
|
assert f"Disallow: {path}" in text
|
|
# The séance is the critical one: crawling it mints guest accounts.
|
|
assert "Disallow: /seance" in text
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_robots_advertises_an_absolute_sitemap_url(client):
|
|
text = (await client.get("/robots.txt")).text
|
|
line = next(ln for ln in text.splitlines() if ln.startswith("Sitemap:"))
|
|
url = line.split(" ", 1)[1]
|
|
# Must be absolute and public — behind the tunnel the app sees an
|
|
# internal host, and a relative or internal URL is useless to a crawler.
|
|
assert url.startswith("https://")
|
|
assert url.endswith("/sitemap.xml")
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_sitemap_is_valid_xml_with_the_public_pages(client):
|
|
response = await client.get("/sitemap.xml")
|
|
assert response.status_code == 200
|
|
assert "xml" in response.headers["content-type"]
|
|
root = ET.fromstring(response.text)
|
|
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
|
|
assert any(loc.endswith("/") for loc in locs)
|
|
assert any(loc.endswith("/codex") for loc in locs)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_sitemap_never_lists_an_interactive_route(client):
|
|
root = ET.fromstring((await client.get("/sitemap.xml")).text)
|
|
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
|
|
for loc in locs:
|
|
for path in DISALLOWED:
|
|
assert not loc.endswith(path.rstrip("/")), f"{loc} advertises {path}"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_sitemap_lists_discovered_spirits(client, db_session):
|
|
from app.models.entity import Entity
|
|
|
|
entity = Entity(
|
|
name="Sitemap Test Spirit",
|
|
epithet="the Indexed",
|
|
persona="A spirit that exists to be crawled.",
|
|
signature="sitemap-sig-1",
|
|
)
|
|
db_session.add(entity)
|
|
await db_session.commit()
|
|
await db_session.refresh(entity)
|
|
|
|
root = ET.fromstring((await client.get("/sitemap.xml")).text)
|
|
locs = [el.text for el in root.findall(".//sm:loc", SITEMAP_NS)]
|
|
assert any(loc.endswith(f"/codex/{entity.id}") for loc in locs)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_sitemap_is_valid_with_an_empty_codex(client):
|
|
# A brand-new install must still serve a well-formed sitemap.
|
|
root = ET.fromstring((await client.get("/sitemap.xml")).text)
|
|
assert root.tag.endswith("urlset")
|