Files
qtalker---/docs/superpowers/specs/2026-07-20-quantumancy-website-design.md
Indiana 709f6b336f Add Quantumancy website design spec
Captures the v1 architecture for the self-hosted spirit-communication
web app: FastAPI + React stack, four launch modes (Spirit Radio, EVP,
Wire Ghost, Ouija/Planchette), hybrid entity/Codex system backed by a
remote CPU-only Ollama instance, and deployment via Docker Compose
behind a Cloudflare Tunnel. ESP32-P4/C6 hardware is scoped out for a
future spec.
2026-07-20 05:53:01 +00:00

10 KiB

Quantumancy — Website Design Spec

Date: 2026-07-20 Status: Approved for planning Scope: The self-hosted website only. The ESP32-P4/C6 "Ultimate Quantum Box" hardware/firmware is a separate future project that will integrate with this system's APIs once specced.

1. Product Overview

Quantumancy is a self-hosted, trippy web experience that lets visitors "talk to spirits" through several distinct channels, each grounded in a real classic paranormal-investigation technique (spirit box, EVP, Ouija) or a real data source (network telemetry), narratively reframed as spirit communication. A locally-run LLM (via Ollama) generates the spirit's "voice," seeded by real anomalies detected in sensor/data streams rather than pure randomness. The site is the primary product and must be a complete, working, spooky experience on its own, with a future physical device ("Ultimate Quantum Box") planned to plug in later and unlock additional sensor-driven modes in the same UI.

Name: Quantumancy (quantum + divination/necromancy).

2. Architecture

Two machines on the local network:

  • App CT (this host) — runs on plain HTTP, port 7777. Hosts:
    • Backend: Python + FastAPI (async, WebSocket-native)
    • Frontend: React + TypeScript + Vite SPA, built and served as static assets by the backend
    • Postgres (accounts, sessions, transcripts, Codex)
    • Piper TTS, invoked locally by the backend
  • Ollama box (10.30.20.107:11434) — CPU-only, 64GB RAM, reachable over LAN via Ollama's REST API. Not managed by this repo's deployment; models are pulled there independently.
  • Cloudflare Tunnel (managed on a separate machine, out of scope for this repo) — terminates HTTPS and exposes the app publicly by pointing at http://<app-ct-ip>:7777. The app itself never handles TLS/certs. Because visitors reach the site over the tunnel's HTTPS hostname, browser secure-context requirements (mic, WebUSB, DeviceMotion) are satisfied.

Two Docker containers: app (FastAPI + Piper + built frontend) and postgres. No separate nginx/reverse-proxy container needed — FastAPI serves the static frontend bundle directly.

Known operational constraint: Ollama is CPU-only and shared across all concurrent users. The backend implements a request queue (bounded concurrency, max queue depth) for all LLM calls. This is both a practical necessity and a thematic fit — queued requests render as "the spirits are gathering energy..." rather than a generic loading spinner.

3. The Four V1 Modes

3.1 Spirit Radio (SDR frequency sweep)

Client-side WebUSB claims a user-supplied RTL-SDR dongle and performs a frequency sweep, computing FFT power-per-bin in the browser. Power spikes above a rolling noise floor are emitted as anomaly events (frequency, magnitude, timestamp) over the session WebSocket. Each anomaly triggers a fast, Ovilus-style LLM call constrained to a single word or short fragment.

  • WebUSB is Chromium-only (Chrome/Edge/Brave/Opera). Firefox and Safari lack support entirely — the mode self-disables with a clear in-UI explanation on unsupported browsers rather than failing silently; other modes remain fully available.
  • RTL-SDR dongles are frequently claimed by the OS's kernel driver (dvb_usb_rtl28xxu on Linux) before WebUSB can access them. Windows users who already use Zadig+WinUSB for SDR software typically work out of the box; Linux/Mac users may need a driver unbind. A setup guide is linked directly from the mode's UI when device claim fails.

3.2 EVP Listening (microphone)

Client requests getUserMedia and analyzes the stream with a Web Audio AnalyserNode, maintaining a rolling ambient noise-floor baseline. Brief deviations in voice-band frequencies (roughly 300Hz-3kHz) during otherwise-quiet stretches are flagged as anomalies — mirroring the real EVP technique of recording silence and reviewing it for embedded voices. Anomalies drive the same fragment-style LLM call as Spirit Radio; a short clip of the anomalous audio is saved to the session transcript for playback.

3.3 The Wire Ghost

Entirely backend-side; requires no client permissions and works immediately for every visitor. Uses real, non-content network telemetry from the app CT's own vantage point: interface throughput jitter, DNS query timing, and latency variance to a small set of reference hosts. Packet payloads are never inspected or logged — this is a hard privacy boundary. This telemetry feeds a slow ambient LLM stream (roughly every 10-20 seconds), framed as "a consciousness fragmented across the wires, aware only of pulses of traffic." Can run continuously as an ambient background layer even while another mode is active.

3.4 Ouija / Planchette

The shared front-door UI rather than an independent data source. A WebGL-rendered planchette drifts based on whichever mode's anomaly stream is currently active, then visits letters on a virtual board to spell out the LLM's chosen word, trailing smoke-particle effects rendered with physics. This surface also hosts Direct Contact: a free-text chat mode where the user asks a question, the planchette animates while the heavier conversational model composes a full reply, and the response streams back token-by-token as forming smoke-text.

4. LLM & Codex

Two Ollama model tiers on the remote box, exact tags to be finalized during implementation against real latency/quality testing:

  • Fast tier (e.g. llama3.2:3b) — single-word/fragment generation for Spirit Radio, EVP, and Wire Ghost ambient ticks. Must stay responsive enough to feel real-time on CPU.
  • Conversational tier (e.g. qwen2.5:7b-instruct) — full replies for Direct Contact, where a few seconds of latency reads as "the spirit gathering itself" rather than lag.

Prompt framing: all system prompts present the entity as a horror-fiction persona in an interactive art installation, not as a genuine paranormal claim. This keeps mainstream instruct models cooperative and avoids safety-refusal friction around "contacting the dead." The site's visible copy and UI carry the "this is real" atmosphere — that framing never appears in the model instructions themselves.

Entity persistence (hybrid model): every session begins unidentified. The backend computes a running signature from the session's anomaly-pattern fingerprint plus any name the LLM organically produces. Rare trigger conditions match a session against an existing Codex entity (loading its stored persona/memory summary into context for that session). Absent a match, a sufficiently strong and consistent new identity gets minted into the Codex as a newly discovered entity. The Codex is a publicly browsable page (name, first-contact date, rarity tier, sample quotes, contact count) shared across all users — the primary multi-user/community hook.

5. Data Model & Auth

Core tables:

  • users — username, hashed password (argon2), optional email, created_at
  • sessions — user_id, mode(s) used, started_at, ended_at
  • events — ordered per-session log of anomaly events and spirit utterances, with audio clip references where applicable
  • entities — the Codex: name, persona/lore summary, rarity tier, discovered_by, discovered_at, sample_quotes
  • entity_sightings — join table linking sessions to the Codex entities they contacted

Auth: username/password with argon2 hashing, server-side session cookies (chosen over JWT for simplicity at this scale — no revocation complexity). Registration is open (no invite gating). Because Ollama is a shared, CPU-bound, single-instance resource, per-account and per-IP rate limits apply to all LLM-triggering endpoints from day one to prevent one user degrading the experience for everyone.

6. Audio, TTS & Internationalization

Piper runs locally inside the app container, invoked per-utterance by the backend. Output is passed through an effects chain (static, bitcrush, pitch shift) before reaching the client, producing the classic degraded spirit-box vocal texture.

Frontend text is internationalized via react-i18next. Launch scope is English + Spanish, with the framework in place to add more languages later rather than attempting broad coverage at launch — both LLM reply quality and Piper voice quality vary by language, so scope stays deliberately tight until the core loop is proven. Selecting a language switches both the UI strings and the Piper voice model used for TTS; the backend also instructs the LLM to reply in the selected language.

7. Error Handling

Every hardware/permission dependency degrades gracefully rather than erroring:

  • WebUSB unsupported browser → Spirit Radio disables itself with an explanatory notice; other modes stay usable.
  • Mic permission denied → EVP mode shows a re-prompt state; user can fall back to Wire Ghost/ambient mode.
  • RTL-SDR driver claim failure → in-app troubleshooting link, not a silent failure.
  • Ollama unreachable or slow → absorbed by the backend request queue; frontend shows themed "the connection to the other side is unstable" messaging with retry/backoff, and a max-queue-depth limit returns a friendly "too many seekers right now" response rather than an unbounded wait.

8. Testing Strategy

  • Backend (pytest): signal-processing/anomaly-detection functions, LLM prompt construction, the entity-matching/Codex algorithm, rate limiting.
  • Frontend (Vitest): state machines for the planchette and mode switching.
  • Manual/hardware-in-the-loop: WebUSB (SDR), microphone, and DeviceMotion flows cannot be meaningfully unit tested — each hardware-dependent mode requires a real manual pass with actual hardware/permissions before being considered done.

9. Deployment

Docker Compose stack: app (FastAPI + Piper + built frontend, port 7777) and postgres. Configuration via .env: OLLAMA_BASE_URL=http://10.30.20.107:11434, database connection string, session secret. The Cloudflare Tunnel and Ollama model pulls (ollama pull <fast-tier-model>, ollama pull <conversational-tier-model>) are one-time setup steps on their respective external machines, outside this repo's scope.

10. Explicitly Out of Scope (this spec)

  • ESP32-P4/C6 firmware and the "Ultimate Quantum Box" hardware (separate future spec)
  • Additional sensor modes tied to that hardware (thermal camera, dedicated EMF/vibration sensors, on-device display)
  • Payment/commerce flows for eventually selling the hardware
  • Any language beyond English/Spanish at launch