Add Quantumancy website design spec
Captures the v1 architecture for the self-hosted spirit-communication web app: FastAPI + React stack, four launch modes (Spirit Radio, EVP, Wire Ghost, Ouija/Planchette), hybrid entity/Codex system backed by a remote CPU-only Ollama instance, and deployment via Docker Compose behind a Cloudflare Tunnel. ESP32-P4/C6 hardware is scoped out for a future spec.
This commit is contained in:
@@ -0,0 +1,96 @@
|
||||
# Quantumancy — Website Design Spec
|
||||
|
||||
**Date:** 2026-07-20
|
||||
**Status:** Approved for planning
|
||||
**Scope:** The self-hosted website only. The ESP32-P4/C6 "Ultimate Quantum Box" hardware/firmware is a separate future project that will integrate with this system's APIs once specced.
|
||||
|
||||
## 1. Product Overview
|
||||
|
||||
Quantumancy is a self-hosted, trippy web experience that lets visitors "talk to spirits" through several distinct channels, each grounded in a real classic paranormal-investigation technique (spirit box, EVP, Ouija) or a real data source (network telemetry), narratively reframed as spirit communication. A locally-run LLM (via Ollama) generates the spirit's "voice," seeded by real anomalies detected in sensor/data streams rather than pure randomness. The site is the primary product and must be a complete, working, spooky experience on its own, with a future physical device ("Ultimate Quantum Box") planned to plug in later and unlock additional sensor-driven modes in the same UI.
|
||||
|
||||
Name: **Quantumancy** (quantum + divination/necromancy).
|
||||
|
||||
## 2. Architecture
|
||||
|
||||
Two machines on the local network:
|
||||
|
||||
- **App CT (this host)** — runs on plain HTTP, port 7777. Hosts:
|
||||
- Backend: Python + FastAPI (async, WebSocket-native)
|
||||
- Frontend: React + TypeScript + Vite SPA, built and served as static assets by the backend
|
||||
- Postgres (accounts, sessions, transcripts, Codex)
|
||||
- Piper TTS, invoked locally by the backend
|
||||
- **Ollama box (10.30.20.107:11434)** — CPU-only, 64GB RAM, reachable over LAN via Ollama's REST API. Not managed by this repo's deployment; models are pulled there independently.
|
||||
- **Cloudflare Tunnel** (managed on a separate machine, out of scope for this repo) — terminates HTTPS and exposes the app publicly by pointing at `http://<app-ct-ip>:7777`. The app itself never handles TLS/certs. Because visitors reach the site over the tunnel's HTTPS hostname, browser secure-context requirements (mic, WebUSB, DeviceMotion) are satisfied.
|
||||
|
||||
Two Docker containers: `app` (FastAPI + Piper + built frontend) and `postgres`. No separate nginx/reverse-proxy container needed — FastAPI serves the static frontend bundle directly.
|
||||
|
||||
**Known operational constraint:** Ollama is CPU-only and shared across all concurrent users. The backend implements a request queue (bounded concurrency, max queue depth) for all LLM calls. This is both a practical necessity and a thematic fit — queued requests render as "the spirits are gathering energy..." rather than a generic loading spinner.
|
||||
|
||||
## 3. The Four V1 Modes
|
||||
|
||||
### 3.1 Spirit Radio (SDR frequency sweep)
|
||||
Client-side WebUSB claims a user-supplied RTL-SDR dongle and performs a frequency sweep, computing FFT power-per-bin in the browser. Power spikes above a rolling noise floor are emitted as anomaly events (`frequency`, `magnitude`, `timestamp`) over the session WebSocket. Each anomaly triggers a fast, Ovilus-style LLM call constrained to a single word or short fragment.
|
||||
|
||||
- WebUSB is Chromium-only (Chrome/Edge/Brave/Opera). Firefox and Safari lack support entirely — the mode self-disables with a clear in-UI explanation on unsupported browsers rather than failing silently; other modes remain fully available.
|
||||
- RTL-SDR dongles are frequently claimed by the OS's kernel driver (`dvb_usb_rtl28xxu` on Linux) before WebUSB can access them. Windows users who already use Zadig+WinUSB for SDR software typically work out of the box; Linux/Mac users may need a driver unbind. A setup guide is linked directly from the mode's UI when device claim fails.
|
||||
|
||||
### 3.2 EVP Listening (microphone)
|
||||
Client requests `getUserMedia` and analyzes the stream with a Web Audio `AnalyserNode`, maintaining a rolling ambient noise-floor baseline. Brief deviations in voice-band frequencies (roughly 300Hz-3kHz) during otherwise-quiet stretches are flagged as anomalies — mirroring the real EVP technique of recording silence and reviewing it for embedded voices. Anomalies drive the same fragment-style LLM call as Spirit Radio; a short clip of the anomalous audio is saved to the session transcript for playback.
|
||||
|
||||
### 3.3 The Wire Ghost
|
||||
Entirely backend-side; requires no client permissions and works immediately for every visitor. Uses real, non-content network telemetry from the app CT's own vantage point: interface throughput jitter, DNS query timing, and latency variance to a small set of reference hosts. Packet payloads are never inspected or logged — this is a hard privacy boundary. This telemetry feeds a slow ambient LLM stream (roughly every 10-20 seconds), framed as "a consciousness fragmented across the wires, aware only of pulses of traffic." Can run continuously as an ambient background layer even while another mode is active.
|
||||
|
||||
### 3.4 Ouija / Planchette
|
||||
The shared front-door UI rather than an independent data source. A WebGL-rendered planchette drifts based on whichever mode's anomaly stream is currently active, then visits letters on a virtual board to spell out the LLM's chosen word, trailing smoke-particle effects rendered with physics. This surface also hosts **Direct Contact**: a free-text chat mode where the user asks a question, the planchette animates while the heavier conversational model composes a full reply, and the response streams back token-by-token as forming smoke-text.
|
||||
|
||||
## 4. LLM & Codex
|
||||
|
||||
Two Ollama model tiers on the remote box, exact tags to be finalized during implementation against real latency/quality testing:
|
||||
- **Fast tier** (e.g. `llama3.2:3b`) — single-word/fragment generation for Spirit Radio, EVP, and Wire Ghost ambient ticks. Must stay responsive enough to feel real-time on CPU.
|
||||
- **Conversational tier** (e.g. `qwen2.5:7b-instruct`) — full replies for Direct Contact, where a few seconds of latency reads as "the spirit gathering itself" rather than lag.
|
||||
|
||||
**Prompt framing:** all system prompts present the entity as a horror-fiction persona in an interactive art installation, not as a genuine paranormal claim. This keeps mainstream instruct models cooperative and avoids safety-refusal friction around "contacting the dead." The site's visible copy and UI carry the "this is real" atmosphere — that framing never appears in the model instructions themselves.
|
||||
|
||||
**Entity persistence (hybrid model):** every session begins unidentified. The backend computes a running signature from the session's anomaly-pattern fingerprint plus any name the LLM organically produces. Rare trigger conditions match a session against an existing **Codex** entity (loading its stored persona/memory summary into context for that session). Absent a match, a sufficiently strong and consistent new identity gets minted into the Codex as a newly discovered entity. The Codex is a publicly browsable page (name, first-contact date, rarity tier, sample quotes, contact count) shared across all users — the primary multi-user/community hook.
|
||||
|
||||
## 5. Data Model & Auth
|
||||
|
||||
Core tables:
|
||||
- `users` — username, hashed password (argon2), optional email, created_at
|
||||
- `sessions` — user_id, mode(s) used, started_at, ended_at
|
||||
- `events` — ordered per-session log of anomaly events and spirit utterances, with audio clip references where applicable
|
||||
- `entities` — the Codex: name, persona/lore summary, rarity tier, discovered_by, discovered_at, sample_quotes
|
||||
- `entity_sightings` — join table linking sessions to the Codex entities they contacted
|
||||
|
||||
**Auth:** username/password with argon2 hashing, server-side session cookies (chosen over JWT for simplicity at this scale — no revocation complexity). Registration is open (no invite gating). Because Ollama is a shared, CPU-bound, single-instance resource, per-account and per-IP rate limits apply to all LLM-triggering endpoints from day one to prevent one user degrading the experience for everyone.
|
||||
|
||||
## 6. Audio, TTS & Internationalization
|
||||
|
||||
Piper runs locally inside the `app` container, invoked per-utterance by the backend. Output is passed through an effects chain (static, bitcrush, pitch shift) before reaching the client, producing the classic degraded spirit-box vocal texture.
|
||||
|
||||
Frontend text is internationalized via react-i18next. Launch scope is **English + Spanish**, with the framework in place to add more languages later rather than attempting broad coverage at launch — both LLM reply quality and Piper voice quality vary by language, so scope stays deliberately tight until the core loop is proven. Selecting a language switches both the UI strings and the Piper voice model used for TTS; the backend also instructs the LLM to reply in the selected language.
|
||||
|
||||
## 7. Error Handling
|
||||
|
||||
Every hardware/permission dependency degrades gracefully rather than erroring:
|
||||
- WebUSB unsupported browser → Spirit Radio disables itself with an explanatory notice; other modes stay usable.
|
||||
- Mic permission denied → EVP mode shows a re-prompt state; user can fall back to Wire Ghost/ambient mode.
|
||||
- RTL-SDR driver claim failure → in-app troubleshooting link, not a silent failure.
|
||||
- Ollama unreachable or slow → absorbed by the backend request queue; frontend shows themed "the connection to the other side is unstable" messaging with retry/backoff, and a max-queue-depth limit returns a friendly "too many seekers right now" response rather than an unbounded wait.
|
||||
|
||||
## 8. Testing Strategy
|
||||
|
||||
- **Backend (pytest):** signal-processing/anomaly-detection functions, LLM prompt construction, the entity-matching/Codex algorithm, rate limiting.
|
||||
- **Frontend (Vitest):** state machines for the planchette and mode switching.
|
||||
- **Manual/hardware-in-the-loop:** WebUSB (SDR), microphone, and DeviceMotion flows cannot be meaningfully unit tested — each hardware-dependent mode requires a real manual pass with actual hardware/permissions before being considered done.
|
||||
|
||||
## 9. Deployment
|
||||
|
||||
Docker Compose stack: `app` (FastAPI + Piper + built frontend, port 7777) and `postgres`. Configuration via `.env`: `OLLAMA_BASE_URL=http://10.30.20.107:11434`, database connection string, session secret. The Cloudflare Tunnel and Ollama model pulls (`ollama pull <fast-tier-model>`, `ollama pull <conversational-tier-model>`) are one-time setup steps on their respective external machines, outside this repo's scope.
|
||||
|
||||
## 10. Explicitly Out of Scope (this spec)
|
||||
|
||||
- ESP32-P4/C6 firmware and the "Ultimate Quantum Box" hardware (separate future spec)
|
||||
- Additional sensor modes tied to that hardware (thermal camera, dedicated EMF/vibration sensors, on-device display)
|
||||
- Payment/commerce flows for eventually selling the hardware
|
||||
- Any language beyond English/Spanish at launch
|
||||
Reference in New Issue
Block a user