Files
photon---photo-intelligence/README.md
drjones 5208c5ecef Add full-text photo library, bulk select/download, duplicate detection, video support, and zero-config folder detection
- Search & Edit becomes a multi-view library: Tagged/Untagged/Failed/Videos/Duplicates
- Bulk select (click, shift-click range, select-all-matching) with bulk recategorize,
  CSV/JSON export, and original-quality ZIP download (verified byte-identical)
- Video browsing/tagging via ffmpeg frame extraction, native playback with range support
- Duplicate detection via a lightweight perceptual-hash index (no new dependencies)
- Full-resolution photo/video viewer with editable metadata, prev/next navigation
- Date range, file-type, and category quick-filters; saved searches (localStorage)
- Appearance settings: accent color picker, grid density, results-per-page; full dark mode
- Auto-detects its photo folder (CLI arg > env var > last used > parent directory),
  so the app can be dropped into any photo collection and just work
2026-07-18 19:43:26 -07:00

5.5 KiB

PHOTON — local photo intelligence console

A local web app that walks through a photo folder, has an Ollama vision model describe + categorize each photo, and embeds the result as standard metadata inside the photo file so everything becomes searchable (Spotlight, Photos, Lightroom, etc.). Nothing is ever deleted, moved, or renamed.

Run it — zero config

Drop this whole folder inside the photo collection you want to organize, then just run it:

cd "your-photo-library/photon"
python3 server.py
# then open http://localhost:8765

It auto-detects the folder to scan as the parent of wherever this app lives — so if you dropped it into ~/Pictures/Vacation2026/photon, it defaults to scanning ~/Pictures/Vacation2026. The folder it last scanned is also remembered automatically (photon_folder.json), so on every future launch it just picks up where you left off — no retyping paths.

Want to point it somewhere else? Any of these work, in priority order:

python3 server.py /path/to/photos        # one-off CLI override
PHOTON_FOLDER=/path/to/photos python3 server.py   # env var override

Or just type a new path into the folder field in the Console tab and hit Scan Directory — that becomes the new remembered default too. Running two libraries at once? PHOTON_PORT=8766 python3 server.py avoids a port clash.

Requires: Ollama running with a vision model, exiftool (installed via brew), macOS (sips for fast downscaling), and ffmpeg/ffprobe (for video thumbnails and tagging — optional, everything else works without it).

How it works

  1. Scan — recursively finds images (jpg/jpeg/png/heic/tiff/webp/bmp). Videos are counted but skipped. Hidden files and ._* AppleDouble sidecars are never touched.
  2. Analyze — each photo is downscaled with sips to a temp copy (original is only ever read), sent to the chosen Ollama vision model with a JSON schema that forces {description, category} output.
  3. Write — exiftool embeds:
    • EXIF:ImageDescription, IPTC:Caption-Abstract, XMP-dc:Description — the description
    • XMP-dc:Subject + IPTC:Keywords — the category, plus a photon-tagged marker
    • Writes use exiftool's temp-file + atomic-rename mode; file dates preserved with -P.
  4. Journal — every processed photo is appended to photon_journal.jsonl (path, description, category, model, timing). Restarting the app resumes where it left off ("skip already tagged").

The 10 categories

People · Animals · Food & Drink · Nature & Outdoors · City & Buildings · Vehicles · Screenshots & Documents · Events & Parties · Objects & Stuff · Art & Miscellaneous

With the smart router toggle on, a fast scout model (glm-ocr, 1.1B) first classifies each image as screenshot or photo, then hands it to the right describer with a specialized prompt. Screenshots also get an OCR text embed: glm-ocr transcribes the visible words and they're appended to the description (… | text: …), so you can find a screenshot by searching the exact words in it.

Tested defaults: scout glm-ocr (6/6 routing accuracy) → describer qwen3.5:4b for both branches (8/8 accuracy, reads product labels correctly).

Settings that affect speed

Setting Effect
Vision model glm-ocr (1.1B) ≈ 7 s/photo; qwen3.5:9b slower but smarter
Image feed resolution 512 px is fastest; originals are untouched either way
Description length brief/standard/detailed — caps the model's output tokens
Keep-alive "forever" keeps the model in RAM between photos (fastest)
Dry run full pipeline but no metadata written
Keep _original backups exiftool keeps a backup copy of every file (doubles disk usage)

Searching afterwards

Spotlight: just type a word from a description in Finder search. Or from terminal: mdfind -onlyin "/path/to/your/photos" "scooter" Or grep the journal: grep -i scooter photon_journal.jsonl

Or use the app itself — the Search & Edit tab is a full photo library browser:

  • Tagged — instant multi-word search across description/category/filename/path, filter by category (click a chip), date range, or file type (photos/videos), sort by newest/name/category. Click any result for a full-resolution viewer + editor with Save / AI Redo / Reveal-in-Finder.
  • Untagged / Failed — see what's left to do or what errored, tag or retry one at a time without leaving the grid.
  • Videos — thumbnails via ffmpeg frame-grab, native playback, AI tagging off an extracted frame.
  • Duplicates — one-time background scan builds a perceptual-hash index and groups exact visual matches (re-saves, burst duplicates). No auto-delete — just Reveal-in-Finder per copy so you decide.

Selecting & downloading photos in bulk: click Select to enter select mode, then:

  • click a photo to toggle it, shift-click to select a range
  • Select all loaded / Select all N matching (grabs everything matching your current search, not just what's rendered) / Clear selection
  • Download selected (ZIP) or Download all matches (ZIP) — streams a real zip of the original files, byte-for-byte, no re-compression (capped at 3000 files per zip)
  • Apply category to selected, or Export the set as CSV/JSON

Appearance

Settings & Safety → Appearance: pick an accent color (native color picker), grid density (compact/comfortable/large), and results-per-page. Saved per-browser, doesn't touch anything on disk.