drjones fc1e8b9c33 Add real video tagging via minicpm-v4.6:1b, fix broken video keyword writes
Videos were never part of the automated pipeline before -- only reachable
one at a time via a manual "Tag Now" button, and even that just grabbed
one static frame and ran it through the photo model. Now videos are
first-class:

- New extract_video_frames_b64(): samples up to 6 frames spread across
  the clip's duration and passes them all to the video model in one
  call, so it sees actual motion/progression instead of one snapshot.
  Verified live: two different real test videos got distinct,
  content-aware descriptions that correctly named what was actually
  happening in each, not generic placeholders.
- ollama_generate() now accepts a list of images (photos still pass a
  single one, unchanged) so the same call path serves both.
- process_loop merges S.video_files into the same pending queue as
  photos, routes videos to a separate configurable video model
  (default minicpm-v4.6:1b, a small dedicated vision model) and prompt,
  skipping the photo/screenshot router entirely. redo_single (the
  lightbox "AI Redo" button) updated the same way for consistency.
- S.total_images is now set to the actual combined pending count for
  the run so the progress bar/ETA reflect videos too, not just photos.
- FOUND WHILE TESTING: write_metadata's video branch used "-Keywords"
  for the category/photon-tagged marker, which is a silent no-op on
  QuickTime/.mov files (exiftool has no mapping for it there) --
  confirmed by direct testing. Every video "tagged" before this would
  have gotten a description but never an actual category keyword
  embedded. Switched to XMP-dc:Subject (the same tag family used for
  photos, which QuickTime containers do support via an embedded XMP
  packet) -- verified the category now lands and stays idempotent
  across repeat writes.
- New frontend controls: video model dropdown and a "tag videos too"
  toggle (on by default) in Engine Configurations.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 18:28:04 -07:00
2026-07-23 03:59:58 -07:00

PHOTON — local photo intelligence console

A local web application that scans your photo collection, uses local Ollama vision models to describe + categorize each photo, and embeds standard metadata inside photo files so everything becomes searchable across Spotlight, Finder, Photos, Lightroom, etc.

Original photos remain completely untouched unless you explicitly choose to edit or delete them.


Quick Start (Zero-Config)

Drop this directory inside the photo collection you want to organize, then run:

cd "your-photo-library/photon"
python3 server.py
# Open http://localhost:8765 in your browser

Automatic Folder Detection & Overrides

  • Auto-Detection: Scans the parent directory of wherever server.py lives.
  • Persistence: Remembers your last scanned folder (photon_folder.json).
  • CLI & Environment Overrides:
    python3 server.py /path/to/photos                       # CLI argument override
    PHOTON_FOLDER=/path/to/photos python3 server.py         # Environment variable override
    PHOTON_PORT=8766 python3 server.py                      # Custom port
    

Prerequisites

  • Ollama running locally with a vision model (e.g. ollama pull qwen3.5:9b or ollama pull qwen3.5:4b).
  • exiftool (installed via Homebrew: brew install exiftool).
  • macOS (sips built-in for fast downscaling & pixel integrity verification).
  • ffmpeg / ffprobe (optional: for video frame extraction & video thumbnails).

Key Features & Capabilities

1. Multithreaded 3-Stage Pipeline

  1. Downscale Stage: Uses macOS sips to create fast temp copies (original files are read-only).
  2. Inference Stage: Calls local Ollama vision models to determine {description, category}.
  3. Write Stage: Uses exiftool to embed metadata with atomic renames and date preservation (-P).

2. Search & Interactive Library Browser

  • Instant Search: Full-text keyword search across descriptions, categories, filenames, and paths.
  • Filter Chips: 1-click taxonomy pills (People, Animals, Screenshots, Vehicles, Objects, etc.).
  • Sub-Views:
    • Tagged: Explore and filter all processed photos.
    • Untagged: Browse photos awaiting tagging.
    • Failed: View and retry failed operations.
    • Videos: Frame extraction, AI tagging, and native video player.
    • Duplicates: Perceptual hash index (pHash) visual duplicate grouping.

3. Full-Screen Lightbox & Organic Browsing

  • Snappy Viewer: Full-resolution image/video lightbox with metadata inspector.
  • 0ms Image Prefetching: Pre-caches adjacent images in memory for instant switching.
  • Keyboard Shortcuts:
    • ← / → : Navigate previous / next photo.
    • Esc : Close Lightbox.
    • Delete / Backspace : Delete current photo on disk.

4. Disk Photo Deletion & Management

  • Single & Bulk Deletion: Click "Delete Photo" or select multiple photos to permanently delete them on disk (or send to macOS Trash).
  • Automated Cleanup: Deleting a photo purges its entry from photon_journal.jsonl, removes _organized/ symlinks, and clears thumbnail & view caches.

5. Bulk Operations & Export

  • Select Mode: Range selection via Shift-Click or "Select All Matching".
  • Exporting: Export catalog metadata to CSV or JSON.
  • ZIP Downloads: Stream original-quality files into a single ZIP archive.

6. Trust & Safety Safeguards

  • Verify Pixel Integrity: Option to double-hash image pixels via raw BMP conversions before and after writes. Guarantees 100% zero image corruption.
  • Organized Symlinks: Generates relative portable Finder aliases in [folder]/_organized/[category]/[name].
  • Undo All Tags: One-click exiftool pass to cleanly remove all photon-tagged metadata and categories.
  • Persistent Failure Retries: Saves errored paths to photon_failures.jsonl with 1-click retry.

7. Quality Control & Custom Taxonomies

  • Blind Accuracy Grader: Interactive 100-photo audit mode to score description quality.
  • Custom Categories: Edit, add, or customize category schemas (photon_categories.json).

Standard Metadata Specifications

exiftool embeds the following tags into image files:

  • EXIF:ImageDescription, IPTC:Caption-Abstract, XMP-dc:Description — Description string
  • XMP-dc:Subject, IPTC:Keywords — Category string + photon-tagged keyword

License & Safety Notice

Original photo files are never deleted or modified unless you explicitly trigger Delete Photo, Metadata Writing, or Undo All Tags.

Description
Ollama photo organizer runs locally with ocr model and a vision model.
Readme 1 MiB
Languages
Python 53%
HTML 46.7%
Shell 0.3%