drjones adbe4d8f24 Initial commit of PHOTON photo intelligence app
Ollama-based photo tagging with exiftool metadata writes.
2026-07-18 09:21:10 -07:00

PHOTON — local photo intelligence console

A local web app that walks through a photo folder, has an Ollama vision model describe + categorize each photo, and embeds the result as standard metadata inside the photo file so everything becomes searchable (Spotlight, Photos, Lightroom, etc.). Nothing is ever deleted, moved, or renamed.

Run it

cd "/Users/drjones/photo ollama app organizer"
python3 server.py
# then open http://localhost:8765

Requires: Ollama running with a vision model, exiftool (installed via brew), macOS (sips is used for fast downscaling).

How it works

  1. Scan — recursively finds images (jpg/jpeg/png/heic/tiff/webp/bmp). Videos are counted but skipped. Hidden files and ._* AppleDouble sidecars are never touched.
  2. Analyze — each photo is downscaled with sips to a temp copy (original is only ever read), sent to the chosen Ollama vision model with a JSON schema that forces {description, category} output.
  3. Write — exiftool embeds:
    • EXIF:ImageDescription, IPTC:Caption-Abstract, XMP-dc:Description — the description
    • XMP-dc:Subject + IPTC:Keywords — the category, plus a photon-tagged marker
    • Writes use exiftool's temp-file + atomic-rename mode; file dates preserved with -P.
  4. Journal — every processed photo is appended to photon_journal.jsonl (path, description, category, model, timing). Restarting the app resumes where it left off ("skip already tagged").

The 10 categories

People · Animals · Food & Drink · Nature & Outdoors · City & Buildings · Vehicles · Screenshots & Documents · Events & Parties · Objects & Stuff · Art & Miscellaneous

With the smart router toggle on, a fast scout model (glm-ocr, 1.1B) first classifies each image as screenshot or photo, then hands it to the right describer with a specialized prompt. Screenshots also get an OCR text embed: glm-ocr transcribes the visible words and they're appended to the description (… | text: …), so you can find a screenshot by searching the exact words in it.

Tested defaults: scout glm-ocr (6/6 routing accuracy) → describer qwen3.5:4b for both branches (8/8 accuracy, reads product labels correctly).

Settings that affect speed

Setting Effect
Vision model glm-ocr (1.1B) ≈ 7 s/photo; qwen3.5:9b slower but smarter
Image feed resolution 512 px is fastest; originals are untouched either way
Description length brief/standard/detailed — caps the model's output tokens
Keep-alive "forever" keeps the model in RAM between photos (fastest)
Dry run full pipeline but no metadata written
Keep _original backups exiftool keeps a backup copy of every file (doubles disk usage)

Searching afterwards

Spotlight: just type a word from a description in Finder search. Or from terminal: mdfind -onlyin "/Volumes/sanD/allphotos from phone" "scooter" Or grep the journal: grep -i scooter photon_journal.jsonl

Description
Ollama photo organizer runs locally with ocr model and a vision model.
Readme 1 MiB
Languages
Python 53%
HTML 46.7%
Shell 0.3%