3.0 KiB
PHOTON — local photo intelligence console
A local web app that walks through a photo folder, has an Ollama vision model describe + categorize each photo, and embeds the result as standard metadata inside the photo file so everything becomes searchable (Spotlight, Photos, Lightroom, etc.). Nothing is ever deleted, moved, or renamed.
Run it
cd "/Users/drjones/photo ollama app organizer"
python3 server.py
# then open http://localhost:8765
Requires: Ollama running with a vision model, exiftool (installed via brew),
macOS (sips is used for fast downscaling).
How it works
- Scan — recursively finds images (
jpg/jpeg/png/heic/tiff/webp/bmp). Videos are counted but skipped. Hidden files and._*AppleDouble sidecars are never touched. - Analyze — each photo is downscaled with
sipsto a temp copy (original is only ever read), sent to the chosen Ollama vision model with a JSON schema that forces{description, category}output. - Write —
exiftoolembeds:EXIF:ImageDescription,IPTC:Caption-Abstract,XMP-dc:Description— the descriptionXMP-dc:Subject+IPTC:Keywords— the category, plus aphoton-taggedmarker- Writes use exiftool's temp-file + atomic-rename mode; file dates preserved with
-P.
- Journal — every processed photo is appended to
photon_journal.jsonl(path, description, category, model, timing). Restarting the app resumes where it left off ("skip already tagged").
The 10 categories
People · Animals · Food & Drink · Nature & Outdoors · City & Buildings · Vehicles · Screenshots & Documents · Events & Parties · Objects & Stuff · Art & Miscellaneous
Smart router (recommended)
With the smart router toggle on, a fast scout model (glm-ocr, 1.1B) first
classifies each image as screenshot or photo, then hands it to the right
describer with a specialized prompt. Screenshots also get an OCR text embed:
glm-ocr transcribes the visible words and they're appended to the description
(… | text: …), so you can find a screenshot by searching the exact words in it.
Tested defaults: scout glm-ocr (6/6 routing accuracy) → describer
qwen3.5:4b for both branches (8/8 accuracy, reads product labels correctly).
Settings that affect speed
| Setting | Effect |
|---|---|
| Vision model | glm-ocr (1.1B) ≈ 7 s/photo; qwen3.5:9b slower but smarter |
| Image feed resolution | 512 px is fastest; originals are untouched either way |
| Description length | brief/standard/detailed — caps the model's output tokens |
| Keep-alive | "forever" keeps the model in RAM between photos (fastest) |
| Dry run | full pipeline but no metadata written |
Keep _original backups |
exiftool keeps a backup copy of every file (doubles disk usage) |
Searching afterwards
Spotlight: just type a word from a description in Finder search.
Or from terminal: mdfind -onlyin "/Volumes/sanD/allphotos from phone" "scooter"
Or grep the journal: grep -i scooter photon_journal.jsonl