Add full-text photo library, bulk select/download, duplicate detection, video support, and zero-config folder detection

- Search & Edit becomes a multi-view library: Tagged/Untagged/Failed/Videos/Duplicates
- Bulk select (click, shift-click range, select-all-matching) with bulk recategorize,
  CSV/JSON export, and original-quality ZIP download (verified byte-identical)
- Video browsing/tagging via ffmpeg frame extraction, native playback with range support
- Duplicate detection via a lightweight perceptual-hash index (no new dependencies)
- Full-resolution photo/video viewer with editable metadata, prev/next navigation
- Date range, file-type, and category quick-filters; saved searches (localStorage)
- Appearance settings: accent color picker, grid density, results-per-page; full dark mode
- Auto-detects its photo folder (CLI arg > env var > last used > parent directory),
  so the app can be dropped into any photo collection and just work
This commit is contained in:
drjones
2026-07-18 19:43:26 -07:00
parent adbe4d8f24
commit 5208c5ecef
4 changed files with 1731 additions and 208 deletions

6
.gitignore vendored
View File

@@ -1,3 +1,9 @@
__pycache__/ __pycache__/
*.pyc *.pyc
photon_journal.jsonl photon_journal.jsonl
photon_failures.jsonl
photon_audits.jsonl
photon_phash.json
photon_folder.json
photon_categories.json
.claude/

View File

@@ -5,16 +5,35 @@ describe + categorize each photo, and embeds the result as standard metadata
inside the photo file so everything becomes searchable (Spotlight, Photos, inside the photo file so everything becomes searchable (Spotlight, Photos,
Lightroom, etc.). Nothing is ever deleted, moved, or renamed. Lightroom, etc.). Nothing is ever deleted, moved, or renamed.
## Run it ## Run it — zero config
Drop this whole folder *inside* the photo collection you want to organize,
then just run it:
```bash ```bash
cd "/Users/drjones/photo ollama app organizer" cd "your-photo-library/photon"
python3 server.py python3 server.py
# then open http://localhost:8765 # then open http://localhost:8765
``` ```
Requires: Ollama running with a vision model, `exiftool` (installed via brew), It auto-detects the folder to scan as **the parent of wherever this app
macOS (`sips` is used for fast downscaling). lives** — so if you dropped it into `~/Pictures/Vacation2026/photon`, it
defaults to scanning `~/Pictures/Vacation2026`. The folder it last scanned is
also remembered automatically (`photon_folder.json`), so on every future
launch it just picks up where you left off — no retyping paths.
Want to point it somewhere else? Any of these work, in priority order:
```bash
python3 server.py /path/to/photos # one-off CLI override
PHOTON_FOLDER=/path/to/photos python3 server.py # env var override
```
Or just type a new path into the folder field in the Console tab and hit
**Scan Directory** — that becomes the new remembered default too. Running
two libraries at once? `PHOTON_PORT=8766 python3 server.py` avoids a port clash.
Requires: Ollama running with a vision model, `exiftool` (installed via
brew), macOS (`sips` for fast downscaling), and `ffmpeg`/`ffprobe` (for video
thumbnails and tagging — optional, everything else works without it).
## How it works ## How it works
@@ -63,5 +82,35 @@ Tested defaults: scout `glm-ocr` (6/6 routing accuracy) → describer
## Searching afterwards ## Searching afterwards
Spotlight: just type a word from a description in Finder search. Spotlight: just type a word from a description in Finder search.
Or from terminal: `mdfind -onlyin "/Volumes/sanD/allphotos from phone" "scooter"` Or from terminal: `mdfind -onlyin "/path/to/your/photos" "scooter"`
Or grep the journal: `grep -i scooter photon_journal.jsonl` Or grep the journal: `grep -i scooter photon_journal.jsonl`
Or use the app itself — the **Search & Edit** tab is a full photo library browser:
- **Tagged** — instant multi-word search across description/category/filename/path,
filter by category (click a chip), date range, or file type (photos/videos),
sort by newest/name/category. Click any result for a full-resolution
viewer + editor with Save / AI Redo / Reveal-in-Finder.
- **Untagged** / **Failed** — see what's left to do or what errored, tag or
retry one at a time without leaving the grid.
- **Videos** — thumbnails via `ffmpeg` frame-grab, native playback, AI tagging
off an extracted frame.
- **Duplicates** — one-time background scan builds a perceptual-hash index and
groups exact visual matches (re-saves, burst duplicates). No auto-delete —
just Reveal-in-Finder per copy so you decide.
**Selecting & downloading photos in bulk:** click **Select** to enter select
mode, then:
- click a photo to toggle it, **shift-click** to select a range
- **Select all loaded** / **Select all N matching** (grabs everything matching
your current search, not just what's rendered) / **Clear selection**
- **Download selected (ZIP)** or **Download all matches (ZIP)** — streams a
real zip of the original files, byte-for-byte, no re-compression (capped at
3000 files per zip)
- **Apply category to selected**, or **Export** the set as CSV/JSON
## Appearance
Settings & Safety → **Appearance**: pick an accent color (native color
picker), grid density (compact/comfortable/large), and results-per-page.
Saved per-browser, doesn't touch anything on disk.

1148
index.html

File diff suppressed because it is too large Load Diff

726
server.py
View File

@@ -23,12 +23,56 @@ import time
import urllib.request import urllib.request
import urllib.error import urllib.error
import urllib.parse import urllib.parse
import zipfile
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
APP_DIR = os.path.dirname(os.path.abspath(__file__)) APP_DIR = os.path.dirname(os.path.abspath(__file__))
OLLAMA = "http://localhost:11434" OLLAMA = "http://localhost:11434"
PORT = 8765 PORT = int(os.environ.get("PHOTON_PORT", 8765))
DEFAULT_FOLDER = "/Volumes/sanD/allphotos from phone"
FOLDER_CONFIG_FILE = os.path.join(APP_DIR, "photon_folder.json")
def load_saved_folder():
if os.path.exists(FOLDER_CONFIG_FILE):
try:
with open(FOLDER_CONFIG_FILE, "r", encoding="utf-8") as f:
folder = json.load(f).get("folder")
if folder and os.path.isdir(folder):
return folder
except Exception:
pass
return None
def save_folder_config(folder):
try:
with open(FOLDER_CONFIG_FILE, "w", encoding="utf-8") as f:
json.dump({"folder": folder}, f)
except Exception:
pass
def resolve_default_folder():
"""Zero-config folder detection, in priority order:
1. `python3 server.py /path/to/photos` — explicit CLI argument
2. PHOTON_FOLDER env var
3. the last folder scanned in a previous session (remembered automatically)
4. the parent directory of wherever this app folder lives — the intended
workflow is dropping the PHOTON folder directly inside a photo
collection, so its parent IS that collection.
Always overridable from the Console tab's folder field + Scan button.
"""
if len(sys.argv) > 1:
arg = os.path.abspath(sys.argv[1])
if os.path.isdir(arg):
return arg
env = os.environ.get("PHOTON_FOLDER")
if env and os.path.isdir(env):
return os.path.abspath(env)
saved = load_saved_folder()
if saved:
return saved
return os.path.dirname(APP_DIR)
DEFAULT_FOLDER = resolve_default_folder()
JOURNAL = os.path.join(APP_DIR, "photon_journal.jsonl") JOURNAL = os.path.join(APP_DIR, "photon_journal.jsonl")
CATEGORIES_FILE = os.path.join(APP_DIR, "photon_categories.json") CATEGORIES_FILE = os.path.join(APP_DIR, "photon_categories.json")
@@ -124,9 +168,12 @@ class State:
self.status = "idle" # idle | scanning | running | paused | stopping | done self.status = "idle" # idle | scanning | running | paused | stopping | done
self.folder = DEFAULT_FOLDER self.folder = DEFAULT_FOLDER
self.files = [] # pending image paths (after scan) self.files = [] # pending image paths (after scan)
self.video_files = [] # video paths found on last scan
self.total_images = 0 self.total_images = 0
self.skipped_videos = 0 self.skipped_videos = 0
self.skipped_sidecars = 0 self.skipped_sidecars = 0
self.dedupe_status = "idle" # idle | running
self.dedupe_progress = {"done": 0, "total": 0}
self.already_done = 0 self.already_done = 0
self.processed_session = 0 self.processed_session = 0
self.failed_session = 0 self.failed_session = 0
@@ -195,6 +242,203 @@ def journal_write(rec):
with open(JOURNAL, "a", encoding="utf-8") as f: with open(JOURNAL, "a", encoding="utf-8") as f:
f.write(json.dumps(rec, ensure_ascii=False) + "\n") f.write(json.dumps(rec, ensure_ascii=False) + "\n")
SEARCH_CACHE = {"mtime": None, "records": []}
def load_search_index():
"""Journal records with a precomputed lowercase haystack, cached by file mtime."""
global SEARCH_CACHE
if not os.path.exists(JOURNAL):
SEARCH_CACHE = {"mtime": None, "records": []}
return []
mtime = os.path.getmtime(JOURNAL)
if SEARCH_CACHE["mtime"] == mtime:
return SEARCH_CACHE["records"]
records = []
with open(JOURNAL, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
rec = json.loads(line)
path = rec.get("path", "")
name = os.path.basename(path)
rec["name"] = name
rec["_ext"] = os.path.splitext(name)[1].lower()
rec["_hay"] = " ".join([
path.lower(), (rec.get("desc") or "").lower(),
(rec.get("category") or "").lower(), name.lower()
])
try:
st = os.stat(path)
rec["_mtime"] = st.st_mtime
rec["sizeBytes"] = st.st_size
except OSError:
rec["_mtime"] = rec.get("ts", 0)
rec["sizeBytes"] = None
records.append(rec)
except Exception:
pass
SEARCH_CACHE = {"mtime": mtime, "records": records}
return records
def parse_date_bound(s, end_of_day=False):
if not s:
return None
try:
t = time.strptime(s, "%Y-%m-%d")
epoch = time.mktime(t)
return epoch + 86399 if end_of_day else epoch
except Exception:
return None
def filtered_search_results(q, cat, date_from, date_to, sort, ftype=""):
records = load_search_index()
tokens = (q or "").strip().lower().split()
df = parse_date_bound(date_from)
dt = parse_date_bound(date_to, end_of_day=True)
def matches(rec):
if cat and rec.get("category") != cat:
return False
if ftype == "photo" and rec.get("_ext") not in IMAGE_EXTS:
return False
if ftype == "video" and rec.get("_ext") not in VIDEO_EXTS:
return False
if df is not None and rec.get("_mtime", 0) < df:
return False
if dt is not None and rec.get("_mtime", 0) > dt:
return False
if not tokens:
return True
hay = rec.get("_hay", "")
return all(t in hay for t in tokens)
filtered = [r for r in records if matches(r)]
if sort == "name":
filtered.sort(key=lambda r: r.get("name", "").lower())
elif sort == "category":
filtered.sort(key=lambda r: ((r.get("category") or ""), r.get("name", "").lower()))
else:
filtered.sort(key=lambda r: r.get("ts", 0), reverse=True)
return filtered
# ---------------------------------------------------------------- duplicate detection
PHASH_FILE = os.path.join(APP_DIR, "photon_phash.json")
PHASH_CACHE = {}
PHASH_LOCK = threading.Lock()
def load_phash_cache():
global PHASH_CACHE
with PHASH_LOCK:
if PHASH_CACHE:
return PHASH_CACHE
if os.path.exists(PHASH_FILE):
try:
with open(PHASH_FILE, "r", encoding="utf-8") as f:
PHASH_CACHE = json.load(f)
except Exception:
PHASH_CACHE = {}
return PHASH_CACHE
def save_phash_cache():
with PHASH_LOCK:
try:
with open(PHASH_FILE, "w", encoding="utf-8") as f:
json.dump(PHASH_CACHE, f)
except Exception:
pass
def compute_phash(path):
"""8x8 grayscale average-hash via sips + manual BMP parsing. No external deps."""
tmp_path = os.path.join(tempfile.gettempdir(), f"photon_ph_{os.getpid()}_{threading.get_ident()}.bmp")
try:
cmd = ["sips", "-s", "format", "bmp", "-z", "8", "8", path, "--out", tmp_path]
r = subprocess.run(cmd, capture_output=True, timeout=20)
if r.returncode != 0 or not os.path.exists(tmp_path):
return None
with open(tmp_path, "rb") as f:
data = f.read()
if data[0:2] != b"BM":
return None
pixel_offset = int.from_bytes(data[10:14], "little")
width = int.from_bytes(data[18:22], "little")
height = int.from_bytes(data[22:26], "little")
bpp = int.from_bytes(data[28:30], "little")
if bpp not in (24, 32) or width <= 0 or height <= 0:
return None
row_bytes = width * (bpp // 8)
row_padded = (row_bytes + 3) & ~3
grays = []
for y in range(height):
row_start = pixel_offset + y * row_padded
for x in range(width):
px = row_start + x * (bpp // 8)
if px + 2 >= len(data):
return None
b, g, rr = data[px], data[px + 1], data[px + 2]
grays.append((rr * 299 + g * 587 + b * 114) // 1000)
if not grays:
return None
avg = sum(grays) / len(grays)
bits = "".join("1" if p >= avg else "0" for p in grays)
return f"{int(bits, 2):0{len(bits)//4}x}"
except Exception:
return None
finally:
try:
os.remove(tmp_path)
except OSError:
pass
def build_dedupe_index():
"""Background worker: hash every tagged image not yet in the phash cache."""
with S.lock:
if S.dedupe_status == "running":
return
S.dedupe_status = "running"
try:
records = load_search_index()
paths = [r["path"] for r in records
if os.path.splitext(r["path"])[1].lower() in IMAGE_EXTS]
cache = load_phash_cache()
todo = [p for p in paths if p not in cache]
total = len(todo)
with S.lock:
S.dedupe_progress = {"done": 0, "total": total}
log("info", f"duplicate scan: hashing {total} photos …")
for i, p in enumerate(todo):
if not os.path.exists(p):
continue
h = compute_phash(p)
if h:
cache[p] = h
if i % 50 == 0:
save_phash_cache()
with S.lock:
S.dedupe_progress = {"done": i + 1, "total": total}
if i % 25 == 0:
broadcast("dedupe", {"done": i + 1, "total": total})
save_phash_cache()
broadcast("dedupe", {"done": total, "total": total, "finished": True})
log("ok", f"duplicate scan complete: {total} photos hashed")
finally:
with S.lock:
S.dedupe_status = "idle"
def find_duplicate_groups():
"""Groups of exact perceptual-hash matches — visually identical / re-saved copies."""
records = {r["path"]: r for r in load_search_index()}
cache = load_phash_cache()
by_hash = {}
for path, h in cache.items():
if path in records:
by_hash.setdefault(h, []).append(records[path])
groups = [g for g in by_hash.values() if len(g) > 1]
groups.sort(key=len, reverse=True)
return groups
# ---------------------------------------------------------------- SSE # ---------------------------------------------------------------- SSE
def broadcast(kind, data): def broadcast(kind, data):
@@ -302,7 +546,7 @@ def ollama_generate(model, prompt, img_b64, opts, keep_alive, think=None, schema
# ---------------------------------------------------------------- pipeline # ---------------------------------------------------------------- pipeline
def scan_folder(folder): def scan_folder(folder):
images, videos, sidecars = [], 0, 0 images, videos, sidecars = [], [], 0
for root, dirs, files in os.walk(folder): for root, dirs, files in os.walk(folder):
dirs[:] = [d for d in dirs if not d.startswith(".") and not d.startswith("_")] dirs[:] = [d for d in dirs if not d.startswith(".") and not d.startswith("_")]
for name in sorted(files): for name in sorted(files):
@@ -313,7 +557,7 @@ def scan_folder(folder):
if ext in IMAGE_EXTS: if ext in IMAGE_EXTS:
images.append(os.path.join(root, name)) images.append(os.path.join(root, name))
elif ext in VIDEO_EXTS: elif ext in VIDEO_EXTS:
videos += 1 videos.append(os.path.join(root, name))
return images, videos, sidecars return images, videos, sidecars
def organize_alias(folder, file_path, category): def organize_alias(folder, file_path, category):
@@ -418,6 +662,38 @@ def downscale(path, max_px, tmpdir):
with open(out, "rb") as f: with open(out, "rb") as f:
return f.read() return f.read()
def extract_video_frame(path, tmpdir):
"""Grab one representative frame from a video via ffmpeg (read-only). Returns a jpeg file path."""
out = os.path.join(tmpdir, "photon_vframe.jpg")
dur = 3.0
try:
pr = subprocess.run(
["ffprobe", "-v", "error", "-show_entries", "format=duration",
"-of", "default=noprint_wrappers=1:nokey=1", path],
capture_output=True, timeout=15, text=True)
dur = float(pr.stdout.strip())
except Exception:
pass
ts = max(0.5, min(dur * 0.3, max(dur - 0.2, 0.5))) if dur > 1 else 0.1
cmd = ["ffmpeg", "-y", "-ss", str(ts), "-i", path, "-frames:v", "1", "-q:v", "3", out]
r = subprocess.run(cmd, capture_output=True, timeout=60)
if r.returncode != 0 or not os.path.exists(out):
raise RuntimeError(f"ffmpeg frame extraction failed: {r.stderr.decode(errors='replace')[:200]}")
return out
def video_thumbnail(path, out_path):
"""Cheap ffmpeg frame-grab thumbnail for the browsing grid (read-only)."""
cmd = ["ffmpeg", "-y", "-ss", "1", "-i", path, "-frames:v", "1",
"-vf", "scale=256:-1", "-q:v", "5", out_path]
r = subprocess.run(cmd, capture_output=True, timeout=30)
if r.returncode != 0 or not os.path.exists(out_path):
# very short clips: fall back to the first frame
cmd2 = ["ffmpeg", "-y", "-i", path, "-frames:v", "1",
"-vf", "scale=256:-1", "-q:v", "5", out_path]
r = subprocess.run(cmd2, capture_output=True, timeout=30)
if r.returncode != 0 or not os.path.exists(out_path):
raise RuntimeError(r.stderr.decode(errors="replace")[:200])
def salvage_json(raw): def salvage_json(raw):
"""Parse model output, surviving truncated/unterminated JSON.""" """Parse model output, surviving truncated/unterminated JSON."""
try: try:
@@ -465,28 +741,39 @@ def build_prompt(length_key, mode="photo"):
) )
def write_metadata(path, desc, category, keep_backup, preserve_date): def write_metadata(path, desc, category, keep_backup, preserve_date):
is_video = os.path.splitext(path)[1].lower() in VIDEO_EXTS
args = ["exiftool", "-m", "-q", "-codedcharacterset=utf8"] args = ["exiftool", "-m", "-q", "-codedcharacterset=utf8"]
if preserve_date: if preserve_date:
args.append("-P") args.append("-P")
if not keep_backup: if not keep_backup:
args.append("-overwrite_original") # writes temp file then atomic rename args.append("-overwrite_original") # writes temp file then atomic rename
args += [ if is_video:
f"-EXIF:ImageDescription={desc}", # mp4/mov containers don't carry EXIF/IPTC — use exiftool's generic
f"-IPTC:Caption-Abstract={desc}", # tag names so it resolves to QuickTime/Keys groups automatically.
f"-XMP-dc:Description={desc}", args += [
f"-XMP-dc:Subject+={category}", f"-Description={desc}",
f"-IPTC:Keywords+={category}", f"-Keywords+={category}",
"-XMP-dc:Subject+=photon-tagged", "-Keywords+=photon-tagged",
path, ]
] else:
args += [
f"-EXIF:ImageDescription={desc}",
f"-IPTC:Caption-Abstract={desc}",
f"-XMP-dc:Description={desc}",
f"-XMP-dc:Subject+={category}",
f"-IPTC:Keywords+={category}",
"-XMP-dc:Subject+=photon-tagged",
]
args.append(path)
r = subprocess.run(args, capture_output=True, timeout=120) r = subprocess.run(args, capture_output=True, timeout=120)
if r.returncode != 0: if r.returncode != 0:
raise RuntimeError(f"exiftool: {r.stderr.decode(errors='replace')[:300]}") raise RuntimeError(f"exiftool: {r.stderr.decode(errors='replace')[:300]}")
def sync_metadata_from_folder(folder): def sync_metadata_from_folder(folder):
log("info", f"Syncing journal with folder metadata: {folder} ...") log("info", f"Syncing journal with folder metadata: {folder} ...")
cmd = ["exiftool", "-r", "-json", "-if", "$Subject =~ /photon-tagged/", cmd = ["exiftool", "-r", "-json", "-if",
"-EXIF:ImageDescription", "-XMP-dc:Description", "-IPTC:Caption-Abstract", "$Subject =~ /photon-tagged/ or $Keywords =~ /photon-tagged/",
"-EXIF:ImageDescription", "-XMP-dc:Description", "-IPTC:Caption-Abstract", "-Description",
"-XMP-dc:Subject", "-IPTC:Keywords", folder] "-XMP-dc:Subject", "-IPTC:Keywords", folder]
r = subprocess.run(cmd, capture_output=True, timeout=300) r = subprocess.run(cmd, capture_output=True, timeout=300)
if r.returncode != 0: if r.returncode != 0:
@@ -826,6 +1113,40 @@ class Handler(BaseHTTPRequestHandler):
self.end_headers() self.end_headers()
self.wfile.write(body) self.wfile.write(body)
def _serve_file_with_range(self, path, ctype):
"""Basic HTTP Range support, needed for native <video> seeking."""
try:
size = os.path.getsize(path)
range_hdr = self.headers.get("Range")
start, end = 0, size - 1
status = 200
if range_hdr and range_hdr.startswith("bytes="):
status = 206
rng = range_hdr[6:].split("-")
if rng[0]:
start = int(rng[0])
if len(rng) > 1 and rng[1]:
end = int(rng[1])
end = min(end, size - 1)
length = end - start + 1
with open(path, "rb") as f:
f.seek(start)
data = f.read(length)
self.send_response(status)
self.send_header("Content-Type", ctype)
self.send_header("Accept-Ranges", "bytes")
self.send_header("Content-Length", str(len(data)))
if status == 206:
self.send_header("Content-Range", f"bytes {start}-{end}/{size}")
self.send_header("Cache-Control", "no-store")
self.end_headers()
self.wfile.write(data)
except (BrokenPipeError, ConnectionResetError):
pass
except Exception as e:
self.send_response(500); self.end_headers()
self.wfile.write(str(e).encode())
def do_GET(self): def do_GET(self):
if self.path == "/" or self.path.startswith("/index"): if self.path == "/" or self.path.startswith("/index"):
with open(os.path.join(APP_DIR, "index.html"), "rb") as f: with open(os.path.join(APP_DIR, "index.html"), "rb") as f:
@@ -872,11 +1193,15 @@ class Handler(BaseHTTPRequestHandler):
os.makedirs(cache_dir, exist_ok=True) os.makedirs(cache_dir, exist_ok=True)
h = hashlib.md5(abs_path.encode()).hexdigest() h = hashlib.md5(abs_path.encode()).hexdigest()
thumb_path = os.path.join(cache_dir, f"{h}.jpg") thumb_path = os.path.join(cache_dir, f"{h}.jpg")
is_video = os.path.splitext(abs_path)[1].lower() in VIDEO_EXTS
if not os.path.exists(thumb_path): if not os.path.exists(thumb_path):
cmd = ["sips", "-s", "format", "jpeg", "-s", "formatOptions", "70", "-Z", "256", abs_path, "--out", thumb_path] if is_video:
r = subprocess.run(cmd, capture_output=True, timeout=15) video_thumbnail(abs_path, thumb_path)
if r.returncode != 0: else:
raise RuntimeError(f"sips failed: {r.stderr.decode(errors='replace')}") cmd = ["sips", "-s", "format", "jpeg", "-s", "formatOptions", "70", "-Z", "256", abs_path, "--out", thumb_path]
r = subprocess.run(cmd, capture_output=True, timeout=15)
if r.returncode != 0:
raise RuntimeError(f"sips failed: {r.stderr.decode(errors='replace')}")
with open(thumb_path, "rb") as f: with open(thumb_path, "rb") as f:
data = f.read() data = f.read()
self.send_response(200) self.send_response(200)
@@ -888,35 +1213,258 @@ class Handler(BaseHTTPRequestHandler):
except Exception as e: except Exception as e:
self.send_response(500); self.end_headers() self.send_response(500); self.end_headers()
self.wfile.write(str(e).encode()) self.wfile.write(str(e).encode())
elif self.path.startswith("/api/photo"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
img_path = params.get("path", [""])[0]
if not img_path or not os.path.exists(img_path):
self.send_response(404); self.end_headers(); return
abs_path = os.path.abspath(img_path)
with S.lock:
folder_ok = abs_path.startswith(os.path.abspath(S.folder))
journal_ok = abs_path in S.done_paths
if not (folder_ok or journal_ok):
self.send_response(403); self.end_headers(); return
ext = os.path.splitext(abs_path)[1].lower()
serve_path = abs_path
ctype = {
".jpg": "image/jpeg", ".jpeg": "image/jpeg", ".png": "image/png",
".webp": "image/webp", ".heic": "image/heic", ".heif": "image/heif",
".tif": "image/tiff", ".tiff": "image/tiff",
".mov": "video/quicktime", ".mp4": "video/mp4", ".m4v": "video/mp4",
".avi": "video/x-msvideo", ".3gp": "video/3gpp", ".mkv": "video/x-matroska",
}.get(ext, "application/octet-stream")
if ext in VIDEO_EXTS:
self._serve_file_with_range(abs_path, ctype)
return
if ext in (".heic", ".heif", ".tif", ".tiff"):
# browsers other than Safari can't render these — serve a read-only
# full-size jpeg conversion via sips, cached, original untouched.
try:
cache_dir = os.path.join(tempfile.gettempdir(), "photon_view")
os.makedirs(cache_dir, exist_ok=True)
h = hashlib.md5(abs_path.encode()).hexdigest()
view_path = os.path.join(cache_dir, f"{h}.jpg")
if not os.path.exists(view_path) or os.path.getmtime(abs_path) > os.path.getmtime(view_path):
cmd = ["sips", "-s", "format", "jpeg", "-s", "formatOptions", "92",
"-Z", "2200", abs_path, "--out", view_path]
r = subprocess.run(cmd, capture_output=True, timeout=30)
if r.returncode != 0 or not os.path.exists(view_path):
raise RuntimeError(r.stderr.decode(errors="replace")[:200])
serve_path = view_path
ctype = "image/jpeg"
except Exception as e:
self.send_response(500); self.end_headers()
self.wfile.write(str(e).encode())
return
try:
with open(serve_path, "rb") as f:
data = f.read()
self.send_response(200)
self.send_header("Content-Type", ctype)
self.send_header("Content-Length", str(len(data)))
self.send_header("Cache-Control", "no-store")
self.end_headers()
self.wfile.write(data)
except Exception as e:
self.send_response(500); self.end_headers()
self.wfile.write(str(e).encode())
elif self.path.startswith("/api/search/paths"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
q = params.get("q", [""])[0]
cat = params.get("cat", [""])[0]
sort = params.get("sort", ["recent"])[0]
date_from = params.get("date_from", [""])[0]
date_to = params.get("date_to", [""])[0]
ftype = params.get("ftype", [""])[0]
filtered = filtered_search_results(q, cat, date_from, date_to, sort, ftype)
MAX_SELECT = 3000
paths = [r["path"] for r in filtered[:MAX_SELECT]]
self._json({"paths": paths, "total": len(filtered), "capped": len(filtered) > MAX_SELECT})
elif self.path.startswith("/api/search"): elif self.path.startswith("/api/search"):
parsed_url = urllib.parse.urlparse(self.path) parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query) params = urllib.parse.parse_qs(parsed_url.query)
q = params.get("q", [""])[0].strip().lower() q = params.get("q", [""])[0]
cat = params.get("cat", [""])[0]
sort = params.get("sort", ["recent"])[0]
date_from = params.get("date_from", [""])[0]
date_to = params.get("date_to", [""])[0]
ftype = params.get("ftype", [""])[0]
try:
limit = max(1, min(200, int(params.get("limit", ["60"])[0])))
except ValueError:
limit = 60
try:
offset = max(0, int(params.get("offset", ["0"])[0]))
except ValueError:
offset = 0
re_read = params.get("re_read", ["0"])[0] == "1" re_read = params.get("re_read", ["0"])[0] == "1"
if re_read: if re_read:
with S.lock: with S.lock:
folder = S.folder folder = S.folder
sync_metadata_from_folder(folder) sync_metadata_from_folder(folder)
push_stats() push_stats()
results = []
if os.path.exists(JOURNAL): filtered = filtered_search_results(q, cat, date_from, date_to, sort, ftype)
with open(JOURNAL, "r", encoding="utf-8") as f: total = len(filtered)
for line in f: page = [{k: v for k, v in r.items() if k not in ("_hay", "_mtime", "_ext")} for r in filtered[offset:offset + limit]]
line = line.strip() self._json({"results": page, "total": total, "offset": offset, "limit": limit})
if not line:
continue elif self.path.startswith("/api/export"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
q = params.get("q", [""])[0]
cat = params.get("cat", [""])[0]
sort = params.get("sort", ["recent"])[0]
date_from = params.get("date_from", [""])[0]
date_to = params.get("date_to", [""])[0]
ftype = params.get("ftype", [""])[0]
fmt = params.get("format", ["json"])[0]
paths_only = params.get("paths", [""])[0].split(",") if params.get("paths", [""])[0] else None
filtered = filtered_search_results(q, cat, date_from, date_to, sort, ftype)
if paths_only:
wanted = set(paths_only)
filtered = [r for r in filtered if r["path"] in wanted]
if fmt == "csv":
import csv, io
buf = io.StringIO()
w = csv.writer(buf)
w.writerow(["path", "name", "category", "description", "model", "route", "tagged_at"])
for r in filtered:
w.writerow([r.get("path", ""), r.get("name", ""), r.get("category", ""),
r.get("desc", ""), r.get("model", ""), r.get("route", ""),
time.strftime("%Y-%m-%d %H:%M:%S", time.localtime(r.get("ts", 0))) if r.get("ts") else ""])
body = buf.getvalue().encode("utf-8")
ctype = "text/csv"
fname = "photon_export.csv"
else:
page = [{k: v for k, v in r.items() if k not in ("_hay", "_mtime", "_ext")} for r in filtered]
body = json.dumps(page, ensure_ascii=False, indent=2).encode("utf-8")
ctype = "application/json"
fname = "photon_export.json"
self.send_response(200)
self.send_header("Content-Type", ctype)
self.send_header("Content-Disposition", f'attachment; filename="{fname}"')
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
elif self.path.startswith("/api/untagged"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
try:
limit = max(1, min(200, int(params.get("limit", ["60"])[0])))
except ValueError:
limit = 60
try:
offset = max(0, int(params.get("offset", ["0"])[0]))
except ValueError:
offset = 0
with S.lock:
pending = [p for p in S.files if p not in S.done_paths]
total = len(pending)
page = [{"path": p, "name": os.path.basename(p)} for p in pending[offset:offset + limit]]
self._json({"results": page, "total": total, "offset": offset, "limit": limit})
elif self.path.startswith("/api/failed"):
with S.lock:
paths = sorted(S.failed_paths)
results = [{"path": p, "name": os.path.basename(p)} for p in paths]
self._json({"results": results, "total": len(results)})
elif self.path.startswith("/api/videos"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
try:
limit = max(1, min(200, int(params.get("limit", ["40"])[0])))
except ValueError:
limit = 40
try:
offset = max(0, int(params.get("offset", ["0"])[0]))
except ValueError:
offset = 0
with S.lock:
vids = list(S.video_files)
done = set(S.done_paths)
total = len(vids)
page = [{"path": p, "name": os.path.basename(p), "tagged": p in done}
for p in vids[offset:offset + limit]]
self._json({"results": page, "total": total, "offset": offset, "limit": limit})
elif self.path.startswith("/api/dedupe/groups"):
groups = find_duplicate_groups()
payload = [
{"hash": None, "count": len(g),
"items": [{k: v for k, v in r.items() if k not in ("_hay", "_mtime", "_ext")} for r in g]}
for g in groups
]
cache = load_phash_cache()
with S.lock:
total_images = len([p for p in S.done_paths if os.path.splitext(p)[1].lower() in IMAGE_EXTS])
self._json({"groups": payload, "hashed": len(cache), "totalImages": total_images,
"dedupeStatus": S.dedupe_status, "dedupeProgress": S.dedupe_progress})
elif self.path.startswith("/api/download_zip"):
parsed_url = urllib.parse.urlparse(self.path)
params = urllib.parse.parse_qs(parsed_url.query)
paths_param = params.get("paths", [""])[0]
if paths_param:
paths = paths_param.split(",")
else:
q = params.get("q", [""])[0]
cat = params.get("cat", [""])[0]
sort = params.get("sort", ["recent"])[0]
date_from = params.get("date_from", [""])[0]
date_to = params.get("date_to", [""])[0]
ftype = params.get("ftype", [""])[0]
paths = [r["path"] for r in filtered_search_results(q, cat, date_from, date_to, sort, ftype)]
paths = [p for p in paths if p and os.path.exists(p)]
if not paths:
self._json({"error": "no files to download"}, 400); return
if len(paths) > 3000:
self._json({"error": f"too many files for one zip ({len(paths)}, max 3000) — narrow your selection"}, 400); return
tmp_zip = tempfile.NamedTemporaryFile(suffix=".zip", delete=False)
tmp_zip.close()
try:
used_names = set()
with zipfile.ZipFile(tmp_zip.name, "w", zipfile.ZIP_STORED, allowZip64=True) as zf:
for p in paths:
name = os.path.basename(p)
base, ext = os.path.splitext(name)
n, i = name, 1
while n in used_names:
n = f"{base}_{i}{ext}"; i += 1
used_names.add(n)
# original bytes, untouched — this is a read-only archive of the source files
zf.write(p, arcname=n)
size = os.path.getsize(tmp_zip.name)
self.send_response(200)
self.send_header("Content-Type", "application/zip")
self.send_header("Content-Disposition", f'attachment; filename="photon_{len(paths)}_photos.zip"')
self.send_header("Content-Length", str(size))
self.end_headers()
with open(tmp_zip.name, "rb") as f:
while True:
chunk = f.read(1024 * 1024)
if not chunk:
break
try: try:
rec = json.loads(line) self.wfile.write(chunk)
path = rec.get("path", "") except (BrokenPipeError, ConnectionResetError):
desc = rec.get("desc", "") break
category = rec.get("category", "") finally:
name = os.path.basename(path) try:
if (not q) or (q in path.lower()) or (q in desc.lower()) or (q in category.lower()) or (q in name.lower()): os.remove(tmp_zip.name)
rec["name"] = name except OSError:
results.append(rec) pass
except Exception:
pass
self._json({"results": results})
elif self.path == "/api/categories": elif self.path == "/api/categories":
self._json({"categories": CATEGORIES}) self._json({"categories": CATEGORIES})
elif self.path == "/api/audit/start": elif self.path == "/api/audit/start":
@@ -935,6 +1483,29 @@ class Handler(BaseHTTPRequestHandler):
sampled = records sampled = records
random.shuffle(sampled) random.shuffle(sampled)
self._json({"results": sampled}) self._json({"results": sampled})
elif self.path == "/api/audit/history":
audit_file = os.path.join(APP_DIR, "photon_audits.jsonl")
entries = []
if os.path.exists(audit_file):
with open(audit_file, "r", encoding="utf-8") as f:
for line in f:
try:
entries.append(json.loads(line))
except Exception:
pass
passed = sum(1 for e in entries if e.get("grade") == "pass")
failed = sum(1 for e in entries if e.get("grade") == "fail")
by_day = {}
for e in entries:
day = time.strftime("%Y-%m-%d", time.localtime(e.get("ts", 0)))
d = by_day.setdefault(day, {"pass": 0, "fail": 0})
d[e.get("grade", "fail")] += 1
days = sorted(by_day.keys())[-14:]
self._json({
"total": len(entries), "passed": passed, "failed": failed,
"accuracy": round(passed / len(entries) * 100, 1) if entries else 0,
"byDay": [{"day": d, **by_day[d]} for d in days],
})
elif self.path == "/api/events": elif self.path == "/api/events":
self.send_response(200) self.send_response(200)
self.send_header("Content-Type", "text/event-stream") self.send_header("Content-Type", "text/event-stream")
@@ -974,6 +1545,7 @@ class Handler(BaseHTTPRequestHandler):
if S.status == "running": if S.status == "running":
self._json({"error": "stop the run before rescanning"}, 400); return self._json({"error": "stop the run before rescanning"}, 400); return
S.status = "scanning"; S.folder = folder S.status = "scanning"; S.folder = folder
save_folder_config(folder)
broadcast("state", {"status": "scanning"}) broadcast("state", {"status": "scanning"})
log("info", f"scanning {folder} …") log("info", f"scanning {folder} …")
images, videos, sidecars = scan_folder(folder) images, videos, sidecars = scan_folder(folder)
@@ -981,14 +1553,15 @@ class Handler(BaseHTTPRequestHandler):
with S.lock: with S.lock:
S.files = images S.files = images
S.total_images = len(images) S.total_images = len(images)
S.skipped_videos = videos S.video_files = videos
S.skipped_videos = len(videos)
S.skipped_sidecars = sidecars S.skipped_sidecars = sidecars
S.status = "idle" S.status = "idle"
log("ok", f"scan complete: {len(images)} images | {videos} videos skipped | " log("ok", f"scan complete: {len(images)} images | {len(videos)} videos found | "
f"{sidecars} hidden/sidecar files ignored | {done} already tagged") f"{sidecars} hidden/sidecar files ignored | {done} already tagged")
broadcast("state", {"status": "idle"}) broadcast("state", {"status": "idle"})
push_stats() push_stats()
self._json({"images": len(images), "videos": videos, self._json({"images": len(images), "videos": len(videos),
"sidecars": sidecars, "alreadyDone": done}) "sidecars": sidecars, "alreadyDone": done})
elif self.path == "/api/start": elif self.path == "/api/start":
@@ -1040,12 +1613,13 @@ class Handler(BaseHTTPRequestHandler):
with S.lock: with S.lock:
folder = S.folder folder = S.folder
log("warn", f"UNDO ALL: Removing all PHOTON metadata from files in {folder} ...") log("warn", f"UNDO ALL: Removing all PHOTON metadata from files in {folder} ...")
args = ["exiftool", "-r", "-P", "-overwrite_original", "-if", "$Subject =~ /photon-tagged/", args = ["exiftool", "-r", "-P", "-overwrite_original", "-if",
"-EXIF:ImageDescription=", "-IPTC:Caption-Abstract=", "-XMP-dc:Description=", "$Subject =~ /photon-tagged/ or $Keywords =~ /photon-tagged/",
"-XMP-dc:Subject-=photon-tagged", "-IPTC:Keywords-=photon-tagged"] "-EXIF:ImageDescription=", "-IPTC:Caption-Abstract=", "-XMP-dc:Description=", "-Description=",
"-XMP-dc:Subject-=photon-tagged", "-Keywords-=photon-tagged"]
for cat in CATEGORIES: for cat in CATEGORIES:
args.append(f"-XMP-dc:Subject-={cat}") args.append(f"-XMP-dc:Subject-={cat}")
args.append(f"-IPTC:Keywords-={cat}") args.append(f"-Keywords-={cat}")
args.append(folder) args.append(folder)
r = subprocess.run(args, capture_output=True, timeout=600) r = subprocess.run(args, capture_output=True, timeout=600)
log("info", f"Exiftool undo completed: {r.stdout.decode(errors='replace')[:200]}") log("info", f"Exiftool undo completed: {r.stdout.decode(errors='replace')[:200]}")
@@ -1140,6 +1714,10 @@ class Handler(BaseHTTPRequestHandler):
with open(JOURNAL, "w", encoding="utf-8") as f: with open(JOURNAL, "w", encoding="utf-8") as f:
for rec in existing_recs: for rec in existing_recs:
f.write(json.dumps(rec, ensure_ascii=False) + "\n") f.write(json.dumps(rec, ensure_ascii=False) + "\n")
with S.lock:
if path in S.failed_paths:
S.failed_paths.remove(path)
write_failures()
load_journal() load_journal()
push_stats() push_stats()
self._json({"ok": True, "desc": desc, "category": category}) self._json({"ok": True, "desc": desc, "category": category})
@@ -1181,7 +1759,9 @@ class Handler(BaseHTTPRequestHandler):
think = False think = False
tmpdir = tempfile.mkdtemp(prefix="photon_redo_") tmpdir = tempfile.mkdtemp(prefix="photon_redo_")
img = downscale(path, max_px, tmpdir) is_video = os.path.splitext(path)[1].lower() in VIDEO_EXTS
frame_source = extract_video_frame(path, tmpdir) if is_video else path
img = downscale(frame_source, max_px, tmpdir)
b64 = base64.b64encode(img).decode() b64 = base64.b64encode(img).decode()
route = None route = None
@@ -1249,6 +1829,10 @@ class Handler(BaseHTTPRequestHandler):
shutil.rmtree(tmpdir) shutil.rmtree(tmpdir)
except Exception: except Exception:
pass pass
with S.lock:
if path in S.failed_paths:
S.failed_paths.remove(path)
write_failures()
load_journal() load_journal()
push_stats() push_stats()
self._json({"ok": True, "desc": desc, "category": category, "route": route}) self._json({"ok": True, "desc": desc, "category": category, "route": route})
@@ -1280,13 +1864,63 @@ class Handler(BaseHTTPRequestHandler):
except Exception as e: except Exception as e:
self._json({"error": str(e)}, 500) self._json({"error": str(e)}, 500)
elif self.path == "/api/dedupe/build":
with S.lock:
already_running = S.dedupe_status == "running"
if already_running:
self._json({"error": "duplicate scan already running"}, 400); return
threading.Thread(target=build_dedupe_index, daemon=True).start()
self._json({"ok": True})
elif self.path == "/api/bulk_recategorize":
paths = body.get("paths") or []
category = body.get("category")
if not paths or category not in CATEGORIES:
self._json({"error": "bad request"}, 400); return
records = {r["path"]: r for r in load_search_index()}
with S.lock:
keep_backup = bool(S.settings.get("keepBackup", False))
preserve_date = bool(S.settings.get("preserveDate", True))
ok, failed = 0, []
for p in paths:
rec = records.get(p)
desc = rec.get("desc", "") if rec else ""
try:
write_metadata(p, desc, category, keep_backup, preserve_date)
ok += 1
except Exception as e:
failed.append({"path": p, "error": str(e)})
existing_recs = []
wanted = set(paths)
if os.path.exists(JOURNAL):
with open(JOURNAL, "r", encoding="utf-8") as f:
for line in f:
try:
rec = json.loads(line)
if rec["path"] in wanted and not any(fe["path"] == rec["path"] for fe in failed):
rec["category"] = category
rec["ts"] = time.time()
existing_recs.append(rec)
except Exception:
pass
with open(JOURNAL, "w", encoding="utf-8") as f:
for rec in existing_recs:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
load_journal()
push_stats()
log("ok", f"Bulk recategorized {ok} photos to '{category}'" + (f" ({len(failed)} failed)" if failed else ""))
self._json({"ok": True, "updated": ok, "failed": failed})
else: else:
self.send_response(404); self.end_headers() self.send_response(404); self.end_headers()
def main(): def main():
n = load_journal() n = load_journal()
load_failures() load_failures()
load_phash_cache()
print(f"PHOTON console → http://localhost:{PORT}") print(f"PHOTON console → http://localhost:{PORT}")
print(f"photo folder → {DEFAULT_FOLDER}"
+ (" (auto-detected — change it anytime from the Console tab)" if len(sys.argv) <= 1 and not os.environ.get("PHOTON_FOLDER") else ""))
if n: if n:
print(f"journal loaded: {n} photos already tagged (will be skipped on resume)") print(f"journal loaded: {n} photos already tagged (will be skipped on resume)")
ThreadingHTTPServer(("127.0.0.1", PORT), Handler).serve_forever() ThreadingHTTPServer(("127.0.0.1", PORT), Handler).serve_forever()