The first full run passed every stage, and three of those passes were worth less than they looked. Diffusion was reported as 0.67 it/s. That single run included loading the SDXL checkpoint from disk, so it understated throughput roughly tenfold against a steady-state 5.49 it/s. Cold and warm are now timed and labelled separately. The idle-purge stage checked the flag immediately, racing the ComfyUI websocket event that sets it, and reported "no purge pending" as a warning about ComfyUI rather than about its own timing. It now waits for the event. The reclaim stage passed while proving nothing: the model chosen was small enough to fit alongside ComfyUI's checkpoint, so the reclaim path never ran. It now picks a model that genuinely will not fit, and reports a warning rather than a pass when the path is not exercised. Sizing that model correctly took two corrections, both real. On-disk weight size is not the VRAM footprint -- a 12.87 GB blob occupies 14.9 GB once context and KV cache are allocated -- and unmanaged VRAM is not reclaimable, so it cannot count toward what a reclaim will free. Ignoring the second picked a model that failed even after a correct reclaim: the service returned 507 and logged "could not fit with ComfyUI holding 7.03 GB -- reclaiming and retrying", which was right. On this box an 842 MB third-party process is the difference between a 14.9 GB model fitting and not. The verifier also died on the 507 instead of reporting it, since a helper called raise_for_status() on responses a stage deliberately provokes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
14 KiB
Executable File
14 KiB
Executable File