The confirm barrier surfaced two yields that left 8.2 GB allocated after 3s. Two separate issues behind that class of failure: - The yield only unloaded loaded_models[0]. Ollama can hold several models resident (OLLAMA_MAX_LOADED_MODELS), so releasing the first left the rest allocated. It now unloads every resident model concurrently. This box runs with the limit at 1, so the change is defensive here rather than a fix for the observed case. - The observed 8.2 GB stalls happened while ComfyUI was starting and Ollama had a generation in flight; Ollama will not unload mid-request. A 3s ceiling reported a timeout for a model that was simply busy finishing. Raised to 10s -- waiting longer is the safer failure mode, since the alternative is diffusion allocating into VRAM that is still occupied. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
36 KiB
36 KiB