Find the stage that ran out of memory
Do not start with a random launch flag. Save the earliest OOM, return to the last workload that passed, and determine whether Forge failed while loading, sampling, upscaling, decoding, switching models, or using host memory.
Recover capacity before you optimize speed
An OOM means one allocation failed in one environment. It is not a verdict on the whole GPU and it does not prove a memory leak.
Restart Forge once, close other GPU-heavy apps, select the correct preset, use batch size 1, disable Hires. fix and every add-on, restore Automatic · Queue · CPU, and lower GPU Weights if Forge reports too little free GPU memory. Generate the same small baseline twice.
- 01Model weightWhat is loaded
- 02Inference headroomSpace to generate
- 03OffloadWhat moves outside GPU memory
This page diagnoses a failure. For a passing workload that you want to tune or benchmark, use GPU Weights, swap and memory controls. The two pages intentionally do not share the same job.
Where did the first allocation fail?
Choose the earliest stage you can prove from the console. The panel gives one recovery state and one controlled next test.
- Signal
- The traceback appears before sampling progress. Nearby lines may name the checkpoint, text encoder, T5, CLIP, VAE, “Loading Model”, or “Moving model(s)”.
- Do now
- Return Diffusion in Low Bits to Automatic, use the correct model-family preset, remove optional modules, lower GPU Weights, and try the smallest supported package for that exact model family.
- Controlled test
- Keep the same file set for the retest. A different quantization or checkpoint is a different workload, not proof that the original file was broken.
- Do not cross
- If the base file set still fails before sampling, record every component filename and the load log. Do not compensate with larger dimensions or launch flags.
- Signal
- Progress begins but stops before a completed latent. Forge may print its Low GPU VRAM Warning or the traceback may pass through the sampler, UNet, attention, ControlNet, or conditioning code.
- Do now
- Use batch size 1, return to the last passing dimensions, disable Hires. fix, ControlNet and LoRAs, and lower GPU Weights while leaving Queue and CPU unchanged.
- Controlled test
- Repeat the same seed twice. Then restore one workload component at a time: dimensions, one LoRA, one ControlNet unit, and only then Hires. fix.
- Do not cross
- VAE tiling targets encode/decode; it does not repair an allocation that already failed during sampling.
- Signal
- The first pass completes. The traceback then names sample_hr_pass, an upscaler, a second sampler, VAE encode, or VAE decode.
- Do now
- Disable Hires. fix and save the passing base receipt. Re-enable it with batch 1, a smaller target, one upscaler, and no ControlNet or LoRA.
- Controlled test
- Record whether failure happens before the second sampling pass, during that pass, or during the final VAE operation. Those branches need different tests.
- Do not cross
- A 512 × 512 base image followed by a larger second pass is not a 512 × 512 memory test.
- Signal
- Progress reaches 100%, a preview may exist, and the console then names VAE, decode_first_stage, encode, or the current tiled-VAE retry warning.
- Do now
- Return to the last passing output size and verify the intended VAE. As one isolated experiment, enable Never OOM “Enabled for VAE (always tiled)”.
- Controlled test
- Current Forge code catches its OOM exception during regular VAE encode/decode and retries the tiled path. Preserve both the original error and the result of that retry.
- Do not cross
- Do not enable UNet maximum offload at the same time unless sampling also failed; it changes another stage.
- Signal
- The same base model, seed, dimensions and batch pass until one add-on is enabled. Current sampling code reserves additional inference memory for linked ControlNet and online LoRA data.
- Do now
- Prove the base image, then add one exact LoRA or one ControlNet unit. For an extension, retest once with all third-party extensions disabled.
- Controlled test
- Keep the add-on file, weight, model family and preprocessor fixed. “Any LoRA” or “some extension” is not a reproducible boundary.
- Do not cross
- Do not diagnose the add-on by reinstalling Python, changing the model and adding multiple memory flags in one run.
- Signal
- A cold run passes, but identical repeats or a named A → B → A switch increase RAM/VRAM until a later failure. One high reading after model load is not enough to call a leak.
- Do now
- Restart once to reset the process, then run a fixed sequence with no add-ons: cold A, warm A, warm A, B, A. Record RAM, dedicated VRAM and shared memory after every completed stage.
- Controlled test
- Look for a plateau versus monotonic growth. Reproduce in a separate clean install on the same commit before filing a leak or regression report.
- Do not cross
- Restarting proves that process state mattered; it does not identify which object, extension, driver, or cache retained memory.
- Signal
- Task Manager or htop shows host RAM climbing; the OS may become unresponsive, print “Killed”, report a CPU allocation failure, or terminate Forge without a CUDA OOM.
- Do now
- Stop the queue, close other heavy applications, return Swap Location to CPU and Swap Method to Queue, remove add-ons, and use the smallest last-passing file set.
- Controlled test
- Record system RAM, committed/page-file use, shared GPU memory, free disk and the exact stage. Offloaded weights still need host capacity.
- Do not cross
- A larger page file may delay termination but does not prove the workload is practical, and random pinned/shared-memory flags can make this failure harsher.
Return to one image that passes
Each step removes ambiguity. When the minimum workload passes, restore costs in the order you can measure.
- 01
Stop automatic repeats
Cancel the queue. Do not press Generate until you have saved the earliest useful error and the last passing settings.
- 02
Preserve the first failure
Copy the complete traceback and the console memory lines above it. Later cleanup errors can be secondary.
- 03
Reset the process once
Close Forge normally if possible, then start the same inspected installation. A restart creates a clean comparison; it is not the final diagnosis.
- 04
Build the minimum workload
Correct preset and base model, batch 1, family-appropriate base dimensions, no Hires. fix, ControlNet, LoRA, custom VAE, or unrelated extension.
- 05
Use the neutral memory state
Automatic · Queue · CPU · preset GPU Weights. If Forge warns about low free GPU memory, lower GPU Weights and repeat without changing swap.
- 06
Add one cost at a time
Restore dimensions first, then batch or Hires. fix, then one LoRA or ControlNet unit. Keep the same seed and record the first failing change.
Match the traceback to the next test
Read the last successful stage, then the first failing call. The most useful line is often above the final repeated CUDA error.
| When it fails | Likely pipeline area | Evidence to keep | Smallest next test |
|---|---|---|---|
| Before progress | Checkpoint, VAE, text encoder, T5 or other model load | Exact file set; preset; storage/quantization; first load traceback | Smaller supported file set; Automatic; lower GPU Weights; no optional modules |
| At prompt setup | Conditioning or text encoder | Last named encoder; prompt; model package; allocation request | Base prompt; correct companion encoders; one fixed package |
| During progress | UNet / sampler / ControlNet / online LoRA | Step reached; dimensions; batch; add-ons; Low GPU VRAM Warning | Batch 1; last passing size; no add-ons; lower GPU Weights |
| After base pass | Hires. fix, upscale or second sampler | Base size; target size; upscaler; second-pass stage | Prove base, then smaller target with no add-ons |
| After 100% | VAE encode/decode or save path | VAE name; regular and tiled retry messages; final size | Correct VAE; smaller output; isolated VAE tiling test |
| After several runs | Retained process state or model switching | Fixed sequence and a memory reading after each stage | Clean restart; identical repeats; clean-install comparison |
| Whole machine stalls | System RAM, shared memory, page file or another process | RAM/commit/shared values; free disk; OS kill line | Queue + CPU; close other apps; smaller package; no add-ons |
VRAM, reserved memory, and RAM are not interchangeable
Use the message to choose the measurement. Do not translate every crash into “not enough VRAM.”
CUDA out of memory · torch.cuda.OutOfMemoryError
GPU allocation failedStart with the earliest CUDA OOM, its requested allocation, and the pipeline stage—not the cleanup traceback that follows.
Allocated versus reserved in a PyTorch error
Live tensor memory versus caching-allocator memoryReserved memory can exceed active tensor memory. The numbers help describe the failure but do not choose a Forge setting by themselves.
High VRAM in nvidia-smi after a run
May include PyTorch cached blocks and other processesPyTorch documents that unused cached memory can remain visible. A high display alone is not proof of a leak.
CPUAllocator · fatal: Memory allocation failure · Killed
Host/system memory pathCheck RAM, commit/page file, shared GPU memory, free disk, and the stage. Do not label it CUDA OOM without CUDA evidence.
Browser disconnect · Error 1006 · Press any key
Secondary UI symptomDetermine whether the Forge process died. The console or OS event is the primary evidence.
A completed sampler can still be followed by an OOM
Hires. fix and VAE work are later stages with their own image-sized allocations.
Hires. fix fails
Save the base receipt, then record base size, target size, upscaler and whether the failure happens in VAE decode, the second sampler, or final decode. Reduce the target and keep batch 1.
VAE fails
Current Forge code retries regular VAE encode/decode with a tiled path after its OOM exception. If that retry also fails, keep both messages and test a smaller final size or the correct VAE.
Never OOM for VAE
The integrated option forces VAE tiling. Use it only when encode/decode is the proven stage. It may make a constrained output possible, but it is not a sampler-memory control.
Prove the file set before blaming the GPU
FLUX, companion encoders, LoRAs and ControlNet change what must be loaded or preserved. The filename and format belong in every result.
Model package
Record diffusion model, storage format or quantization, CLIP/T5 files and VAE. “FLUX” alone does not describe the memory contract.
GPU Weights
Do not set it to the full card capacity. It is weight allocation; Forge still needs free GPU memory for conditioning and matrix computation.
Host capacity
CPU offload moves pressure, not cost. Watch RAM, committed memory, page-file activity and free disk separately from dedicated VRAM.
Add-on boundary
If the base passes and one exact add-on fails, that is useful evidence. Preserve the file, family, weight and clean-extension result.
No universal 4/6/8/12/24 GB table: the local videos expose useful variables, but their fixed hardware claims and timings are dated single-system snapshots, not compatibility requirements.
Measure a sequence, not a feeling
Cached memory, retained live state, model switching and host pressure can look similar in one screenshot.
- Use one commit, one extension-free installation and the same fixed generation receipt.
- After every completed stage, record system RAM, dedicated VRAM, shared memory and whether memory reaches a plateau.
- Repeat the exact sequence after a clean restart; then repeat it in a separate clean install if growth returns.
- File a report only with the smallest sequence that reproduces monotonic growth or the switch-specific OOM.
Do not stack five memory fixes
Every added flag changes the experiment. The safest diagnosis begins in the UI and changes one variable.
--always-offload-from-vram
Evidence: The dated maintainer post describes maximum offload as slower but less risky for unusual OOM or coexistence cases.
Boundary: Test only after the baseline names persistent model residency as the constraint.
--cuda-stream
Evidence: The dated maintainer post records possible OOM on older devices or large resolutions and black/NaN output.
Boundary: Do not add while diagnosing stability. Return to Queue first.
--pin-shared-memory
Evidence: The dated maintainer post warns shared-memory OOM can be more severe and may crash without dynamic recovery.
Boundary: Do not use as a generic system-RAM fix. Return Swap Location to CPU.
--disable-gpu-warning
Evidence: Current sampler code says this suppresses the warning and accepts fallback risk.
Boundary: It does not create inference headroom; do not use it as a remedy.
PYTORCH_ALLOC_CONF / PYTORCH_CUDA_ALLOC_CONF
Evidence: PyTorch exposes advanced allocator configuration; the older name remains an alias in current docs.
Boundary: Do not paste a generic value before a reproducible Forge baseline. It can change the allocator, not the workload.
Empty cache is not extra VRAM
PyTorch separates memory occupied by tensors from memory reserved by its caching allocator.
Live tensor memory
memory_allocated reports memory occupied by tensors. Live tensors cannot be freed by emptying unused cache blocks.
Allocator-managed memory
memory_reserved reports memory managed by the caching allocator. Some unused reserved blocks may remain visible in nvidia-smi.
Releases unoccupied blocks
PyTorch explicitly says empty_cache() does not increase memory available to PyTorch. It may reduce fragmentation in some cases.
Snapshot with limits
PyTorch memory snapshots can capture allocator OOM events, but only allocations visible to the PyTorch allocator. Use them for an escalated reproducible case, not as the first user step.
Make the failure reproducible
Check only fields you actually captured. The complete state becomes the compact input for a useful issue report.
Before publishing logs or Sysinfo, remove usernames, local paths, prompts, tokens, remote URLs and other private data.
Current mechanics stay separate from user reports
The original repository defines behavior. Issues and videos reveal questions and failure shapes, but they do not become universal product facts.
VERIFIED Primary implementation and upstream allocator docs
- Original Forge commit dfdcbabExact technical snapshot inspected for this page.
- Forge memory managerWeight placement, inference reservation, offload, free-memory and cache mechanics.
- Forge sampling functionLow-VRAM warning plus ControlNet and online-LoRA memory preservation.
- Forge VAE patcherRegular encode/decode and tiled retry after the current OOM exception.
- Forge memory-control UIExact preset, GPU Weights, swap and low-bits labels.
- Never OOM implementationSeparate UNet maximum-offload and VAE always-tiled behavior.
- Maintainer NeverOOM postDated risks and trade-offs for offload, stream and pinned memory.
- Maintainer FLUX guideDated FLUX packaging, low-bits and swap context.
- Maintainer performance-reporting guideControlled comparison requirements.
- PyTorch CUDA memory managementAllocated, reserved and caching-allocator semantics.
- PyTorch empty_cache referenceWhat unused cache release can and cannot do.
- PyTorch memory snapshotsAdvanced allocator trace and visibility boundary.
COMMUNITY-REPORTED Original-repository issue patterns
- Model switching OOMOpen report on commit-sensitive model switching; Never OOM changed the reporter’s result.
- FLUX fills system RAMOpen report separating 99% host RAM pressure from VRAM and noting the final stage.
- LoRA fills system RAMOpen report where the base workflow passes and adding any LoRA changes host RAM behavior.
- RAM grows across operationsClosed report spanning generation, model changes, merging and Linux OOM killing; no universal cause established.
- CUDA OOM during FLUX conditioningOpen report whose traceback fails in T5 conditioning after a high weight allocation.
- Fatal allocation at VAE decodeOpen report where sampling completes and the later upscale/decode path fails.
- Hires. fix VAE CUDA OOMClosed report whose traceback reaches the Hires pass and VAE decode before CUDA OOM.
Issue state does not mean root cause or repair was verified. Hardware, commit, extension state and workload differ.
STALE SNAPSHOT Dated video transcripts
- 8 GB comparisonA dated benchmark that changes dimensions and tiled VAE behavior; its timings and percentage claims are not carried forward.
- FLUX low-VRAM walkthroughA dated single-system walkthrough using a 24 GB allocation and fixed RAM/VRAM advice; useful as vocabulary, not a requirement.
Subtitles were reviewed for the page. Fixed timings, speed percentages and minimum-memory claims are intentionally not generalized.
Forge out-of-memory answers
Short answers for the questions that usually appear after the first crash.
How do I fix CUDA out of memory in Stable Diffusion WebUI Forge?
Find the first failing stage, then return to batch 1, the last passing dimensions, no Hires. fix, ControlNet or LoRA, and Automatic · Queue · CPU. Lower GPU Weights if the sampler has too little free GPU memory. Add one workload component back at a time.
Why does Forge run out of memory when my GPU has 12 GB or 24 GB?
VRAM capacity alone does not define the workload. Model and encoder files, storage precision, GPU Weights, dimensions, batch, conditioning, ControlNet, LoRAs, Hires. fix, VAE work, extensions and other GPU processes all affect the first failing allocation.
Should I set GPU Weights to my full VRAM amount?
No. GPU Weights allocates memory to model weights while generation still needs inference headroom. Current Forge code explicitly tells users to lower GPU Weights when free memory for a diffusion iteration is below its warning boundary.
Can lowering GPU Weights fix Forge CUDA OOM?
It can return inference headroom and is a good controlled test for a sampling or conditioning failure. It cannot by itself prove that a model package, VAE, add-on, host RAM, or retained-state problem is fixed.
Why does the base image work but Hires. fix runs out of memory?
Hires. fix adds an upscale and another processing path at a larger target. Record whether the failure occurs during VAE work, the second sampling pass, or final decode; a passing base image does not test those allocations.
How do I fix Forge VAE decode out of memory?
Return to the last passing output size and verify the intended VAE. If sampling completed and the error names VAE encode/decode, test Never OOM “Enabled for VAE (always tiled)” as the only change and preserve the regular and tiled results.
What is Never OOM in Forge?
The current integrated script has separate switches: UNet forces the NO_VRAM state and maximum offload, while VAE forces tiled encode/decode. Both trade normal behavior for a constrained path and should be chosen by the failing stage.
Should I enable both Never OOM options?
Not for diagnosis. Enable only the switch that addresses the proven stage. VAE tiling does not fix sampling, and UNet offload changes model movement even when only final decode failed.
Does batch size use more VRAM than Batch count?
The active batch increases the tensors processed together and is a direct memory variable. Batch count schedules separate work; begin diagnosis with batch size 1 and one queued result so the failing run is unambiguous.
Why do larger image dimensions cause CUDA OOM?
Larger spatial tensors require more runtime memory during sampling and especially image-space VAE work. Return to the last passing dimensions, then increase one dimension target at a time with batch 1.
Why does ControlNet cause Forge out of memory?
Current Forge sampling code adds linked ControlNet model and inference-memory requirements before sampling. Prove the base run, then test one compatible ControlNet unit with a fixed preprocessor, model, dimensions and control image.
Why does a LoRA fill RAM or VRAM in Forge?
A LoRA changes the model patching workload, and current code reserves memory for online LoRA data. User reports also describe host-RAM failures, but they do not establish a universal Forge bug. Reproduce one exact LoRA against a passing base model.
Why does FLUX use so much system RAM in Forge?
FLUX packages can include large diffusion and text-encoder components, while offloaded pieces need host capacity. Use the correct preset and documented file set, keep GPU Weights below the no-headroom boundary, remove add-ons, and record RAM separately from VRAM.
How much RAM or VRAM does Forge need for FLUX?
There is no universal number supported by the inspected evidence. File format, quantization, encoders, dimensions, batch, add-ons, Forge commit, Torch/CUDA build and swap behavior change the result. Test a specific package and publish the receipt.
Why is system RAM at 99% while VRAM is not full?
Offloaded weights, encoders, add-ons, shared-memory behavior and other applications can consume host memory even when dedicated VRAM has headroom. Record the stage, committed memory, page-file activity and exact file set.
Is shared GPU memory the same as dedicated VRAM?
No. On supported systems, shared GPU memory uses host-backed capacity and can coincide with heavy system-RAM pressure and slower transfers. Forge’s Shared swap path is a separate, environment-sensitive choice.
Should I increase the Windows page file to fix Forge OOM?
A page file can delay an operating-system allocation failure when disk space is available, but it does not reduce the workload or guarantee usable speed. First prove the smallest file set and measure committed memory and paging.
How do I prove Forge has a memory leak?
Run a fixed, extension-free sequence on one commit: cold A, repeated warm A runs, then a named model-switch sequence. Record RAM, dedicated VRAM and shared memory after every stage. Look for reproducible monotonic growth rather than one high cached reading.
Why does Forge OOM when switching checkpoints?
The repository has an open user report of model-switch OOM, but that does not make every switch failure the same bug. Record the exact A → B → A sequence, file formats, clean-install result and memory after each load.
Does restarting Forge fix a memory leak?
Restarting clears process state and creates a clean baseline. If it helps, that proves only that process state mattered; it does not identify whether Forge core, an extension, a driver, an allocator cache, or a model patch retained memory.
Will torch.cuda.empty_cache() fix CUDA out of memory?
PyTorch states that empty_cache releases only unoccupied cached memory and does not increase the memory available to PyTorch. It may help fragmentation in some cases, but it cannot free live tensors or replace a smaller workload.
Why does nvidia-smi still show high VRAM after generation?
PyTorch uses a caching allocator, so unused cached blocks may remain visible in nvidia-smi. Compare allocated, reserved and workload-stage evidence before calling the reading a leak.
Should I paste a PYTORCH_CUDA_ALLOC_CONF max_split_size_mb fix?
Not as a first step. Allocator configuration is an advanced PyTorch control and a generic value is not an official Forge repair. Preserve the failing workload first; otherwise you cannot tell whether allocator behavior or workload reduction changed the result.
What should I include in a Forge OOM bug report?
Include the original repository identity, commit, install route, GPU/VRAM, driver, RAM and page file, exact files, preset and memory controls, complete workload, add-ons, earliest traceback, last passing case, clean-install result and a repeatable failure sequence.
What does the Forge Low GPU VRAM Warning mean?
Current sampler code prints it when free GPU memory for that diffusion iteration is below its internal safe boundary. Preserve the full warning, lower GPU Weights, keep Queue and CPU fixed, and repeat the same workload. Suppressing the warning does not add headroom.
Why does CUDA OOM say “Tried to allocate 64 MiB” when some VRAM looks free?
The failed request is the next block, not the workload’s total cost. Live tensors, allocator-reserved blocks, other processes, fragmentation, and allocations outside PyTorch’s visible pool can make one apparently small request fail. Save allocated, reserved, free and stage evidence together.
Will lowering sampling steps fix Forge out of memory?
It is not the primary capacity test. Current Forge estimates sampling memory from the active tensor shape and additional models, while fewer steps mainly shorten how long that workload runs. Start with batch, dimensions, add-ons and inference headroom; keep steps fixed during diagnosis.
Should I reinstall Forge, delete the venv, or reboot Windows after an OOM?
Restart Forge once to clear process state and close other GPU-heavy applications. Do not rebuild the Python environment for a generation-stage allocation failure unless separate environment evidence points there. A clean install is useful later as a controlled comparison, not as the first repair.