Give model weights room without starving generation
Start with Automatic · Queue · CPU and the preset GPU Weights value. If Forge warns, slows down, or runs out of memory, lower the weight allocation and retest the exact same workload before changing swap behavior.
lllyasviel/stable-diffusion-webui-forge onlydfdcbab26 Jun 2025 · checked 24 Aug 2026
This guide tunes a passing or measurable workload
It does not certify a GPU, invent a recommended MB value, or replace diagnosis of a repeatable crash.
You can reach the UI and select a model
You need to understand the current controls, preserve headroom, or test one safer memory change.
Forge exits, the environment fails, or the model is incomplete
Start with troubleshooting by symptom or repair the FLUX file set before tuning memory.
Do not maximize GPU Weights. Keep a known-good baseline, read the earliest failing stage, and move one boundary at a time.
Three resources share one job
A model fitting in dedicated VRAM does not prove that sampling, decoding, or a larger workflow will fit.
- 01Model weightWhat is loaded
- 02Inference headroomSpace to generate
- 03OffloadWhat moves outside GPU memory
Persistent model storage
Checkpoint, diffusion model, text encoders, VAE and patches do not all have the same lifetime or placement.
Active work space
Sampling tensors, attention, conditioning, resolution, batch and workflow add-ons need memory beyond the stored weights.
Transfer instead of disappearance
Moving weights outside dedicated VRAM uses another location and transport path. It can trade capacity for transfer cost.
The rail is explanatory. It does not calculate whether a model will run on a named GPU because model format, resolution, batch, add-ons, driver, Torch build, system RAM, commit and swap behavior change the result.
Know what each switch changes
The labels and current defaults below come from the inspected original Forge code.
| Control | Current choices | What it changes | Safe interpretation |
|---|---|---|---|
| UI preset | sd · xl · flux · all | Shows a workload-oriented subset and applies preset defaults. | Changing preset can change visibility and reset memory-related values. Use all if a referenced control is hidden. |
| GPU Weights (MB) | 0 → detected total VRAM | Sets the model-weight allocation; Forge derives inference memory as detected total VRAM minus this value. | Maximum leaves essentially no compute headroom. The preset value is a starting observation, not a universal optimum. |
| Diffusion in Low Bits | Automatic plus explicit NF4, FP4 and float8 choices | Changes diffusion-weight storage/loading behavior; “fp16 LoRA” variants also change LoRA handling. | Automatic is the neutral baseline. Do not force a second format onto an already quantized file without a documented reason. |
| Swap Method | Queue · Async | Queue moves then computes; Async attempts to overlap movement and compute with separate workers/streams. | Queue is the current default. Async can help or starve computation of headroom; measure it on the same workload. |
| Swap Location | CPU · Shared | Chooses ordinary CPU offload or pinned/shared GPU-memory behavior. | CPU is the current default. Shared has dated reports of gains and crashes; it is not a safe universal recommendation. |
| Never OOM Integrated | UNet offload · VAE tiled | Forces the UNet path to NO_VRAM or always tiles VAE encode/decode. | Use for a named constraint or failure stage. Maximum offload and tiling can trade speed for capacity. |
Prove one fixed job before tuning
The baseline creates a known return point and separates model loading from repeated generation.
- 01FIXED
Record identity first
Save the original Forge commit, install route, launch arguments, exact model URL, filename, format and hash.
- 02VISIBLE
Select the workload preset
Choose
sd,xl, orfluxbefore recording controls. Useallonly when you need to expose a hidden control. - 03BASELINE
Restore the neutral transport
Use Diffusion in Low Bits: Automatic, Swap Method: Queue, Swap Location: CPU, and the preset GPU Weights value.
- 04ISOLATED
Remove optional pressure
Batch size 1; no Hires. fix, ControlNet, LoRA, refiner, face restoration or unrelated extensions. Use the model family’s documented dimensions.
- 05MEASURED
Run cold, then warm
Generate the same seed at least twice. Record model-load time separately from later warm sampling.
- 06EVIDENCE
Read the console before changing anything
Capture memory-management lines, the first warning or traceback, output result, dedicated/shared memory and system RAM.
Current code generally derives preset GPU Weights from detected VRAM while preserving a nominal margin. The target workload can require more headroom, so the console and repeat test decide whether to lower it.
Inspect preset behavior ↗Choose the state you actually have
The desk gives one next move, the settings to hold constant, and the evidence needed to accept or reject the change.
- Do now
- Select the correct UI preset. Leave Diffusion in Low Bits on Automatic, Swap Method on Queue, Swap Location on CPU, and keep the preset GPU Weights value.
- Hold constant
- Use one documented model, batch size 1, no Hires. fix, LoRA, ControlNet, or unrelated extensions. Generate twice so model loading is not confused with warm sampling.
- Record
- Exact model file and format, Forge commit, preset, dimensions, console memory lines, dedicated VRAM, shared memory, system RAM, cold time, and warm time.
- Stop / route
- Do not optimize a setup that has not completed one clean base image.
Use NeverOOM by failure stage
The integrated controls remain in current code, but their 2024 announcement is a dated capability demonstration—not a command to enable both.
Enabled for UNet (always maximize offload)
Current code switches the memory manager to NO_VRAM and unloads models when this state changes.
- Use when
- Sampling-stage capacity is the isolated constraint.
- Cost
- More weight movement and potentially much slower generation.
Enabled for VAE (always tiled)
Current code forces tiled VAE encode/decode. It changes the image conversion stage, not diffusion sampling.
- Use when
- The failure is specifically at VAE encode or final decode.
- Cost
- Tiling overhead and a different processing path.
--always-offload-from-vram, --cuda-malloc, --cuda-stream, and forced VRAM states exist, but they change the baseline globally. The 2024 maintainer note calls always-offload slower and cudaMallocAsync riskier.
--disable-gpu-warning suppresses the current warning; it does not create headroom. Current code explicitly marks that route as highly discouraged.
Route the earliest memory signal
Later browser symptoms are less useful than the first console warning and the pipeline stage where it appeared.
“Low GPU VRAM Warning” appears before or during sampling
- What it means
- The current sampler found less free GPU memory than its internal warning threshold for that iteration.
- Next controlled check
- Lower GPU Weights, keep Queue/CPU, repeat the identical workload, and preserve the complete warning. Do not use --disable-gpu-warning as a fix.
Model loads, then generation becomes about ten times slower
- What it means
- A common official failure pattern is too much weight allocation leaving too little compute headroom and triggering fallback/shared-memory pressure.
- Next controlled check
- Lower GPU Weights and compare warm runs. Confirm whether shared GPU memory activates; do not assume that “fully loaded” means fast.
CUDA out of memory while the model is loading
- What it means
- The selected file set or weight allocation does not fit the current load path.
- Next controlled check
- Verify the package, remove extra modules, use Automatic, lower GPU Weights, and test a supported lower-storage distribution if needed.
CUDA out of memory during sampling
- What it means
- The model loaded, but dimensions, batch, conditioning, ControlNet, LoRA, or another runtime allocation consumed the remaining headroom.
- Next controlled check
- Return to batch 1 and baseline dimensions; disable add-ons; lower GPU Weights; then add one workload component at a time.
The preview completes, but VAE decode fails
- What it means
- The failure is later than diffusion sampling and may be specific to VAE encode/decode memory.
- Next controlled check
- Test the documented VAE and smaller output first. Then test Never OOM “Enabled for VAE (always tiled)” as one isolated change.
Shared works once, then Forge crashes or freezes
- What it means
- Shared is a device- and environment-sensitive swap location, not a compatibility guarantee.
- Next controlled check
- Return to CPU, keep Queue, capture system RAM/shared-memory behavior and the last console lines, then retest.
Changing a preset moved or hid memory controls
- What it means
- Current code changes control visibility and values for sd, xl, flux, and all.
- Next controlled check
- Choose the correct preset first. Use all to inspect the controls; record values again after every preset change.
Memory grows across repeated runs or model switches
- What it means
- This is not proven to be normal tuning behavior. The repository contains unresolved and historical user reports with different causes.
- Next controlled check
- Reproduce in a clean install with the same model and no extensions; capture each run and switch. Route the evidence to OOM/update diagnosis instead of stacking flags.
Keep the run reproducible
A time or memory screenshot without the workload and control state cannot establish which change helped.
Exact questions users ask while tuning Forge
These answers preserve the boundary between current code behavior, dated maintainer experiments and hardware-specific community reports.
Where is GPU Weights in Stable Diffusion WebUI Forge?
It is a top-level Forge control labeled GPU Weights (MB). The flux and all presets expose it in the inspected code; if you cannot see it, select all in the UI preset area and confirm you are using the original Forge repository.
What does GPU Weights mean in Forge?
It is the memory allocation for model weights, not a cap on total VRAM use. Forge calculates inference memory as detected total VRAM minus the GPU Weights value, so generation still needs the remainder.
Should I set GPU Weights to maximum VRAM?
No. The current UI and official performance guidance warn that maximum weight allocation can leave no memory for matrix computation, causing fallback, OOM, or severe slowdown.
What is the best GPU Weights setting for 4 GB, 6 GB, 8 GB, 12 GB, or 24 GB VRAM?
There is no universal number. Start from the preset value, run a fixed baseline, and lower the value if Forge reports low free VRAM, uses shared memory unexpectedly, becomes much slower, or fails at the target workload.
What is inference headroom?
It is GPU memory left for the active computation and temporary runtime allocations after weight placement. Resolution, batch, conditioning and add-ons can increase the required headroom.
Why did Forge become extremely slow after I increased GPU Weights?
The model may have occupied memory needed for computation. Restore Queue and CPU, lower GPU Weights, and compare the same warm run while watching the low-VRAM warning and shared-memory use.
What is the difference between Queue and Async swap in Forge?
Queue moves a layer and then computes it. Async attempts to move layers and compute concurrently with separate workers/streams. Async can be faster on a passing setup but can also move too much data and starve computation.
Should I use Queue or Async in Forge?
Use Queue for the first reproducible baseline. Test Async only after Queue is stable, changing no other variable and reverting if time, memory, stability, or output regresses.
What is the difference between CPU and Shared swap location?
CPU keeps the offloaded portion in ordinary CPU memory. Shared enables Forge’s pinned/shared-memory path. Shared may improve transfer behavior on some systems but has dated reports of crashes on others.
Why does Shared swap crash Forge?
The official 2024 tuning post explicitly notes that some devices crash with Shared. Treat that as an environment-specific incompatibility signal and return to CPU rather than forcing the option.
What does Diffusion in Low Bits: Automatic do?
Automatic lets the current loader choose its default storage behavior for the selected model instead of forcing NF4, FP4, or float8 from the UI. It is the cleanest first-test state.
Should I choose bnb-nf4 for a GGUF or already quantized checkpoint?
Not by default. GGUF or an NF4-packaged checkpoint already has a storage contract. Keep Automatic unless the exact model documentation and a controlled test require an explicit Forge override.
Does Diffusion in Low Bits improve image quality?
It is a storage/loading choice, not a quality control. Precision can affect compatibility, memory, speed, LoRA behavior and sometimes output, so compare it as a named experiment rather than a quality slider.
Should I enable Never OOM all the time?
No. The current integrated script forces maximum UNet offload or tiled VAE operation. Those can make a constrained job possible but can add transfers or tiling overhead when the normal path already fits.
What is the difference between Never OOM for UNet and VAE?
UNet changes the VRAM state to NO_VRAM and maximizes model offload. VAE always tiles image encode/decode. Choose by the stage that fails; VAE tiling does not solve a sampling-stage allocation.
Can lowering GPU Weights fix CUDA out of memory?
It can return headroom to inference and is an appropriate controlled test, but it is not the only cause. Record whether the OOM occurred at model load, sampling, VAE decode, or a model switch.
Why does a larger image fail when the smaller image works?
A successful model load proves only that workload. Larger dimensions require more runtime memory; Hires. fix, batch, ControlNet and other add-ons create additional allocations.
Do I need more system RAM or a larger page file for Forge swap?
Offload needs available system memory, and the official FLUX post mentions system swap as a crash fallback. Neither is a substitute for a workload that exceeds practical memory limits; measure RAM, page-file activity and disk capacity.
Why is the first Forge generation slower than the second?
The first result can include model loading, conversion, patching and cache setup. Record cold load plus first generation separately from repeated warm generation before calling a setting faster or slower.
How should I benchmark Forge memory settings?
Keep the exact model, format, commit, seed, sampler, scheduler, steps, dimensions, batch, precision and extensions fixed. Change one memory control, run multiple warm passes, save full console logs, and report failures as well as times.
Why did my GPU Weights value change after selecting another preset?
The inspected preset handler applies separate stored or calculated values for sd, xl, flux and all. Select the workload preset before recording or tuning memory settings.
Should I use --disable-gpu-warning?
Not as a remedy. Current code describes it as a risky way to suppress the warning while leaving the underlying headroom problem unchanged. Fix and measure the allocation instead.
Current mechanics, dated experiments, explicit unknowns
We use tutorials to discover questions and vocabulary; product behavior is tied to original Forge code and maintainer material.
Current original Forge code
Exact labels, choices, preset behavior, inference-memory calculation, stream state, warning text and NeverOOM actions were inspected at commit dfdcbab.
Official 2024 tuning notes
The maintainer explains GPU Weights, Queue/Async, CPU/Shared and NeverOOM with dated device experiments. Mechanisms still present in code are retained; percentages and hardware outcomes are not generalized.
Inspect discussion #981 ↗Crashes and repeated-run symptoms
Issues show real user language around model switching, Shared memory and NeverOOM, but unresolved reports do not prove one universal cause or fix.
Inspect issue #2127 ↗