GPU WEIGHTS · HEADROOM · OFFLOAD · ONE CHANGE AT A TIME

Give model weights room without starving generation

Start with Automatic · Queue · CPU and the preset GPU Weights value. If Forge warns, slows down, or runs out of memory, lower the weight allocation and retest the exact same workload before changing swap behavior.

Current code inspected · original lllyasviel/stable-diffusion-webui-forge onlydfdcbab
26 Jun 2025 · checked 24 Aug 2026
01 · Scope gate

This guide tunes a passing or measurable workload

It does not certify a GPU, invent a recommended MB value, or replace diagnosis of a repeatable crash.

USE THIS PAGE

You can reach the UI and select a model

You need to understand the current controls, preserve headroom, or test one safer memory change.

The operating rule

Do not maximize GPU Weights. Keep a known-good baseline, read the earliest failing stage, and move one boundary at a time.

02 · Memory model

Three resources share one job

A model fitting in dedicated VRAM does not prove that sampling, decoding, or a larger workflow will fit.

RESOURCE RAILCONCEPTUAL · NOT A BENCHMARK
  1. 01Model weightWhat is loaded
  2. 02Inference headroomSpace to generate
  3. 03OffloadWhat moves outside GPU memory
WEIGHTS

Persistent model storage

Checkpoint, diffusion model, text encoders, VAE and patches do not all have the same lifetime or placement.

HEADROOM

Active work space

Sampling tensors, attention, conditioning, resolution, batch and workflow add-ons need memory beyond the stored weights.

OFFLOAD

Transfer instead of disappearance

Moving weights outside dedicated VRAM uses another location and transport path. It can trade capacity for transfer cost.

The rail is explanatory. It does not calculate whether a model will run on a named GPU because model format, resolution, batch, add-ons, driver, Torch build, system RAM, commit and swap behavior change the result.

03 · Exact controls

Know what each switch changes

The labels and current defaults below come from the inspected original Forge code.

Control → current choices → mechanism → safe interpretation
ControlCurrent choicesWhat it changesSafe interpretation
UI presetsd · xl · flux · allShows a workload-oriented subset and applies preset defaults.Changing preset can change visibility and reset memory-related values. Use all if a referenced control is hidden.
GPU Weights (MB)0 → detected total VRAMSets the model-weight allocation; Forge derives inference memory as detected total VRAM minus this value.Maximum leaves essentially no compute headroom. The preset value is a starting observation, not a universal optimum.
Diffusion in Low BitsAutomatic plus explicit NF4, FP4 and float8 choicesChanges diffusion-weight storage/loading behavior; “fp16 LoRA” variants also change LoRA handling.Automatic is the neutral baseline. Do not force a second format onto an already quantized file without a documented reason.
Swap MethodQueue · AsyncQueue moves then computes; Async attempts to overlap movement and compute with separate workers/streams.Queue is the current default. Async can help or starve computation of headroom; measure it on the same workload.
Swap LocationCPU · SharedChooses ordinary CPU offload or pinned/shared GPU-memory behavior.CPU is the current default. Shared has dated reports of gains and crashes; it is not a safe universal recommendation.
Never OOM IntegratedUNet offload · VAE tiledForces the UNet path to NO_VRAM or always tiles VAE encode/decode.Use for a named constraint or failure stage. Maximum offload and tiling can trade speed for capacity.
04 · Safe baseline

Prove one fixed job before tuning

The baseline creates a known return point and separates model loading from repeated generation.

  1. 01

    Record identity first

    Save the original Forge commit, install route, launch arguments, exact model URL, filename, format and hash.

    FIXED
  2. 02

    Select the workload preset

    Choose sd, xl, or flux before recording controls. Use all only when you need to expose a hidden control.

    VISIBLE
  3. 03

    Restore the neutral transport

    Use Diffusion in Low Bits: Automatic, Swap Method: Queue, Swap Location: CPU, and the preset GPU Weights value.

    BASELINE
  4. 04

    Remove optional pressure

    Batch size 1; no Hires. fix, ControlNet, LoRA, refiner, face restoration or unrelated extensions. Use the model family’s documented dimensions.

    ISOLATED
  5. 05

    Run cold, then warm

    Generate the same seed at least twice. Record model-load time separately from later warm sampling.

    MEASURED
  6. 06

    Read the console before changing anything

    Capture memory-management lines, the first warning or traceback, output result, dedicated/shared memory and system RAM.

    EVIDENCE
The preset is a starting observation, not a hardware guarantee.

Current code generally derives preset GPU Weights from detected VRAM while preserving a nominal margin. The target workload can require more headroom, so the console and repeat test decide whether to lower it.

Inspect preset behavior ↗
05 · Tuning desk

Choose the state you actually have

The desk gives one next move, the settings to hold constant, and the evidence needed to accept or reject the change.

Current state
SAFE NEXT MOVEEstablish the neutral baseline
Do now
Select the correct UI preset. Leave Diffusion in Low Bits on Automatic, Swap Method on Queue, Swap Location on CPU, and keep the preset GPU Weights value.
Hold constant
Use one documented model, batch size 1, no Hires. fix, LoRA, ControlNet, or unrelated extensions. Generate twice so model loading is not confused with warm sampling.
Record
Exact model file and format, Forge commit, preset, dimensions, console memory lines, dedicated VRAM, shared memory, system RAM, cold time, and warm time.
Stop / route
Do not optimize a setup that has not completed one clean base image.
06 · Capacity tools

Use NeverOOM by failure stage

The integrated controls remain in current code, but their 2024 announcement is a dated capability demonstration—not a command to enable both.

UNET / DIFFUSION

Enabled for UNet (always maximize offload)

Current code switches the memory manager to NO_VRAM and unloads models when this state changes.

Use when
Sampling-stage capacity is the isolated constraint.
Cost
More weight movement and potentially much slower generation.
VAE / IMAGE

Enabled for VAE (always tiled)

Current code forces tiled VAE encode/decode. It changes the image conversion stage, not diffusion sampling.

Use when
The failure is specifically at VAE encode or final decode.
Cost
Tiling overhead and a different processing path.
Do not start with launch flags

--always-offload-from-vram, --cuda-malloc, --cuda-stream, and forced VRAM states exist, but they change the baseline globally. The 2024 maintainer note calls always-offload slower and cudaMallocAsync riskier.

Never hide the diagnostic first

--disable-gpu-warning suppresses the current warning; it does not create headroom. Current code explicitly marks that route as highly discouraged.

07 · Signal decoder

Route the earliest memory signal

Later browser symptoms are less useful than the first console warning and the pipeline stage where it appeared.

“Low GPU VRAM Warning” appears before or during sampling
What it means
The current sampler found less free GPU memory than its internal warning threshold for that iteration.
Next controlled check
Lower GPU Weights, keep Queue/CPU, repeat the identical workload, and preserve the complete warning. Do not use --disable-gpu-warning as a fix.
Model loads, then generation becomes about ten times slower
What it means
A common official failure pattern is too much weight allocation leaving too little compute headroom and triggering fallback/shared-memory pressure.
Next controlled check
Lower GPU Weights and compare warm runs. Confirm whether shared GPU memory activates; do not assume that “fully loaded” means fast.
CUDA out of memory while the model is loading
What it means
The selected file set or weight allocation does not fit the current load path.
Next controlled check
Verify the package, remove extra modules, use Automatic, lower GPU Weights, and test a supported lower-storage distribution if needed.
CUDA out of memory during sampling
What it means
The model loaded, but dimensions, batch, conditioning, ControlNet, LoRA, or another runtime allocation consumed the remaining headroom.
Next controlled check
Return to batch 1 and baseline dimensions; disable add-ons; lower GPU Weights; then add one workload component at a time.
The preview completes, but VAE decode fails
What it means
The failure is later than diffusion sampling and may be specific to VAE encode/decode memory.
Next controlled check
Test the documented VAE and smaller output first. Then test Never OOM “Enabled for VAE (always tiled)” as one isolated change.
Shared works once, then Forge crashes or freezes
What it means
Shared is a device- and environment-sensitive swap location, not a compatibility guarantee.
Next controlled check
Return to CPU, keep Queue, capture system RAM/shared-memory behavior and the last console lines, then retest.
Changing a preset moved or hid memory controls
What it means
Current code changes control visibility and values for sd, xl, flux, and all.
Next controlled check
Choose the correct preset first. Use all to inspect the controls; record values again after every preset change.
Memory grows across repeated runs or model switches
What it means
This is not proven to be normal tuning behavior. The repository contains unresolved and historical user reports with different causes.
Next controlled check
Reproduce in a clean install with the same model and no extensions; capture each run and switch. Route the evidence to OOM/update diagnosis instead of stacking flags.
08 · Test receipt

Keep the run reproducible

A time or memory screenshot without the workload and control state cannot establish which change helped.

Memory-control test record
RECORDED EVIDENCEOne variable per comparison
0 / 8 recorded
Complete records can be compared; partial ones remain observations.
09 · Memory FAQ

Exact questions users ask while tuning Forge

These answers preserve the boundary between current code behavior, dated maintainer experiments and hardware-specific community reports.

Where is GPU Weights in Stable Diffusion WebUI Forge?

It is a top-level Forge control labeled GPU Weights (MB). The flux and all presets expose it in the inspected code; if you cannot see it, select all in the UI preset area and confirm you are using the original Forge repository.

What does GPU Weights mean in Forge?

It is the memory allocation for model weights, not a cap on total VRAM use. Forge calculates inference memory as detected total VRAM minus the GPU Weights value, so generation still needs the remainder.

Should I set GPU Weights to maximum VRAM?

No. The current UI and official performance guidance warn that maximum weight allocation can leave no memory for matrix computation, causing fallback, OOM, or severe slowdown.

What is the best GPU Weights setting for 4 GB, 6 GB, 8 GB, 12 GB, or 24 GB VRAM?

There is no universal number. Start from the preset value, run a fixed baseline, and lower the value if Forge reports low free VRAM, uses shared memory unexpectedly, becomes much slower, or fails at the target workload.

What is inference headroom?

It is GPU memory left for the active computation and temporary runtime allocations after weight placement. Resolution, batch, conditioning and add-ons can increase the required headroom.

Why did Forge become extremely slow after I increased GPU Weights?

The model may have occupied memory needed for computation. Restore Queue and CPU, lower GPU Weights, and compare the same warm run while watching the low-VRAM warning and shared-memory use.

What is the difference between Queue and Async swap in Forge?

Queue moves a layer and then computes it. Async attempts to move layers and compute concurrently with separate workers/streams. Async can be faster on a passing setup but can also move too much data and starve computation.

Should I use Queue or Async in Forge?

Use Queue for the first reproducible baseline. Test Async only after Queue is stable, changing no other variable and reverting if time, memory, stability, or output regresses.

What is the difference between CPU and Shared swap location?

CPU keeps the offloaded portion in ordinary CPU memory. Shared enables Forge’s pinned/shared-memory path. Shared may improve transfer behavior on some systems but has dated reports of crashes on others.

Why does Shared swap crash Forge?

The official 2024 tuning post explicitly notes that some devices crash with Shared. Treat that as an environment-specific incompatibility signal and return to CPU rather than forcing the option.

What does Diffusion in Low Bits: Automatic do?

Automatic lets the current loader choose its default storage behavior for the selected model instead of forcing NF4, FP4, or float8 from the UI. It is the cleanest first-test state.

Should I choose bnb-nf4 for a GGUF or already quantized checkpoint?

Not by default. GGUF or an NF4-packaged checkpoint already has a storage contract. Keep Automatic unless the exact model documentation and a controlled test require an explicit Forge override.

Does Diffusion in Low Bits improve image quality?

It is a storage/loading choice, not a quality control. Precision can affect compatibility, memory, speed, LoRA behavior and sometimes output, so compare it as a named experiment rather than a quality slider.

Should I enable Never OOM all the time?

No. The current integrated script forces maximum UNet offload or tiled VAE operation. Those can make a constrained job possible but can add transfers or tiling overhead when the normal path already fits.

What is the difference between Never OOM for UNet and VAE?

UNet changes the VRAM state to NO_VRAM and maximizes model offload. VAE always tiles image encode/decode. Choose by the stage that fails; VAE tiling does not solve a sampling-stage allocation.

Can lowering GPU Weights fix CUDA out of memory?

It can return headroom to inference and is an appropriate controlled test, but it is not the only cause. Record whether the OOM occurred at model load, sampling, VAE decode, or a model switch.

Why does a larger image fail when the smaller image works?

A successful model load proves only that workload. Larger dimensions require more runtime memory; Hires. fix, batch, ControlNet and other add-ons create additional allocations.

Do I need more system RAM or a larger page file for Forge swap?

Offload needs available system memory, and the official FLUX post mentions system swap as a crash fallback. Neither is a substitute for a workload that exceeds practical memory limits; measure RAM, page-file activity and disk capacity.

Why is the first Forge generation slower than the second?

The first result can include model loading, conversion, patching and cache setup. Record cold load plus first generation separately from repeated warm generation before calling a setting faster or slower.

How should I benchmark Forge memory settings?

Keep the exact model, format, commit, seed, sampler, scheduler, steps, dimensions, batch, precision and extensions fixed. Change one memory control, run multiple warm passes, save full console logs, and report failures as well as times.

Why did my GPU Weights value change after selecting another preset?

The inspected preset handler applies separate stored or calculated values for sd, xl, flux and all. Select the workload preset before recording or tuning memory settings.

Should I use --disable-gpu-warning?

Not as a remedy. Current code describes it as a risky way to suppress the warning while leaving the underlying headroom problem unchanged. Fix and measure the allocation instead.

10 · Evidence and next task

Current mechanics, dated experiments, explicit unknowns

We use tutorials to discover questions and vocabulary; product behavior is tied to original Forge code and maintainer material.

VERIFIED

Current original Forge code

Exact labels, choices, preset behavior, inference-memory calculation, stream state, warning text and NeverOOM actions were inspected at commit dfdcbab.

Inspect the UI code ↗
STALE SNAPSHOT

Official 2024 tuning notes

The maintainer explains GPU Weights, Queue/Async, CPU/Shared and NeverOOM with dated device experiments. Mechanisms still present in code are retained; percentages and hardware outcomes are not generalized.

Inspect discussion #981 ↗
COMMUNITY-REPORTED

Crashes and repeated-run symptoms

Issues show real user language around model switching, Shared memory and NeverOOM, but unresolved reports do not prove one universal cause or fix.

Inspect issue #2127 ↗
Transcripts reviewedV27 ↗V33 ↗
NEXT DECISION

Memory baseline passes. Now return to the actual workflow.

For FLUX, continue with the model baseline. If the smallest clean job still fails, capture the receipt and route the symptom instead of stacking more settings.