Diagnose Forge performance before you tune it
A useful speed test starts by naming what is slow. Separate the cold model load, warm sampling loop, VAE decode, and browser response—then change one variable and keep the full console record.
VERIFIED Tested by source and code audit at original Forge commit dfdcbab. No universal speed or VRAM claim is made.
There is no single “fastest Forge setting”
Performance belongs to a complete workload and environment, not to one slider or GPU model.
Measure before optimizing
Lock the model hash, seed, sampler, scheduler, steps, dimensions and add-ons. Change one variable only.
Route the failure stage
Model load, sampling and VAE decode do not share one fix. Keep the earliest error, not only the final traceback line.
Early announcements and videos reported large gains on selected hardware. We retain their variables and vocabulary, but not their percentages as a current promise.
Time the phase, not the whole mystery
The same “slow image” complaint can originate before, during, or after diffusion. The console is the timeline that separates them.
Load
Checkpoint, text encoders, VAE and patches are prepared or moved.
- Keep
- Console model-load and movement lines
- Why it matters
- A slow first run may be a load problem, not slow sampling.
Condition
Prompts and optional networks are prepared for the selected model family.
- Keep
- Model family, LoRAs, ControlNet and prompt state
- Why it matters
- An add-on can change both memory demand and output behavior.
Sample
The denoising loop performs the repeated diffusion steps.
- Keep
- Warm run time plus s/it or it/s
- Why it matters
- This is the phase most people mean by generation speed.
Decode & save
The VAE turns latents into pixels; previews, upscaling and file writing may follow.
- Keep
- Last console stage and output timestamp
- Why it matters
- A late stall or OOM needs a different fix from slow sampling.
Seconds per iteration: lower is faster.
it/s ↑Iterations per second: higher is faster.
Check the unit before comparing numbers.Choose the symptom you actually have
The router gives one first test and sends you to the page that owns the next job. It does not guess a GPU score.
- What the evidence supports
- Model loading, patching, file reads and cache setup can be charged to the first request. Historical issue #1630 also describes a FLUX first-run-only slowdown, but it is not a universal explanation.
- Do now
- Restart once, run the exact workload, then repeat it without changing any control. Record model-load time separately from both generations.
Build a receipt another person can reproduce
A prompt screenshot is not a benchmark. Lock the artifact, workload and environment before interpreting time.
- 01Build
Repository identity, full commit, install type and whether the working tree is modified.
- 02Environment
OS, GPU model, dedicated VRAM, driver, Forge-reported Torch/CUDA, system RAM and active launch flags.
- 03Model contract
Exact filenames, hash shown by Forge, family, quantization/storage format, VAE and text encoders.
- 04Workload
Prompt, seed, sampler, scheduler, steps, width, height, batch, Hires. fix, ControlNet and LoRAs.
- 05Memory state
UI preset, Low Bits, GPU Weights, swap method/location, NeverOOM state, dedicated/shared memory evidence.
- 06Timing
Cold load, first generation, repeated warm generations, displayed s/it or it/s, decode/save delay and any failure.
- One documented cold run after restart.
- One unchanged warm-up run.
- At least three unchanged warm runs; retain every result and failure.
- Change one variable, then repeat the same sequence.
This run count is our editorial protocol, not an official Forge minimum. The maintainer’s verified requirement is comparable before/after commits, models and full console logs.
Open the complete benchmark protocol and run sheet →Do not let a proxy become the conclusion
These signals may be useful observations. None identifies performance by itself.
One fast screenshot
It omits model load, later runs, errors, shared-memory pressure and the exact environment.
Different model formats
NF4, GGUF, FP8 and full-precision files can follow different load and compute paths.
Same filename
Names can be changed. Preserve the hash reported by Forge and the complete companion-file set.
100% VRAM used
Full allocation is not proof of efficiency; computation still needs headroom.
GPU utilization alone
It does not separate transfer, compute, decode, browser latency or the monitoring engine being shown.
A cold run vs a warm run
The cold side may include file reads, model movement and initialization absent from the warm side.
Use the narrowest next page
This hub classifies the job. The linked guides carry the settings and failure procedures.
Forge performance FAQ
Short answers for the queries that usually hide a missing test variable.
Why is Stable Diffusion WebUI Forge slow?
First identify whether the delay is model load, every sampling step, VAE decode/save, or the browser interface. Then reproduce with the exact model, commit and extensions-off workload. “Forge is slow” is not one diagnosable state.
Why is the first Forge generation slower than the second?
The first request can include model loading, patching, file reads and cache setup. Measure cold load and first generation separately, then compare repeated warm runs. If only the first FLUX run stalls, keep the full console log because historical reports describe that distinct symptom.
How can I make Forge generate faster?
Begin with the correct preset and a passing base workload. Remove extensions and add-ons, keep the model and generation settings fixed, and find the slow phase. Only tune GPU Weights or swap when memory evidence points there.
What is the best performance setting for Forge?
There is no universal setting across GPUs, model families, formats, dimensions and environments. The safest baseline is the preset value with Automatic, Queue and CPU, followed by one controlled change.
Should GPU Weights equal all available VRAM?
No. Original Forge code warns that maximum weight allocation can leave no GPU memory for matrix computation and trigger fallback, OOM or a large slowdown.
Why is Forge using shared GPU memory?
The model or workload may not fit the chosen dedicated-memory allocation, or Shared was selected as the swap location. Shared memory is not extra dedicated VRAM and can make transfer or system-memory pressure part of the run.
Does low GPU utilization mean Forge is using the CPU?
Not by itself. Measure the phase, console memory lines, warm iteration rate, dedicated/shared memory and CPU activity together. A single utilization percentage cannot distinguish transfer, decode, waiting or fallback.
What do s/it and it/s mean in Forge?
s/it is seconds per diffusion iteration, so lower is faster. it/s is iterations per second, so higher is faster. Do not compare the raw numbers as if both units move in the same direction.
How many runs should I use for a Forge benchmark?
Our editorial protocol is one documented cold run, one warm-up, then at least three unchanged warm runs with every result retained. This is a repeatability rule for this guide, not an official Forge minimum.
Can I compare Forge and A1111 with the same prompt?
A prompt alone is insufficient. Lock the model hash, VAE, sampler, scheduler, steps, dimensions, batch, seed, precision, extensions and environment, then report cold and warm phases separately.
Why did Forge become slow after an update?
Record both commits and reproduce with the same model hash, settings and environment after disabling extensions. The official reporting post requests full before/after console logs; a date or “latest” label is not enough.
Should I add --cuda-malloc or --cuda-stream to fix speed?
Not as the first move. The flags exist in the inspected code, but the original announcements describe device-specific gains and failure risks. Establish a default baseline, change one flag, and keep the result only if stability and output also pass.
Does more VRAM always make Forge faster?
More headroom can reduce offload for a fixed workload, but speed still depends on GPU architecture, model format, precision, software environment and the actual task. VRAM capacity alone is not a benchmark.
How fast should Forge be with 6 GB, 8 GB, 12 GB, or 24 GB VRAM?
A VRAM tier does not determine one expected speed. Publish the GPU model, exact model files, dimensions, steps, memory settings and warm-run unit; otherwise a time from another 6 GB or 12 GB system is not a valid target.
Why does Forge take so long to load a checkpoint?
Keep model-load time separate from sampling. Record the exact checkpoint and companion files, storage format, disk location, previous loaded model, console movement lines and whether the next unchanged warm generation is normal.
Is Forge always faster than AUTOMATIC1111?
No universal claim is supported. Local research contains both faster and conflicting community results. A meaningful comparison requires the same artifacts, workload, environment and multiple warm runs.
Where can I find Forge performance statistics?
The generation interface and console expose timing and memory evidence. Preserve the full console, the generation parameters and screenshots together; a cropped iteration rate is not a complete performance record.
What this page was allowed to conclude
Primary code and maintainer guidance establish mechanics. Community reports establish symptoms and vocabulary, not universal speed.
- Commit
dfdcbab685e57677014f05a3309b48cc87383167- Inspected
- UI controls ↗ · memory manager ↗ · launch flags ↗
- Runtime boundary
- Source-audited; no local NVIDIA benchmark was run for this page.
Official 2024 NeverOOM announcement; establishes the UNet/VAE capacity trade-off and risk boundary.
Maintainer performance-reporting note; requires comparable models, commits, screenshots, and full console logs.
Original Forge UI code; verifies UI presets, Low Bits, Queue/Async, CPU/Shared, and GPU Weights labels.
Launch-era reports about low-VRAM behavior and speed; useful for intent, not a current guarantee.
Conflicting Forge/A1111 performance reports; supports controlled testing rather than a universal winner claim.
Dated 8 GB Forge/A1111 test; variables were reviewed, headline timing was not generalized.
Dated low-VRAM FLUX walkthrough; fixed hardware settings were not reused as universal advice.