LOAD · SAMPLE · DECODE · MEASURE THE RIGHT PHASE

Diagnose Forge performance before you tune it

A useful speed test starts by naming what is slow. Separate the cold model load, warm sampling loop, VAE decode, and browser response—then change one variable and keep the full console record.

VERIFIED Tested by source and code audit at original Forge commit dfdcbab. No universal speed or VRAM claim is made.

01 · Direct answer

There is no single “fastest Forge setting”

Performance belongs to a complete workload and environment, not to one slider or GPU model.

START HERERun the correct model-family preset with no third-party extensions, batch size 1, and unchanged defaults. Generate once, repeat warm, and preserve the console.
IF IT PASSES

Measure before optimizing

Lock the model hash, seed, sampler, scheduler, steps, dimensions and add-ons. Change one variable only.

IF IT FAILS

Route the failure stage

Model load, sampling and VAE decode do not share one fix. Keep the earliest error, not only the final traceback line.

STALE SNAPSHOT

Early announcements and videos reported large gains on selected hardware. We retain their variables and vocabulary, but not their percentages as a current promise.

02 · Run anatomy

Time the phase, not the whole mystery

The same “slow image” complaint can originate before, during, or after diffusion. The console is the timeline that separates them.

01

Load

Checkpoint, text encoders, VAE and patches are prepared or moved.

Keep
Console model-load and movement lines
Why it matters
A slow first run may be a load problem, not slow sampling.
02

Condition

Prompts and optional networks are prepared for the selected model family.

Keep
Model family, LoRAs, ControlNet and prompt state
Why it matters
An add-on can change both memory demand and output behavior.
03

Sample

The denoising loop performs the repeated diffusion steps.

Keep
Warm run time plus s/it or it/s
Why it matters
This is the phase most people mean by generation speed.
04

Decode & save

The VAE turns latents into pixels; previews, upscaling and file writing may follow.

Keep
Last console stage and output timestamp
Why it matters
A late stall or OOM needs a different fix from slow sampling.
s/it ↓

Seconds per iteration: lower is faster.

it/s ↑

Iterations per second: higher is faster.

Check the unit before comparing numbers.
03 · Performance triage

Choose the symptom you actually have

The router gives one first test and sends you to the page that owns the next job. It does not guess a GPU score.

Observed behavior
FIRST CONTROLLED MOVESeparate cold setup from warm generation
What the evidence supports
Model loading, patching, file reads and cache setup can be charged to the first request. Historical issue #1630 also describes a FLUX first-run-only slowdown, but it is not a universal explanation.
Do now
Restart once, run the exact workload, then repeat it without changing any control. Record model-load time separately from both generations.
Build a comparable receipt
04 · Comparable evidence

Build a receipt another person can reproduce

A prompt screenshot is not a benchmark. Lock the artifact, workload and environment before interpreting time.

  1. 01
    Build

    Repository identity, full commit, install type and whether the working tree is modified.

  2. 02
    Environment

    OS, GPU model, dedicated VRAM, driver, Forge-reported Torch/CUDA, system RAM and active launch flags.

  3. 03
    Model contract

    Exact filenames, hash shown by Forge, family, quantization/storage format, VAE and text encoders.

  4. 04
    Workload

    Prompt, seed, sampler, scheduler, steps, width, height, batch, Hires. fix, ControlNet and LoRAs.

  5. 05
    Memory state

    UI preset, Low Bits, GPU Weights, swap method/location, NeverOOM state, dedicated/shared memory evidence.

  6. 06
    Timing

    Cold load, first generation, repeated warm generations, displayed s/it or it/s, decode/save delay and any failure.

ASSUMPTIONForge Field Guide repeatability protocol
  1. One documented cold run after restart.
  2. One unchanged warm-up run.
  3. At least three unchanged warm runs; retain every result and failure.
  4. Change one variable, then repeat the same sequence.

This run count is our editorial protocol, not an official Forge minimum. The maintainer’s verified requirement is comparable before/after commits, models and full console logs.

Open the complete benchmark protocol and run sheet →
05 · Interpretation

Do not let a proxy become the conclusion

These signals may be useful observations. None identifies performance by itself.

01

One fast screenshot

It omits model load, later runs, errors, shared-memory pressure and the exact environment.

02

Different model formats

NF4, GGUF, FP8 and full-precision files can follow different load and compute paths.

03

Same filename

Names can be changed. Preserve the hash reported by Forge and the complete companion-file set.

04

100% VRAM used

Full allocation is not proof of efficiency; computation still needs headroom.

05

GPU utilization alone

It does not separate transfer, compute, decode, browser latency or the monitoring engine being shown.

06

A cold run vs a warm run

The cold side may include file reads, model movement and initialization absent from the warm side.

07 · Questions users ask

Forge performance FAQ

Short answers for the queries that usually hide a missing test variable.

Why is Stable Diffusion WebUI Forge slow?

First identify whether the delay is model load, every sampling step, VAE decode/save, or the browser interface. Then reproduce with the exact model, commit and extensions-off workload. “Forge is slow” is not one diagnosable state.

Why is the first Forge generation slower than the second?

The first request can include model loading, patching, file reads and cache setup. Measure cold load and first generation separately, then compare repeated warm runs. If only the first FLUX run stalls, keep the full console log because historical reports describe that distinct symptom.

How can I make Forge generate faster?

Begin with the correct preset and a passing base workload. Remove extensions and add-ons, keep the model and generation settings fixed, and find the slow phase. Only tune GPU Weights or swap when memory evidence points there.

What is the best performance setting for Forge?

There is no universal setting across GPUs, model families, formats, dimensions and environments. The safest baseline is the preset value with Automatic, Queue and CPU, followed by one controlled change.

Should GPU Weights equal all available VRAM?

No. Original Forge code warns that maximum weight allocation can leave no GPU memory for matrix computation and trigger fallback, OOM or a large slowdown.

Why is Forge using shared GPU memory?

The model or workload may not fit the chosen dedicated-memory allocation, or Shared was selected as the swap location. Shared memory is not extra dedicated VRAM and can make transfer or system-memory pressure part of the run.

Does low GPU utilization mean Forge is using the CPU?

Not by itself. Measure the phase, console memory lines, warm iteration rate, dedicated/shared memory and CPU activity together. A single utilization percentage cannot distinguish transfer, decode, waiting or fallback.

What do s/it and it/s mean in Forge?

s/it is seconds per diffusion iteration, so lower is faster. it/s is iterations per second, so higher is faster. Do not compare the raw numbers as if both units move in the same direction.

How many runs should I use for a Forge benchmark?

Our editorial protocol is one documented cold run, one warm-up, then at least three unchanged warm runs with every result retained. This is a repeatability rule for this guide, not an official Forge minimum.

Can I compare Forge and A1111 with the same prompt?

A prompt alone is insufficient. Lock the model hash, VAE, sampler, scheduler, steps, dimensions, batch, seed, precision, extensions and environment, then report cold and warm phases separately.

Why did Forge become slow after an update?

Record both commits and reproduce with the same model hash, settings and environment after disabling extensions. The official reporting post requests full before/after console logs; a date or “latest” label is not enough.

Should I add --cuda-malloc or --cuda-stream to fix speed?

Not as the first move. The flags exist in the inspected code, but the original announcements describe device-specific gains and failure risks. Establish a default baseline, change one flag, and keep the result only if stability and output also pass.

Does more VRAM always make Forge faster?

More headroom can reduce offload for a fixed workload, but speed still depends on GPU architecture, model format, precision, software environment and the actual task. VRAM capacity alone is not a benchmark.

How fast should Forge be with 6 GB, 8 GB, 12 GB, or 24 GB VRAM?

A VRAM tier does not determine one expected speed. Publish the GPU model, exact model files, dimensions, steps, memory settings and warm-run unit; otherwise a time from another 6 GB or 12 GB system is not a valid target.

Why does Forge take so long to load a checkpoint?

Keep model-load time separate from sampling. Record the exact checkpoint and companion files, storage format, disk location, previous loaded model, console movement lines and whether the next unchanged warm generation is normal.

Is Forge always faster than AUTOMATIC1111?

No universal claim is supported. Local research contains both faster and conflicting community results. A meaningful comparison requires the same artifacts, workload, environment and multiple warm runs.

Where can I find Forge performance statistics?

The generation interface and console expose timing and memory evidence. Preserve the full console, the generation parameters and screenshots together; a cropped iteration rate is not a complete performance record.

08 · Evidence record

What this page was allowed to conclude

Primary code and maintainer guidance establish mechanics. Community reports establish symptoms and vocabulary, not universal speed.

VERIFIEDOriginal Forge code audit
Commit
dfdcbab685e57677014f05a3309b48cc87383167
Runtime boundary
Source-audited; no local NVIDIA benchmark was run for this page.
O11 STALE SNAPSHOT

Official 2024 NeverOOM announcement; establishes the UNet/VAE capacity trade-off and risk boundary.

O18 VERIFIED

Maintainer performance-reporting note; requires comparable models, commits, screenshots, and full console logs.

O23 VERIFIED

Original Forge UI code; verifies UI presets, Low Bits, Queue/Async, CPU/Shared, and GPU Weights labels.

R01 COMMUNITY-REPORTED

Launch-era reports about low-VRAM behavior and speed; useful for intent, not a current guarantee.

R04 COMMUNITY-REPORTED

Conflicting Forge/A1111 performance reports; supports controlled testing rather than a universal winner claim.

V27 COMMUNITY-REPORTED

Dated 8 GB Forge/A1111 test; variables were reviewed, headline timing was not generalized.

V33 COMMUNITY-REPORTED

Dated low-VRAM FLUX walkthrough; fixed hardware settings were not reused as universal advice.

AuthorForge Field Guide editorial teamReviewerTechnical editorial reviewUpdated1 Sep 2026EnvironmentOriginal Forge · source audit