LOCK THE JOB · SEPARATE COLD · KEEP EVERY RUN

Benchmark the whole job, not one flattering number

A Forge benchmark is useful only when another person can replay it. Name both builds, freeze every workload variable, measure cold and warm phases separately, and publish the slow runs and failures beside the fast ones.

VERIFIED Source-audited against original Forge commit dfdcbab. No local GPU result or universal winner is claimed.

01 · Valid result

A benchmark compares two identities

The identity includes the software, artifacts, workflow and machine state—not merely the UI name.

THE RULEIf more than one intended variable changed, the result cannot tell you which change caused the difference.
VERIFIED

Project requirement

The maintainer’s reporting guide asks for comparable model files, named before/after commits and complete console logs.

ASSUMPTION

Our repeatability layer

One cold observation, one excluded warm-up and five retained warm runs per condition. This is the site’s protocol, not an official Forge rule.

Different job? GGUF versus NF4, base versus Hires. fix, batch 1 versus batch 4, or extensions-on versus extensions-off can all be worthwhile tests. Label the changed item as the experiment; do not present it as a pure Forge-versus-A1111 result.

02 · Claim builder

Decide what the test can prove

Select the user job. The card gives the correct outcome, lock boundary and reason to stop.

Comparison type
VALID CLAIMCompare the same completed job in two named builds
Measure
Primary: end-to-end warm wall time. Secondary: sampling rate, peak memory, failures and output agreement.
Lock
Model hash, VAE, prompt, seed, sampler, scheduler, steps, dimensions, batch, precision, add-ons and extensions.
Stop when
Stop if either UI changes the effective workflow, cannot load the same artifacts, or produces a different kind of output.
03 · Identity lock

Make both sides replayable

Check a field only when its value is written down for both conditions.

0 / 12 lockedComparison is not replayable

12 / 12 does not make the benchmark true. It makes the conditions inspectable. Output validation and retained raw runs still decide whether the comparison survives.

04 · Run protocol

Cold once. Warm consistently. Retain everything.

The sequence prevents initialization cost and order effects from becoming invisible advantages.

  1. 01

    Write the claim first

    Name the two conditions and the primary outcome before running either. “Feels faster” is not a measurable claim.

  2. 02

    Capture the environment

    Save full commits, System info, launch command, GPU/driver, Python, Torch/CUDA, RAM and extension state. Remove credentials before sharing.

  3. 03

    Freeze one workload

    Copy the complete generation parameters and artifact hashes. A matching prompt with a different VAE, scheduler or quantization is a different job.

  4. 04

    Choose the clock boundary

    Use Forge’s Time taken for the UI job, console phase lines for diagnosis, and an external clock only when startup or load is the declared outcome.

  5. 05

    Record cold separately

    After a clean process start, record startup, model load and first generation. Never mix this observation into the warm series.

  6. 06

    Warm, then alternate

    Our house protocol uses one unchanged warm-up, then five recorded runs per condition in alternating A/B order when both can coexist fairly. Keep every result.

  7. 07

    Validate each output

    Mark errors, black images, wrong dimensions, missing add-ons and fallback warnings. A fast invalid output is a failed run, not a win.

  8. 08

    Publish the complete receipt

    Report every time, the median and range, memory evidence, failures, full settings, commits and the raw console—not only the best screenshot.

ASSUMPTION

Why five and alternating order? It is a practical house standard for exposing obvious variation without turning a user test into a laboratory study. Publish the sequence so readers can see warm-up, drift and outliers for themselves.

05 · Run sheet

Enter the five warm wall times

Use one unit per series. The calculator retains all values and summarizes the median and observed range in your browser only.

CONDITION AWarm Time taken
Median
—
Range
—
Valid runs
0 / 5
CONDITION BWarm Time taken
Median
—
Range
—
Valid runs
0 / 5
SERIES STATUSEnter all ten valid warm times

Cold and warm-up observations belong in the report, not these fields.

06 · Measurement map

Use one primary outcome and keep the diagnostic signals

No single GPU percentage can represent the load, denoising, decode and save path.

SignalWhat it recordsUse it forBoundary
End-to-end wall timeForge’s “Time taken” covers the queued generation call around processing and cleanup.Use as the primary user-facing latency for an already loaded UI job.Do not mix with launch-to-ready or model load.
Sampling rateThe progress display may use s/it or it/s.Use to inspect the denoising loop while model and step count are identical.Lower s/it is faster; higher it/s is faster. Never average the two units together.
Cold startupExternal timestamp from process launch to the local URL being ready.Use only for a declared startup test.OS cache, antivirus, disk and dependency checks can dominate it.
Model-load timeConsole timestamps around loading and patching the exact model package.Use when switching or first-image delay is the claim.A warm, already loaded model is not a comparable side.
VRAM evidenceActive, reserved and system peaks shown when Forge memory monitoring is enabled.Use with dedicated/shared-memory observations to explain a boundary.Allocated, reserved and system values are not interchangeable.
ReliabilityCompleted valid outputs divided by attempted measured runs.Keep failures visible, especially near a memory limit.A crash, OOM or black output is not a missing data point.
Output agreementSame dimensions and intended pipeline; optionally compare metadata and images.Confirms both conditions performed the job you claimed.Speed alone cannot prove equivalent output behavior.
s/it ↓lower is faster
it/s ↑higher is faster

Convert to a shared unit before summarizing. End-to-end seconds remain separate from per-iteration rate.

07 · Invalid results

Eight comparisons that cannot support their headline

Keep the observation, correct the claim, then rerun only if the user job still matters.

01

Different model bytes

A shared filename does not prove a shared artifact. Preserve the Forge-reported hash and companion files.

02

One cold, one warm

File reads, model movement, conversion and cache setup are charged to only one condition.

03

Changed extensions

An extension can alter callbacks, memory, scripts and effective generation behavior.

04

Best run only

Selecting one flattering number hides variance, failures and thermal or background-process effects.

05

Different batch semantics

Batch size changes simultaneous work and peak memory; batch count repeats jobs. They are not substitutes.

06

Crossed the VRAM boundary

A small setting change can trigger offload or shared-memory traffic, creating a nonlinear result. Report it; do not generalize it.

07

Recording overhead on one side

Screen capture, browser tools and monitoring must be identical or absent. A reviewed video estimated its own recording overhead.

08

Different output contract

Missing LoRA effect, failed ControlNet, wrong precision behavior or a black image invalidates a performance comparison.

08 · Publish the receipt

Let the reader replay or reject it

Copy the structure, fill every field, attach both full consoles, and redact credentials or private paths before posting.

System info boundary: the inspected code hides Gradio and API auth arguments, but exports local paths, configuration, environment fields and package data. Review the file yourself.

Result boundary: the calculator does not prove significance or hardware-wide performance. It only summarizes the values you entered.

Nothing copied yet
CLAIM
Condition A:
Condition B:
Primary outcome:

BUILD & ENVIRONMENT
Repository: lllyasviel/stable-diffusion-webui-forge
Commit A:
Commit B:
OS / GPU / dedicated VRAM / RAM:
Driver / Python / Torch / CUDA:
Launch arguments:
Extensions:

LOCKED WORKLOAD
Checkpoint + hash / VAE / encoders:
Prompt / negative / seed:
Sampler / scheduler / steps / CFG:
Width / height / batch size / batch count:
Hires / ControlNet / LoRAs / other scripts:
Low Bits / GPU Weights / Swap Method / Swap Location:

RUNS
Cold A (startup / load / first image):
Cold B (startup / load / first image):
Warm-up excluded: yes / no
A warm times:
B warm times:
Run order:
Median + range:
Memory evidence:
Failed or invalid outputs:

ATTACHMENTS
Full console A / B:
System info A / B:
Generation metadata and output samples:
09 · Questions users ask

Forge benchmark FAQ

Short answers for real comparison, regression and memory-test questions.

How do I benchmark Stable Diffusion WebUI Forge?

Define one claim, capture the full environment, freeze the exact model and workload, record cold behavior separately, run one warm-up, retain five unchanged measured runs per condition, validate every output, then publish all values and logs. Five runs is this site’s house protocol, not an official Forge minimum.

Is Forge faster than AUTOMATIC1111?

There is no hardware-independent answer supported by the reviewed evidence. Compare named commits on your actual workflow; community reports include faster, similar and slower outcomes.

Can I compare Forge and A1111 using the same prompt and seed?

Not yet. Also lock model and VAE hashes, sampler, scheduler, steps, CFG, dimensions, batch, precision, add-ons, extensions and memory behavior. Confirm both outputs represent the same job.

How many benchmark runs should I perform?

This guide uses one cold observation, one warm-up and five retained warm runs per condition. The count is an ASSUMPTION in our repeatability protocol, not a requirement from the Forge project.

Should I use average or median generation time?

Report every measured value. This worksheet emphasizes the median because one unusually slow run moves it less than the arithmetic mean; also publish the range so variation remains visible. This reporting choice is our protocol.

Should the first generation count in a Forge speed test?

Record it, but do not place it in the warm series. The first request can include model loading, patching, conversion, file reads and cache setup.

What is a warm-up run?

It is one unchanged generation after the cold observation that prepares the loaded state. This guide excludes it from the measured warm series and applies the same rule to every condition.

What does Time taken measure in Forge?

At the inspected commit, modules/call_queue.py measures the queued generation function from just before processing until cleanup and appends “Time taken” to the result. It is not process startup time.

Should I compare it/s or total generation time?

Use end-to-end warm wall time as the primary user-facing measure and sampling rate as a diagnostic secondary measure. A faster denoising loop can still coexist with slower loading, decode or save.

What is the difference between s/it and it/s?

s/it means seconds per iteration, where lower is faster. it/s means iterations per second, where higher is faster. Convert to one unit before summarizing a series.

How do I benchmark Forge GPU Weights?

Keep the entire workload fixed, establish a passing baseline, change only GPU Weights, reload equally if required, then repeat the cold/warm protocol. Record active, reserved, system and shared-memory evidence plus failures.

How do I test Queue versus Async swap?

Treat Queue and Async as two conditions while keeping Swap Location, GPU Weights, Low Bits, model and workload fixed. A result applies only to the recorded device and environment.

Can I compare GGUF, NF4, FP8 and full-precision checkpoints?

You may benchmark them as different artifact choices, but not as proof that one UI is faster. State that model format or precision is the independent variable and validate output behavior.

Does maximum GPU utilization prove the fastest setup?

No. Utilization does not identify load, transfer, compute or decode, and it does not show whether the workload entered shared memory or failed. Use phase timing and memory evidence together.

Why do two identical GPUs get different Forge results?

Driver, power and thermal state, CPU/RAM, storage, background processes, runtime packages, commit, extensions, model bytes and memory settings can differ. A GPU name is only one field in the receipt.

How do I benchmark a Forge update regression?

Save the current full commit and System info, reproduce with extensions disabled, then test the last known-good commit with the same environment, artifact hashes and workload. Provide full before/after console logs.

Should failed or black-image runs be removed?

No. Mark them as failed measured attempts. They are part of reliability and may prove that the faster condition did not complete the intended job.

Can I publish only the fastest run?

No. Publish all retained runs, the median, range, errors and run order. A best-only screenshot cannot show repeatability.

What information should I include in a Forge performance report?

Include repository, full commits, OS, GPU/VRAM, driver, Python/Torch/CUDA, RAM, launch flags, extensions, exact model hashes, complete generation settings, cold observation, all warm times, memory evidence, output validation and full console logs.

Where do I find Forge System info?

The inherited WebUI System info function exists in the inspected original Forge source and exports build, runtime, command-line, RAM, extension, config and package data. Review exported text for local paths or sensitive values before posting it.

10 · Evidence record

Mechanics verified; hardware conclusions withheld

Official code and maintainer guidance define what Forge records. Dated community material reveals failure modes, not universal performance.

VERIFIEDOriginal Forge source audit
Commit
dfdcbab685e57677014f05a3309b48cc87383167
Runtime
Source and transcript audit only. No NVIDIA benchmark was run for this page.
O18 VERIFIED

Maintainer performance-reporting guidance asks for comparable model files, commits and full before/after console logs.

O23 VERIFIED

Current original Forge UI code verifies Low Bits, Queue/Async, CPU/Shared and GPU Weights as independent variables.

CODE VERIFIED

Current call wrapper verifies Time taken and optional active, reserved and system VRAM output.

CODE VERIFIED

Current System info code records commit, command line, runtime, RAM, extensions, config and packages; credentials are hidden, but review before sharing.

O11 STALE SNAPSHOT

Historical NeverOOM announcement documents device-specific speed and stability trade-offs; no old percentage becomes a current promise.

R01 COMMUNITY-REPORTED

Launch-era users reported gains, ties and regressions around VRAM limits and batch sizes.

R04 COMMUNITY-REPORTED

A reported large gain shrank after A1111 was reinstalled; replies range from equal speed to slower or faster behavior.

V27 STALE SNAPSHOT

A dated RTX 3060 Ti comparison disclosed default-oriented settings and recording overhead, but lacked commits, repetitions and raw logs.

V33 STALE SNAPSHOT

A dated FLUX tutorial makes cold-versus-warm behavior visible, but is not a controlled cross-product benchmark.

#575 COMMUNITY-REPORTED

A user first attributed a complex-workflow slowdown to Forge, then expressed uncertainty; cache and disk were proposed, not proven.

#681 COMMUNITY-REPORTED

Users described similar inference, different switching behavior and extension-dependent results.

#685 COMMUNITY-REPORTED

A 3090 report lacked enough reproducibility to explain a large gap from other claimed rates.

#479 COMMUNITY-REPORTED

Community replies explicitly tie potential benefit to whether a workflow crosses a VRAM threshold.

#373 COMMUNITY-REPORTED

A GTX 1060 discussion shows settings, xformers, batch and restart state changing the observed result.

AuthorForge Field Guide editorial teamReviewerTechnical editorial reviewUpdated1 Sep 2026EnvironmentOriginal Forge · commit dfdcbab
NEXT CONTROLLED MOVE

Run the test or return to the symptom

Use the performance hub when the slow phase is still unknown. Use memory controls only after the workload passes and memory is the intended variable.