Benchmark the whole job, not one flattering number
A Forge benchmark is useful only when another person can replay it. Name both builds, freeze every workload variable, measure cold and warm phases separately, and publish the slow runs and failures beside the fast ones.
VERIFIED Source-audited against original Forge commit dfdcbab. No local GPU result or universal winner is claimed.
A benchmark compares two identities
The identity includes the software, artifacts, workflow and machine state—not merely the UI name.
Project requirement
The maintainer’s reporting guide asks for comparable model files, named before/after commits and complete console logs.
Our repeatability layer
One cold observation, one excluded warm-up and five retained warm runs per condition. This is the site’s protocol, not an official Forge rule.
Different job? GGUF versus NF4, base versus Hires. fix, batch 1 versus batch 4, or extensions-on versus extensions-off can all be worthwhile tests. Label the changed item as the experiment; do not present it as a pure Forge-versus-A1111 result.
Decide what the test can prove
Select the user job. The card gives the correct outcome, lock boundary and reason to stop.
- Measure
- Primary: end-to-end warm wall time. Secondary: sampling rate, peak memory, failures and output agreement.
- Lock
- Model hash, VAE, prompt, seed, sampler, scheduler, steps, dimensions, batch, precision, add-ons and extensions.
- Stop when
- Stop if either UI changes the effective workflow, cannot load the same artifacts, or produces a different kind of output.
Make both sides replayable
Check a field only when its value is written down for both conditions.
12 / 12 does not make the benchmark true. It makes the conditions inspectable. Output validation and retained raw runs still decide whether the comparison survives.
Cold once. Warm consistently. Retain everything.
The sequence prevents initialization cost and order effects from becoming invisible advantages.
- 01
Write the claim first
Name the two conditions and the primary outcome before running either. “Feels faster” is not a measurable claim.
- 02
Capture the environment
Save full commits, System info, launch command, GPU/driver, Python, Torch/CUDA, RAM and extension state. Remove credentials before sharing.
- 03
Freeze one workload
Copy the complete generation parameters and artifact hashes. A matching prompt with a different VAE, scheduler or quantization is a different job.
- 04
Choose the clock boundary
Use Forge’s Time taken for the UI job, console phase lines for diagnosis, and an external clock only when startup or load is the declared outcome.
- 05
Record cold separately
After a clean process start, record startup, model load and first generation. Never mix this observation into the warm series.
- 06
Warm, then alternate
Our house protocol uses one unchanged warm-up, then five recorded runs per condition in alternating A/B order when both can coexist fairly. Keep every result.
- 07
Validate each output
Mark errors, black images, wrong dimensions, missing add-ons and fallback warnings. A fast invalid output is a failed run, not a win.
- 08
Publish the complete receipt
Report every time, the median and range, memory evidence, failures, full settings, commits and the raw console—not only the best screenshot.
Why five and alternating order? It is a practical house standard for exposing obvious variation without turning a user test into a laboratory study. Publish the sequence so readers can see warm-up, drift and outliers for themselves.
Enter the five warm wall times
Use one unit per series. The calculator retains all values and summarizes the median and observed range in your browser only.
- Median
- Range
- Valid runs
- Median
- Range
- Valid runs
Cold and warm-up observations belong in the report, not these fields.
Use one primary outcome and keep the diagnostic signals
No single GPU percentage can represent the load, denoising, decode and save path.
| Signal | What it records | Use it for | Boundary |
|---|---|---|---|
| End-to-end wall time | Forge’s “Time taken” covers the queued generation call around processing and cleanup. | Use as the primary user-facing latency for an already loaded UI job. | Do not mix with launch-to-ready or model load. |
| Sampling rate | The progress display may use s/it or it/s. | Use to inspect the denoising loop while model and step count are identical. | Lower s/it is faster; higher it/s is faster. Never average the two units together. |
| Cold startup | External timestamp from process launch to the local URL being ready. | Use only for a declared startup test. | OS cache, antivirus, disk and dependency checks can dominate it. |
| Model-load time | Console timestamps around loading and patching the exact model package. | Use when switching or first-image delay is the claim. | A warm, already loaded model is not a comparable side. |
| VRAM evidence | Active, reserved and system peaks shown when Forge memory monitoring is enabled. | Use with dedicated/shared-memory observations to explain a boundary. | Allocated, reserved and system values are not interchangeable. |
| Reliability | Completed valid outputs divided by attempted measured runs. | Keep failures visible, especially near a memory limit. | A crash, OOM or black output is not a missing data point. |
| Output agreement | Same dimensions and intended pipeline; optionally compare metadata and images. | Confirms both conditions performed the job you claimed. | Speed alone cannot prove equivalent output behavior. |
Convert to a shared unit before summarizing. End-to-end seconds remain separate from per-iteration rate.
Eight comparisons that cannot support their headline
Keep the observation, correct the claim, then rerun only if the user job still matters.
Different model bytes
A shared filename does not prove a shared artifact. Preserve the Forge-reported hash and companion files.
One cold, one warm
File reads, model movement, conversion and cache setup are charged to only one condition.
Changed extensions
An extension can alter callbacks, memory, scripts and effective generation behavior.
Best run only
Selecting one flattering number hides variance, failures and thermal or background-process effects.
Different batch semantics
Batch size changes simultaneous work and peak memory; batch count repeats jobs. They are not substitutes.
Crossed the VRAM boundary
A small setting change can trigger offload or shared-memory traffic, creating a nonlinear result. Report it; do not generalize it.
Recording overhead on one side
Screen capture, browser tools and monitoring must be identical or absent. A reviewed video estimated its own recording overhead.
Different output contract
Missing LoRA effect, failed ControlNet, wrong precision behavior or a black image invalidates a performance comparison.
Let the reader replay or reject it
Copy the structure, fill every field, attach both full consoles, and redact credentials or private paths before posting.
System info boundary: the inspected code hides Gradio and API auth arguments, but exports local paths, configuration, environment fields and package data. Review the file yourself.
Result boundary: the calculator does not prove significance or hardware-wide performance. It only summarizes the values you entered.
CLAIM
Condition A:
Condition B:
Primary outcome:
BUILD & ENVIRONMENT
Repository: lllyasviel/stable-diffusion-webui-forge
Commit A:
Commit B:
OS / GPU / dedicated VRAM / RAM:
Driver / Python / Torch / CUDA:
Launch arguments:
Extensions:
LOCKED WORKLOAD
Checkpoint + hash / VAE / encoders:
Prompt / negative / seed:
Sampler / scheduler / steps / CFG:
Width / height / batch size / batch count:
Hires / ControlNet / LoRAs / other scripts:
Low Bits / GPU Weights / Swap Method / Swap Location:
RUNS
Cold A (startup / load / first image):
Cold B (startup / load / first image):
Warm-up excluded: yes / no
A warm times:
B warm times:
Run order:
Median + range:
Memory evidence:
Failed or invalid outputs:
ATTACHMENTS
Full console A / B:
System info A / B:
Generation metadata and output samples:Forge benchmark FAQ
Short answers for real comparison, regression and memory-test questions.
How do I benchmark Stable Diffusion WebUI Forge?
Define one claim, capture the full environment, freeze the exact model and workload, record cold behavior separately, run one warm-up, retain five unchanged measured runs per condition, validate every output, then publish all values and logs. Five runs is this site’s house protocol, not an official Forge minimum.
Is Forge faster than AUTOMATIC1111?
There is no hardware-independent answer supported by the reviewed evidence. Compare named commits on your actual workflow; community reports include faster, similar and slower outcomes.
Can I compare Forge and A1111 using the same prompt and seed?
Not yet. Also lock model and VAE hashes, sampler, scheduler, steps, CFG, dimensions, batch, precision, add-ons, extensions and memory behavior. Confirm both outputs represent the same job.
How many benchmark runs should I perform?
This guide uses one cold observation, one warm-up and five retained warm runs per condition. The count is an ASSUMPTION in our repeatability protocol, not a requirement from the Forge project.
Should I use average or median generation time?
Report every measured value. This worksheet emphasizes the median because one unusually slow run moves it less than the arithmetic mean; also publish the range so variation remains visible. This reporting choice is our protocol.
Should the first generation count in a Forge speed test?
Record it, but do not place it in the warm series. The first request can include model loading, patching, conversion, file reads and cache setup.
What is a warm-up run?
It is one unchanged generation after the cold observation that prepares the loaded state. This guide excludes it from the measured warm series and applies the same rule to every condition.
What does Time taken measure in Forge?
At the inspected commit, modules/call_queue.py measures the queued generation function from just before processing until cleanup and appends “Time taken” to the result. It is not process startup time.
Should I compare it/s or total generation time?
Use end-to-end warm wall time as the primary user-facing measure and sampling rate as a diagnostic secondary measure. A faster denoising loop can still coexist with slower loading, decode or save.
What is the difference between s/it and it/s?
s/it means seconds per iteration, where lower is faster. it/s means iterations per second, where higher is faster. Convert to one unit before summarizing a series.
How do I benchmark Forge GPU Weights?
Keep the entire workload fixed, establish a passing baseline, change only GPU Weights, reload equally if required, then repeat the cold/warm protocol. Record active, reserved, system and shared-memory evidence plus failures.
How do I test Queue versus Async swap?
Treat Queue and Async as two conditions while keeping Swap Location, GPU Weights, Low Bits, model and workload fixed. A result applies only to the recorded device and environment.
Can I compare GGUF, NF4, FP8 and full-precision checkpoints?
You may benchmark them as different artifact choices, but not as proof that one UI is faster. State that model format or precision is the independent variable and validate output behavior.
Does maximum GPU utilization prove the fastest setup?
No. Utilization does not identify load, transfer, compute or decode, and it does not show whether the workload entered shared memory or failed. Use phase timing and memory evidence together.
Why do two identical GPUs get different Forge results?
Driver, power and thermal state, CPU/RAM, storage, background processes, runtime packages, commit, extensions, model bytes and memory settings can differ. A GPU name is only one field in the receipt.
How do I benchmark a Forge update regression?
Save the current full commit and System info, reproduce with extensions disabled, then test the last known-good commit with the same environment, artifact hashes and workload. Provide full before/after console logs.
Should failed or black-image runs be removed?
No. Mark them as failed measured attempts. They are part of reliability and may prove that the faster condition did not complete the intended job.
Can I publish only the fastest run?
No. Publish all retained runs, the median, range, errors and run order. A best-only screenshot cannot show repeatability.
What information should I include in a Forge performance report?
Include repository, full commits, OS, GPU/VRAM, driver, Python/Torch/CUDA, RAM, launch flags, extensions, exact model hashes, complete generation settings, cold observation, all warm times, memory evidence, output validation and full console logs.
Where do I find Forge System info?
The inherited WebUI System info function exists in the inspected original Forge source and exports build, runtime, command-line, RAM, extension, config and package data. Review exported text for local paths or sensitive values before posting it.
Mechanics verified; hardware conclusions withheld
Official code and maintainer guidance define what Forge records. Dated community material reveals failure modes, not universal performance.
- Commit
dfdcbab685e57677014f05a3309b48cc87383167- Inspected
- UI controls ↗ · timing/memory output ↗ · System info ↗ · memory manager ↗
- Runtime
- Source and transcript audit only. No NVIDIA benchmark was run for this page.
Maintainer performance-reporting guidance asks for comparable model files, commits and full before/after console logs.
Current original Forge UI code verifies Low Bits, Queue/Async, CPU/Shared and GPU Weights as independent variables.
Current call wrapper verifies Time taken and optional active, reserved and system VRAM output.
Current System info code records commit, command line, runtime, RAM, extensions, config and packages; credentials are hidden, but review before sharing.
Historical NeverOOM announcement documents device-specific speed and stability trade-offs; no old percentage becomes a current promise.
Launch-era users reported gains, ties and regressions around VRAM limits and batch sizes.
A reported large gain shrank after A1111 was reinstalled; replies range from equal speed to slower or faster behavior.
A dated RTX 3060 Ti comparison disclosed default-oriented settings and recording overhead, but lacked commits, repetitions and raw logs.
A dated FLUX tutorial makes cold-versus-warm behavior visible, but is not a controlled cross-product benchmark.
A user first attributed a complex-workflow slowdown to Forge, then expressed uncertainty; cache and disk were proposed, not proven.
Users described similar inference, different switching behavior and extension-dependent results.
A 3090 report lacked enough reproducibility to explain a large gap from other claimed rates.
Community replies explicitly tie potential benefit to whether a workflow crosses a VRAM threshold.
A GTX 1060 discussion shows settings, xformers, batch and restart state changing the observed result.