Choose GGUF or NF4 by the job—not the label
NF4 and GGUF take different routes through Forge. Pick the exact runnable package first, keep the model and settings fixed, then let your own warm-run evidence decide.
Start with the constraint that actually decides your route
This chooser gives a first test candidate, not a universal winner. Every result ends in a proof you can reproduce.
It gives you the shortest documented route when that exact distribution identifies its contents and your inspected Forge environment can load bitsandbytes.
- Required proof
- Record the publisher page, filename, hash, bundled roles, Forge commit and a clean base generation.
- Stop signal
- Do not call every NF4 file complete. Package contents belong to the exact distribution, not the NF4 label.
Availability can decide the route. Treat the GGUF diffusion file as one component unless its publisher explicitly documents a different package.
- Required proof
- Identify the base variant, GGUF encoding, companion VAE, CLIP-L and T5, plus their exact sources.
- Stop signal
- A .gguf extension proves a container—not a complete FLUX set, license, speed or quality level.
GGUF has several encodings and an NF4 upload may bundle more components. Total the diffusion file and every required companion before deciding.
- Required proof
- Record downloaded bytes, complete package bytes and the components included in each candidate.
- Stop signal
- Smaller on disk does not prove lower peak VRAM, lower RAM use, faster loading or faster generation.
The base format is only half the decision. Forge exposes normal and “(fp16 LoRA)” low-bit modes with different patching and per-step work.
- Required proof
- Prove base-only first, then add one exact family-matched LoRA at fixed weight and compare patch time, warm generation and visible effect.
- Stop signal
- Do not infer LoRA compatibility from .safetensors, GGUF or NF4 alone.
At the inspected commit, Forge only auto-installs bitsandbytes when PyTorch reports CUDA and CUDA 11.7 or newer. Broader upstream bitsandbytes support does not prove this integration.
- Required proof
- Use a platform-specific original-Forge baseline and capture the first console result from the exact file.
- Stop signal
- GGUF parsing also does not certify that the whole FLUX workflow works on an untested platform.
Your exact GPU, offload state, companion modules and LoRA plan matter more than an uncited winner from another computer.
- Required proof
- Hold the user task and all non-format variables fixed; separate cold load from repeated warm generations and inspect several seeds.
- Stop signal
- One image or one first run is not enough to rank speed, memory or output behavior.
VERIFIED Current Forge code dispatches NF4 and GGUF to different operation classes. The recommendation above is a testing order, not a ranking claim. Inspect operations.py ↗
Separate the model, package, encoding and runtime
Most bad comparisons change several of these at once. If the NF4 file is dev and the GGUF is schnell—or their T5 files differ—you are not isolating the format.
Model identity
FLUX.1-dev, FLUX.1-schnell or a fine-tuneThe format does not change the base variant, training or model terms.
Package layout
Bundled checkpoint or split component setNF4 may arrive as a documented full checkpoint; the official GGUF guide uses separate diffusion, VAE, CLIP-L and T5 roles.
Weight encoding
NF4 or a specific GGUF tensor encodingQ4, Q5, Q8 and NF4 are not interchangeable labels and do not prove identical outputs.
Forge compute path
bitsandbytes 4-bit operation or GGUF dequantization into PyTorch operationsThis is why file size alone cannot predict speed on your environment.
What original Forge actually does with each file
This is a commit-bound description of dfdcbab, not a promise for forks or future backends.
state dict→ForgeOperationsBNB4bits→bnb.matmul_4bitGGUFReader→ForgeOperationsGGUF→dequantize + linear| Question | NF4 | GGUF |
|---|---|---|
| How Forge recognizes it | Prequantized state-dict keys include `bitsandbytes__nf4`. | A `.gguf` file is read by GGUFReader; tensors carry GGUF quantization classes. |
| Current operation path | The inspected backend calls `bnb.matmul_4bit` for 4-bit linear work. | The inspected backend dequantizes supported GGUF tensors and runs PyTorch linear operations. |
| Manual low-bit control | `bnb-nf4` and `bnb-nf4 (fp16 LoRA)` are selectable. | There is no manual `gguf` item in the current dropdown; Auto detects the loaded file. |
| Typical official 2024 layout | The dated NF4 guide recommended a specific packaged `.safetensors` checkpoint. | The dated GGUF guide pairs the diffusion file with separate VAE, CLIP-L and T5. |
| What is not guaranteed | Every NF4 upload being complete, supported or faster. | Every GGUF encoding being supported, faster, lower-VRAM or self-contained. |
Current bitsandbytes documentation lists broader backends than the dated Forge integration; that does not prove this integration supports them. At inspected commit dfdcbab, Forge’s automatic BNB installer returns false without PyTorch CUDA or below CUDA 11.7, and its on-load BNB quantization path requires a CUDA computation device. We therefore mark other original-Forge platform combinations UNKNOWN until tested.
Load the encoding you downloaded; do not force a second one
Both routes begin with the publisher’s exact manifest and Automatic under Diffusion in Low Bits.
Documented packaged NF4
- Confirm the model variant, publisher, license and bundled roles.
- Put the finished checkpoint in
models/Stable-diffusion/. - Select it under Checkpoint and leave Diffusion in Low Bits → Automatic.
- Leave VAE / Text Encoder empty only when the package says those roles are bundled.
- Generate once with LoRAs and extensions off; preserve the first console result.
Accept when: the exact package loads, the base image completes and the console identifies the intended storage path.
Documented GGUF split set
- Confirm the FLUX variant and the exact GGUF encoding supplied by the publisher.
- Put the diffusion GGUF in
models/Stable-diffusion/. - Put the documented VAE in
models/VAE/and CLIP-L plus T5 inmodels/text_encoder/. - Select all required roles and leave Diffusion in Low Bits → Automatic.
- Generate once with LoRAs and extensions off; preserve the first console result.
Accept when: all four roles are recognized and one clean base generation finishes.
A filename label is a candidate—not a result
The GGUF specification can carry encoding information, but the reader and model loader still determine whether Forge can use the actual tensors.
Q2_KQ3_KQ4_0Q4_KQ4_1Q5_0Q5_1Q5_KQ6_KQ8_0BF16Q4_0 ≠ Q4_K ≠ NF4
They have different structures and runtime paths. The inspected mapping also does not prove that every publisher filename, mixed-tensor file or future variant will load.
Forge prints GGUF tensor-type counts when it reads the state dictionary.
A higher-storage candidate may preserve more behavior for your prompts; do not turn that expectation into a universal ranking.
Which candidate is fastest, fits, or looks best with your complete pipeline.
Decide how the LoRA is applied—not only how the base is stored
The official explanation is dated August 2024, while the paired UI choices still exist at the inspected commit. Treat timing and compatibility as test results.
Precompute into base precision
The dated guide describes normal low-bit mode as patching the LoRA into the base model’s precision. Up-front patching may take time; after that, the guide says diffusion speed should not change from the patch itself.
Measure: model/LoRA patch timeCompute the LoRA on each step
The LoRA stays at higher precision and is evaluated during each diffusion iteration. This avoids the same prepatch route, but one LoRA can add work and several can add much more.
Measure: repeated warm generation timeOne-variable LoRA check
- Save a base-only image and generation metadata.
- Add one LoRA declared for the exact FLUX base.
- Keep seed, prompt and LoRA weight fixed.
- Record load/patch time and three warm generations.
- Confirm a reproducible visible effect before adding another adapter.
UNKNOWN An adapter’s extension and base format do not establish that its key layout is supported.
Open the full FLUX LoRA diagnosticMake the comparison survive a restart
A first generation includes loading and setup costs. A single output can hide random or prompt-specific differences. Record both phases and several seeds.
Keep these ten fields fixed or recorded
Record the environment before declaring a winner.
Load + first image
Restart, select the candidate, time until the first completed image, and record model/T5/VAE loading plus any one-time LoRA patch.
Repeated generation
Run several fixed seeds after loading. Report median or the full short sequence—not the fastest isolated run.
Peak resources
Record dedicated VRAM, shared GPU memory, system RAM and errors. Do not substitute download size for these measurements.
Blinded grid
Review several seeds without format labels visible. Note task failures and consistent differences, not “looks better” from one image.
VERIFIED The official performance post asks for reproducible environment and settings; the GGUF announcement warns against one- or two-image comparisons. Read #1181 ↗
Use the first failed layer to choose the next action
Do not answer a loader failure by randomly changing sampler, prompt, GPU Weights and format together.
01NF4 is absent or bitsandbytes fails during startup+
At the inspected commit, Forge’s installer gate requires PyTorch CUDA and CUDA 11.7+. Capture the first installation/import error and verify the actual Forge environment; current upstream bitsandbytes platform support is not proof of this older integration.
02“BNB Must Use CUDA as Computation Device”+
The inspected on-load BNB quantization path asserts a CUDA compute device. Return to the exact prequantized package and environment contract instead of forcing another file through NF4.
03GGUF file appears but fails while loading+
The filename extension was recognized, but its architecture, tensor encoding or download may still be unsupported. The inspected mapping is limited to the tensor types listed on this page; preserve the first console error.
04GGUF loads, then Forge asks for model, CLIP, T5 or VAE state dict+
You have an incomplete or mismatched split set. Repair the named role using the publisher’s manifest and the FLUX file map.
05Loading suddenly takes much longer after choosing bnb-nf4 manually+
Return Diffusion in Low Bits to Automatic. The official guide warns that forcing an already quantized file through another precision can cause dequantization and requantization.
06The smaller file is slower+
That is possible. Transfer, dequantization, kernels, offload and memory pressure are different costs. Compare warm generations with the same memory controls before changing formats again.
07A LoRA takes a long time to patch+
Normal low-bit mode may precompute the LoRA into the base precision. If you test an “(fp16 LoRA)” mode, record the trade: less up-front patching can mean extra work every diffusion step.
08A LoRA has no visible effect or throws a version mismatch+
Prove the base model, verify the LoRA’s exact FLUX family/format and reproduce with one LoRA and a fixed seed. Format alone does not establish adapter compatibility.
09Generation is extremely slow or falls into shared memory+
Keep the chosen format fixed and tune memory separately. An excessive GPU Weights value can create a severe slowdown; use the memory-controls guide and change one variable per run.
What supports this decision guide
Product facts are tied to the original repository. Community material supplied questions and vocabulary; it did not override the code.
Original Forge source
Loader, operation dispatch, UI choices and official maintainer announcements.
Dated walkthroughs
Useful for user journeys, but their hardware thresholds, exact files and speed claims are not generalized.
Question evidence
Confirms comparison and documentation gaps, not feature support or a winning format.
Your exact result
Fit, speed, output behavior and LoRA compatibility until the controlled receipt passes.
GGUF vs NF4 questions, answered without folklore
These answers are scoped to the original lllyasviel/stable-diffusion-webui-forge repository and the inspected commit.
What is the difference between GGUF and NF4 in Forge?+
In the inspected original Forge backend, NF4 uses a bitsandbytes 4-bit operation path, while GGUF tensors are read from the GGUF container, dequantized by supported tensor classes and passed to PyTorch operations. They are different encodings and runtime paths, not two names for the same switch.
Is GGUF or NF4 better for FLUX in Stable Diffusion WebUI Forge?+
There is no universal winner. Start from exact model availability and package documentation, verify that the runtime works on your platform, then compare the complete candidates under the same controlled job.
Is NF4 always faster than GGUF in Forge?+
No universal speed promise is supported. The 2024 maintainer posts describe reasons NF4 could be faster in tested CUDA/offload cases, but also say the measured speedups were random across a small device sample. Measure cold and warm phases locally.
Does GGUF use less VRAM than NF4?+
A smaller quantized diffusion file can reduce weight storage, but peak VRAM also depends on encoding, dequantization, resolution, batch, text encoders, VAE, LoRAs and offload. File size is not a VRAM certificate.
Is a GGUF file smaller than an NF4 checkpoint?+
Sometimes, depending on the GGUF encoding and what the NF4 checkpoint bundles. Compare exact byte sizes for the complete runnable sets rather than comparing extensions.
Does GGUF have better image quality than NF4?+
The format names alone do not settle output behavior. Use the same base variant and companion modules, generate multiple fixed seeds and review a blinded grid; the official GGUF announcement explicitly warns against trusting one or two images.
Which FLUX GGUF should I choose: Q4, Q5, Q6 or Q8?+
Choose only among files documented for your exact model and supported by your Forge build. Higher-number labels usually represent different storage trade-offs, but this page does not claim a universal quality or speed ranking. Begin with the publisher’s tested recommendation, then measure.
Is NF4 the same as Q4 GGUF?+
No. Both may be described casually as 4-bit approaches, but they use different quantization structures and different Forge operation paths. Do not treat Q4_0, Q4_K and NF4 as interchangeable files.
Does a FLUX GGUF need T5, CLIP-L and VAE in Forge?+
For the official Forge GGUF route documented in discussion #1050, yes: the diffusion GGUF is paired with separate VAE, CLIP-L and T5 files. Follow the exact distribution manifest if it documents a different package.
Does an NF4 checkpoint need separate T5, CLIP-L and VAE?+
It depends on the exact package. The specific NF4 checkpoint in the dated official guide was presented as a full checkpoint, but the NF4 label itself does not prove that another upload bundles the same roles.
Where do I put a FLUX GGUF model in Forge?+
Put a documented FLUX diffusion GGUF under `models/Stable-diffusion/`. For the official split layout, put `ae.safetensors` under `models/VAE/` and CLIP-L plus the documented T5 under `models/text_encoder/`.
Where do I put an NF4 FLUX checkpoint in Forge?+
Put the documented checkpoint under `models/Stable-diffusion/`. Add separate modules only when that exact package’s manifest requires them.
What should Diffusion in Low Bits be for a GGUF model?+
Use Automatic first. At the inspected commit, GGUF is detected from the loaded tensors and is not a manual dropdown choice.
What should Diffusion in Low Bits be for an NF4 checkpoint?+
Use Automatic first for a checkpoint already stored as NF4. The inspected loader detects prequantized NF4 state-dict keys; forcing another storage choice can add conversion work or compound quantization loss.
Why is there no GGUF option in Diffusion in Low Bits?+
Because the current UI dropdown lists Automatic, BNB NF4/FP4 and float8 choices, while the loader recognizes GGUF from the file and tensor metadata. Select the GGUF checkpoint and leave the control on Automatic.
Can I load a GGUF model with the bnb-nf4 option?+
Do not use that as the normal route. It mixes a detected GGUF file with a manual BNB storage override and no source in this review validates the result as an equivalent conversion. Use Automatic for the packaged encoding.
Can I use NF4 on GTX 10-series, RTX 20-series, AMD or macOS?+
Do not rely on a GPU-generation slogan. The inspected Forge integration’s automatic installer requires PyTorch CUDA and CUDA 11.7+, while its on-load BNB path requires CUDA. Test the exact original-Forge environment; current upstream bitsandbytes support does not retroactively prove this commit’s integration.
Does GGUF make FLUX work on any low-VRAM GPU?+
No. GGUF support removes no requirement to fit or offload the complete workload, and it does not prove support for every GPU, operating system, encoding, resolution or companion set.
Do FLUX LoRAs work with both GGUF and NF4?+
Some exact combinations can work, but the base format is not a compatibility guarantee. Verify the LoRA’s declared FLUX base, Forge support and visible effect with one adapter before stacking more.
What does “Automatic (fp16 LoRA)” mean?+
The dated Forge LoRA guide says the higher-precision LoRA stays separate and is computed on each diffusion iteration instead of being precomputed into the low-bit base. This can reduce patching delay but add per-step work, especially with multiple LoRAs.
Can Forge load a GGUF T5 as well as a GGUF diffusion model?+
At inspected commit dfdcbab, the current loader has a GGUF path for both diffusion and T5 state dictionaries. That code capability does not guarantee every GGUF T5 upload or model combination; follow the exact distribution.
Should I convert an NF4 checkpoint to GGUF or GGUF to NF4 in Forge?+
The inspected Forge UI is a loader, not a provenance-preserving conversion workflow. Prefer an original publisher’s documented artifact. If you convert externally, treat the output as a new unverified distribution and record tool, source hash and settings.
Can I compare GGUF and NF4 with the same seed?+
Yes, keep prompts, seeds and generation settings fixed, but compare several seeds and record the complete environment. A matching seed helps control the job; it does not guarantee pixel-identical paths across quantizations.
Does GGUF or NF4 change the FLUX license?+
No. Encoding does not replace the model’s own terms. Verify the base variant and the exact distribution license; a renamed or converted file does not grant new usage rights.