FLUX · FORMAT DECISION

Choose GGUF or NF4 by the job—not the label

NF4 and GGUF take different routes through Forge. Pick the exact runnable package first, keep the model and settings fixed, then let your own warm-run evidence decide.

NF4bitsandbytes 4-bit opsusually a prequantized .safetensors package
GGUFdequantize → PyTorch ops.gguf diffusion weights + documented companions
Inspected original Forge · commit dfdcbab · code dated 26 Jun 2025Reviewed 01 Sep 2026
01 · Decision desk

Start with the constraint that actually decides your route

This chooser gives a first test candidate, not a universal winner. Every result ends in a proof you can reproduce.

What matters in this run?
FIRST TESTTest the packaged NF4 checkpoint first

It gives you the shortest documented route when that exact distribution identifies its contents and your inspected Forge environment can load bitsandbytes.

Required proof
Record the publisher page, filename, hash, bundled roles, Forge commit and a clean base generation.
Stop signal
Do not call every NF4 file complete. Package contents belong to the exact distribution, not the NF4 label.
Open the NF4 route

VERIFIED Current Forge code dispatches NF4 and GGUF to different operation classes. The recommendation above is a testing order, not a ranking claim. Inspect operations.py ↗

02 · Four axes

Separate the model, package, encoding and runtime

Most bad comparisons change several of these at once. If the NF4 file is dev and the GGUF is schnell—or their T5 files differ—you are not isolating the format.

01

Model identity

FLUX.1-dev, FLUX.1-schnell or a fine-tune

The format does not change the base variant, training or model terms.

02

Package layout

Bundled checkpoint or split component set

NF4 may arrive as a documented full checkpoint; the official GGUF guide uses separate diffusion, VAE, CLIP-L and T5 roles.

03

Weight encoding

NF4 or a specific GGUF tensor encoding

Q4, Q5, Q8 and NF4 are not interchangeable labels and do not prove identical outputs.

04

Forge compute path

bitsandbytes 4-bit operation or GGUF dequantization into PyTorch operations

This is why file size alone cannot predict speed on your environment.

03 · Runtime path

What original Forge actually does with each file

This is a commit-bound description of dfdcbab, not a promise for forks or future backends.

NF4state dictForgeOperationsBNB4bitsbnb.matmul_4bit
GGUFGGUFReaderForgeOperationsGGUFdequantize + linear
Behavior observed in the inspected original Forge source
QuestionNF4GGUF
How Forge recognizes itPrequantized state-dict keys include `bitsandbytes__nf4`.A `.gguf` file is read by GGUFReader; tensors carry GGUF quantization classes.
Current operation pathThe inspected backend calls `bnb.matmul_4bit` for 4-bit linear work.The inspected backend dequantizes supported GGUF tensors and runs PyTorch linear operations.
Manual low-bit control`bnb-nf4` and `bnb-nf4 (fp16 LoRA)` are selectable.There is no manual `gguf` item in the current dropdown; Auto detects the loaded file.
Typical official 2024 layoutThe dated NF4 guide recommended a specific packaged `.safetensors` checkpoint.The dated GGUF guide pairs the diffusion file with separate VAE, CLIP-L and T5.
What is not guaranteedEvery NF4 upload being complete, supported or faster.Every GGUF encoding being supported, faster, lower-VRAM or self-contained.
PLATFORM BOUNDARY
Upstream support and Forge integration are two different facts.

Current bitsandbytes documentation lists broader backends than the dated Forge integration; that does not prove this integration supports them. At inspected commit dfdcbab, Forge’s automatic BNB installer returns false without PyTorch CUDA or below CUDA 11.7, and its on-load BNB quantization path requires a CUDA computation device. We therefore mark other original-Forge platform combinations UNKNOWN until tested.

04 · Load routes

Load the encoding you downloaded; do not force a second one

Both routes begin with the publisher’s exact manifest and Automatic under Diffusion in Low Bits.

ROUTE A

Documented packaged NF4

  1. Confirm the model variant, publisher, license and bundled roles.
  2. Put the finished checkpoint in models/Stable-diffusion/.
  3. Select it under Checkpoint and leave Diffusion in Low Bits → Automatic.
  4. Leave VAE / Text Encoder empty only when the package says those roles are bundled.
  5. Generate once with LoRAs and extensions off; preserve the first console result.

Accept when: the exact package loads, the base image completes and the console identifies the intended storage path.

ROUTE B

Documented GGUF split set

  1. Confirm the FLUX variant and the exact GGUF encoding supplied by the publisher.
  2. Put the diffusion GGUF in models/Stable-diffusion/.
  3. Put the documented VAE in models/VAE/ and CLIP-L plus T5 in models/text_encoder/.
  4. Select all required roles and leave Diffusion in Low Bits → Automatic.
  5. Generate once with LoRAs and extensions off; preserve the first console result.

Accept when: all four roles are recognized and one clean base generation finishes.

05 · Quant labels

A filename label is a candidate—not a result

The GGUF specification can carry encoding information, but the reader and model loader still determine whether Forge can use the actual tensors.

GGUF TENSOR TYPES MAPPED IN CURRENT CODE
Q2_KQ3_KQ4_0Q4_KQ4_1Q5_0Q5_1Q5_KQ6_KQ8_0BF16
DO NOT COLLAPSE THESE

Q4_0 ≠ Q4_K ≠ NF4

They have different structures and runtime paths. The inspected mapping also does not prove that every publisher filename, mixed-tensor file or future variant will load.

VERIFIED

Forge prints GGUF tensor-type counts when it reads the state dictionary.

ASSUMPTION TO TEST

A higher-storage candidate may preserve more behavior for your prompts; do not turn that expectation into a universal ranking.

UNKNOWN UNTIL RUN

Which candidate is fastest, fits, or looks best with your complete pipeline.

06 · LoRA path

Decide how the LoRA is applied—not only how the base is stored

The official explanation is dated August 2024, while the paired UI choices still exist at the inspected commit. Treat timing and compatibility as test results.

AUTOMATIC / BNB-NF4

Precompute into base precision

The dated guide describes normal low-bit mode as patching the LoRA into the base model’s precision. Up-front patching may take time; after that, the guide says diffusion speed should not change from the patch itself.

Measure: model/LoRA patch time
AUTOMATIC (FP16 LORA)

Compute the LoRA on each step

The LoRA stays at higher precision and is evaluated during each diffusion iteration. This avoids the same prepatch route, but one LoRA can add work and several can add much more.

Measure: repeated warm generation time

One-variable LoRA check

  1. Save a base-only image and generation metadata.
  2. Add one LoRA declared for the exact FLUX base.
  3. Keep seed, prompt and LoRA weight fixed.
  4. Record load/patch time and three warm generations.
  5. Confirm a reproducible visible effect before adding another adapter.

UNKNOWN An adapter’s extension and base format do not establish that its key layout is supported.

Open the full FLUX LoRA diagnostic
07 · Controlled test

Make the comparison survive a restart

A first generation includes loading and setup costs. A single output can hide random or prompt-specific differences. Record both phases and several seeds.

A/B EVIDENCE RECEIPT

Keep these ten fields fixed or recorded

0 / 10 recorded

Record the environment before declaring a winner.

COLD

Load + first image

Restart, select the candidate, time until the first completed image, and record model/T5/VAE loading plus any one-time LoRA patch.

WARM

Repeated generation

Run several fixed seeds after loading. Report median or the full short sequence—not the fastest isolated run.

FIT

Peak resources

Record dedicated VRAM, shared GPU memory, system RAM and errors. Do not substitute download size for these measurements.

OUTPUT

Blinded grid

Review several seeds without format labels visible. Note task failures and consistent differences, not “looks better” from one image.

VERIFIED The official performance post asks for reproducible environment and settings; the GGUF announcement warns against one- or two-image comparisons. Read #1181 ↗

08 · Failure map

Use the first failed layer to choose the next action

Do not answer a loader failure by randomly changing sampler, prompt, GPU Weights and format together.

01NF4 is absent or bitsandbytes fails during startup

At the inspected commit, Forge’s installer gate requires PyTorch CUDA and CUDA 11.7+. Capture the first installation/import error and verify the actual Forge environment; current upstream bitsandbytes platform support is not proof of this older integration.

02“BNB Must Use CUDA as Computation Device”

The inspected on-load BNB quantization path asserts a CUDA compute device. Return to the exact prequantized package and environment contract instead of forcing another file through NF4.

03GGUF file appears but fails while loading

The filename extension was recognized, but its architecture, tensor encoding or download may still be unsupported. The inspected mapping is limited to the tensor types listed on this page; preserve the first console error.

04GGUF loads, then Forge asks for model, CLIP, T5 or VAE state dict

You have an incomplete or mismatched split set. Repair the named role using the publisher’s manifest and the FLUX file map.

05Loading suddenly takes much longer after choosing bnb-nf4 manually

Return Diffusion in Low Bits to Automatic. The official guide warns that forcing an already quantized file through another precision can cause dequantization and requantization.

06The smaller file is slower

That is possible. Transfer, dequantization, kernels, offload and memory pressure are different costs. Compare warm generations with the same memory controls before changing formats again.

07A LoRA takes a long time to patch

Normal low-bit mode may precompute the LoRA into the base precision. If you test an “(fp16 LoRA)” mode, record the trade: less up-front patching can mean extra work every diffusion step.

08A LoRA has no visible effect or throws a version mismatch

Prove the base model, verify the LoRA’s exact FLUX family/format and reproduce with one LoRA and a fixed seed. Format alone does not establish adapter compatibility.

09Generation is extremely slow or falls into shared memory

Keep the chosen format fixed and tune memory separately. An excessive GPU Weights value can create a severe slowdown; use the memory-controls guide and change one variable per run.

09 · Evidence register

What supports this decision guide

Product facts are tied to the original repository. Community material supplied questions and vocabulary; it did not override the code.

VERIFIED

Original Forge source

Loader, operation dispatch, UI choices and official maintainer announcements.

STALE SNAPSHOT

Dated walkthroughs

Useful for user journeys, but their hardware thresholds, exact files and speed claims are not generalized.

COMMUNITY-REPORTED

Question evidence

Confirms comparison and documentation gaps, not feature support or a winning format.

UNKNOWN

Your exact result

Fit, speed, output behavior and LoRA compatibility until the controlled receipt passes.

PACKAGE QUESTION?

Map every FLUX file

Use the separate file guide for exact VAE, CLIP-L, T5 and diffusion folders.

Open files and folders →
MEMORY QUESTION?

Tune after the format loads

Keep format fixed while changing Swap Method, Swap Location or GPU Weights.

Open memory controls →
MODEL QUESTION?

Return to the FLUX run guide

Choose dev versus schnell, package layout and a clean first-image contract.

Open the FLUX guide →
10 · Direct answers

GGUF vs NF4 questions, answered without folklore

These answers are scoped to the original lllyasviel/stable-diffusion-webui-forge repository and the inspected commit.

What is the difference between GGUF and NF4 in Forge?

In the inspected original Forge backend, NF4 uses a bitsandbytes 4-bit operation path, while GGUF tensors are read from the GGUF container, dequantized by supported tensor classes and passed to PyTorch operations. They are different encodings and runtime paths, not two names for the same switch.

Is GGUF or NF4 better for FLUX in Stable Diffusion WebUI Forge?

There is no universal winner. Start from exact model availability and package documentation, verify that the runtime works on your platform, then compare the complete candidates under the same controlled job.

Is NF4 always faster than GGUF in Forge?

No universal speed promise is supported. The 2024 maintainer posts describe reasons NF4 could be faster in tested CUDA/offload cases, but also say the measured speedups were random across a small device sample. Measure cold and warm phases locally.

Does GGUF use less VRAM than NF4?

A smaller quantized diffusion file can reduce weight storage, but peak VRAM also depends on encoding, dequantization, resolution, batch, text encoders, VAE, LoRAs and offload. File size is not a VRAM certificate.

Is a GGUF file smaller than an NF4 checkpoint?

Sometimes, depending on the GGUF encoding and what the NF4 checkpoint bundles. Compare exact byte sizes for the complete runnable sets rather than comparing extensions.

Does GGUF have better image quality than NF4?

The format names alone do not settle output behavior. Use the same base variant and companion modules, generate multiple fixed seeds and review a blinded grid; the official GGUF announcement explicitly warns against trusting one or two images.

Which FLUX GGUF should I choose: Q4, Q5, Q6 or Q8?

Choose only among files documented for your exact model and supported by your Forge build. Higher-number labels usually represent different storage trade-offs, but this page does not claim a universal quality or speed ranking. Begin with the publisher’s tested recommendation, then measure.

Is NF4 the same as Q4 GGUF?

No. Both may be described casually as 4-bit approaches, but they use different quantization structures and different Forge operation paths. Do not treat Q4_0, Q4_K and NF4 as interchangeable files.

Does a FLUX GGUF need T5, CLIP-L and VAE in Forge?

For the official Forge GGUF route documented in discussion #1050, yes: the diffusion GGUF is paired with separate VAE, CLIP-L and T5 files. Follow the exact distribution manifest if it documents a different package.

Does an NF4 checkpoint need separate T5, CLIP-L and VAE?

It depends on the exact package. The specific NF4 checkpoint in the dated official guide was presented as a full checkpoint, but the NF4 label itself does not prove that another upload bundles the same roles.

Where do I put a FLUX GGUF model in Forge?

Put a documented FLUX diffusion GGUF under `models/Stable-diffusion/`. For the official split layout, put `ae.safetensors` under `models/VAE/` and CLIP-L plus the documented T5 under `models/text_encoder/`.

Where do I put an NF4 FLUX checkpoint in Forge?

Put the documented checkpoint under `models/Stable-diffusion/`. Add separate modules only when that exact package’s manifest requires them.

What should Diffusion in Low Bits be for a GGUF model?

Use Automatic first. At the inspected commit, GGUF is detected from the loaded tensors and is not a manual dropdown choice.

What should Diffusion in Low Bits be for an NF4 checkpoint?

Use Automatic first for a checkpoint already stored as NF4. The inspected loader detects prequantized NF4 state-dict keys; forcing another storage choice can add conversion work or compound quantization loss.

Why is there no GGUF option in Diffusion in Low Bits?

Because the current UI dropdown lists Automatic, BNB NF4/FP4 and float8 choices, while the loader recognizes GGUF from the file and tensor metadata. Select the GGUF checkpoint and leave the control on Automatic.

Can I load a GGUF model with the bnb-nf4 option?

Do not use that as the normal route. It mixes a detected GGUF file with a manual BNB storage override and no source in this review validates the result as an equivalent conversion. Use Automatic for the packaged encoding.

Can I use NF4 on GTX 10-series, RTX 20-series, AMD or macOS?

Do not rely on a GPU-generation slogan. The inspected Forge integration’s automatic installer requires PyTorch CUDA and CUDA 11.7+, while its on-load BNB path requires CUDA. Test the exact original-Forge environment; current upstream bitsandbytes support does not retroactively prove this commit’s integration.

Does GGUF make FLUX work on any low-VRAM GPU?

No. GGUF support removes no requirement to fit or offload the complete workload, and it does not prove support for every GPU, operating system, encoding, resolution or companion set.

Do FLUX LoRAs work with both GGUF and NF4?

Some exact combinations can work, but the base format is not a compatibility guarantee. Verify the LoRA’s declared FLUX base, Forge support and visible effect with one adapter before stacking more.

What does “Automatic (fp16 LoRA)” mean?

The dated Forge LoRA guide says the higher-precision LoRA stays separate and is computed on each diffusion iteration instead of being precomputed into the low-bit base. This can reduce patching delay but add per-step work, especially with multiple LoRAs.

Can Forge load a GGUF T5 as well as a GGUF diffusion model?

At inspected commit dfdcbab, the current loader has a GGUF path for both diffusion and T5 state dictionaries. That code capability does not guarantee every GGUF T5 upload or model combination; follow the exact distribution.

Should I convert an NF4 checkpoint to GGUF or GGUF to NF4 in Forge?

The inspected Forge UI is a loader, not a provenance-preserving conversion workflow. Prefer an original publisher’s documented artifact. If you convert externally, treat the output as a new unverified distribution and record tool, source hash and settings.

Can I compare GGUF and NF4 with the same seed?

Yes, keep prompts, seeds and generation settings fixed, but compare several seeds and record the complete environment. A matching seed helps control the job; it does not guarantee pixel-identical paths across quantizations.

Does GGUF or NF4 change the FLUX license?

No. Encoding does not replace the model’s own terms. Verify the base variant and the exact distribution license; a renamed or converted file does not grant new usage rights.