DEV · SCHNELL · PACKAGED · SPLIT · MEASURED

Run FLUX in Forge without guessing the package

Choose the exact variant first, identify whether its VAE and text encoders are bundled or separate, then prove one base image before tuning memory or adding a LoRA.

Evidence checked24 Aug 2026Code snapshotdfdcbab · 26 Jun 2025ScopeOriginal Forge only
01 · Readiness gate

FLUX is a dependency set, not one magic filename

A successful load requires the model identity, complete layout, correct controls, accepted terms, and enough headroom for the actual workload.

  1. 01

    Original Forge identified

    Record the repository and commit. A fork tutorial can expose different loaders, controls, defaults, and supported formats.

    SCOPE
  2. 02

    Variant identified

    Record dev, schnell, or the fine-tune’s declared upstream. The variant determines guidance, step, and license decisions.

    BEHAVIOR
  3. 03

    Package identified

    Separate model identity from FP16, FP8, NF4, GGUF, packaged, and split-file labels.

    LAYOUT
  4. 04

    Terms recorded

    Save the exact model page and current license before the file loses its source context.

    PASS / FAIL
  5. 05

    Acceptance limit written

    Define acceptable load success, time, memory, stability, and output before changing controls.

    MEASURE
VRAM capacity alone does not decide whether a FLUX workflow fits.

Text encoders, VAE, precision, resolution, batch, LoRA patching, system RAM, shared memory, driver, and Forge settings all contribute to the measured path.

02 · Package router

What exactly did you download?

Select the strongest description. The output tells you the minimum load contract and the first reason to stop.

Choose a FLUX package type
LOAD CONTRACTOne checkpoint firstVariant: dev · Layout: packaged
Files
Use the exact Forge-targeted checkpoint package. Do not add separate CLIP-L, T5, or VAE unless that package page explicitly requires them.
Forge controls
Select flux, the checkpoint, Diffusion in Low Bits: Automatic, and no VAE / Text Encoder modules for the first load unless documented otherwise.
First baseline
Use CFG 1. For a distilled dev build, start with Distilled CFG 3.5 and the exact distribution’s documented steps and dimensions.
Stop signal
Stop if the download page does not say what is bundled, which Forge build it targets, or which base license applies.
03 · Two decision axes

Variant controls behavior; format controls storage

Do not rank a dev NF4 file against schnell as though only one property changed.

AXIS A · MODEL

FLUX.1-dev

Guidance-distilled 12B model. Current Forge applies Distilled CFG to recognized dev models. The BFL page currently gates files behind acceptance of its dev non-commercial license.

Forge start state
CFG 1 · Distilled CFG 3.5 · Euler · Simple
Exact steps
Follow the selected distribution, not a copied schnell recipe.
Read the publisher model card ↗
AXIS A · MODEL

FLUX.1-schnell

Separate 12B distilled variant. Current Forge ignores Distilled CFG when it recognizes schnell. The BFL card currently lists Apache-2.0 and a 1–4-step reference.

Forge behavior
Schnell predictor · Distilled CFG ignored
Exact steps
Follow the selected schnell distribution.
Read the publisher model card ↗
AXIS B · WEIGHT FORMATFP16 · FP8 · NF4 · GGUF

These labels affect storage, precision, loading, computation, and package layout. They do not replace the variant, fine-tune identity, license, or benchmark conditions. Choose and test GGUF vs NF4 → Install and troubleshoot a FLUX LoRA →

04 · File contract

Select files by role, not by extension

This is the launch-level checklist. The complete path resolver remains in the dedicated model-folder guide.

PRIMARY

Diffusion weights

models/Stable-diffusion

Packaged checkpoint, raw model, or GGUF. It appears under Checkpoint.

SPLIT ONLY

Image decoder

models/VAE

The documented FLUX VAE, commonly ae.safetensors, selected in VAE / Text Encoder.

SPLIT ONLY

Prompt encoders

models/text_encoder

The documented CLIP-L and T5-XXL files, selected together in VAE / Text Encoder.

PACKAGEDCheckpoint contains the promised componentsDo not add modules by habit
SPLITDiffusion model + VAE + CLIP-L + T5All roles must match

The inspected scanner accepts .ckpt, .pt, .bin, .safetensors, and .gguf for additional-module inventories. Discovery does not prove that the selected state dict is the expected FLUX component.

05 · First baseline

Prove the base package before optimizing it

The first run answers one question: can this exact FLUX set load, condition, sample, and decode on this machine?

  1. 01

    Start current original Forge

    Record the commit and launch arguments. Watch the console from model selection through final decode.

  2. 02

    Select flux

    The preset exposes FLUX controls and starts txt2img at 896 × 1152, CFG 1, Distilled CFG 3.5, Euler, and Simple in the inspected code. These are UI defaults—not a contract for every fine-tune.

  3. 03

    Select the exact checkpoint and modules

    Packaged: choose the checkpoint and leave modules empty unless documented. Split: select the model, VAE, CLIP-L, and T5 as one set.

  4. 04

    Keep loading neutral

    Begin with Diffusion in Low Bits: Automatic, Swap Method: Queue, and Swap Location: CPU. Preserve the preset GPU Weights value for the first observation.

  5. 05

    Apply the variant’s recipe

    Dev and schnell do not share one step/guidance contract. Use the exact distribution’s documented dimensions and steps; keep batch size 1.

  6. 06

    Remove optional load

    No LoRA, ControlNet, Hires fix, refiner, face restoration, styles, or unrelated extensions in the acceptance run.

  7. 07

    Generate twice

    Separate cold model load plus first generation from the next warm generation. A slow first run alone is not a stable performance result.

06 · Evidence receipt

Keep enough evidence to reproduce the image

A screenshot without identity, memory, and console context cannot diagnose a FLUX setup.

First FLUX baseline record
RECORDED EVIDENCE0 / 7 recorded
07 · Memory controls

Move one boundary at a time

Memory tuning starts only after the base package is known and the initial attempt is recorded.

GPU WEIGHTS

Weights allocation is not compute headroom

Increasing the value keeps more weights in dedicated VRAM but leaves less room for computation. Maximum is not the fastest setting.

Change in measured increments
LOW BITS

Automatic preserves the file’s first contract

Forcing another dtype can add conversion, double-quantization, or LoRA behavior you did not intend. Change it only as a named experiment.

Never stack guesses
SWAP

Queue / CPU is the neutral baseline

Async and Shared can help some environments and hurt others. Record system RAM and shared-memory use before switching.

Compare same seed and workload
WORKLOAD

Resolution and add-ons need headroom too

A base image loading does not prove Hires fix, LoRA, ControlNet, batch, or a larger aspect ratio fits.

Add one dependency per pass
Do not set GPU Weights to the maximum.

Both the current code and the official performance note warn that leaving almost no inference memory can force fallback, cause OOM, or make generation dramatically slower.

Inspect the official warning ↗
08 · Acceptance and failures

Route the earliest state that failed

Fix package identity before memory, and memory before optional workflow components.

The FLUX checkpoint does not appear

Confirm the active models/Stable-diffusion path, finished download, supported extension, and top refresh control. A VAE or text encoder belongs to another inventory.

Forge reports a missing CLIP or T5 state dict

Treat this as a split-set failure. Verify CLIP-L—not CLIP-G or CLIP Vision—and the documented T5 file in models/text_encoder; select both in VAE / Text Encoder.

Inspect the current user report ↗
The preview is visible, but the final image is black or grey

Final decoding implicates the VAE/package contract. Remove extra modules from a packaged checkpoint, or verify the exact FLUX ae.safetensors for a split setup. Disable extensions and retest.

Separate preview, Full decode, and save stages →
Generation takes many minutes

Record whether the time is cold load or warm sampling. Restore Automatic/Queue/CPU, move GPU Weights below maximum, remove extra modules, and compare the exact same model, seed, dimensions, and Forge commit.

Forge runs out of VRAM or system RAM

Keep batch 1, remove optional components, return to the documented native baseline, and lower GPU Weights in measured steps. If the smallest valid workload still exceeds written limits, reject that package.

Dev output looks wrong with classic CFG or a negative prompt

Return to the current flux-preset contract: CFG 1 and the dev distribution’s Distilled CFG value. Describe the wanted result directly instead of importing an SDXL negative-prompt recipe.

Schnell output is noisy or unchanged when Distilled CFG moves

Current Forge ignores Distilled CFG for recognized schnell models. Restore the schnell package’s documented low step count and guidance behavior.

The base works, but a FLUX LoRA breaks or changes nothing

Fix the seed and add one LoRA declared for that exact FLUX variant. Compare with and without it; record trigger, strength, console patches, low-bit mode, load time, VRAM, and RAM.

Troubleshoot a FLUX LoRA step by step →Read the dated official LoRA note ↗
09 · FLUX FAQ

Questions to settle before downloading another quant

The answers separate variant, format, package, memory, and license decisions.

Can Stable Diffusion WebUI Forge run FLUX?

Yes. The inspected original Forge code contains a FLUX diffusion engine, FLUX and FLUX Schnell model detection, a flux UI preset, separate-module loading, and low-bit controls. Support for one architecture does not guarantee every conversion or add-on.

What do I need to run FLUX in Forge?

You need a current original Forge installation, one exact FLUX variant, its complete packaged or split file set, accepted model terms, enough local storage and memory for a measured baseline, and the matching flux interface controls.

Should I download FLUX dev or FLUX schnell?

Choose by workflow and terms. Dev uses distilled guidance and generally a longer sampling contract; schnell is a separate distilled variant whose publisher reference uses 1–4 steps. Their current publisher licenses also differ.

What is the difference between FLUX dev and FLUX schnell in Forge?

Current Forge recognizes both through its FLUX engine, but applies Distilled CFG to dev and ignores it for schnell. Use each exact distribution’s step, guidance, file, and license contract.

Is NF4, FP8, or GGUF a different FLUX model?

Not by itself. Dev or schnell is the model variant; NF4, FP8, GGUF, and full precision describe storage or computation choices and packaging. A fine-tune can add another identity layer.

Which FLUX format is best for low VRAM in Forge?

There is no universal winner. Format, GPU generation, CUDA/PyTorch path, offload, encoders, VAE, resolution, batch, LoRA, and system RAM change the result. Shortlist a supported distribution and measure it locally.

Can FLUX run in Forge with 4 GB, 6 GB, or 8 GB VRAM?

Dated demonstrations report some low-memory configurations, but they are not a support guarantee. Test the smallest valid package at batch size 1 and record load success, dedicated VRAM, shared memory, RAM, time, and console output.

Does a packaged FLUX checkpoint need separate T5, CLIP-L, and VAE files?

Not when the exact package documents those components as bundled. Do not attach separate modules just because another raw/GGUF tutorial uses them. Follow the package page.

Does a FLUX GGUF file include T5, CLIP-L, and VAE?

Do not assume so. Official Forge’s split-layout guidance treats GGUF diffusion weights as one role and separately loads VAE, CLIP-L, and T5. Verify the exact distribution.

Where do FLUX files go in Forge?

Primary packaged checkpoints and diffusion-model/GGUF files use models/Stable-diffusion by default; separate VAE uses models/VAE; CLIP-L and T5 use models/text_encoder. Path arguments can override these roots.

Which Forge preset should I use for FLUX?

Use flux. It exposes VAE / Text Encoder, Diffusion in Low Bits, Swap Method, Swap Location, GPU Weights, and Distilled CFG controls with FLUX-oriented defaults. It does not convert or complete the selected model.

What CFG and Distilled CFG should I use for FLUX dev?

The current flux preset starts CFG at 1 and Distilled CFG at 3.5. Treat those as Forge interface defaults, then follow the exact dev distribution’s tested recipe and keep the values in saved infotext.

Why does Distilled CFG do nothing with FLUX schnell?

Current Forge’s FLUX engine explicitly ignores Distilled CFG for models recognized as schnell. Use the schnell distribution’s own guidance and step contract.

Can I use a negative prompt with FLUX in Forge?

With the common dev setup at CFG 1, do not expect the classic SD negative-prompt workflow. Describe the wanted scene directly and follow the exact variant documentation rather than copying an SDXL recipe.

Why does Forge say “You do not have CLIP state dict” with FLUX?

A split layout is missing the expected CLIP-L component or has the wrong encoder selected. CLIP-G and a CLIP vision encoder are not substitutes. Rebuild the documented set and select CLIP-L plus T5 in VAE / Text Encoder.

Why does my FLUX preview turn black at the final step?

The preview can exist before VAE decoding. Return to the exact package contract: remove unnecessary modules for a packaged checkpoint, or verify the documented ae.safetensors for a split layout. Retest without extensions or LoRAs.

Why is FLUX extremely slow in Forge?

First record the exact model file, format, Forge commit, settings, and console. Do not set GPU Weights to the maximum: weights need VRAM, but computation also needs headroom. Compare cold load separately from warm generation.

Should GPU Weights equal all of my VRAM?

No. Current Forge warns that using essentially all VRAM for weights leaves no compute headroom and can cause fallback, OOM, or severe slowdown. Start from the preset value and tune from measurements.

Can I add a FLUX LoRA before the first image?

Do not. Prove the base package first, fix the seed, then add one LoRA declared compatible with the exact variant. Low-bit loading and LoRA precision can change patch time and memory behavior.

Can I use FLUX commercially?

A family name is not legal clearance. The current BFL pages list FLUX.1-dev under the dev non-commercial license and FLUX.1-schnell under Apache-2.0; fine-tunes and conversions may add terms. Review the exact current licenses and obtain legal advice when needed.

10 · Evidence scope

Current code defines controls; your package defines the recipe

Dated official experiments and community demonstrations locate risks, not permanent hardware promises.

VERIFIED

Current original code

FLUX/Schnell recognition, encoder roles, Distilled CFG behavior, preset defaults, file scanning, and memory-control warnings at dfdcbab.

VERIFIED

Publisher model cards

Variant identity, reference usage, access conditions, displayed license, and limitations for BFL dev and schnell.

STALE SNAPSHOT

August 2024 Forge guides

Packaged NF4/FP8, split raw/GGUF, LoRA, and performance posts document then-current experiments and layouts.

COMMUNITY-REPORTED

User intent and symptoms

Five catalogued video transcripts and issue reports map setup confusion; their GPU thresholds, timing, and quality opinions are not universalized.

Author
Forge Field Guide editorial team
Technical review
Forge code + publisher cards
Inspected snapshot
dfdcbab · 26 Jun 2025
Refresh trigger
FLUX engine, preset, loader, model card, or license change