COMPONENT GUIDE · ORIGINAL FORGE

VAE and text encoders: select the files your checkpoint is actually missing

The combined VAE / Text Encoder field is not a required-files checklist. A complete checkpoint may need nothing there; a split package may need several files. Start with the package manifest, then let the first console line—not the filename suffix—identify what is absent.

Evidence boundary: current product behavior is tied to original Forge commit dfdcbab. No GPU generation was performed for this review, so machine-specific compatibility and memory outcomes remain UNKNOWN.

01PROMPTtokens
02CLIP / T5conditioning
03DIFFUSIONlatents
04VAERGB image
TEXT SIDEIMAGE SIDE
01 · Roles

Three jobs, three failure boundaries

Forge puts two component types in one control for convenience. Their work is still separate, and the stage where a run fails helps identify which one to inspect.

PROMPT

Text encoder

Turns tokens into conditioning the diffusion model can use.

Wrong family or missing weights can stop loading or make prompt behavior invalid.
LATENT

Diffusion model

Builds the latent representation during sampling.

This is the Checkpoint selection—not a VAE or text encoder.
PIXELS

VAE

Encodes images to latent space and decodes latents back to RGB.

Wrong or unstable decoding can surface at the final image step.
02 · Package router

Is the checkpoint integrated, split or still unknown?

Choose the statement you can prove now. Each route ends with one acceptance condition before you add another file.

Checkpoint evidence
NEXT DECISIONStart with the field empty

A checkpoint can already contain the diffusion model, VAE and every required text encoder. Adding separate files without a publisher instruction replaces components and creates a new, unproven package.

Do next
Select the checkpoint, leave VAE / Text Encoder empty, run one fixed-seed baseline and save its infotext plus console. Add nothing unless the artifact manifest or first error proves a component is missing.
Accept when
The model loads, the console reports component state-dict counts, and the image completes without a missing VAE, CLIP or T5 assertion.
Open the relevant section
03 · Package contract

A suffix tells you how a file is stored—not which role it fulfills

Write the component manifest before moving files. The checkpoint page should name the model family, packaging and exact companion artifacts.

MINIMUM MANIFEST
  1. Model familySD 1.5, SDXL, FLUX or another explicit architecture
  2. Diffusion fileexact checkpoint URL, filename, bytes and hash
  3. Image codecbuilt-in VAE or exact external artifact
  4. Prompt encodersbuilt-in or every required CLIP/T5 architecture
CONTAINER ≠ COMPONENT
.safetensors.gguf.ckpt.bin

Current Forge discovers all of these in its additional-module roots, then inspects tensor keys and shapes. Copying, renaming or selecting a file cannot change its internal role.

FamilyEngine componentsPackage ruleCommit-bound evidence
SD 1.5CLIP-L + VAECommon checkpoints bundle both. A separate VAE is an explicit replacement, not a routine requirement.Current sd15 engine maps one text_encoder and one vae component.
SDXL baseCLIP-L + CLIP-G + VAECommon checkpoints bundle the set. A refiner is a different model role and must not be treated as a drop-in base package.Current sdxl engine maps text_encoder, text_encoder_2 and vae.
FLUX.1CLIP-L + T5-XXL + VAESeparated diffusion-only files need all three; some packaged checkpoints include some or all components. Verify the exact artifact.Current flux engine maps CLIP-L to pooled conditioning, T5 to cross-attention and a VAE to image encode/decode.
04 · Install and select

Move one documented package from files to a reproducible run

This procedure works for an integrated package with zero additions and a split package with several. The difference is the manifest—not a different installation ritual.

  1. 01

    Identify the package

    Save the artifact page, direct file URL and its declared component list. Do not infer “complete” from size, suffix or marketplace category.

  2. 02

    Prove the active Forge build

    Record original repository commit dfdcbab, launch arguments and the installation directory that actually owns the running process.

  3. 03

    Place files by role

    Use models/VAE for VAE files and models/text_encoder for CLIP or T5 files. The diffusion checkpoint stays in models/Stable-diffusion.

  4. 04

    Refresh the inventory

    Use the top Refresh control. At the inspected commit it refreshes checkpoints and both additional-module roots; a full browser restart is not the first diagnostic step.

  5. 05

    Select the exact component set

    Leave the multiselect empty for a proven integrated package. For a split package, choose exactly one artifact for each required role—no speculative extras.

  6. 06

    Read the first console failure

    Preserve the first “You do not have … state dict” assertion or model-recognition error before secondary memory and connection messages obscure it.

  7. 07

    Save a controlled receipt

    Generate one fixed-seed job and retain the checkpoint, Module N fields, prompt, dimensions, sampler, steps, commit and console state-dict counts.

05 · Discovery

Make the file visible before asking whether it is compatible

The inspected inventory is filename-based and recursive. This makes subfolders convenient, but duplicate basenames can hide which physical file the UI resolves.

Default image-codec rootmodels/VAE
Default prompt-encoder rootmodels/text_encoder
Accepted discovery suffixes.ckpt · .pt · .bin · .safetensors · .gguf
Optional directory overrides--vae-dir · --text-encoder-dir
Search depthRecursive below every active root
UI identityFilename only; duplicate basenames are ambiguous
VISIBLEFolder + suffix + Refresh

The file appears in the combined field.

VALIDFamily + keys + complete load

The required role is recognized and the controlled run passes.

06 · Failure decoder

Let the first failed stage choose the repair

Open the matching symptom. Preserve the first assertion and full traceback; later browser disconnects often describe the consequence, not the cause.

01The selector is empty after I copied the file

Confirm you are editing the installation that owns the running process. Use models/VAE or models/text_encoder, verify the download finished with an accepted suffix, then press Refresh. A listed file proves discovery only.

02“You do not have VAE state dict!”

The selected checkpoint plus additional modules did not yield a usable VAE state dict. Verify the exact VAE required by that package, its direct source and selection. Renaming a random model to ae.safetensors cannot satisfy the keys.

03“You do not have CLIP state dict!”

A required CLIP text encoder is absent or the selected file has the wrong CLIP architecture. For FLUX, current code expects CLIP-L—not CLIP-G and not an IP-Adapter CLIP Vision file.

04“You do not have T5 state dict!”

The package expects T5 weights and none were recognized. Select the documented T5-XXL artifact; do not substitute a CLIP file just because both appear in one UI field.

05“Failed to recognize model type!”

The checkpoint could not be identified before component construction. Recheck that Checkpoint contains the intended diffusion model and is not a VAE, encoder, partial download or unsupported export.

06The image turns grey or black at the final step

That timing points toward decode, but it is not proof by itself. Test the publisher VAE with the same seed and settings, preserve console NaN/OOM messages, and use the dedicated black-image diagnosis if the failure remains.

07The prompt is ignored or composition changes unexpectedly

Confirm the required encoder family and remove speculative alternatives. A loaded-but-wrong encoder can change conditioning; compare the documented package against one fixed prompt and seed.

08Forge runs out of memory while loading T5 or VAE

Treat this as a resource failure, not evidence that the component is incompatible. Record whether it fails during model load, text encoding, VAE encode or VAE decode, then use the memory-controls route.

09Changing checkpoints leaves the old modules selected

Selections are persisted in Forge settings. Re-read the new checkpoint manifest and clear or replace the multiselect before the first generation; do not assume the old package matches the new family.

10Hires fix uses a different VAE or encoder set

The current UI has a separate Hires VAE / Text Encoder control. “Use same choices” inherits the first pass; an empty Hires selection is recorded as Built-in. Audit both passes in the saved infotext.

11Two different files have the same basename

The inspected inventory uses the basename as its key. Rename neither artifact blindly: preserve source identity and instead keep only the intended duplicate in active roots or use unique publisher filenames.

07 · Controlled proof

Build a component receipt another user can replay

A successful image without component identity is not a reusable answer. Record the artifact chain, UI selection, first console evidence and one unchanged generation job.

A · BASELINE

Use the published package

Integrated field empty or exact split set selected.

B · ONE CHANGE

Test one component only

Keep seed, prompt and every generation setting fixed.

C · ATTRIBUTE

Match symptom to stage

Load, conditioning, encode or final decode.

REPRODUCTION RECEIPT0 / 9

Complete every field before calling the package compatible or filing a component bug.

08 · Evidence

Product behavior stays attached to original Forge

Official source and commit-bound code determine what this page calls Forge behavior. Community failures identify questions; dated tutorials supply vocabulary, not universal fixes.

VERIFIED

Code and primary records

Roles, roots, discovery, selection, reload behavior, errors and metadata.

COMMUNITY-REPORTED

Issues and discussions

Real symptoms whose individual fixes may not generalize.

STALE SNAPSHOT

Articles and video

Dated workflows reviewed for user language; hardware advice excluded.

UNKNOWN

Your exact package

Compatibility, memory and output until its receipt passes locally.

MEMORY FAILURE?

Locate the pressure stage

Record load, text encoding, VAE encode or VAE decode before tuning memory.

Open memory controls →
VERIFIEDO17 · Official split-model announcementCombined multiselect, official FLUX component roles and folders.VERIFIEDCurrent module inventoryRoots, suffixes, recursive discovery, refresh and persisted selection.VERIFIEDCurrent component loaderState-dict replacement, recognition and exact missing-component assertions.VERIFIEDCurrent FLUX engineCLIP-L, T5-XXL and VAE component mapping.VERIFIEDCurrent SD 1.5 engineSingle CLIP-L and VAE component map.VERIFIEDCurrent SDXL engineCLIP-L, CLIP-G and VAE component map.VERIFIEDCurrent VAE optionsEncode/decode role, precision safeguards and managed settings.VERIFIEDCurrent infotext parserModule N preservation and Built-in/Use same choices semantics.VERIFIEDCurrent Hires UISeparate Hires VAE / Text Encoder multiselect.VERIFIEDCurrent directory arguments--vae-dir and --text-encoder-dir.COMMUNITY-REPORTEDMissing CLIP discussion #1923Wrong-role and unselected-file diagnostic language.COMMUNITY-REPORTEDWrong CLIP issue #3064Real CLIP-G/CLIP Vision confusion; issue remains unverified as a universal fix.COMMUNITY-REPORTEDFinal-step image discussion #2029VAE replacement as one reported black/grey-image diagnostic.VERIFIEDHugging Face FLUX pipeline docsUpstream component roles: VAE, CLIP-L and T5.VERIFIEDBlack Forest Labs FLUX.1-dev filesPrimary artifact inventory for official FLUX package.STALE SNAPSHOTA10 · Forge FLUX setupFolder, selection and missing-state-dict user journey.STALE SNAPSHOTA11 · GGUF walkthroughSeparated component workflow and user vocabulary.STALE SNAPSHOTA24 · Regional FLUX guideIntegrated versus separated package distinction.STALE SNAPSHOTV03 · Full transcript reviewedRefresh and separate-file workflow; hardware claims excluded.STALE SNAPSHOTV04 · Full transcript reviewedThree-component FLUX selection and encoder confusion; settings treated as dated.
REVIEW RECORDSource/code inspection only
Product
lllyasviel/stable-diffusion-webui-forge
Commit
dfdcbab685e57677014f05a3309b48cc87383167
Reviewed
24 Aug 2026
Author
Forge Field Guide editorial team
Reviewer
Forge Field Guide technical review
Runtime
GPU generation not performed
09 · Direct answers

VAE and text encoder questions, answered by component boundary

Each answer distinguishes what the inspected code proves from what your exact checkpoint package still needs to establish.

What does the VAE do in Stable Diffusion WebUI Forge?

It converts between RGB images and the latent representation used during sampling. Txt2img needs it to decode the final latent; img2img and inpainting also use it to encode an input image.

What does a text encoder do in Forge?

It turns prompt tokens into conditioning for the diffusion model. The required encoder architecture belongs to the model family and package; it is not interchangeable with every file called CLIP.

What is the difference between a VAE and a text encoder?

A VAE operates between pixels and latents. A text encoder operates between prompt text and conditioning. They solve different jobs even though Forge selects both from one combined field.

Why are VAE and Text Encoder in the same Forge dropdown?

Modern split model packages can require several additional state dictionaries. Forge’s combined multiselect keeps the A1111-shaped top bar while allowing zero, one or several VAE/encoder files.

Can I leave VAE / Text Encoder empty?

Yes, when the selected checkpoint is a verified integrated package containing every required component. Empty is not automatically an error; prove the package with its manifest and a clean baseline.

When do I need a separate VAE in Forge?

Use one when the exact model publisher requires it, when you deliberately test an explicit replacement, or when a missing-VAE assertion proves the package lacks it. Do not add one by habit.

Where do VAE files go in Forge?

The default current root is models/VAE. Forge also scans a directory supplied with --vae-dir. Use the active installation and one evidence-backed location rather than duplicating files.

Where do text encoder files go in Forge?

The default current root is models/text_encoder. A --text-encoder-dir launch argument can add another root. CLIP Vision files for IP-Adapter are a different component category.

Why is my VAE not showing in Forge?

Check the active installation, models/VAE or --vae-dir, completed filename, accepted suffix and basename collision, then use Refresh. Discovery does not prove the file matches your checkpoint.

Why is my text encoder not showing in Forge?

Check models/text_encoder or --text-encoder-dir in the running installation, then Refresh. The inspected scanner accepts .ckpt, .pt, .bin, .safetensors and .gguf recursively.

Do I need to restart Forge after adding a VAE or text encoder?

Usually start with the top Refresh control, which rebuilds the checkpoint and additional-module choices at the inspected commit. Restart only after active-path and refresh checks fail.

Which files does a separated FLUX model need?

The official separated route uses the diffusion checkpoint plus ae/VAE, CLIP-L and T5-XXL. Exact filenames and precision must come from the artifact publisher, not from a generic list.

Does every FLUX checkpoint need separate CLIP, T5 and VAE files?

No. Packaging differs. Some checkpoints include components; diffusion-only packages do not. Determine the manifest for the exact file and test the integrated route empty before adding replacements.

Does FLUX use CLIP-L or CLIP-G in Forge?

The inspected original Forge FLUX engine expects CLIP-L together with T5-XXL. Selecting CLIP-G or a CLIP Vision encoder does not satisfy that role.

What does “You do not have CLIP state dict” mean?

The loader reached a CLIP component but did not receive a recognized state dictionary with enough keys. The file may be missing, unselected, incomplete or the wrong CLIP architecture.

What does “You do not have T5 state dict” mean?

The selected package requires T5 and Forge did not recognize usable T5 weights. Verify the documented T5 artifact, folder and UI selection before changing unrelated generation settings.

What does “You do not have VAE state dict” mean?

The package did not provide a recognized VAE state dictionary. Verify the publisher’s VAE artifact and selection; a .safetensors suffix alone does not establish the role.

Can I use any .safetensors file as a VAE or encoder?

No. Safetensors is a container format. Forge identifies component roles from tensor keys and shapes, so the wrong content remains wrong after copying or renaming.

Can Forge load a GGUF text encoder?

The inspected discovery and T5 loader include GGUF handling. That is code capability, not a guarantee for every GGUF export; use a publisher-documented encoder and validate its state-dict load.

Why does the image look black, grey or washed out?

A VAE mismatch, precision/NaN problem or decode memory failure can cause final-stage symptoms, but the image alone cannot distinguish them. Preserve the console and run one documented-VAE A/B test.

Can the wrong text encoder make prompts behave badly?

Yes, conditioning comes from the encoder set. First prove that the exact required architectures loaded; then compare a fixed prompt and seed without changing sampler, guidance or model.

Does Clip skip choose a different text encoder?

No. Clip skip changes which layer output is used by a compatible CLIP processing path. It does not supply missing weights or convert CLIP-G, CLIP-L, T5 and CLIP Vision into one another.

Why did Forge keep my VAE after switching checkpoints?

The additional-module selection is saved in Forge settings. Clear or replace it whenever the next checkpoint has a different package contract, then confirm Module N entries in infotext.

How can I prove which VAE and encoders generated an image?

Save the PNG infotext and console. Current Forge records selected additional files as Module 1, Module 2 and so on; add hashes and source URLs to make the receipt unambiguous.

What should I include in a VAE or text encoder bug report?

Include original Forge commit, launch arguments, checkpoint and component URLs/hashes, ordered UI selections, StateDict Keys output, the first assertion, complete traceback and one fixed-seed reproduction.

Should Hires fix use the same VAE and text encoders?

Use same choices is the controlled default when both passes belong to the same package. If you deliberately override Hires VAE / Text Encoder, record that second component set separately.