Text encoder
Turns tokens into conditioning the diffusion model can use.
Wrong family or missing weights can stop loading or make prompt behavior invalid.The combined VAE / Text Encoder field is not a required-files checklist. A complete checkpoint may need nothing there; a split package may need several files. Start with the package manifest, then let the first console line—not the filename suffix—identify what is absent.
Evidence boundary: current product behavior is tied to original Forge commit dfdcbab. No GPU generation was performed for this review, so machine-specific compatibility and memory outcomes remain UNKNOWN.
Forge puts two component types in one control for convenience. Their work is still separate, and the stage where a run fails helps identify which one to inspect.
Turns tokens into conditioning the diffusion model can use.
Wrong family or missing weights can stop loading or make prompt behavior invalid.Builds the latent representation during sampling.
This is the Checkpoint selection—not a VAE or text encoder.Encodes images to latent space and decodes latents back to RGB.
Wrong or unstable decoding can surface at the final image step.Choose the statement you can prove now. Each route ends with one acceptance condition before you add another file.
A checkpoint can already contain the diffusion model, VAE and every required text encoder. Adding separate files without a publisher instruction replaces components and creates a new, unproven package.
A split package needs external components. For the official separated FLUX route, the documented set is a VAE, CLIP-L and T5-XXL alongside the diffusion checkpoint.
.safetensors and .gguf are containers, not a component manifest. A file can be complete, diffusion-only, a VAE, a CLIP encoder or T5 depending on its state-dict keys and publisher packaging.
At the inspected commit, Forge scans two recursive roots plus any --vae-dir and --text-encoder-dir overrides. An empty list means the active installation did not discover an accepted filename there.
Write the component manifest before moving files. The checkpoint page should name the model family, packaging and exact companion artifacts.
.safetensors.gguf.ckpt.binCurrent Forge discovers all of these in its additional-module roots, then inspects tensor keys and shapes. Copying, renaming or selecting a file cannot change its internal role.
| Family | Engine components | Package rule | Commit-bound evidence |
|---|---|---|---|
| SD 1.5 | CLIP-L + VAE | Common checkpoints bundle both. A separate VAE is an explicit replacement, not a routine requirement. | Current sd15 engine maps one text_encoder and one vae component. |
| SDXL base | CLIP-L + CLIP-G + VAE | Common checkpoints bundle the set. A refiner is a different model role and must not be treated as a drop-in base package. | Current sdxl engine maps text_encoder, text_encoder_2 and vae. |
| FLUX.1 | CLIP-L + T5-XXL + VAE | Separated diffusion-only files need all three; some packaged checkpoints include some or all components. Verify the exact artifact. | Current flux engine maps CLIP-L to pooled conditioning, T5 to cross-attention and a VAE to image encode/decode. |
This procedure works for an integrated package with zero additions and a split package with several. The difference is the manifest—not a different installation ritual.
Save the artifact page, direct file URL and its declared component list. Do not infer “complete” from size, suffix or marketplace category.
Record original repository commit dfdcbab, launch arguments and the installation directory that actually owns the running process.
Use models/VAE for VAE files and models/text_encoder for CLIP or T5 files. The diffusion checkpoint stays in models/Stable-diffusion.
Use the top Refresh control. At the inspected commit it refreshes checkpoints and both additional-module roots; a full browser restart is not the first diagnostic step.
Leave the multiselect empty for a proven integrated package. For a split package, choose exactly one artifact for each required role—no speculative extras.
Preserve the first “You do not have … state dict” assertion or model-recognition error before secondary memory and connection messages obscure it.
Generate one fixed-seed job and retain the checkpoint, Module N fields, prompt, dimensions, sampler, steps, commit and console state-dict counts.
The inspected inventory is filename-based and recursive. This makes subfolders convenient, but duplicate basenames can hide which physical file the UI resolves.
The file appears in the combined field.
The required role is recognized and the controlled run passes.
Open the matching symptom. Preserve the first assertion and full traceback; later browser disconnects often describe the consequence, not the cause.
Confirm you are editing the installation that owns the running process. Use models/VAE or models/text_encoder, verify the download finished with an accepted suffix, then press Refresh. A listed file proves discovery only.
The selected checkpoint plus additional modules did not yield a usable VAE state dict. Verify the exact VAE required by that package, its direct source and selection. Renaming a random model to ae.safetensors cannot satisfy the keys.
A required CLIP text encoder is absent or the selected file has the wrong CLIP architecture. For FLUX, current code expects CLIP-L—not CLIP-G and not an IP-Adapter CLIP Vision file.
The package expects T5 weights and none were recognized. Select the documented T5-XXL artifact; do not substitute a CLIP file just because both appear in one UI field.
The checkpoint could not be identified before component construction. Recheck that Checkpoint contains the intended diffusion model and is not a VAE, encoder, partial download or unsupported export.
That timing points toward decode, but it is not proof by itself. Test the publisher VAE with the same seed and settings, preserve console NaN/OOM messages, and use the dedicated black-image diagnosis if the failure remains.
Confirm the required encoder family and remove speculative alternatives. A loaded-but-wrong encoder can change conditioning; compare the documented package against one fixed prompt and seed.
Treat this as a resource failure, not evidence that the component is incompatible. Record whether it fails during model load, text encoding, VAE encode or VAE decode, then use the memory-controls route.
Selections are persisted in Forge settings. Re-read the new checkpoint manifest and clear or replace the multiselect before the first generation; do not assume the old package matches the new family.
The current UI has a separate Hires VAE / Text Encoder control. “Use same choices” inherits the first pass; an empty Hires selection is recorded as Built-in. Audit both passes in the saved infotext.
The inspected inventory uses the basename as its key. Rename neither artifact blindly: preserve source identity and instead keep only the intended duplicate in active roots or use unique publisher filenames.
A successful image without component identity is not a reusable answer. Record the artifact chain, UI selection, first console evidence and one unchanged generation job.
Integrated field empty or exact split set selected.
Keep seed, prompt and every generation setting fixed.
Load, conditioning, encode or final decode.
Complete every field before calling the package compatible or filing a component bug.
Official source and commit-bound code determine what this page calls Forge behavior. Community failures identify questions; dated tutorials supply vocabulary, not universal fixes.
Roles, roots, discovery, selection, reload behavior, errors and metadata.
Real symptoms whose individual fixes may not generalize.
Dated workflows reviewed for user language; hardware advice excluded.
Compatibility, memory and output until its receipt passes locally.
Each answer distinguishes what the inspected code proves from what your exact checkpoint package still needs to establish.
It converts between RGB images and the latent representation used during sampling. Txt2img needs it to decode the final latent; img2img and inpainting also use it to encode an input image.
It turns prompt tokens into conditioning for the diffusion model. The required encoder architecture belongs to the model family and package; it is not interchangeable with every file called CLIP.
A VAE operates between pixels and latents. A text encoder operates between prompt text and conditioning. They solve different jobs even though Forge selects both from one combined field.
Modern split model packages can require several additional state dictionaries. Forge’s combined multiselect keeps the A1111-shaped top bar while allowing zero, one or several VAE/encoder files.
Yes, when the selected checkpoint is a verified integrated package containing every required component. Empty is not automatically an error; prove the package with its manifest and a clean baseline.
Use one when the exact model publisher requires it, when you deliberately test an explicit replacement, or when a missing-VAE assertion proves the package lacks it. Do not add one by habit.
The default current root is models/VAE. Forge also scans a directory supplied with --vae-dir. Use the active installation and one evidence-backed location rather than duplicating files.
The default current root is models/text_encoder. A --text-encoder-dir launch argument can add another root. CLIP Vision files for IP-Adapter are a different component category.
Check the active installation, models/VAE or --vae-dir, completed filename, accepted suffix and basename collision, then use Refresh. Discovery does not prove the file matches your checkpoint.
Check models/text_encoder or --text-encoder-dir in the running installation, then Refresh. The inspected scanner accepts .ckpt, .pt, .bin, .safetensors and .gguf recursively.
Usually start with the top Refresh control, which rebuilds the checkpoint and additional-module choices at the inspected commit. Restart only after active-path and refresh checks fail.
The official separated route uses the diffusion checkpoint plus ae/VAE, CLIP-L and T5-XXL. Exact filenames and precision must come from the artifact publisher, not from a generic list.
No. Packaging differs. Some checkpoints include components; diffusion-only packages do not. Determine the manifest for the exact file and test the integrated route empty before adding replacements.
The inspected original Forge FLUX engine expects CLIP-L together with T5-XXL. Selecting CLIP-G or a CLIP Vision encoder does not satisfy that role.
The loader reached a CLIP component but did not receive a recognized state dictionary with enough keys. The file may be missing, unselected, incomplete or the wrong CLIP architecture.
The selected package requires T5 and Forge did not recognize usable T5 weights. Verify the documented T5 artifact, folder and UI selection before changing unrelated generation settings.
The package did not provide a recognized VAE state dictionary. Verify the publisher’s VAE artifact and selection; a .safetensors suffix alone does not establish the role.
No. Safetensors is a container format. Forge identifies component roles from tensor keys and shapes, so the wrong content remains wrong after copying or renaming.
The inspected discovery and T5 loader include GGUF handling. That is code capability, not a guarantee for every GGUF export; use a publisher-documented encoder and validate its state-dict load.
A VAE mismatch, precision/NaN problem or decode memory failure can cause final-stage symptoms, but the image alone cannot distinguish them. Preserve the console and run one documented-VAE A/B test.
Yes, conditioning comes from the encoder set. First prove that the exact required architectures loaded; then compare a fixed prompt and seed without changing sampler, guidance or model.
No. Clip skip changes which layer output is used by a compatible CLIP processing path. It does not supply missing weights or convert CLIP-G, CLIP-L, T5 and CLIP Vision into one another.
The additional-module selection is saved in Forge settings. Clear or replace it whenever the next checkpoint has a different package contract, then confirm Module N entries in infotext.
Save the PNG infotext and console. Current Forge records selected additional files as Module 1, Module 2 and so on; add hashes and source URLs to make the receipt unambiguous.
Include original Forge commit, launch arguments, checkpoint and component URLs/hashes, ordered UI selections, StateDict Keys output, the first assertion, complete traceback and one fixed-seed reproduction.
Use same choices is the controlled default when both passes belong to the same package. If you deliberately override Hires VAE / Text Encoder, record that second component set separately.