ONE REFERENCE · ONE JOB · ONE A/B PROOF

Use a reference image in Forge

A reference can supply source pixels, structure, visual appearance, or a person’s identity. Choose that job first; the correct Forge route follows from it.

Product boundaryOriginal Forge onlySource snapshotdfdcbab · 26 Jun 2025GPU executionNOT RUNReviewed2 Sep 2026
00 · Direct answer

Do not ask one tool to preserve everything.

Start with the smallest signal that solves the actual job.

PIXELSImg2img

Keep the source canvas and control change with denoising.

STRUCTUREControlNet

Carry pose, edges, depth, or another explicit map.

APPEARANCEReference / IP-Adapter

Guide content or style without starting from the source pixels.

IDENTITYPhotoMaker V2

Use one or more face photos in a separate SDXL Space.

VERIFIED
Page boundary

This guide chooses and proves a reference-image method. It does not promise exact identity, certify arbitrary third-party weights, or turn an old A1111 or Forge-fork tutorial into an Original Forge workflow.

01 · Method picker

What must survive the change?

Select the property you would call a failure if it disappeared.

Should the result remain recognizably close to this exact image?

Img2img

Open
Img2img → img2img
Input
One source image plus a text prompt
Main control
Denoising strength
Preserves
Broad layout, color, shapes, and source pixels to the degree allowed by denoising.
Boundary
It does not independently lock pose, style, or identity.
Use this route
02 · Signal contract

Similar inputs, different influence.

The reference image is not evidence that the method understands the property you care about.

Reference-image methods in the inspected Original Forge snapshot
MethodStarts from source pixels?Needs separate weights?Best first proof
Img2imgYesNo adapterTwo denoising values, same seed
ControlNetNo, unless combined with Img2imgUsually a map- and family-matched control modelUseful detected map + enabled/disabled A/B
Reference-onlyNoNo separate control modelreference_only at weight 1, fidelity 0.5
IP-AdapterNoAdapter weight + matching image encoderCentered square reference + same-seed A/B
PhotoMaker V2NoSeparate fixed Space pipeline downloadsClear face set + one img token

VERIFIED Current code separates the integrated ControlNet routes from the PhotoMaker V2 Space. “Reference image” is the shared input, not a shared implementation.

03 · Appearance workflows

Prove Reference-only before adding a model file.

It is the shortest current route from one image to visual guidance.

VERIFIED

Reference-only baseline

NO CONTROL MODEL REQUIRED
  1. 01

    Prove txt2img first

    Choose the intended checkpoint and preset, use batch size 1 and a fixed seed, and save a passing image before adding any reference unit.

  2. 02

    Add one reference

    Open one ControlNet Integrated unit under Txt2img, upload the image under Single Image, and leave every other unit disabled.

  3. 03

    Choose the Reference type

    Select Control Type: Reference. Start with reference_only. Current code also exposes reference_adain and reference_adain+attn.

  4. 04

    Leave Model empty

    The current Reference preprocessors set do_not_need_model. Forge hides the Model control for this route; downloading an IP-Adapter file is not part of this test.

  5. 05

    Start from neutral settings

    Use Enable, Control Weight 1, Timestep Range 0–1, and the inspected Style Fidelity default of 0.5. Keep source and output aspect ratios close for the first comparison.

  6. 06

    Run the same-seed proof

    Generate once enabled and once disabled without changing anything else. Keep the method only if the enabled result moves toward the intended look without unacceptable composition lock or artifacts.

PASS CONDITIONThe enabled image moves toward the intended visual qualities, the disabled image does not, and only the reference unit changed.
04 · IP-Adapter workflow

Match three files before touching strength.

The checkpoint family, adapter variant, and image encoder form one contract.

01Base checkpointSD 1.5 or SDXL
02IP-Adapter weightsame family, exact variant
03Image encoderViT-H, ViT-bigG, or face-aware path
  1. 01

    Name the target

    Decide whether the reference supplies broad image features or a face-oriented signal. Do not select a FaceID file merely because the filename contains “face”.

  2. 02

    Match the base family

    Use an IP-Adapter weight explicitly published for the active checkpoint family. SD 1.5 and SDXL files are not interchangeable.

  3. 03

    Match the image encoder

    Use the maintainer’s model-to-encoder matrix and the exact labels in your build: CLIP-ViT-H, CLIP-ViT-bigG, or InsightFace+CLIP-H.

  4. 04

    Place and refresh the weight

    Put the IP-Adapter model under the effective models/ControlNet path, then refresh the Model list. Encoder files download separately into models/ControlNetPreprocessor when first requested.

  5. 05

    Prepare one clean image

    For a first proof, crop close to square and center the important subject. Current code warns that a non-square image is resized and center-cropped by the CLIP processor.

  6. 06

    Build one integrated unit

    Select Control Type: IP-Adapter, choose the paired preprocessor and model, enable the unit, and keep weight 1 and range 0–1 for the initial A/B run.

  7. 07

    Compare before combining

    Generate enabled and disabled with the same seed. Only after that passes should you add pose ControlNet, img2img, Hires. fix, a LoRA, a mask, or another reference.

05 · Person identity

PhotoMaker V2 runs beside Forge, not inside the main pipeline.

The separate boundary explains why its model, prompt token, downloads, port, and controls differ.

SPACESPHOTOMAKER V2NEW LOCAL PORT
Fixed pipeline in the inspected wrapperSG161222/RealVisXL_V4.0 + PhotoMaker V2 + SDXL T2I Adapter

The checkpoint selected in the main Forge header does not control this app.

Inspect the Space lifecycle and downloads →
  1. 01

    Open the separate app

    Go to Spaces, find PhotoMaker V2, install its pinned Space snapshot, then Launch it. The resulting app opens on another local port.

  2. 02

    Let first Launch finish

    Keep the console visible while the wrapper downloads its PhotoMaker V2 adapter, SDXL T2I Adapter, RealVisXL base, and face-analysis assets. Install status alone is not proof.

  3. 03

    Upload clear identity photos

    Use one or more images where the intended face is visible. The embedded usage tip says that more photos can improve identity fidelity; the PhotoMaker paper reports the largest gain from one to two inputs.

  4. 04

    Write the required token once

    Place img after the person class in the prompt—for example, “a photo of a woman img”. The wrapper rejects a missing token and more than one token.

  5. 05

    Generate a simple portrait

    Leave the optional doodle control off for the baseline. Use one known style and seed before changing aspect ratio, style strength, steps, or guidance.

  6. 06

    Judge identity and editability separately

    Compare face traits, then check whether the prompt still controls clothing, scene, age, or style. More identity inputs can trade some text alignment for fidelity.

  7. 07

    Return to Spaces when done

    Closing the browser tab does not stop a Space. Use Terminate in Forge, and export wanted files before any Uninstall.

VALID BASELINEa photo of a woman img, soft window light, neutral backgroundOne trigger token, placed after the person class. This is a syntax example, not a guaranteed likeness recipe.
06 · Versioned matrix

Know what is verified—and what is still open.

This table belongs to commit dfdcbab; it is not a promise for forks or future builds.

Compatibility boundary checked 2 Sep 2026
RouteCurrent locationBase contractCompanion contractEvidence
Img2imgCore Img2img tabThe active checkpoint itselfNo companion adapter requiredVERIFIED
Reference-onlyControl Type: ReferenceNo separate control modelreference_only, reference_adain, reference_adain+attnVERIFIED
IP-Adapter · SD 1.5Control Type: IP-AdapterSD 1.5 IP-Adapter weightUsually ViT-H; the sd15_vit-G variant uses ViT-bigG. FaceID variants have a separate contract.VERIFIED
IP-Adapter · SDXLControl Type: IP-AdapterSDXL IP-Adapter weightEncoder depends on the exact variant; use the maintainer matrix, not a guessed name.VERIFIED
IP-Adapter FaceIDIntegrated ControlNetFaceID weight, and for some variants a matching LoRACurrent source can request InsightFace, but open reports show unresolved variant and dependency failures.COMMUNITY-REPORTED
PhotoMaker V2Separate Spaces appWrapper-fixed RealVisXL V4.0 SDXL pipelinePhotoMaker V2 adapter, T2I Adapter, face analysis, and the img tokenVERIFIED
FLUX reference workflowNo universal route certified hereExact publisher contract requiredDo not reuse SD 1.5/SDXL adapter recipesUNKNOWN
07 · Combination order

Add one influence at a time.

A strong final image is not a useful test if five controls changed together.

  1. A

    Base image

    Checkpoint + prompt + fixed seed

    Proves generation before reference control.

  2. B

    Appearance

    Reference-only or one IP-Adapter unit

    Proves what the image prompt adds.

  3. C

    Structure

    One pose, edge, or depth ControlNet unit

    Adds composition without hiding whether B works.

  4. D

    Source pixels

    Img2img and one denoising strength

    Adds the original canvas as a third influence.

  5. E

    Finish

    Hires. fix, upscale, inpaint, or LoRA

    Added only after the lower-complexity receipt passes.

Keep one receipt through A–E

Record commit, checkpoint, adapter and encoder filenames, reference crop, prompt, seed, dimensions, sampler, scheduler, steps, every unit, weight, timestep range, denoising strength, and enabled/disabled outputs.

08 · Failure router

Fix the earliest broken contract.

Prompt tuning cannot repair a missing adapter, wrong family, cropped-away subject, or failed face detector.

The result copies the composition too strongly

Use a reference with a different crop, simplify its background, or reduce one influence at a time. If you only need a palette or texture, Reference-only may be easier to diagnose than a fine-grained IP-Adapter.

The reference seems ignored

Confirm Enable, a nonzero weight, an active 0–1 range, a selected non-None IP-Adapter model where required, and an enabled-versus-disabled comparison using the same seed.

The important subject is missing or off-center

Crop a square test image with the subject centered. Current IP-Adapter code warns that non-square input is resized and center-cropped before CLIP encoding.

The IP-Adapter model does not appear

Confirm models/ControlNet, refresh the Model list, and verify that the filename passes the IP-Adapter filter. Discovery still does not prove family or encoder compatibility.

InsightFace says no face detected

Use a clear, front-biased face with enough resolution and little occlusion. If detection still fails, keep the exact image, model, preprocessor, platform, and first console traceback.

FaceID throws a DLL, ONNX, shape, or dict error

Stop tuning prompts. Open Original Forge reports show environment- and variant-specific failures. Reproduce on a clean build and treat the pairing as unresolved instead of applying a random A1111 patch.

PhotoMaker rejects the prompt

Include img exactly once and place it after the person class, such as “a man img”. The separate Space enforces this token contract.

PhotoMaker is missing or will not launch

Check the Spaces inventory and console. The current wrapper is PhotoMaker V2, downloads multiple remote dependencies, and does not run through the normal ControlNet accordion.

An old tutorial shows PhotoMaker V1 in ControlNet

That is a dated interface. The inspected current Original Forge commit ships PhotoMaker V2 as a separate Space with a fixed pipeline; do not move V1 files until your exact build’s code proves that route.

Two controls fight each other

Return to the last passing stage in the A–E order. Compare appearance alone, structure alone, and then both. Change one weight or range per run and record the result.

09 · Reference FAQ

Questions that appear after the first image.

Each answer stays within the Original Forge evidence boundary.

How do I use a reference image in Stable Diffusion WebUI Forge?

First decide what must be preserved. Use Img2img for source pixels, ControlNet for structure, Reference-only or IP-Adapter for visual guidance, and the separate PhotoMaker V2 Space for a person’s identity.

What is the difference between img2img and IP-Adapter?

Img2img starts diffusion from an encoded source image, so denoising governs how much of that canvas remains. IP-Adapter supplies image-prompt features alongside the text prompt without using the reference as the starting canvas.

What is the difference between ControlNet and IP-Adapter?

A conventional ControlNet pair carries an explicit map such as pose, edges, or depth. IP-Adapter carries image features such as content or style. They can be combined, but they solve different jobs.

What is Reference-only in Forge?

It is a built-in ControlNet preprocessor family that uses the reference during sampling and requires no separate control-model file. Current Original Forge exposes reference_only, reference_adain, and reference_adain+attn.

Does Reference-only need a model download?

No separate ControlNet model is required by the inspected Reference preprocessors. They set do_not_need_model, so the integrated UI hides its Model control for this route.

What does Style Fidelity do in Forge Reference-only?

It controls the balance used by the Reference preprocessor; the inspected default is 0.5. Treat it as a comparison variable, not a quality score, and change it only after one same-seed baseline.

Where is IP-Adapter in Forge?

Open ControlNet Integrated in Txt2img or Img2img, then choose the IP-Adapter Control Type. The current labels include CLIP-ViT-H (IPAdapter), CLIP-ViT-bigG (IPAdapter), and InsightFace+CLIP-H (IPAdapter).

Where do IP-Adapter models go in Forge?

The integrated scanner uses the effective models/ControlNet directory and any configured extra ControlNet path. Image-encoder downloads use models/ControlNetPreprocessor.

Which IP-Adapter preprocessor should I use?

Use the preprocessor required by the exact adapter variant. The maintainer matrix maps standard SD 1.5 variants mainly to ViT-H, sd15_vit-G to ViT-bigG, and SDXL variants to either encoder depending on the filename. FaceID is a separate and less reliable boundary.

Can I use an SD 1.5 IP-Adapter with SDXL?

No general cross-family support is established. Use adapter weights published for the same base family as the active checkpoint.

Can I use IP-Adapter with FLUX in Original Forge?

This audit does not establish a universal FLUX reference-image route in Original Forge. Do not reuse SD 1.5 or SDXL files; follow an exact publisher contract and label the result unknown until it passes.

Why does a non-square IP-Adapter reference work badly?

Current code logs that the CLIP image processor resizes and center-crops non-square images. A subject near an edge may be cropped away, so use a centered square crop for the diagnostic baseline.

Can IP-Adapter preserve the exact same person?

Do not treat it as an exact identity guarantee. The IP-Adapter paper says the base method produces resemblance in content and style rather than the high subject consistency of personalization methods. Face-oriented variants add different dependencies and failure modes.

Can I use IP-Adapter and ControlNet together?

Yes in principle, and the IP-Adapter paper demonstrates image prompting with structural controls. In Forge, prove the IP-Adapter unit and the structural unit separately with the same base receipt before combining them.

Can I use multiple reference images with Forge IP-Adapter?

Prove one Single Image first. The maintainer explicitly declined the old WebUI multi-input mixture and recommended separate ControlNet units for similar logic; an open report says batch input used only its first image. Test each unit alone, then combine them. Use PhotoMaker V2 when several photos of one person are the actual requirement.

How do I keep the same face in Forge?

For a dedicated identity workflow, use the separate PhotoMaker V2 Space with clear photos and the required img token. IP-Adapter FaceID exists in the code path, but open reports show unresolved variant and dependency problems, so it is not presented as the default reliable route.

Where is PhotoMaker in current Forge?

At the inspected commit it appears under Spaces as PhotoMaker V2. Install and Launch open a separate local Gradio app; it is not a normal ControlNet preprocessor.

Does PhotoMaker use my selected Forge checkpoint?

No. The inspected wrapper constructs its own SDXL pipeline with SG161222/RealVisXL_V4.0 and separate adapters, independently of the checkpoint selected in the main Forge header.

What prompt does PhotoMaker V2 require?

Use the img trigger exactly once, after the person class—for example, “a photo of a woman img”. The wrapper checks for a missing token and rejects multiple occurrences.

How many photos should I use with PhotoMaker?

Start with two clear, varied views if available. The paper found the largest identity-fidelity improvement when moving from one image to two, with diminishing gains and a possible tradeoff in text alignment as inputs increase.

Is PhotoMaker a face swap?

No. It generates a new portrait through a personalized SDXL pipeline. Identity similarity can vary, and neither the paper nor the wrapper promises pixel-exact face replacement.

Can PhotoMaker V2 use a pose or sketch?

The current Forge wrapper has an optional doodle control using an SDXL T2I Adapter, but its own interface warns that quality may decrease. Prove identity with doodle off before adding that condition.

Why does my reference workflow change style and composition together?

Image features are not perfectly separable. Use a simpler crop, compare Reference-only with IP-Adapter, and add an explicit pose or depth ControlNet only after the appearance route passes.

How do I reproduce a reference-image result?

Record Original Forge commit, checkpoint or Space pipeline, exact adapter and encoder filenames, source image crop, prompt, seed, dimensions, sampler, scheduler, steps, every unit, weight, range, and same-seed enabled/disabled outputs.

10 · Evidence ledger

Current code outranks remembered tutorials.

External sources explain the product and user intent. The workspace SEO book alone controls page methodology.

VERIFIED Original Forge code inspected
COMMUNITY-REPORTED Open Original Forge reports
  • Issue #1854 — reported FaceID preprocessor and InsightFace load failure.
  • Issue #3057 — reported IP-Adapter runtime type failure in one environment.
  • Issue #1896 — reported that IP-Adapter batch input used only its first image.
  • Discussion #2679 — recurring confusion between a model name and Forge’s encoder labels.

No single universal fix is inferred from these reports.

STALE SNAPSHOT Local research catalog
O06COMMUNITY-REPORTED

Original Forge issue #1854

Reported SDXL FaceID preprocessor and InsightFace dependency failure; retained as an unresolved report.

VERIFIED Fresh primary research
  • IP-Adapter paper — Defines image prompting, text compatibility, structural-control composition, and the base method’s identity limit.
  • IP-Adapter repository — Primary model-family and variant context; Forge-specific steps remain grounded in Original Forge.
  • PhotoMaker paper — Defines stacked identity embeddings, multi-image behavior, and the fidelity/text-control tradeoff.
  • PhotoMaker repository — Primary trigger-word contract and pipeline behavior checked against the bundled wrapper.
UNKNOWN

Universal compatibility for arbitrary FLUX adapters, every FaceID variant, every multi-reference input mode, or fork-specific files is not established by this audit.

Written by the Forge Field Guide editorial team · reviewed against Original Forge commit dfdcbab · source-audited, not GPU-run here · updated 2 Sep 2026.