AMD · ROCm · DIRECTML · ORIGINAL FORGE

Choose the AMD compute path before you install Forge

Original Forge has AMD-related code, but it does not publish a current AMD support matrix or one-click package. First prove that your exact GPU works with ROCm or DirectML; then test Forge in a disposable clone.

UNKNOWN

There is no single “AMD install command.” A valid route is a chain: exact GPU → supported driver/OS → supported backend → compatible PyTorch → clean Forge generation. A failure at an earlier link cannot be repaired with a later flag.

Which environment are you actually running?
COMMUNITY-REPORTED

The most coherent AMD route, but not a maintained Forge recipe.

Use it only when AMD’s current matrix lists your exact GPU, Linux release and ROCm/PyTorch pair. Original Forge contains ROCm-aware code, but its automatic shell branch still points to ROCm 5.7 and does not represent the current AMD stack.

Audit the Linux route
01 · Version gap

Four “supported” layers can still disagree

Forge, PyTorch, AMD ROCm and DirectML have independent release cycles. Record all four rather than asking only whether the GPU is “supported.”

01Radeon GPUExact SKU + architecture
02Compute backendROCm or DirectML
03PyTorch buildBackend-specific package
04Original ForgeCommit + clean workflow
Source snapshots checked on 29 August 2026
LayerObserved version boundaryEvidenceWhat it means
Original Forge launcherPyTorch 2.3.1 · torchvision 0.18.1 · CUDA 12.1 indexVERIFIEDThis is the inspected default, not an AMD compatibility claim.
Original Forge Linux shellROCm 5.7 wheels; older Navi 1 branches use ROCm 5.2 nightliesSTALE SNAPSHOTDo not let an old auto-detection branch choose a current Radeon environment.
Current AMD Radeon stackHardware, OS, ROCm and framework versions are selected as one matrixVERIFIED UPSTREAMAMD support for a stack does not prove Forge supports that stack.
Torch-DirectMLMicrosoft documents support up to PyTorch 2.3.1; DirectML is in sustained engineeringVERIFIED UPSTREAMThe version aligns with Forge’s Torch pin, but Forge does not install or certify the backend.
The current mismatch has no universal winner.

PyTorch 2.3.1 + ROCm 6.0 is closest to Forge’s pin, but may be outside support for a current Radeon. A current AMD-validated PyTorch stack fits the hardware, but is outside Forge’s inspected pin. That is why this page starts with a matrix and an isolated lab—not a copy-paste upgrade.

02 · Evidence gate

Build a receipt before changing packages

Check each item yourself. The receipt captures what is known and what is still missing; it does not claim that this machine was tested by our team.

AMD FORGE PREFLIGHT0 / 8 checks recorded
REPRODUCTION RECEIPTFill the blanks after copying.
route: native Linux / ROCm
gpu: [exact model]
os-kernel: [version]
backend: [ROCm or DirectML version]
python-torch: [versions]
forge: dfdcbab (source audit)
extensions: disabled
checks-recorded: 0/8
result: [backend test / localhost / generation / first error]
03 · Native Linux

Prove ROCm outside Forge first

A native Linux route is valid only when AMD’s current matrix connects the GPU, distribution, kernel, ROCm release and PyTorch build.

Open AMD’s current Linux guide ↗
  1. Find the exact GPU

    Do not continue with a family label such as “RX 7000.” Record the full marketing name and architecture identifier reported by the compute runtime.

    lspci -nn | grep -Ei 'VGA|3D|Display'
  2. Match the entire AMD matrix row

    Check the GPU, supported Linux release, kernel, Radeon software/ROCm release and framework version together. A check mark for a different operating system or PyTorch version is a different environment.

  3. Install ROCm from AMD’s current procedure

    Use AMD’s commands for the selected matrix row. This page does not freeze a driver URL that will become stale, and it does not reuse the Fedora 40 video’s one-command package install on another distribution.

  4. Verify the compute device

    rocminfo must list a GPU agent. Then the selected ROCm PyTorch build must report a HIP version and return True from torch.cuda.is_available(). Under ROCm, PyTorch intentionally keeps the CUDA API name.

    python3 -c "import torch; print('torch', torch.__version__); print('hip', torch.version.hip); print('available', torch.cuda.is_available()); print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'NO GPU')"
  5. Compare that build with Forge’s pin

    If the validated PyTorch is not 2.3.1, you are outside the inspected Forge runtime. Keep it in the receipt. Do not downgrade a supported Radeon stack merely to silence Forge’s package check.

  6. Test only in a separate clone

    Clone original Forge, disable extensions, and let the first error identify whether the gap is package resolution, startup, model load or inference. A community Fedora 40 / RX 6900 XT case succeeded after manually replacing Torch, but that 2024 environment is not a current matrix.

    git clone https://github.com/lllyasviel/stable-diffusion-webui-forge.git forge-amd-lab
STALE SNAPSHOT
Why we do not recommend the automatic ROCm branch

webui.sh still selects ROCm 5.7 for a generic AMD detection and ROCm 5.2 nightlies for older Navi paths. Those lines prove historical intent, not current validation. Override experiments need their own tested version record and should never be presented as a universal install.

Inspect the exact shell lines ↗
04 · WSL 2

A supported bridge is not a supported Forge stack

AMD’s current ROCDXG route makes ROCm compute available inside WSL for selected hardware. Forge still sits above that bridge as a separate, unverified application layer.

WINDOWS HOSTCompatible Adrenalin driverAuthoritative display driver
WSL BRIDGEDXCore + ROCDXGUser-mode compute interface
LINUX USERSPACESupported Ubuntu + ROCm + PyTorchExact current AMD matrix row
APPLICATIONOriginal Forge lab cloneUnknown until reproduced
  • Use the current AMD WSL page.The modern ROCDXG flow supersedes older tightly coupled WSL recipes. Do not mix commands from ROCm 6.1 preview documentation, a native-Linux driver install and the current bridge.
  • Prove the framework before cloning.Run AMD’s PyTorch verification inside WSL. Record the device, ROCm and PyTorch versions. A passing tensor test is the handoff point to Forge—not the final result.
  • Keep the lab in the Linux filesystem.Use a clean WSL directory for the checkout and venv. Avoid adding Windows path semantics to an already experimental package stack.
  • Expect the Forge version gap.The current AMD WSL matrix validates a newer PyTorch line than Forge’s inspected 2.3.1 pin. Do not hide dependency errors or imply that upstream production support covers Forge.
05 · Windows lab

DirectML is a reversible experiment

This sequence is source-audited, not hardware-verified by our team. It modifies only a disposable clone and stops before model complexity can hide a backend failure.

COMMUNITY-REPORTED
Do not use the NVIDIA one-click archive.Its CUDA environment is not an AMD package. Start from a Git clone so every package change stays visible and removable.
  1. 01

    Create a separate lab clone

    Clone the original repository into a new forge-amd-lab folder. Do not convert a working CUDA or third-party AMD environment in place.

    git clone https://github.com/lllyasviel/stable-diffusion-webui-forge.git forge-amd-lab
  2. 02

    Record the source snapshot

    Save the origin URL and full HEAD. This guide inspected dfdcbab; a later commit can change the backend or dependency boundary.

    git -C forge-amd-lab remote get-url origin && git -C forge-amd-lab rev-parse HEAD
  3. 03

    Let Forge create only its local environment

    Run webui-user.bat once. On an AMD-only system it may stop at the CUDA availability test; that expected failure still creates the local venv needed for the next isolated check.

    webui-user.bat
  4. 04

    Install Torch-DirectML in that venv

    Use the venv interpreter, never a global pip. Microsoft’s current package limit is PyTorch 2.3.1. Save the package output so an unexpected Torch replacement is visible.

    .\venv\Scripts\python.exe -m pip install torch-directml
  5. 05

    Use the flag that exists in current source

    Set COMMANDLINE_ARGS to --directml --skip-torch-cuda-test in webui-user.bat. Current backend/args.py accepts --directml. The --use-directml alias is still proposed in an open PR and is not in the inspected commit.

    set COMMANDLINE_ARGS=--directml --skip-torch-cuda-test
  6. 06

    Prove DirectML before loading a model

    Import torch_directml with the Forge venv and print the selected device. If this fails, stop: Forge cannot repair the DirectML layer.

    .\venv\Scripts\python.exe -c "import torch_directml; d=torch_directml.device(); print(torch_directml.device_name(0)); print(d)"
  7. 07

    Start clean and read every fallback

    Launch webui-user.bat with no extensions. Reaching localhost is only server success; a small SD 1.5 generation and the console are needed to prove an inference path. Stop on unsupported-operator, mixed-device or repeated CPU-fallback errors.

What the CUDA skip flag does—and does not do.

It bypasses Forge’s NVIDIA/ROCm-style startup check because DirectML is a different backend. Use it only after the standalone DirectML device command succeeds. It does not install the package, select a GPU, implement missing operators or prevent CPU fallback.

06 · Failure map

Stop at the first failed layer

Repeated reinstalling makes the environment harder to explain. Preserve the first failure, identify its owner, and change one layer only.

GPU MATRIX

Exact GPU is absent

Stop. Do not use an environment-variable architecture override as proof of support. Choose software with a documented path for that hardware or wait for a validated matrix update.

ROCm DEVICE

rocminfo has no GPU agent

Return to AMD’s driver, OS and kernel documentation. Forge has not started yet and cannot fix this layer.

PYTORCH DEVICE

torch.cuda.is_available() is false

Confirm this is a ROCm build, inspect torch.version.hip, and match AMD’s framework matrix. Do not add Forge skip flags.

DIRECTML IMPORT

torch_directml is missing

Confirm the package was installed by the disposable venv interpreter. Do not install it globally or into a CUDA environment you need to preserve.

FORGE STARTUP

Server never prints localhost

Capture the first traceback, exact package list, commit and flags. Remove extensions and route Python/Torch errors through the environment troubleshooting guide.

INFERENCE

Localhost works; generation fails

Use a small SD 1.5 baseline. Unsupported operators, mixed-device tensors, dtype failures and VAE decode errors mean the backend/application path is not yet proven.

07 · Project boundary

A working AMD fork is still a different project

Forks and containers may solve real problems. Their patches, binaries, owners and update paths must remain correctly attributed.

VERIFIED

Original Forge

Repository owner lllyasviel. The claims and source audit on this page refer only to that codebase and commit.

Open original source ↗
COMMUNITY-REPORTED

AMD-specific forks and ZLUDA

Separate projects can add DirectML, ZLUDA or other patches. Use their own instructions and issue trackers; never copy their features into an original-Forge claim.

COMMUNITY-REPORTED

Containers and launchers

AI-Dock and other images own their base OS, ROCm packages, services and security model. They are not official Forge downloads and are not interchangeable with a native install.

08 · Questions

AMD, ROCm and DirectML answers

These answers separate current upstream support, original-Forge source facts and community experiments.

Does original Stable Diffusion WebUI Forge officially support AMD GPUs?

The inspected original README does not publish a current AMD support matrix or AMD one-click package. The source contains ROCm and DirectML-related paths, while open issues still request installation and support guidance. We therefore label AMD routes community-reported or unknown, not official parity with Windows and NVIDIA.

Should I use ROCm or DirectML for Forge on an AMD GPU?

Native Linux ROCm is the most coherent upstream compute path when AMD’s current matrix lists your exact GPU and OS. On native Windows, DirectML is a source-visible but manual experiment. ROCm in WSL 2 is now an upstream AMD route for selected hardware, but its current framework stack is not verified against original Forge.

Can Forge run on AMD in Windows 11?

A DirectML experiment is possible because current Forge source accepts --directml, but the launcher does not install torch-directml and the original project provides no tested Windows AMD package. Use a disposable clone and prove DirectML before treating localhost as success.

Does --use-directml work in original Forge?

Not at the inspected commit. The existing parser accepts --directml. An open, unmerged PR #3101 proposes --use-directml as an alias and describes it being silently ignored in the current source.

Why does Forge say “No module named torch_directml”?

The Forge launcher does not automatically install torch-directml. Open issue #749 reproduces the sequence: --directml passes control to code that imports torch_directml, but the package is absent. Install it only inside a disposable Forge venv and verify the device independently.

Why does Forge say “Torch is not able to use GPU” on AMD?

The launcher calls torch.cuda.is_available(). ROCm PyTorch uses that same API and should return true when its stack works; DirectML does not use CUDA semantics and needs the skip only after torch-directml itself passes a device test. The skip flag hides the launcher test—it does not make a GPU usable.

Can I use the Forge one-click Windows archive with AMD?

The documented archive bundles CUDA/PyTorch environments for NVIDIA. It is not an AMD package. Replacing packages inside it creates an untracked hybrid, so use a separate Git clone for any AMD experiment.

Which AMD GPUs work with Forge?

There is no maintained original-Forge AMD GPU matrix. First use AMD’s current Radeon/Ryzen matrix for the compute layer, then treat Forge as a separate compatibility gate. A GPU absent from the AMD matrix is not made supported by a Forge flag.

Can an RX 5000, RX 6000, RX 7000 or RX 9000 card run Forge?

The family name is insufficient. Support changes by exact SKU, architecture, operating system, ROCm release and PyTorch build. Check the current AMD matrix rather than applying a broad “RX 5000–9000” promise from a community tutorial.

Why should I not trust Forge webui.sh to pick ROCm automatically?

At the inspected commit, its generic AMD branch still selects ROCm 5.7 and some older Navi branches select ROCm 5.2 nightlies. Those are historical launcher defaults, not the current AMD compatibility matrix.

Which PyTorch version does original Forge expect?

The inspected launcher defaults to PyTorch 2.3.1 and torchvision 0.18.1. PyTorch’s archive provides a ROCm 6.0 build for that pair on Linux, but AMD’s current hardware support may require a newer stack. Matching one layer can therefore break another.

Can I install the newest ROCm and newest PyTorch for Forge?

You can test them in a disposable environment, but that combination is outside Forge’s inspected pin. First prove the AMD/PyTorch layer, record exact versions, then run a clean Forge clone. Do not call the result supported until generation and the required workflow are reproduced.

Can Forge use ROCm on WSL 2 with an AMD GPU?

AMD now documents a production ROCDXG WSL path for selected Radeon and Ryzen hardware. That is upstream compute support, not Forge validation. Check the current WSL matrix and expect a version gap between its validated PyTorch and Forge’s older pin.

Should I install a Linux AMD display driver inside WSL?

Follow the current AMD WSL procedure for your exact release. The modern ROCDXG route uses the Windows Adrenalin driver as the authoritative display driver and a user-mode bridge inside WSL; do not mix it with a native-Linux driver recipe.

Does Forge support native ROCm on Windows?

AMD’s upstream Windows ROCm capabilities do not automatically create a Forge integration. The inspected original Forge code exposes DirectML, not a documented native-Windows ROCm installation path. This page does not claim one.

Does ZLUDA work with original Forge?

The inspected original source contains no maintained ZLUDA installation path. ZLUDA recipes and AMD-specific forks are separate projects with separate owners, patches and risks; do not relabel them as original Forge.

Can I use FLUX or NF4 on AMD Forge?

Do not make FLUX the first hardware test. First prove the backend with a small SD 1.5 baseline. FLUX adds text encoders, much larger memory demand and quantization backends; open original-Forge issues show AMD-specific NF4 and bitsandbytes uncertainty.

Why does AMD Forge open in the browser but fail during generation?

Localhost proves only that the web server started. Model load and inference can still fail on an unsupported operator, dtype, mixed CPU/GPU tensor, VAE decode or memory allocation. Keep the console open and record the first error.

Should I add --lowvram, --always-normal-vram or performance flags?

Not for the first test. Community guides disagree on these flags, and Forge manages memory dynamically. Prove one small clean generation before changing one variable at a time.

How do I undo an AMD Forge experiment?

Stop the process and delete only the dedicated forge-amd-lab checkout after saving its receipt and error log. Because the clone, venv and configuration were isolated, no global Python rollback or repair of a working installation is required.

09 · Evidence ledger

What this guide is built on

Product claims come from original source and issues. Backend compatibility comes from AMD, PyTorch and Microsoft. Community sources supply dated workflows and user language, never universal support claims.

Primary and direct evidence

20 SOURCES
  1. 01Original Forge READMEDefines the original project and its Windows/NVIDIA-first published installation boundary.
  2. 02Inspected original Forge commitExact source snapshot audited locally on 29 August 2026.
  3. 03Current webui.shContains automatic AMD detection and dated ROCm 5.2/5.7 wheel branches.
  4. 04Current launch_utils.pyDefines the PyTorch 2.3.1 default and torch.cuda.is_available startup gate.
  5. 05Current backend argumentsConfirms --directml exists and --use-directml does not exist at the inspected commit.
  6. 06Current memory managementImports torch_directml when the DirectML backend is selected and includes ROCm-aware branches.
  7. 07DirectML launch issue #58Owner response and community reproduction show manual package and operator limitations; remains open.
  8. 08AMD CUDA-test issue #369Open AMD report for the launcher GPU availability gate.
  9. 09Missing torch-directml issue #749Clean-install reproduction of the missing package after selecting --directml.
  10. 10ROCm regression issue #1348Dated RX Vega / ROCm 6.1 / PyTorch 2.4 failure showing why one success cannot become a matrix.
  11. 11Ubuntu AMD documentation gap #2376Open unanswered request for an original-Forge AMD installation path.
  12. 12AMD support issue #2698Open support request; no maintained matrix or resolution.
  13. 13Open DirectML alias PR #3101Proposes --use-directml and no-NVIDIA fallback; unmerged and test boxes remain unchecked.
  14. 14AMD ROCm compatibility matricesPrimary hardware, OS and framework gate for current Radeon/Ryzen environments.
  15. 15AMD native Linux ROCm guideCurrent upstream sequence for Radeon compute on Linux.
  16. 16AMD ROCm on WSL guideCurrent ROCDXG architecture, supported distro boundary and host-driver model.
  17. 17PyTorch local installationExplains ROCm selection and device verification through torch.cuda.is_available.
  18. 18PyTorch previous versionsConfirms the Linux ROCm 6.0 wheel pair for PyTorch 2.3.1.
  19. 19Microsoft DirectML overviewStates DirectML is supported in sustained engineering while new feature development moved.
  20. 20Microsoft Torch-DirectML setupDocuments package installation, device verification and the PyTorch 2.3.1 ceiling.

Local research catalog

6 SOURCES
  1. O01 · Original Forge repository ↗

    Product identity and published install boundary.

  2. A18 · AI-Dock Forge image ↗

    Third-party Docker project reviewed; its ROCm 6.0 image is not an official Forge binary or a native install recipe.

  3. A35 · AMD Forge community guide ↗

    Full page reviewed. Its --use-directml instruction conflicts with current original source and its broad GPU/performance claims were excluded.

  4. V34 · ROCm 6 + Fedora 40 video ↗

    Full automatic transcript reviewed. RX 6900 XT success used as a dated case, not universal instructions; manual Torch replacement and custom flags remain community evidence.

  5. R12 · FLUX on Forge with AMD ↗

    Catalogued high-volatility workflow; inaccessible during this check and used only as a FLUX/AMD intent signal.

  6. R13 · Forge on AMD/Linux Mint ↗

    Catalogued ROCm, skip-install and fork-confusion discussion; inaccessible during this check and not used for technical claims.

Author: Forge Field Guide editorial teamTechnical review: source and dependency auditUpdated: 29 Aug 2026Runtime: no AMD GPU generation run