
Stable Diffusion has no single required CUDA version. On an NVIDIA GPU, use the PyTorch CUDA build supported by your chosen application, your GPU architecture, and your installed driver. ComfyUI, AUTOMATIC1111, and a custom Diffusers pipeline can require different combinations.
This guide helps you identify the versions actually in use, choose a compatible installation, and distinguish driver or package problems from insufficient GPU memory. It focuses on NVIDIA setups on Windows and Linux. Apple, AMD, and Intel acceleration use other software paths.
CUDA is NVIDIA's platform for running general-purpose computations on its GPUs. A Stable Diffusion model performs repeated tensor calculations to generate an image; PyTorch's CUDA support lets those calculations run on a compatible NVIDIA GPU. GPU memory, also called VRAM, holds model weights and the data needed during execution.
The CUDA version does not determine image quality or generation speed on its own. The model, hardware, precision, resolution, and application also matter. First establish a working installation, then measure performance with a repeatable workflow.
Start with the application you intend to run, then check its supported PyTorch package, Python version, and GPU requirements. A model checkpoint does not prescribe a universal CUDA release for every interface.
The compatibility chain is: GPU architecture → NVIDIA driver → PyTorch CUDA build → application → compiled extensions. Each part must support the next. Installing the newest CUDA toolkit alone does not repair an incompatible Python environment.
Exact package recommendations below were checked on October 6, 2026. Recheck the linked installation instructions before changing a working environment, because application requirements and available packages change.
The driver provides access to the GPU. It must support both the hardware and the CUDA features used by the application. A recent runtime paired with an older driver can fail before a model loads.
NVIDIA's CUDA compatibility guidance lists minimum driver families of 450 for CUDA 11.x, 525 for CUDA 12.x, and 580 for CUDA 13.x. These are broad compatibility families, not sufficient installation specifications for every GPU, operating system, and application.
Check the full driver release requirements. New hardware, newer features, or code compiled at runtime may require a newer driver than the family minimum. NVIDIA also documents limitations when relying on minor-version compatibility.
The nvidia-smi documentation describes the CUDA field as the version supported by the driver. Newer documentation calls it the CUDA UMD version. It does not identify the CUDA build imported by your Stable Diffusion Python process.
For example, a driver reporting CUDA 13.0 can run a supported PyTorch build packaged for CUDA 12.6. Those numbers differ without establishing a fault. Do not downgrade a compatible newer driver simply to make the displayed numbers identical.
The toolkit includes development tools such as the nvcc compiler. It matters when building CUDA code, PyTorch, or certain extensions from source. A prebuilt PyTorch installation provides its runtime dependencies and normally does not need a separate local toolkit to execute GPU operations.
A PyTorch maintainer explains this distinction. A missing nvcc command can coexist with a working GPU application. Installing a toolkit does not turn a CPU-only PyTorch package into a CUDA-enabled package.
torch.__version__ identifies the installed PyTorch release. torch.version.cuda reports the CUDA version used to build that package. For this NVIDIA workflow, a value of None indicates that the imported package has no CUDA build version.
Check both values inside the application's actual Python environment. A globally installed PyTorch package may differ from the one inside a virtual environment, notebook kernel, portable distribution, or container.
The Stable Diffusion interface may manage or pin its own dependencies. Custom nodes and compiled extensions can add requirements for Python, PyTorch, CUDA, or specific GPU architectures.
A working base application followed by an extension import failure points to a narrower problem than a broken driver. Save the error and the versions before changing packages.
Run this in a terminal on the machine that will execute the model:
nvidia-smi
Record the GPU model, driver information, and memory usage. If it fails to communicate with the driver or cannot see the expected GPU, investigate the driver or the environment's GPU access first. Reinstalling Python packages will not fix every failure at this level.
On a remote instance, run the check there rather than on your laptop. For a container, compare the host and container results if you administer both.
Save the following as check_gpu.py. Activate the application's environment, or use its documented Python executable, before running it.
import sys
import torch
print("Python:", sys.executable)
print("PyTorch:", torch.__version__)
print("PyTorch CUDA:", torch.version.cuda)
available = torch.cuda.is_available()
print("CUDA available:", available)
if available:
print("GPU:", torch.cuda.get_device_name(0))
print("Compute capability:", torch.cuda.get_device_capability(0))
value = torch.ones(1, device="cuda")
print("CUDA test:", (value + value).item())Run the saved file with the Python interpreter used by the application:
python check_gpu.py
Here, python means that interpreter, not necessarily the first Python installation on your system. A portable UI may have an embedded interpreter; a notebook uses its selected kernel. If the printed executable points somewhere unexpected, correct that before installing anything.
The PyTorch CUDA API provides these availability and device checks. On a working NVIDIA setup, availability should be True, the GPU should be the intended device, and the small tensor calculation should print 2.0.
This is a basic execution check. It does not prove that the entire model, every extension, or the intended image resolution will work. If importing PyTorch fails, preserve that error; if the tiny CUDA calculation fails, preserve its traceback.
These commands inspect packages in the same Python environment:
python -m pip show torch torchvision
python -m pip checkThe first shows installed package metadata and location. The second checks declared dependency compatibility, but cannot verify GPU kernel support or guarantee that an extension works.
If you need to compile CUDA code, check the toolkit compiler separately:
nvcc --version
Its absence is not a reason to repair a prebuilt installation that already passes the GPU test. The output from nvidia-smi, nvcc, and PyTorch describes different layers; the numbers do not all need to match.
Use this sequence for a new installation or a deliberate repair:
The PyTorch package matrix currently lists CUDA 12.6, 13.0, and 13.2 packages for PyTorch 2.13.0 on Linux and Windows. That is an example of available builds, not a recommendation to replace an application's pinned dependencies with PyTorch 2.13.0.
Keep related packages, such as torchvision, aligned with the documented combination. An independently upgraded extension can change the dependency requirements even when you leave the CUDA driver untouched.
ComfyUI's current system requirements point NVIDIA users toward PyTorch with CUDA 13.0. They describe the current Windows portable build as using Python 3.13 and PyTorch CUDA 13.0. The page also warns that its instructions may lag behind the repository.
The official ComfyUI repository separates its portable NVIDIA downloads: the main build supports RTX 20-series and newer GPUs, while a CUDA 12.6/Python 3.12 build is offered for 10-series and older GPUs. Follow the download's hardware guidance rather than choosing solely by the largest version number.
For manual installation, use the repository's current instructions and the Python environment you created for ComfyUI. For a portable installation, diagnose its embedded environment. Changing a separate system Python installation may have no effect on the UI.
Test a built-in workflow before introducing custom nodes. If only a custom workflow fails, check its nodes and dependencies before replacing the base stack.
AUTOMATIC1111's NVIDIA installation wiki provides application-specific setup instructions. At this review, the page still directs RTX 50-series users to its development-branch instructions for PyTorch 2.7 support. That is guidance for this project, not a universal minimum for every Stable Diffusion interface.
Use the documented installation method for your operating system and GPU. Avoid replacing the virtual environment's PyTorch packages with an unrelated “latest CUDA” command: a newer package can conflict with the WebUI or its extensions.
Options that change memory use or numerical precision address other parts of execution. Skipping an availability check does not create GPU support, and a precision flag does not repair an unsupported binary. Diagnose the reported error first.
The Diffusers installation guide tells users to install PyTorch for their system. At this review it requires Python 3.10 or newer and says it is tested with PyTorch 2.6 or newer. It does not define one CUDA version for every model and device.
Use an isolated environment for the pipeline, install a supported PyTorch CUDA build, and verify GPU execution before adding the remaining model dependencies. The Python packaging guide explains virtual environments and installing packages into the active environment.
Use a model's documented pipeline and a modest supported test configuration. A model download, access-permission error, or missing dependency is not automatically a CUDA failure. Keep the first working example simple so that later changes have a clear baseline.
A supported driver is only one part of running on a newer GPU. PyTorch and any compiled extensions must also contain compatible GPU code. An older package may recognize the device while failing when a kernel executes.
NVIDIA's Blackwell compatibility guide distinguishes native GPU code from PTX code that can be compiled at runtime. It describes native Blackwell support in CUDA 12.8 and explains why some older applications can remain compatible when they include suitable PTX.
For a Stable Diffusion user, the practical step is to follow the application's explicit GPU support instructions. Do not assume every old package fails solely because its CUDA number is smaller, or that every new package supports every GPU. The packaged kernels and extension builds matter.
The imported PyTorch package lacks the CUDA support requested by the operation. Check the printed Python executable and torch.version.cuda. On an intended NVIDIA setup, install the application's supported CUDA-enabled package into that exact environment.
Check GPU visibility, the driver, the PyTorch build, and the interpreter path. A working nvidia-smi result does not prove that the application imported the right package. A failed availability check also does not, by itself, identify which layer is wrong.
Compare the runtime requirements with the installed driver. Update the driver through the supported process, or use a compatible package combination supported by the application and GPU. On managed infrastructure, the host driver may be the provider's responsibility.
Confirm that the intended environment exposes a GPU. For containers, inspect GPU access as well as host visibility. Also check whether device-visibility settings intentionally hide the GPU from the process.
Check architecture support in both PyTorch and the component named in the traceback. If the basic tensor test succeeds but a custom node fails, investigate that node's compiled dependencies. Rebuilding or replacing the failing component may be appropriate; changing every package at once makes the cause harder to isolate.
Record the extension's expected Python, PyTorch, and CUDA combination. Restore a documented compatible set or follow its rebuild instructions. Keep the previous environment available until the updated workflow passes the same test.
A CUDA out of memory error means an allocation could not be satisfied. Start with memory pressure: model size, image resolution, batch size, optional networks, and other GPU processes. The fact that the message contains “CUDA” does not establish a version mismatch.
Reduce the workload and use the application's supported memory-saving options. Precision changes, offloading, and attention implementations have compatibility and performance trade-offs; do not apply a universal flag to every UI.
PyTorch uses a caching allocator. Its memory-management documentation distinguishes live tensor memory from reserved cache. Clearing unused cache cannot free memory still held by active tensors, so it is not a general fix for a model that exceeds capacity.
For sizing rather than software compatibility, see our Stable Diffusion hardware and VRAM guide. The CUDA version alone cannot determine how much memory a particular model and workflow will need.
A container can package an application and its Python and CUDA runtime dependencies. The host still needs a compatible NVIDIA driver and a supported way to expose the GPU. NVIDIA's Container Toolkit installation guide covers that host-side setup.
Check the GPU from inside the execution environment, not only from the host. Reinstalling a driver inside an application container is not the default remedy for missing host access. If you do not administer the host, use the provider's supported GPU image or ask support about the driver.
Record the image version or digest, application version, and installed packages when a workflow succeeds. Treat community images as third-party software and check their source and maintenance. ComfyUI's documentation states that it does not provide an official Docker image.
Compute with Hivenet offers GPU environments to evaluate against the selected application's requirements. The Compute FAQ documents container and virtual-machine options, SSH access, and custom templates. It does not establish one permanent CUDA version for every instance.
After connecting, identify the actual GPU and driver, then run the Python checks in the chosen application environment. Use a documented compatible package set, test a small supported generation, and add custom nodes or extensions incrementally. Verify the current image and runtime rather than assuming an older tutorial describes the instance.
Keep a record of a working environment and separate copies of important models, workflows, and outputs. A reusable template can help reproduce software; it does not replace a data backup. Check current GPU availability and pricing when choosing capacity for the workload.
Usually not with a supported prebuilt PyTorch setup. You may need development tools when compiling CUDA extensions or building from source. Follow the component's requirements.
Yes. One describes driver support and the other describes the imported PyTorch build. Check compatibility instead of trying to make the numbers equal.
No. It appears in current ComfyUI guidance, but other applications, GPU generations, and package combinations have their own requirements. Use the configuration supported by the application you run.
Yes, provided their hardware and driver requirements can be met. Separate their Python environments to reduce dependency conflicts, and account for GPU memory if they run concurrently.
Update for a supported feature, hardware requirement, bug fix, or security fix. Save the working configuration, review the project's update guidance, and test the new environment before replacing the old one.
Pick one AI, compute, or storage workload and see the difference for yourself. Spin it up in minutes, or let our team map your fastest path to production.