Hardware Platforms¶
TeleFuser abstracts execution hardware behind the platform layer in telefuser/platforms/. Each process resolves a single current_platform object at import time, and the Ops dispatch layer uses it to select operator implementations. Pipelines and examples configure a torch.device-style device and do not branch on the vendor.
Platform selection¶
_resolve_current_platform() in telefuser/platforms/__init__.py probes the environment in a fixed order and instantiates the first match:
- ROCm — a PyTorch
+rocm(HIP) build with at least one visible AMD GPU - CUDA — a CUDA build with at least one visible NVIDIA GPU
- NPU — a
torch_npuinstallation with a visible Ascend device - CPU — the fallback when no accelerator is detected
HIP builds expose the torch.cuda API, so the ROCm platform reuses it (device_type stays cuda); check torch.version.hip to distinguish a ROCm host from an NVIDIA one. A HIP or CUDA build without a visible GPU falls back to CPU. CUDA_VISIBLE_DEVICES controls device visibility on both GPU platforms, and ASCEND_RT_VISIBLE_DEVICES on NPU.
Platform matrix¶
| Platform | device_type | Distributed backend | Operator dispatch | tf-kernel | torch.compile |
|---|---|---|---|---|---|
| CUDA (NVIDIA) | cuda | NCCL | Optimized forward_cuda paths | Supported | Experimental (details) |
| ROCm (AMD) | cuda | RCCL (nccl) | Triton forward_cuda paths and native fallbacks | Not available | Not validated |
| NPU (Ascend) | npu | HCCL | Native fallbacks | Not available | Not validated |
| CPU | cpu | Gloo | Native fallbacks | Not available | Native path |
Attention backends by platform¶
Attention backend availability is resolved at import time (see Attention). A backend whose dependencies are missing falls back to TORCH_SDPA with a one-time warning.
- CUDA: all dense backends —
TORCH_SDPA, FlashAttention ⅔/4, SageAttention throughtf-kernelorsageattention, and cuDNN — subject to GPU architecture; sparse backends likewise. - ROCm:
TORCH_SDPAonly.flash_attnhas no official ROCm wheels for consumer RDNA GPUs, andtf-kernel, SageAttention, and SpargeAttn are CUDA-only. The AOTriton-backed SDPA path is the fast attention kernel on RDNA4. - NPU / CPU:
TORCH_SDPAthrough the native fallback paths; no vendor attention kernels are integrated.
CUDA¶
CUDA is the primary validated path: Python 3.10–3.13, PyTorch 2.6 or newer, CUDA toolkit 12.8 or newer, with H100 as the validated target for optimized kernels. Optional tf-kernel provides fused elementwise operations, quantized GEMM, SageAttention, and block-sparse attention — see tf-kernel for build and artifact compatibility. Multi-GPU inference uses NCCL; see Parallel Inference.
ROCm¶
ROCm support targets AMD GPUs with ROCm 7.x and a PyTorch +rocm build; see Installation for the setup path.
- Attention uses
TORCH_SDPA; notf-kernel,flash_attn, orsageattentioninstallation is required. - The ops layer selects
forward_rocmwhere a kernel defines one and otherwise reuses the CUDA Triton path (Triton supports ROCm), falling back to native PyTorch. tf-kernelimports are gated toCudaPlatform, so no CUDA-only extension is loaded on ROCm hosts.- Multi-GPU inference uses RCCL, AMD's NCCL-compatible collectives library. PyTorch's ROCm build exposes it through the
ncclbackend string, so the platform layer requires no special configuration. torch.compileis not validated on ROCm; ROCm examples run eager.- Validated entry points are the
*_rocm.pyexamples, for example Wan2.1 1.3B text-to-video on a Radeon RX 9070 (gfx1201, ROCm 7.2). Multi-GPU branches reuse the_h100.pyparallel configuration but are not yet validated.
NPU and CPU¶
The NPU platform targets Huawei Ascend devices through torch_npu with the HCCL distributed backend. It is wired into the platform and ops dispatch layers, but the maintained examples are validated on CUDA and, for select examples, ROCm — validate on your target NPU before production use. The CPU platform is the fallback when no accelerator is detected; it is intended for tests and for pipelines that explicitly request CPU execution.