tilelang.ascend.torch_exchange¶

Runtime-only Torch NPU stream integration for TVM-FFI execution.

Functions¶

install_torch_npu_stream_exchange()

Make TVM-FFI submit Ascend kernels to Torch's current NPU stream.

is_torch_npu_stream_exchange_installed()

Return whether TileLang owns Torch's current Exchange API table.

npu_current_device()

Torch device for the current NPU, for kernel output allocation.

npu_current_raw_stream()

Raw stream of Torch's current NPU stream.

Module Contents¶

tilelang.ascend.torch_exchange.install_torch_npu_stream_exchange()¶

Make TVM-FFI submit Ascend kernels to Torch’s current NPU stream.

Return type:

bool

tilelang.ascend.torch_exchange.is_torch_npu_stream_exchange_installed()¶

Return whether TileLang owns Torch’s current Exchange API table.

Return type:

bool

tilelang.ascend.torch_exchange.npu_current_device()¶

Torch device for the current NPU, for kernel output allocation.

tilelang.ascend.torch_exchange.npu_current_raw_stream()¶

Raw stream of Torch’s current NPU stream.

Uses the low-level accessor when available: torch.npu.current_stream() goes through a Python wrapper that internally probes torch.cuda.is_available(), which costs ~150us per call on a CUDA-less NPU host; the _C accessor returns the raw stream in <1us.