tilelang.ascend.torch_exchange ============================== .. py:module:: tilelang.ascend.torch_exchange .. autoapi-nested-parse:: Runtime-only Torch NPU stream integration for TVM-FFI execution. Functions --------- .. autoapisummary:: tilelang.ascend.torch_exchange.install_torch_npu_stream_exchange tilelang.ascend.torch_exchange.is_torch_npu_stream_exchange_installed tilelang.ascend.torch_exchange.npu_current_device tilelang.ascend.torch_exchange.npu_current_raw_stream Module Contents --------------- .. py:function:: install_torch_npu_stream_exchange() Make TVM-FFI submit Ascend kernels to Torch's current NPU stream. .. py:function:: is_torch_npu_stream_exchange_installed() Return whether TileLang owns Torch's current Exchange API table. .. py:function:: npu_current_device() Torch device for the current NPU, for kernel output allocation. .. py:function:: npu_current_raw_stream() Raw stream of Torch's current NPU stream. Uses the low-level accessor when available: torch.npu.current_stream() goes through a Python wrapper that internally probes torch.cuda.is_available(), which costs ~150us per call on a CUDA-less NPU host; the _C accessor returns the raw stream in <1us.