tilelang.ascend.language.annotations ==================================== .. py:module:: tilelang.ascend.language.annotations .. autoapi-nested-parse:: Ascend auto-schedule annotations: multi-buffer version control and memory capacity overrides consumed by the Z3 scheduler. Functions --------- .. autoapisummary:: tilelang.ascend.language.annotations.annotate_buffer_versions tilelang.ascend.language.annotations.annotate_manual_multi_buffer tilelang.ascend.language.annotations.annotate_unlimit_memory Module Contents --------------- .. py:function:: annotate_buffer_versions(buffer_versions_map) Control automatic multi-buffer version counts and indexing modes. When using T.Persistent with num_stages > 1, the Z3 auto-scheduler computes optimal buffer version counts for each on-chip buffer. Each mapping value may be one of: - ``1``: opt out of multi-buffer eligibility, including explicit owner claims; - ``num_versions >= 2``: use a fixed version count and automatic mode; - ``(num_versions, mode)``: use a fixed count and explicit mode; - ``mode``: select a mode while leaving the version count to the scheduler. ``mode`` is ``"auto"``, ``"iteration"``, or ``"counter"``. ``"iteration"`` uses the affine flattened-loop index. ``"counter"`` uses a monotonic local counter that advances after a completed active buffer epoch. ``"auto"`` keeps the affine fast path when safe and selects a counter for conditionally executed or non-affine epochs. A fixed count is useful when you want to guarantee double-buffering for a critical buffer regardless of the Z3 solver's choice. ``{buf: 1}`` removes the storage (including aliases) from multi-buffer eligibility, overriding inferred or explicit owner claims and preserving ordinary dependencies instead of owner-exclusion dependencies. Overriding explicit ``multi_buffer_eligible`` claims emits a warning once per storage. Use this when the eligibility heuristic misses reads of previous data. To keep eligibility with one version, explicitly specify a mode: ``{buf: (1, "auto")}``. AutoSchedule consumes this map and replaces it with the selected version counts, including versions chosen automatically by the solver. .. rubric:: Example >>> @T.prim_func ... def my_kernel(x: T.Tensor((M, K), T.float16), ... w: T.Tensor((N, K), T.float16), ... y: T.Tensor((M, N), T.float16)): ... with T.Kernel(T.ceildiv(N, block_N), threads=128) as bx: ... fixed = T.alloc_shared((block_M, block_K), T.float16) ... inferred = T.alloc_shared((block_M, block_K), T.float16) ... T.annotate_buffer_versions({ ... fixed: (2, "counter"), ... inferred: "iteration", ... }) ... with T.Persistent(T.ceildiv(M, block_M), num_stages=4) as i: ... ... .. py:function:: annotate_manual_multi_buffer(*items) Declare buffers that the user has *manually* multi-buffered. Unlike :func:`annotate_buffer_versions` (which asks the Z3 auto-scheduler to version a buffer for you), this annotation tells AutoSchedule that the buffer is *already* multi-buffered by hand: the user writes the version index into every access themselves (e.g. ``buf[l % 2]`` / ``buf[(l + 1) % 2]`` ping-pong), and the pass only needs to analyze dependencies and emit the ``set_flag`` / ``wait_flag`` synchronization. The per-buffer version count ``N`` is used as the ring size / the upper bound on the cross-iteration distance the analysis searches for — it does NOT reshape or reallocate the buffer and does NOT assume the version index lives in dimension 0. Each positional argument is either: - a bare buffer: ``N`` defaults to its leading dimension ``buffer.shape[0]`` (which must then be a compile-time constant) — the common ``T.alloc_shared((N, ...))`` ping-pong case; or - a ``{buffer: N}`` dict: gives ``N`` explicitly, decoupled from the shape. Use this when the buffer has no dedicated leading version dim — e.g. a flattened buffer multi-buffered by disjoint sub-ranges — or when you want a version count that differs from ``shape[0]``. .. rubric:: Example >>> @T.prim_func ... def my_kernel(...): ... with T.Kernel(...) as core_id: ... buf = T.alloc_shared((2, N), T.bfloat16) # leading version dim ... flat = T.alloc_shared((3 * N,), T.bfloat16) # flat, sub-ranges ... T.annotate_manual_multi_buffer(buf, {flat: 3}) ... for l in T.serial(L): ... ... # buf[l % 2] ping-pong; flat[(l % 3) * N : ...] triple-buffer .. py:function:: annotate_unlimit_memory(*scopes) Remove the memory capacity limit for the given scopes during auto-scheduling. When unlimit'd, a scope's capacity is effectively unbounded so Z3 will not constrain multi-buffer versions based on that memory pool. Use this when you know a particular scope has spare headroom. Valid scopes: ``"shared"``, ``"shared.l1"``, ``"shared.l0c"``, ``"shared.l0a"``, ``"shared.l0b"``. .. rubric:: Example >>> @T.prim_func ... def my_kernel(x: T.Tensor((M, K), T.float16), ... w: T.Tensor((N, K), T.float16), ... y: T.Tensor((M, N), T.float16)): ... with T.Kernel(T.ceildiv(N, block_N), threads=128) as bx: ... T.annotate_unlimit_memory("shared") ... ...