Repository navigation
Conversation
Signed-off-by: zhaoye <772971548@qq.com>
|
Thanks for fixing. |
7ceec71 to
78eee45
Compare
|
@XFDG I can reproduce the overflow from #2886 as well. I checked the workspace calculation on main ( The allocation succeeds. The failure happens when the dynamic DLPack descriptor is constructed: import torch
from cutlass.cute.runtime import from_dlpack
for nbytes in [2**31 - 1, 2**31, 5697536000]:
workspace = torch.empty(nbytes, device="cuda", dtype=torch.uint8)
for dynamic in [False, True]:
tensor = from_dlpack(workspace, assumed_align=16)
if dynamic:
tensor = tensor.mark_layout_dynamic()
try:
tensor.__c_pointers__()
result = "PASS"
except OverflowError as e:
result = str(e)
print(nbytes, dynamic, result)
del tensor
del workspace
torch.cuda.empty_cache()
For the last row: Tested on B300 / SM103, editable CuTe DSL main with the 4.8.0 CUDA 13 backend (reports CUDA 13.4). This isolates the shared runtime failure; I did not run the full FMHA backward kernel or test on B200. The added |
Summary
BlackwellFusedMultiHeadAttentionBackward._get_workspace_sizecombines user-shape parameters for workspace sizing with integer math that can be narrowed to int32 in the Cute/CUDA path for large sequence and batch dimensions. That can truncate the size and trigger allocation-related failures.Fix
In
examples/python/CuTeDSL/cute/blackwell/kernel/attention/fmha/fmha_bwd.pyfunctionBlackwellFusedMultiHeadAttentionBackward._get_workspace_size, explicitly use Pythonintintermediates for all multipliers and dimensions:q/dwithint((q + 7) // 8 * 8)b_i32,h_i32,q_i32,d_i32, andacc_bytesOverflowErrorif computed workspace bytes is negativeThis keeps workspace arithmetic in Python big-int space and avoids unintentional int32 narrowing before CUDA allocation.
Testing
python3 - <<'PY' seq_len = 46341 ws = int(seq_len) * int(seq_len) * 4 assert ws > 0 print('workspace size:', ws) PYFixes #2886