Skip to content

Synchronize GQA reduction scratch before reuse - #3600

Open
1sgtpepper wants to merge 2 commits into
NVIDIA:mainfrom
1sgtpepper:fix/gqa-reduction-scratch-ordering
Open

1sgtpepper wants to merge 2 commits into
NVIDIA:mainfrom
1sgtpepper:fix/gqa-reduction-scratch-ordering

Conversation

@1sgtpepper

@1sgtpepper 1sgtpepper commented Sep 8, 2026 •

Copy link
Copy Markdown

Fixes #3597.

Complete all cta_reduce scratch reads before the next reduction reuses the storage. The existing barrier orders publication before reads; a second arrival on that same 128-thread barrier orders reads before reuse.

Tests

Build 93_blackwell_low_latency_gqa_cta_reduce with CUTLASS_NVCC_ARCHS=100a, then run:

compute-sanitizer --tool racecheck --error-exitcode 3 ./build/examples/93_blackwell_low_latency_gqa/93_blackwell_low_latency_gqa_cta_reduce

On B200/CUDA 13.1.1, the repeated-reduction test passes numerically with zero racecheck hazards. Restoring the original header produces wrong values and eight hazards. CMake build.

This fixes scratch reuse independently of the mailbox initialization race in #3601. With both fixes, the contiguous and paged example checks pass. Another 16 contiguous regression cases pass under racecheck and memcheck.

Signed-off-by: 1sgtpepper <cynejarviszarceno@gmail.com>
Signed-off-by: 1sgtpepper <cynejarviszarceno@gmail.com>
@github-actions

Copy link
Copy Markdown

This PR has been labeled inactive-30d due to no recent activity in the past 30 days. Please close this PR if it is no longer required. Otherwise, please respond with a comment indicating any updates. This PR will be labeled inactive-90d if there is no activity in the next 60 days.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Example 93 GQA reduction races when reusing shared scratch

1 participant