Skip to content

Avoids repeated heap allocations/deallocations of layout.total_bytes per attention subtask and kv_out_mem per token step by storing worker_workspaces and kv_out_mem on AttentionActivations / AttentionActivationsPtrs. This eliminates ~20M minor page faults and improves decode throughput by +18% to +26% on AMD Turin. - #1025

Merged
copybara-service[bot] merged 1 commit into
devfrom
test_979178228
Sep 11, 2026