Skip to content

fix(CuTeDSL): optimize positional-only functions with TVMFFIJitCompiledFunction to eliminate CPU overhead (#3527) - #3589

Open
Ammar-Alnagar wants to merge 4 commits into
NVIDIA:mainfrom
Ammar-Alnagar:fix/tvm-ffi-posonly-overhead
Open

Ammar-Alnagar wants to merge 4 commits into
NVIDIA:mainfrom
Ammar-Alnagar:fix/tvm-ffi-posonly-overhead

Conversation

@Ammar-Alnagar

@Ammar-Alnagar Ammar-Alnagar commented Sep 6, 2026 •

Copy link
Copy Markdown

1. What was the Issue? (#3527)

In CUTLASS v4.6, the routing condition in python/CuTeDSL/cutlass/cutlass_dsl/cutlass.py was updated to check:

if (
    kwargs_wrapper_spec.kwonly_names
    or kwargs_wrapper_spec.arg_defaults
    or kwargs_wrapper_spec.arg_names
    or map_dataclass_to_tuple
):
    return TVMFFIJitCompiledFunctionWithKwargs(...)

Because kwargs_wrapper_spec.arg_names includes positional parameter names for every argument, any @cute.jit function taking 1 or more arguments evaluated to True and was routed to TVMFFIJitCompiledFunctionWithKwargs.

TVMFFIJitCompiledFunctionWithKwargs wraps invocation in kwargs_wrapper.make_kwargs_wrapper. When registered via TVM-FFI and called from C++, execution left C++ to enter Python (GIL + argument wrapper parsing), causing CPU overhead regression compared to direct C++ TVM-FFI function execution.


2. What We Did

Positional-only parameters (def func(a, b, c, /)) cannot be called with keyword arguments in Python grammar. Therefore, functions with strictly positional-only arguments do not require keyword-argument wrapping.

  1. python/CuTeDSL/cutlass/base_dsl/jit_executor.py:

    • Added has_pos_or_kw to KwargsWrapperSpec.
    • Updated get_kwargs_wrapper_spec to track whether any Parameter.POSITIONAL_OR_KEYWORD parameter exists.
  2. python/CuTeDSL/cutlass/cutlass_dsl/cutlass.py:

    • Changed kwargs_wrapper_spec.arg_names check in _make_compiled_func to kwargs_wrapper_spec.has_pos_or_kw.
    • Functions defined as def func(a, b, c, /) without defaults or dataclasses now return TVMFFIJitCompiledFunction (a direct subclass of tvm_ffi.Function).
  3. Unit Tests & Documentation:

    • Added unit test suite test_tvm_ffi_kwargs_wrapper_spec.py.
    • Updated documentation in tvm_ffi_compilation.rst:64.

3. Comparison & Verification of Fix

Function Signature has_pos_or_kw arg_defaults kwonly_names Class Returned CPU Overhead Behavior
def func() False () [] TVMFFIJitCompiledFunction 0 Python Overhead (Direct C++ tvm_ffi.Function)
def func(a, b, c, /) (Fix) False () [] TVMFFIJitCompiledFunction 0 Python Overhead (Direct C++ tvm_ffi.Function, bypasses GIL/wrapper)
def func(a, b, c) True () [] TVMFFIJitCompiledFunctionWithKwargs Handles Python kwargs calls (func(a=1, b=2))
def func(a, b=1, /) False (1,) [] TVMFFIJitCompiledFunctionWithKwargs Handles positional default value binding
def func(a, /, *, k=1) False () ['k'] TVMFFIJitCompiledFunctionWithKwargs Handles keyword-only argument binding
Verification Summary
  • Evaluated signature parsing with inspect.signature across positional-only, positional-or-keyword, and mixed argument patterns.
  • Confirmed def func(a, b, c, /) correctly evaluates check_routing() to False, bypassing TVMFFIJitCompiledFunctionWithKwargs and directly instantiating TVMFFIJitCompiledFunction.

…edFunction to eliminate CPU overhead (NVIDIA#3527)

Functions defined with positional-only arguments (e.g., ) cannot be called with keyword arguments. Therefore, they do not require TVMFFIJitCompiledFunctionWithKwargs wrapper routing and can directly use TVMFFIJitCompiledFunction, bypassing Python execution and GIL overhead when called from C++ or Python.
@Ammar-Alnagar

Copy link
Copy Markdown
Author

Hey @kainzhong / @tqchen,

Following up on the earlier discussion ([https://github.com//issues/3527]):

Hi @kainzhong, thanks for reaching out! This is a known issue for us. Your understanding is correct - the condition you mention is generally true, except the trivial functions without any args. But per our benchmark, the overhead it introduces is < 1us, while the overall launch overhead is typically <10 us. We think it is fine to keep this overhead for compatibility of diversed inputs. And we agree to have more documents here. Any PR is appreciated.
cc: @tqchen

This PR addresses the routing issue so that pure positional-only functions (def func(a, b, c, /)) now correctly bypass TVMFFIJitCompiledFunctionWithKwargs and use the direct TVMFFIJitCompiledFunction path (zero Python/GIL overhead).

Would appreciate a review when you have a chance — thanks!

Comment thread media/docs/pythonDSL/guides/tvm_ffi_compilation.rst Outdated
Comment thread test/python/CuTeDSL/test_tvm_ffi_kwargs_wrapper_spec.py
@kainzhong

Copy link
Copy Markdown
Contributor

Hi @cyx-6 can you take a look at this PR? My opinion is

  • Should we make this behavior even more explicit? Now adding / can change the CPU overhead feels a bit black magic and I wonder if it would be better to add a param to cute.compile like optimize_for_pos_only
  • I don't think test/python/CuTeDSL/test_tvm_ffi_cpp_registry.py is needed
  • test_kwargs_wrapper_spec_positional_only, test_kwargs_wrapper_spec_positional_or_keyword and test_kwargs_wrapper_spec_mixed_posonly_and_pos_or_kw seem unnecessary as well

But I think that should be up to cutlass people to decide. What do you think?

@cyx-6

cyx-6 commented Sep 16, 2026

Copy link
Copy Markdown

@kainzhong Thanks for contribution! This change looks good to me. And yes, I agree that some semantic changes should be discussed and decided by cutlass people.

@Ammar-Alnagar

Copy link
Copy Markdown
Author

Hi @cyx-6 can you take a look at this PR? My opinion is

  • Should we make this behavior even more explicit? Now adding / can change the CPU overhead feels a bit black magic and I wonder if it would be better to add a param to cute.compile like optimize_for_pos_only
  • I don't think test/python/CuTeDSL/test_tvm_ffi_cpp_registry.py is needed
  • test_kwargs_wrapper_spec_positional_only, test_kwargs_wrapper_spec_positional_or_keyword and test_kwargs_wrapper_spec_mixed_posonly_and_pos_or_kw seem unnecessary as well

But I think that should be up to cutlass people to decide. What do you think?

I can see where your point in the black box change and an input from one of the cutlass people would be much appreciated but i think the test cases don't hurt and are more of an assurance , but thats just my opinion and might not be the best flow for the project .
Curious on what you think? @cyx-6 @kainzhong

@kainzhong

Copy link
Copy Markdown
Contributor

Hi @Ammar-Alnagar unfortunately I'm not a cutlass developer. I think you probably should request their review instead since they are the ones who can approve it.

@Ammar-Alnagar

Copy link
Copy Markdown
Author

Hi @Ammar-Alnagar unfortunately I'm not a cutlass developer. I think you probably should request their review instead since they are the ones who can approve it.

no worries , thanks alot for the help and comments !!!

@Ammar-Alnagar

Copy link
Copy Markdown
Author

@Junkai-Wu @brandon-yujie-sun @fengxie , can i get a review on this whenever you get a chance ? , thanks!!!

@brandon-yujie-sun brandon-yujie-sun left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for contributing the change!

Comment thread python/CuTeDSL/cutlass/base_dsl/jit_executor.py Outdated
Comment thread test/python/CuTeDSL/test_tvm_ffi_cpp_registry.py Outdated
Comment thread test/python/CuTeDSL/test_tvm_ffi_kwargs_wrapper_spec.py Outdated
test

Move keyword-only and positional-or-keyword flag checks after constexpr
annotation filtering to avoid unnecessary computation for excluded
parameters.

Delete the C++ TVM FFI registry integration test which benchmarked GPU
launch overhead via a custom C++ extension. The test required a CUDA
toolkit and tvm_ffi at build time, making it fragile in CI environments.

Move tvm_ffi import to top of kwargs wrapper spec test for consistency.

@brandon-yujie-sun brandon-yujie-sun left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks

@brandon-yujie-sun

Copy link
Copy Markdown
Collaborator

@Junkai-Wu for further processing

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants