Skip to content
@uccl-project

UCCL

Next-generation GPU communication

Pinned Loading

  1. uccl uccl Public

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    C++ 1.5k 175

  2. mKernel mKernel Public

    mKernel: fast multi-node, multi-GPU fused kernels

    Cuda 276 26

  3. CommBench CommBench Public

    Can LLMs Write Correct and Efficient GPU Communication Code?

    Python 63 2

  4. rdmatop rdmatop Public

    htop-like TUI for real-time RDMA network monitoring.

    Rust 104 9

Repositories

Showing 9 of 9 repositories
  • rdmatop Public

    htop-like TUI for real-time RDMA network monitoring.

    uccl-project/rdmatop's past year of commit activity
    Rust 104 Apache-2.0 9 0 1 Updated Sep 18, 2026
  • mKernel Public

    mKernel: fast multi-node, multi-GPU fused kernels

    uccl-project/mKernel's past year of commit activity
    Cuda 276 MIT 26 1 4 Updated Sep 17, 2026
  • uccl Public

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    uccl-project/uccl's past year of commit activity
    C++ 1,519 Apache-2.0 175 57 (1 issue needs help) 8 Updated Sep 15, 2026
  • CommBench Public

    Can LLMs Write Correct and Efficient GPU Communication Code?

    uccl-project/CommBench's past year of commit activity
    Python 63 2 0 1 Updated Jul 6, 2026
  • uccl-project/uccl-project.github.io's past year of commit activity
    Astro 3 MIT 3 1 1 Updated Jun 14, 2026
  • nixl Public Forked from ai-dynamo/nixl

    NVIDIA Inference Xfer Library (NIXL)

    uccl-project/nixl's past year of commit activity
    C++ 0 Apache-2.0 452 0 0 Updated Nov 24, 2025
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    uccl-project/vllm's past year of commit activity
    Python 0 Apache-2.0 22,641 0 1 Updated Nov 3, 2025
  • ray-uccl Public Forked from ray-project/ray

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    uccl-project/ray-uccl's past year of commit activity
    Python 0 Apache-2.0 8,261 0 0 Updated Jul 22, 2025
  • nccl Public Forked from NVIDIA/nccl

    Optimized primitives for collective multi-GPU communication

    uccl-project/nccl's past year of commit activity
    C++ 0 1,433 0 0 Updated Jul 4, 2025