ReleaseNVIDIANVIDIApublished Jul 24, 2026seen 2d

NVIDIA/nvloom v2.1.0

NVIDIA/nvloom

Open original ↗

Captured source

source ↗
published Jul 24, 2026seen 2dcaptured 2dhttp 200method plain

v2.1.0

Repository: NVIDIA/nvloom

Tag: v2.1.0

Published: 2026-07-24T06:29:34Z

Prerelease: no

Release notes:

Added

  • Added support for CUDA Compute Fabric Transport (CFT) programming model with a coverage of unicast and multicast testcases
  • Added fully random and random permutation (ring, bisect) fabric stress tests
  • Added TMA flavors of all-to-one/bisect/gpu-to-rack/rack-to-rack testcases
  • Added --blockCount argument to override block count of copy kernels
  • Added --tmaChunkSize argument to override TMA staging buffer size
  • Added --richOutput support to all-to-one, rack-aware-all-to-one and multicast_all_to_all testcases
  • Added support for CUDA MPS
  • Testcases now output median value alongside preexisting min/average/max/sum

Changed

  • Increased default TMA staging buffer size to 196 kB