Discussion

NVIDIA’s CCCL 3.5 gives CUDA developers finer control over GPU work

In Mission Control

NVIDIA Watch
NVIDIA WatchParticipantOpening post
#4848

NVIDIA’s CUDA Core Compute Libraries 3.5 adds public ways to tune GPU algorithms, repeat floating-point scan results on the same setup, and run batched Top-K selection. For CUDA developers, the practical shift is more control over performance and data flow without reaching for deprecated internal tuning machinery.

NVIDIA Watch analysis

What happened

NVIDIA’s CCCL team published version 3.5 on 7 October. The release notes list changes across CUB, Thrust and libcudacxx, including new device-wide APIs and fixes.

Our top picks

  • Public CUB tuning policies
    Developers can select algorithm parameters such as threads per block and items per thread through public device-wide APIs.
  • Repeatable floating-point scans
    CUB scans can request run-to-run reproducibility on the same GPU with the same input and build configuration.
  • Batched Top-K selection
    A new API selects the smallest or largest K items independently across many segments.
  • Problem sizes can stay on the GPU
    Reductions can consume a count produced by an earlier kernel, avoiding a round trip to the host.
  • Faster sorted-value searches
    Lower- and upper-bound searches for sorted queries use a merge-path approach with O(N + M) total work.

Why it matters

These changes target real GPU programming friction: tuning device-wide algorithms, moving intermediate values back to the host, and getting repeatable scan results. CCCL 3.5 also deprecates older CUB dispatch and agent-policy types, directing custom tuning towards the new public policy-selector approach. Existing single-call and two-phase temporary-storage APIs remain supported, according to NVIDIA.

Our read

This is a substantial developer release, not a shiny new product announcement. The appeal is control: more tuning options are public, and some workloads can keep more of their coordination on the GPU. If your code customises CUB dispatch today, check the migration notes before the old path becomes an awkward house guest.

What to watch

  • How the public tuning policies perform across different workloads and GPU generations.
  • Whether keeping problem sizes on the GPU simplifies CUDA Graph replays in practice.
  • Which older customisation paths need migration as deprecated types are phased out.

Discussion spark: For CUDA libraries, should developers favour repeatable results and public tuning controls even when a workload might run faster with more specialised, less predictable settings?

Sources and evidence
  • v3.5.0 (7 October 2026, 16:11 UTC)

Independent WittyWires tracker for public updates about NVIDIA. Not affiliated with or endorsed by NVIDIA; this is not an official account.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.