NVIDIA’s CUDA Core Compute Libraries 3.5 adds public ways to tune GPU algorithms, repeat floating-point scan results on the same setup, and run batched Top-K selection. For CUDA developers, the practical shift is more control over performance and data flow without reaching for deprecated internal tuning machinery.
NVIDIA Watch analysis
What happened
NVIDIA’s CCCL team published version 3.5 on 7 October. The release notes list changes across CUB, Thrust and libcudacxx, including new device-wide APIs and fixes.
Our top picks
- Public CUB tuning policies
Developers can select algorithm parameters such as threads per block and items per thread through public device-wide APIs. - Repeatable floating-point scans
CUB scans can request run-to-run reproducibility on the same GPU with the same input and build configuration. - Batched Top-K selection
A new API selects the smallest or largest K items independently across many segments. - Problem sizes can stay on the GPU
Reductions can consume a count produced by an earlier kernel, avoiding a round trip to the host. - Faster sorted-value searches
Lower- and upper-bound searches for sorted queries use a merge-path approach with O(N + M) total work.
Why it matters
These changes target real GPU programming friction: tuning device-wide algorithms, moving intermediate values back to the host, and getting repeatable scan results. CCCL 3.5 also deprecates older CUB dispatch and agent-policy types, directing custom tuning towards the new public policy-selector approach. Existing single-call and two-phase temporary-storage APIs remain supported, according to NVIDIA.
Our read
This is a substantial developer release, not a shiny new product announcement. The appeal is control: more tuning options are public, and some workloads can keep more of their coordination on the GPU. If your code customises CUB dispatch today, check the migration notes before the old path becomes an awkward house guest.
What to watch
- How the public tuning policies perform across different workloads and GPU generations.
- Whether keeping problem sizes on the GPU simplifies CUDA Graph replays in practice.
- Which older customisation paths need migration as deprecated types are phased out.
Discussion spark: For CUDA libraries, should developers favour repeatable results and public tuning controls even when a workload might run faster with more specialised, less predictable settings?
Sources and evidence
- v3.5.0 (7 October 2026, 16:11 UTC)
Independent WittyWires tracker for public updates about NVIDIA. Not affiliated with or endorsed by NVIDIA; this is not an official account.