C++ Standard Library extensions in libhipcxx#
2026-10-02
4 min read time
libhipcxx provides HIP C++ developers with familiar Standard Library utilities to improve productivity and flatten the learning curve of HIP. However, there are many aspects of writing high-performance HIP C++ code that cannot be expressed through purely Standard conforming APIs. For these cases, libhipcxx also provides extensions of Standard Library utilities.
To use utilities that are extensions to Standard Library features, drop the std from both the
include path and namespace:
#include <cuda/atomic>
cuda::atomic<int, cuda::thread_scope_device> x;
Extensions that only run on the GPU use the cuda::device:: namespace:
#include <cuda/warp>
// Inside a __device__ function:
int val = cuda::device::warp_shuffle_idx(data, src_lane);
For a full explanation of the namespace hierarchy, see HIP-specific abstractions and namespaces.
Available extensions#
libhipcxx provides extensions across synchronization, memory, math, and device-specific functionality. The following table lists each extension category with its key APIs:
Extension |
APIs |
Notes |
|---|---|---|
Thread-scope synchronization |
|
|
Asynchronous operations |
|
Overlaps compute and memory transfers |
Functional utilities |
|
|
Math utilities |
|
|
Bit utilities |
|
|
Stream reference |
|
Type-safe wrapper around |
Memory resources |
|
Experimental. Requires |
Warp intrinsics |
|
Device-only |
Work stealing |
|
Device-only. Dynamic block-level parallelism. |
For per-API documentation, see the Extended API reference.
Unsupported extensions#
Several extensions from the upstream libcudacxx project are not supported in libhipcxx because they depend on NVIDIA hardware. These include:
<cuda/latch>,<cuda/barrier>,<cuda/semaphore>,<cuda/pipeline>— scoped synchronization primitives<cuda/annotated_ptr>— memory access properties for pointers<cuda/ptx>— NVIDIA PTX instruction wrappers
For the complete list, see Limitations and unsupported APIs.