Primbench API#
The core types for defining and running GPU benchmarks with Primbench include the state object passed into run(), the JSON builder, and the user-facing macros and utility functions. For command-line option details, see Command-line options. For correctness validation workflows, see Validate benchmark output.
Command-line values override any programmatic settings values passed to the executor constructor.
Flags#
Enumeration and wrapper for combining benchmark flags.
Settings#
All tunable parameters for benchmark execution. Each field has a default that can be overridden programmatically or from the command line.
-
struct settings#
Settings that benchmarks and users can pass.
Public Members
-
bool hot = false#
Hot means not clearing GPU cache between batches.
-
uint32_t seed = 42#
The seed to use for input array generation.
-
std::string json_out = "results.json"#
Output JSON file path.
-
std::string csv_out = ""#
Output CSV file path.
-
std::string filter = ""#
Regex filter of specialization names to benchmark.
-
bool dry = false#
Flag to perform a dry run.
-
double min_gpu_ms_per_batch = 10.0#
Minimum GPU batch duration.
-
double min_secs = 1.0#
Minimum benchmark duration.
-
double noise_timeout_secs = 10.0#
Max duration before noisy benchmark times out.
-
size_t batch_window_size = 10#
Noise window size for early stopping.
-
double noise_tolerance_percent = 1.0#
Noise tolerance for early stopping.
-
uint16_t min_gpu_temp = 50#
Minimum GPU temperature.
-
uint16_t max_gpu_temp = 60#
Maximum GPU temperature.
-
double max_warming_secs = 60.0#
Max GPU warmup time.
-
double max_cooling_secs = 60.0#
Max GPU cooldown time.
-
bool output_batches = false#
Flag to output batch details.
-
uint32_t spaces_per_indent = 4#
JSON indentation spaces.
-
double stream_blocking_timeout_secs = 10.0#
Max duration before stream blocking times out.
-
std::map<std::string, custom_arg_value> custom_args#
Custom user-registered arguments with types.
-
bool hot = false#
Benchmark interface#
Abstract base class that users subclass to define a benchmark specialization. Implement meta() to describe the specialization and run() to execute it.
-
struct benchmark_interface#
Interface for all benchmark specializations.
A benchmark implementation describes:
the algorithm (
meta()["algo"]),a JSON-formatted specialization identifier (other keys in
meta()),and the code that performs the timed measurement (
run()).
The executor uses this interface to:
validate and sort benchmarks,
construct per-benchmark state objects,
run kernels and collect performance data,
and emit structured JSON results.
Public Functions
-
virtual json meta() const = 0#
Returns a JSON object describing the benchmark.
The returned JSON must include:
”algo”: canonical algorithm name,
other keys describing the specialization.
All benchmarks queued for one executor run must have the same “algo”.
-
virtual void run(state &state) = 0#
Executes the benchmark using the provided state.
Implementations allocate input/output data, perform any required setup, and launch the algorithm under test.
-
virtual ~benchmark_interface() = default#
Virtual destructor for polymorphic cleanup.
Executor#
Manages the benchmark lifecycle: parses command-line arguments, queues benchmark specializations, and runs them.
-
class executor#
Executes a suite of GPU benchmarks with configurable parameters.
The executor class handles command-line parsing, benchmark queueing, execution, and logging of results in JSON format. Supports tuning GPU and benchmark parameters, including batch sizes, durations, and temperature limits.
Public Functions
-
inline executor(int argc, char *argv[], primbench::settings settings = {}, detail::flags::FlagTag flags = flags::none, stream_t stream = default_stream)#
Constructs the executor, and runs setup code.
-
inline void run()#
Prepares and runs all queued benchmark specializations.
-
template<typename T>
inline T get(std::string_view name, const T &default_val, std::string_view description)# Parses a command-line argument.
-
inline double get_last_bytes_per_second()#
Returns the bytes per second of the last ran benchmark.
Public Static Functions
-
template<typename Benchmark, typename ...Args>
static inline bool queue(Args&&... args)# Queue a benchmark for execution.
- Template Parameters:
Benchmark – Type of benchmark to queue.
Args – Argument types for benchmark constructor.
- Parameters:
args – Arguments to forward to the benchmark constructor.
- Returns:
true, which allows the function to be called in global scope.
-
template<typename BulkCreateFunction>
static inline bool queue_autotune(BulkCreateFunction &&fn)# Queue benchmarks using an autotune bulk creation function.
- Template Parameters:
BulkCreateFunction – Callable that populates specializations.
- Parameters:
fn – Function that creates benchmarks.
- Returns:
true, which allows the function to be called in global scope.
-
inline executor(int argc, char *argv[], primbench::settings settings = {}, detail::flags::FlagTag flags = flags::none, stream_t stream = default_stream)#
Benchmark state#
The state object serves as the primary interface for declaring throughput metrics, registering the kernel lambda, setting up per-iteration callbacks, and running correctness tests. It is passed to benchmark_interface::run(), and exposes the GPU stream and the current input size.
Public fields#
-
class state#
Manages benchmark execution, GPU warm-up/cool-down, timing, and logging.
Public Functions
-
inline void set_items(size_t items)#
Sets the total number of items processed per iteration.
This must be called exactly once before calling run() or any memory tracking methods such as add_reads() or add_writes().
-
template<typename T>
inline void add_reads(size_t items)# Adds an estimate of global memory reads performed by the benchmark.
Must be called after set_items() and before any call to add_writes(). Multiple calls accumulate total read bytes.
The total number of bytes read (from all calls to this function) is summed together with the total bytes written (added via add_writes()) to compute the reported memory throughput.
-
template<typename T>
inline void add_writes(size_t items)# Adds an estimate of global memory writes performed by the benchmark.
Must be called after set_items(). Multiple calls accumulate total written bytes.
The total number of bytes written (from all calls to this function) is summed together with the total bytes read (added via add_reads()) to compute the reported memory throughput.
-
inline void run_before_every_iteration(std::function<void()> lambda)#
Sets a callback to run before each iteration, which should be used to reset the input data of in-place algorithms.
-
inline void run(std::function<void()> kernel)#
Executes the benchmark loop for the provided kernel.
Handles warm-up, timing, CV-based stopping, and logging.
The benchmark manages all required stream synchronization internally to ensure accurate timing and prevent command queue buildup. Users should not perform any manual synchronization before or during the benchmark run.
-
inline void test(std::function<void()> test_lambda)#
Registers test_lambda, which runs once during warmup.
Must not be called after run().
Define
PRIMBENCH_NO_TESTto disable.
-
inline double get_last_bytes_per_second()#
Returns the bytes per second of the last ran benchmark.
-
inline void set_items(size_t items)#
Throughput declarations#
These methods declare how many logical items, read bytes, and written bytes each kernel invocation processes. The executor uses these values to compute throughput metrics in the output.
-
inline void primbench::detail::state::set_items(size_t items)
Sets the total number of items processed per iteration.
This must be called exactly once before calling run() or any memory tracking methods such as add_reads() or add_writes().
-
template<typename T>
inline void primbench::detail::state::add_reads(size_t items) Adds an estimate of global memory reads performed by the benchmark.
Must be called after set_items() and before any call to add_writes(). Multiple calls accumulate total read bytes.
The total number of bytes read (from all calls to this function) is summed together with the total bytes written (added via add_writes()) to compute the reported memory throughput.
-
template<typename T>
inline void primbench::detail::state::add_writes(size_t items) Adds an estimate of global memory writes performed by the benchmark.
Must be called after set_items(). Multiple calls accumulate total written bytes.
The total number of bytes written (from all calls to this function) is summed together with the total bytes read (added via add_reads()) to compute the reported memory throughput.
Kernel registration#
Register the kernel lambda that the executor times, and an optional callback that runs before every iteration, for example to reset output buffers.
-
inline void primbench::detail::state::run(std::function<void()> kernel)
Executes the benchmark loop for the provided kernel.
Handles warm-up, timing, CV-based stopping, and logging.
The benchmark manages all required stream synchronization internally to ensure accurate timing and prevent command queue buildup. Users should not perform any manual synchronization before or during the benchmark run.
-
inline void primbench::detail::state::run_before_every_iteration(std::function<void()> lambda)
Sets a callback to run before each iteration, which should be used to reset the input data of in-place algorithms.
Correctness testing#
Register a callable that validates kernel output. The callable runs once after the warmup batch, before timed iterations begin. Use PRIMBENCH_ASSERT inside the test callable to check results.
JSON builder#
The json struct is a lightweight builder used inside benchmark_interface::meta() to attach algorithm names, type names, and custom fields to a benchmark specialization. Calls to add() can be chained, and nested json objects are supported.
-
struct json#
Simple JSON-like container.
Stores key-value pairs where values can be nested JSON objects, strings, integers, or doubles. Provides basic serialization to JSON and to a human-readable name.
Public Functions
-
template<typename T>
inline json &add(std::string_view key, T value)# Adds a key-value pair to the JSON object.
-
inline std::string serialize() const#
Serializes the JSON object to a JSON string.
-
inline std::string serialize_name() const#
Serializes the JSON object into a human-readable name.
-
template<typename T>
Size constants#
-
constexpr size_t primbench::KiB = 1024#
Macros#
The following macros configure type names, error checking, correctness assertions, and compile-time behavior.
Type registration#
-
PRIMBENCH_REGISTER_TYPE(TYPE, NAME)#
Registers a custom type name, used by
primbench::name<T>().
Registers a human-readable display name for a C++ type so that primbench::name<T>() returns it. The macro must be invoked at namespace scope.
Error checking#
-
PRIMBENCH_CHECK(condition)#
Exits the program with an error message if the given CUDA API call returns a failure status.
Wraps a HIP API call. If the call returns a failure status, the macro prints the file, line, and error string to stderr and exits the program.
Correctness assertions#
-
PRIMBENCH_ASSERT(input, ...)#
Asserts equality between input and expected. Works for scalar arithmetic types, iterable containers, and brace-enclosed initializer lists. The expected value and an optional tolerance (tol, defaults to 0.0) are captured via VA_ARGS to handle initializer list commas. Prints file:line, an error message to stderr, and exits on mismatch.
Asserts equality between an input value and an expected value. Overloads handle scalar arithmetic types, iterable containers, and brace-enclosed initializer lists. An optional tolerance parameter defaults to 0.0 and controls the maximum allowed difference for floating-point comparisons. On mismatch the macro prints file, line, and a diagnostic message to stderr and exits.
Compile-time configuration#
-
PRIMBENCH_GPU_CACHE_SIZE#
Default GPU cache size used for clearing caches.
This conservative size is currently used to evict cached data before kernel launches. In the future, introducing HSA as a dependency may allow querying the actual largest GPU cache at runtime.
Sets the size of the buffer used to evict GPU caches before kernel launches. Override by defining the macro before including the header or passing it as a compiler flag.
PRIMBENCH_NO_MONITORINGWhen defined, disables GPU temperature monitoring. The library compiles without a monitoring dependency.
PRIMBENCH_NO_TESTWhen defined, disables correctness-test execution.
Version-control metadata#
BRANCH_NAME and COMMIT_HASH are optional compile-time macros. Pass them with -DBRANCH_NAME=... and -DCOMMIT_HASH=.... Their values are embedded in the context.general section of the JSON output so that benchmark results can be traced back to a specific source revision.
Free functions#
-
template<typename ...Args>
void primbench::log(Args&&... args)# This function is primarily used in benchmarks to display progress or setup messages (for example, “Generating matrix of size 32x64”). It accepts any number of arguments of varying types, concatenates them, and prints them as a gray line.
This logging is especially helpful for diagnosing slow setup steps.
Examples:
primbench::log("Loading dataset..."); // Output: Loading dataset... primbench::log("Generating matrix of size ", 32, "x", 64); // Output: Generating matrix of size 32x64
-
template<class T>
std::string primbench::name()# Used to retrieve the name of a type that was registered with PRIMBENCH_REGISTER_TYPE().