Benchmarking with Primbench

Contents

Benchmarking with Primbench#

Primbench is a single-header library. There are no build steps, package installations, or linking against Primbench binaries required. All that’s needed to use Primbench is to import primbench.hpp into your benchmarking code.

Use the copy_benchmark.cpp example and the information in this tutorial to help you use Primbench in your application.

Using Primbench#

PRIMBENCH_REGISTER_TYPE is used to register a name for each variable type, or specialization, that you benchmark. This name is used to identify the type in the output.

For example, in copy_benchmark.cpp the char and long long types are given the names "char" and "long long":

PRIMBENCH_REGISTER_TYPE(char, "char")
PRIMBENCH_REGISTER_TYPE(long long, "long long")

Registering also lets you provide alternate names for your types. For example, you could register long long as "longx2":

PRIMBENCH_REGISTER_TYPE(long long, "longx2")

Both the meta() and run() functions in primbench::benchmark_interface must be implemented.

The meta() function returns metadata as a JSON object.

The returned JSON object must include a value for the algo key. The algo key sets the name of the algorithm being benchmarked. This name is used in the JSON output.

For example, from copy_benchmark.cpp:

template<typename T>
struct copy_benchmark : public primbench::benchmark_interface
{
  primbench::json meta() const override
  {
    return primbench::json{}.add("algo", "copy").add("type", primbench::name<T>());
  }
  [...]
}

Define the algorithm to benchmark. It is passed to state.run() in the implementation of primbench::benchmark_interface::run().

For example, the copy_kernel algorithm is defined in copy_benchmark.cpp:

template<typename T, unsigned int BlockSize, unsigned int ItemsPerThread>
__global__ __launch_bounds__(BlockSize)
void copy_kernel(const T* input, T* output)
{
  unsigned int idx = threadIdx.x + blockIdx.x * BlockSize * ItemsPerThread;
  #pragma unroll
  for(unsigned int i = 0; i < ItemsPerThread; ++i)
      output[idx + i * BlockSize] = input[idx + i * BlockSize];
}

The run() function runs the benchmark. run() must include a call to state.set_items(). state.set_items() sets the number of items processed per iteration. The number of items must be greater than 0.

The state class saves the state of the benchmarking run, including the number of reads and writes.

Depending on the algorithm being benchmarked, run() might call state.add_reads(), state.add_writes(), or both. These functions calculate the number of items or bytes processed per second.

set_items() must be called before add_reads() or add_writes(). If you call both, call add_reads() before add_writes(). Call state.run() after set_items() and any read or write counters you need.

For example, from copy_benchmark.cpp:

state.set_items(items);
state.add_reads<T>(items);
state.add_writes<T>(items);

The kernel call is wrapped in a lambda and passed to state.run(). state.run() runs the kernel as many times as required.

state.run(
        [&] {
            copy_kernel<T, BlockSize, ItemsPerThread>
                <<<grid, block, 0, stream>>>(d_input, d_output);
        });

Benchmark settings and flags can be passed to the executor class constructor as optional parameters. The executor queues and runs the benchmarks.

For more information on settings, see Configure benchmark settings.

executor.queue() is called to queue the benchmark for each specialization. When executor.run() is called, the queued benchmark specializations run in alphabetical order.

For example, from copy_benchmark.cpp:

int main(int argc, char* argv[])
{
  primbench::executor executor(argc, argv);

  executor.queue<copy_benchmark<char>>();
  executor.queue<copy_benchmark<long long>>();

  executor.run();
}

Compile the benchmark using hipcc. For example, on Linux:

hipcc -o copy_benchmark copy_benchmark.cpp -lamd_smi
./copy_benchmark

On Windows:

hipcc -o copy_benchmark.exe examples/hip/copy_benchmark.cpp -I. -DPRIMBENCH_NO_MONITORING -std=c++17 -g --offload-arch=$(amdgpu-arch)
./copy_benchmark.exe

For the complete list of command-line options, see Command-line options.

The output is written to the terminal and to results.json.