RPP audio augmentations#

RppStatus rppt_non_silent_region_detection(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, Rpp32s *srcLengthTensor, Rpp32s *detectedIndexTensor, Rpp32s *detectionLengthTensor, Rpp32f cutOffDB, Rpp32s windowLength, Rpp32f referencePower, Rpp32s resetInterval, rppHandle_t rppHandle, RppBackend executionBackend)#

Non Silent Region Detection augmentation on HIP/HOST backend.

Non Silent Region Detection augmentation for 1D audio buffer

Finds the starting index and length of non silent region in the audio buffer by comparing the calculated short-term power with cutoff value passed

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)

  • srcLengthTensor – [in] source audio buffer length (1D tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • detectedIndexTensor – [out] beginning index of non silent region (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • detectionLengthTensor – [out] length of non silent region (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • cutOffDB – [in] cutOff in dB below which the signal is considered silent

  • windowLength – [in] window length used for computing short-term power of the signal

  • referencePower – [in] reference power that is used to convert the signal to dB

  • resetInterval – [in] number of samples after which the moving mean average is recalculated to avoid precision loss

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_to_decibels(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, RpptImagePatchPtr srcDims, Rpp32f cutOffDB, Rpp32f multiplier, Rpp32f referenceMagnitude, rppHandle_t rppHandle, RppBackend executionBackend)#

To Decibels augmentation on HIP/HOST backend.

To Decibels augmentation for 1D/2D audio buffer converts magnitude values to decibel values

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel/2D audio tensor with 1 channel), offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel/2D audio tensor with 1 channel), offsetInBytes >= 0, dataType = F32)

  • srcDims – [in] source tensor sizes for each element in batch (2D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)

  • cutOffDB – [in] minimum or cut-off ratio in dB

  • multiplier – [in] factor by which the logarithm is multiplied

  • referenceMagnitude – [in] Reference magnitude if not provided maximum value of input used as reference

  • rppHandle – [in] RPP HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_pre_emphasis_filter(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, Rpp32f *coeffTensor, RpptAudioBorderType borderType, rppHandle_t rppHandle, RppBackend executionBackend)#

Pre Emphasis Filter augmentation on HIP/HOST backend.

Pre Emphasis Filter augmentation for audio data

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32)

  • srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • coeffTensor – [in] preemphasis coefficient (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • borderType – [in] border value policy

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_down_mixing(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcDimsTensor, bool normalizeWeights, rppHandle_t rppHandle, RppBackend executionBackend)#

Down Mixing augmentation on HIP/HOST backend.

Down Mixing augmentation for audio data

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)

  • srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)

  • normalizeWeights – [in] bool flag to specify if normalization of weights is needed

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_spectrogram(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, bool centerWindows, bool reflectPadding, Rpp32f *windowFunction, Rpp32s nfft, Rpp32s power, Rpp32s windowLength, Rpp32s windowStep, rppHandle_t rppHandle, RppBackend executionBackend)#

Produces a spectrogram from a 1D audio buffer on HIP/HOST backend.

Spectrogram for 1D audio buffer

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT / NTF)

  • srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • centerWindows – [in] indicates whether extracted windows should be padded so that the window function is centered at multiples of window_step

  • reflectPadding – [in] indicates the padding policy when sampling outside the bounds of the signal

  • windowFunction – [in]

    samples of the window function that will be multiplied to each extracted window when calculating the Short Time Fourier Transform (STFT).

    if windowFunction is a nullptr, then required windowFunction values will be generated inside the kernel

  • nfft – [in] size of the FFT

  • power – [in] exponent of the magnitude of the spectrum

  • windowLength – [in] window size in number of samples

  • windowStep – [in] step between the STFT windows in number of samples

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_mel_filter_bank(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcDims, Rpp32f maxFreq, Rpp32f minFreq, RpptMelScaleFormula melFormula, Rpp32s numFilter, Rpp32f sampleRate, bool normalize, rppHandle_t rppHandle, RppBackend executionBackend)#

Mel filter bank augmentation HIP/HOST backend.

Mel filter bank augmentation for audio data

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT)

  • srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)

  • maxFreq – [in] maximum frequency if not provided maxFreq = sampleRate / 2

  • minFreq – [in] minimum frequency

  • melFormula – [in] formula used to convert frequencies from hertz to mel and from mel to hertz (SLANEY / HTK)

  • numFilter – [in] number of mel filters

  • sampleRate – [in] sampling rate of the audio

  • normalize – [in] boolean variable that determine whether to normalize weights / not

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_resample(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32f *inRateTensor, Rpp32f *outRateTensor, Rpp32s *srcDimsTensor, RpptResamplingWindow &window, rppHandle_t rppHandle, RppBackend executionBackend)#

Resample augmentation on HIP/HOST backend.

Resample augmentation for audio data

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • inRate – [in] Input sampling rate (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • outRate – [in] Output sampling rate (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)

  • window – [in] Resampling window (struct of type RpptRpptResamplingWindow)

  • rppHandle – [in] RPP HOST handle created with rppCreate()

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_audio_tensor_add_tensor(RppPtr_t srcPtr1, RppPtr_t srcPtr2, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, rppHandle_t rppHandle, RppBackend executionBackend)#

Audio Tensor Add Tensor augmentation on HIP/HOST backend.

Audio Tensor Add Tensor augmentation for audio data. Performs scalar broadcasting addition where srcPtr2 contains one scalar value per batch (shape: batchSize) that is broadcasted and added to all elements in the corresponding batch of srcPtr1.

Parameters:
  • srcPtr1 – [in] first source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • srcPtr2 – [in] second source tensor containing one scalar per batch in HIP memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

  • executionBackend – [in] RPP execution backend (HOST or HIP)

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.

RppStatus rppt_audio_tensor_mul_scalar(RppPtr_t srcPtr, Rpp32f scalarValue, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, rppHandle_t rppHandle, RppBackend executionBackend)#

Audio Tensor Multiply Scalar augmentation on HIP/HOST backend.

Audio Tensor Multiply Scalar augmentation for audio data. Multiplies each element in srcPtr by a single scalar value. The scalar value is broadcasted and multiplied to all elements across the entire tensor.

Parameters:
  • srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • scalarValue – [in] scalar value to multiply with all elements in the tensor

  • srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)

  • dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)

  • srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)

  • rppHandle – [in] RPP HIP/HOST handle created with rppCreate()

  • executionBackend – [in] RPP execution backend (HOST or HIP)

Return values:
  • RPP_SUCCESS – Successful completion.

  • RPP_ERROR* – Unsuccessful completion.

Returns:

A RppStatus enumeration.