RPP audio augmentations#
-
RppStatus rppt_non_silent_region_detection(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, Rpp32s *srcLengthTensor, Rpp32s *detectedIndexTensor, Rpp32s *detectionLengthTensor, Rpp32f cutOffDB, Rpp32s windowLength, Rpp32f referencePower, Rpp32s resetInterval, rppHandle_t rppHandle, RppBackend executionBackend)#
Non Silent Region Detection augmentation on HIP/HOST backend.
Non Silent Region Detection augmentation for 1D audio buffer
Finds the starting index and length of non silent region in the audio buffer by comparing the calculated short-term power with cutoff value passed
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)
srcLengthTensor – [in] source audio buffer length (1D tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
detectedIndexTensor – [out] beginning index of non silent region (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
detectionLengthTensor – [out] length of non silent region (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
cutOffDB – [in] cutOff in dB below which the signal is considered silent
windowLength – [in] window length used for computing short-term power of the signal
referencePower – [in] reference power that is used to convert the signal to dB
resetInterval – [in] number of samples after which the moving mean average is recalculated to avoid precision loss
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_to_decibels(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, RpptImagePatchPtr srcDims, Rpp32f cutOffDB, Rpp32f multiplier, Rpp32f referenceMagnitude, rppHandle_t rppHandle, RppBackend executionBackend)#
To Decibels augmentation on HIP/HOST backend.
To Decibels augmentation for 1D/2D audio buffer converts magnitude values to decibel values
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel/2D audio tensor with 1 channel), offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel/2D audio tensor with 1 channel), offsetInBytes >= 0, dataType = F32)
srcDims – [in] source tensor sizes for each element in batch (2D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)
cutOffDB – [in] minimum or cut-off ratio in dB
multiplier – [in] factor by which the logarithm is multiplied
referenceMagnitude – [in] Reference magnitude if not provided maximum value of input used as reference
rppHandle – [in] RPP HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_pre_emphasis_filter(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, Rpp32f *coeffTensor, RpptAudioBorderType borderType, rppHandle_t rppHandle, RppBackend executionBackend)#
Pre Emphasis Filter augmentation on HIP/HOST backend.
Pre Emphasis Filter augmentation for audio data
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32)
srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
coeffTensor – [in] preemphasis coefficient (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
borderType – [in] border value policy
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_down_mixing(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcDimsTensor, bool normalizeWeights, rppHandle_t rppHandle, RppBackend executionBackend)#
Down Mixing augmentation on HIP/HOST backend.
Down Mixing augmentation for audio data
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)
srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)
normalizeWeights – [in] bool flag to specify if normalization of weights is needed
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_spectrogram(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, bool centerWindows, bool reflectPadding, Rpp32f *windowFunction, Rpp32s nfft, Rpp32s power, Rpp32s windowLength, Rpp32s windowStep, rppHandle_t rppHandle, RppBackend executionBackend)#
Produces a spectrogram from a 1D audio buffer on HIP/HOST backend.
Spectrogram for 1D audio buffer
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2, offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT / NTF)
srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
centerWindows – [in] indicates whether extracted windows should be padded so that the window function is centered at multiples of window_step
reflectPadding – [in] indicates the padding policy when sampling outside the bounds of the signal
windowFunction – [in]
samples of the window function that will be multiplied to each extracted window when calculating the Short Time Fourier Transform (STFT).
if windowFunction is a nullptr, then required windowFunction values will be generated inside the kernel
nfft – [in] size of the FFT
power – [in] exponent of the magnitude of the spectrum
windowLength – [in] window size in number of samples
windowStep – [in] step between the STFT windows in number of samples
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_mel_filter_bank(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcDims, Rpp32f maxFreq, Rpp32f minFreq, RpptMelScaleFormula melFormula, Rpp32s numFilter, Rpp32f sampleRate, bool normalize, rppHandle_t rppHandle, RppBackend executionBackend)#
Mel filter bank augmentation HIP/HOST backend.
Mel filter bank augmentation for audio data
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 3, offsetInBytes >= 0, dataType = F32, layout - NFT)
srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)
maxFreq – [in] maximum frequency if not provided maxFreq = sampleRate / 2
minFreq – [in] minimum frequency
melFormula – [in] formula used to convert frequencies from hertz to mel and from mel to hertz (SLANEY / HTK)
numFilter – [in] number of mel filters
sampleRate – [in] sampling rate of the audio
normalize – [in] boolean variable that determine whether to normalize weights / not
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_resample(RppPtr_t srcPtr, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32f *inRateTensor, Rpp32f *outRateTensor, Rpp32s *srcDimsTensor, RpptResamplingWindow &window, rppHandle_t rppHandle, RppBackend executionBackend)#
Resample augmentation on HIP/HOST backend.
Resample augmentation for audio data
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
inRate – [in] Input sampling rate (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
outRate – [in] Output sampling rate (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
srcDimsTensor – [in] source audio buffer length and number of channels (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize * 2)
window – [in] Resampling window (struct of type RpptRpptResamplingWindow)
rppHandle – [in] RPP HOST handle created with
rppCreate()
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_audio_tensor_add_tensor(RppPtr_t srcPtr1, RppPtr_t srcPtr2, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, rppHandle_t rppHandle, RppBackend executionBackend)#
Audio Tensor Add Tensor augmentation on HIP/HOST backend.
Audio Tensor Add Tensor augmentation for audio data. Performs scalar broadcasting addition where srcPtr2 contains one scalar value per batch (shape: batchSize) that is broadcasted and added to all elements in the corresponding batch of srcPtr1.
- Parameters:
srcPtr1 – [in] first source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
srcPtr2 – [in] second source tensor containing one scalar per batch in HIP memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()executionBackend – [in] RPP execution backend (HOST or HIP)
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.
-
RppStatus rppt_audio_tensor_mul_scalar(RppPtr_t srcPtr, Rpp32f scalarValue, RpptDescPtr srcDescPtr, RppPtr_t dstPtr, RpptDescPtr dstDescPtr, Rpp32s *srcLengthTensor, rppHandle_t rppHandle, RppBackend executionBackend)#
Audio Tensor Multiply Scalar augmentation on HIP/HOST backend.
Audio Tensor Multiply Scalar augmentation for audio data. Multiplies each element in srcPtr by a single scalar value. The scalar value is broadcasted and multiplied to all elements across the entire tensor.
- Parameters:
srcPtr – [in] source tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
scalarValue – [in] scalar value to multiply with all elements in the tensor
srcDescPtr – [in] source tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
dstPtr – [out] destination tensor in HIP memory (for HIP backend) or HOST memory (for HOST backend)
dstDescPtr – [in] destination tensor descriptor (Restrictions - numDims = 2 or 3 (for single-channel or multi-channel audio tensor), offsetInBytes >= 0, dataType = F32)
srcLengthTensor – [in] source audio buffer length (1D tensor in pinned memory (for HIP backend) or HOST memory (for HOST backend), of size batchSize)
rppHandle – [in] RPP HIP/HOST handle created with
rppCreate()executionBackend – [in] RPP execution backend (HOST or HIP)
- Return values:
RPP_SUCCESS – Successful completion.
RPP_ERROR* – Unsuccessful completion.
- Returns:
A
RppStatusenumeration.