nvidia/bigvgan_v2_44khz_128band_512x

Name: nvidia/bigvgan_v2_44khz_128band_512x
Rating: 5 (67 reviews)
Author: nvidia

audio to audioPyTorchPyTorchneural-vocoderaudio-generationaudio-to-audioarxiv:2206.04658license:mitmit

67

HuggingFace

765.6K

instantiate the model. You can optionally set use_cuda_kernel=True for faster inference.

model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_44khz_128band_512x', use_cuda_kernel=False)

remove weight norm in the model and set to eval mode

model.remove_weight_norm() model = model.eval().to(device)

load wav file and compute mel spectrogram

wav_path = '/path/to/your/audio.wav' wav, sr = librosa.load(wav_path, sr=model.h.sampling_rate, mono=True) # wav is np.ndarray with shape [T_time] and values in [-1, 1] wav = torch.FloatTensor(wav).unsqueeze(0) # wav is FloatTensor with shape [B(1), T_time]

compute mel spectrogram from the ground truth audio

mel = get_mel_spectrogram(wav, model.h).to(device) # mel is FloatTensor with shape [B(1), C_mel, T_frame]

generate waveform from mel

with torch.inference_mode(): wav_gen = model(mel) # wav_gen is FloatTensor with shape [B(1), 1, T_time] and values in [-1, 1] wav_gen_float = wav_gen.squeeze(0).cpu() # wav_gen is FloatTensor with shape [1, T_time]

you can convert the generated waveform to 16 bit linear PCM

wav_gen_int16 = (wav_gen_float * 32767.0).numpy().astype('int16') # wav_gen is now np.ndarray with shape [1, T_time] and int16 dtype


## Using Custom CUDA Kernel for Synthesis
You can apply the fast CUDA inference kernel by using a parameter `use_cuda_kernel` when instantiating BigVGAN:

```python
import bigvgan
model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_44khz_128band_512x', use_cuda_kernel=True)

When applied for the first time, it builds the kernel using nvcc and ninja. If the build succeeds, the kernel is saved to alias_free_activation/cuda/build and the model automatically loads the kernel. The codebase has been tested using CUDA 12.1.

Please make sure that both are installed in your system and nvcc installed in your system matches the version your PyTorch build is using.

For detail, see the official GitHub repository: https://github.com/NVIDIA/BigVGAN?tab=readme-ov-file#using-custom-cuda-kernel-for-synthesis

Pretrained Models

We provide the pretrained models on Hugging Face Collections. One can download the checkpoints of the generator weight (named bigvgan_generator.pt) and its discriminator/optimizer states (named bigvgan_discriminator_optimizer.pt) within the listed model repositories.

Model Name	Sampling Rate	Mel band	fmax	Upsampling Ratio	Params	Dataset	Steps	Fine-Tuned
bigvgan_v2_44khz_128band_512x	44 kHz	128	22050	512	122M	Large-scale Compilation	5M	No
bigvgan_v2_44khz_128band_256x	44 kHz	128	22050	256	112M	Large-scale Compilation	5M	No
bigvgan_v2_24khz_100band_256x	24 kHz	100	12000	256	112M	Large-scale Compilation	5M	No
bigvgan_v2_22khz_80band_256x	22 kHz	80	11025	256	112M	Large-scale Compilation	5M	No
bigvgan_v2_22khz_80band_fmax8k_256x	22 kHz	80	8000	256	112M	Large-scale Compilation	5M	No
bigvgan_24khz_100band	24 kHz	100	12000	256	112M	LibriTTS	5M	No
bigvgan_base_24khz_100band	24 kHz	100	12000	256	14M	LibriTTS	5M	No
bigvgan_22khz_80band	22 kHz	80	8000	256	112M	LibriTTS + VCTK + LJSpeech	5M	No
bigvgan_base_22khz_80band	22 kHz	80	8000	256	14M	LibriTTS + VCTK + LJSpeech	5M	No

Deploy Model on Runcrate

Run this model on powerful GPU infrastructure. Deploy in 60 seconds.

Pay per second

H100, A100, RTX GPUs

Instant deployment

DEPLOY IN 60 SECONDS

Run bigvgan_v2_44khz_128band_512x on Runcrate

Deploy on H100, A100, or RTX GPUs. Pay only for what you use. No setup required.