Instructions to use ujjwal5454/nepali-voice-engine-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ujjwal5454/nepali-voice-engine-v4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Nepali Voice Engine v0.4
A self-contained Transformers remote-code package for the v4 checkpoint, with Aakriti, Sangeeta, Prabal and Baje persona presets. The release contains the speech decoder, both tokenizers and embedded text-encoder/audio-codec configurations. T5 and DAC run through standard Transformers; no separate TTS framework or codec package is imported or downloaded.
Install and use
pip install torch 'transformers==4.46.1'
This version is pinned because the decoder uses Transformers' internal generation and cache interfaces. Transformers 5.x is not supported by this release. PyTorch and Transformers' normal transitive dependencies are required; “self-contained” means no additional model/framework packages or external model repositories.
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained(
"ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True
)
# Optional on a CUDA machine:
# model = model.to("cuda")
audio = model.generate_speech(prompt="नमस्ते सबैलाई!", persona="aakriti")
sample_rate = model.sampling_rate # 44100
audio is a one-dimensional CPU float32 PyTorch tensor. Personas are aakriti,
sangeeta, prabal, and baje (case insensitive). Generation overrides are
accepted, for example max_new_tokens=1200. The default 800-token budget can
truncate long speech. Presets condition the model with text descriptions; they
are not hard speaker-ID guarantees. No postprocessing filter changes the waveform.
AutoTokenizer.from_pretrained(...) loads the Nepali prompt tokenizer; the engine
loads the bundled description_tokenizer/ automatically from the same revision.
Once the complete release is downloaded, local loading works offline:
model = AutoModel.from_pretrained(
"./complete-release", trust_remote_code=True, local_files_only=True
)
Release contents and verification
config.json is the active Hub config; sanitized_config.json is an identical
review copy. Both use NepaliVoiceConfig, model type nepali_voice, and the
NepaliVoiceEngine architecture. The weights retain their original tensor names.
The two Python modules load submodels from embedded configs, without repository
lookups for T5 or DAC.
This directory is an overlay for the existing Hub repository. Its existing
model.safetensors must remain present in the published repository. To turn the
overlay into a complete local checkpoint, stage the original weights:
python tools/stage_weights.py /path/to/v4-snapshot
python tools/validate_release.py --load-weights --synthesize
To check the configuration, tokenizers, import boundaries, and all original weight shapes without allocating the full checkpoint:
python tools/validate_release.py
python tools/test_tiny_inference.py
The tiny test exercises autoregressive decoding and DAC synthesis with random
weights, plus save/reload in offline mode. It does not measure trained voice quality.
See PACKAGING_PLAN.md and VALIDATION.md for scope and validation status.
The custom decoder is an Apache-2.0 licensed derivative of Hugging Face's
Parler-TTS implementation.
The trained checkpoint derives from
Indic Parler-TTS.
Runtime independence does not erase provenance; see NOTICE and LICENSE.
- Downloads last month
- 744