Nepali Voice Engine v0.4

A self-contained Transformers remote-code package for the v4 checkpoint, with Aakriti, Sangeeta, Prabal and Baje persona presets. The release contains the speech decoder, both tokenizers and embedded text-encoder/audio-codec configurations. T5 and DAC run through standard Transformers; no separate TTS framework or codec package is imported or downloaded.

Install and use

pip install torch 'transformers==4.46.1'

This version is pinned because the decoder uses Transformers' internal generation and cache interfaces. Transformers 5.x is not supported by this release. PyTorch and Transformers' normal transitive dependencies are required; “self-contained” means no additional model/framework packages or external model repositories.

from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained(
    "ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True
)
# Optional on a CUDA machine:
# model = model.to("cuda")
audio = model.generate_speech(prompt="नमस्ते सबैलाई!", persona="aakriti")
sample_rate = model.sampling_rate  # 44100

audio is a one-dimensional CPU float32 PyTorch tensor. Personas are aakriti, sangeeta, prabal, and baje (case insensitive). Generation overrides are accepted, for example max_new_tokens=1200. The default 800-token budget can truncate long speech. Presets condition the model with text descriptions; they are not hard speaker-ID guarantees. No postprocessing filter changes the waveform.

AutoTokenizer.from_pretrained(...) loads the Nepali prompt tokenizer; the engine loads the bundled description_tokenizer/ automatically from the same revision. Once the complete release is downloaded, local loading works offline:

model = AutoModel.from_pretrained(
    "./complete-release", trust_remote_code=True, local_files_only=True
)

Release contents and verification

config.json is the active Hub config; sanitized_config.json is an identical review copy. Both use NepaliVoiceConfig, model type nepali_voice, and the NepaliVoiceEngine architecture. The weights retain their original tensor names. The two Python modules load submodels from embedded configs, without repository lookups for T5 or DAC.

This directory is an overlay for the existing Hub repository. Its existing model.safetensors must remain present in the published repository. To turn the overlay into a complete local checkpoint, stage the original weights:

python tools/stage_weights.py /path/to/v4-snapshot
python tools/validate_release.py --load-weights --synthesize

To check the configuration, tokenizers, import boundaries, and all original weight shapes without allocating the full checkpoint:

python tools/validate_release.py
python tools/test_tiny_inference.py

The tiny test exercises autoregressive decoding and DAC synthesis with random weights, plus save/reload in offline mode. It does not measure trained voice quality. See PACKAGING_PLAN.md and VALIDATION.md for scope and validation status.

The custom decoder is an Apache-2.0 licensed derivative of Hugging Face's Parler-TTS implementation. The trained checkpoint derives from Indic Parler-TTS. Runtime independence does not erase provenance; see NOTICE and LICENSE.

Downloads last month
744
Safetensors
Model size
0.9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support