QTSTRM — Quiet Storm R&B LoRAs for YuE2

Style LoRAs that push YuE2-3B into 1980s–1990s quiet storm: slow, late-night R&B and sophisti-soul with a sweet, sultry female lead. Deep pulsing bass, electric piano, synth pads, slow drum machines, and on the sophisti-pop side a tenor saxophone, fretless bass and Latin percussion. English lyrics, female lead.

The training set mixes several female voices, and the captions describe each one with the same small vocabulary, so the prompt can steer the voice: a low smoky contralto, a light breathy soprano, a deep husky contralto or a warm soulful mezzo, or a blend (see Steering the voice).

Each file patches both halves of YuE2: the autoregressive planner (writes the score and the vocal lines) and the flow-matching decoder (the sound). Trigger word: qtstrm. All five were trained with the experimental YuE2 trainer in Ostris AI Toolkit, in two generations on two datasets (v1 and v2).

Last updated: 28 September 2026.

LoRA File Generation, checkpoint Character
🕯️ QTSTRM Candlelight qtstrm_candlelight.safetensors v1, step 500 Lightest touch; the most varied voices
🌙 QTSTRM Nocturne qtstrm_nocturne.safetensors v1, step 700 Deep smoky slow jams
🥃 QTSTRM Afterhours qtstrm_afterhours.safetensors v1, step 800 Strongest voice character; can grit on long high notes
🎷 QTSTRM Velvet qtstrm_velvet.safetensors v2, step 700 Larger dataset, cleaner; smooth sax-blend sound
🌌 QTSTRM Midnight qtstrm_midnight.safetensors v2, step 800 v2 at full strength; the most polished

Start with clip 1.0 / model 1.0 and score_mode full.

Listen

32 steps dpm_2 / sgm_uniform, cfg 1.0, no post-processing. clip is strength_clip (the planner), model is strength_model (the decoder). Every lyric here is an original: "I Stayed" was written by becausereasons; the others were written for these demos.

Afterhours, "Blue Light": prompts/blend_sax.txt + lyrics/blue_light.txt, seed 7, clip 1.0 / model 1.0, cap 300 s. The sax blend on the peak v1 file: smoky contralto, tenor sax, fretless bass, congas (4:05):

Midnight, "I Stayed": prompts/blend_sax.txt + lyrics/i_stayed.txt, seed 1832754214, clip 1.0 / model 1.0, cap 360 s. The sax blend on the v2 step-800 file, with a lyric written by becausereasons (4:58):

Velvet, "Blue Light": prompts/blend_sax.txt + lyrics/blue_light.txt, seed 7, clip 1.0 / model 1.0, cap 260 s. Same prompt, lyric and seed as the Afterhours demo, on the v2 step-700 file (3:54):

Afterhours, "I Stayed": prompts/blend_sax.txt + lyrics/i_stayed.txt, seed 1148926555, clip 1.0 / model 1.0, cap 360 s. The same lyric on the v1 peak file. This take sings an earlier draft of the chorus ("Slowly disappearing here") (4:30):

Nocturne, "Nothing but Static": prompts/smoky_slowjam.txt + lyrics/nothing_but_static.txt, seed 1691473935, clip 1.0 / model 1.0, cap 260 s. The smoky slow-jam prompt: pulsing bass, electric piano, synth pads, slow drums (4:20):

Afterhours, "Blue Light" (another seed): prompts/blend_sax.txt + lyrics/blue_light.txt, seed 353841619, clip 1.0 / model 1.0, cap 240 s. Same file, prompt and lyric as the first demo, different seed (4:00):

Candlelight, "Velvet Hour" (sax blend): prompts/blend_sax.txt + lyrics/velvet_hour.txt, seed 7, clip 1.0 / model 1.0, cap 300 s. The early v1 file with the reference lyric (4:17):

Candlelight, "Velvet Hour" (breathy soprano): prompts/breathy_slowjam.txt + lyrics/velvet_hour.txt, seed 7, clip 1.0 / model 1.0, cap 300 s. The breathy-soprano voice prompt: whispery delivery, airy stacked harmonies, electric piano, synth bass (4:22):

Quick start (ComfyUI)

The LoRAs use Comfy's fused YuE2 key layout (text_encoders.* for the planner, diffusion_model.* for the decoder). They load through the FS_Audio Suite LoRA loader and through Comfy's standard LoRA loader.

  1. ComfyUI ≥ v0.36.0 (native YuE2 support) and the FS_Audio Suite custom node pack.
  2. Base model: yue2_3b_bf16.safetensors from Comfy-Org/YuE2 in models/checkpoints/.
  3. Drop one qtstrm_*.safetensors into models/loras/.
  4. Chain the nodes:
🧩 FS_Audio Lora Loader  ──loras──▶  🎤 FS_Audio Model Loader  ──pipe──▶  🎵 FS_Audio Sampler  ──▶  💿 FS_Audio Output
   lora_name      = qtstrm_midnight.safetensors     yue2_checkpoint = yue2_3b_bf16       style   = <prompt, starts with "qtstrm,">
   strength_clip  = 1.0   (planner)                 melody_transcriber = none            lyrics  = <tagged lyric blocks>
   strength_model = 1.0   (decoder)                                                      score_mode = full
widget value
steps / sampler / scheduler 32 / dpm_2 / sgm_uniform
score_mode full
song_length_cap 260–360 (a ~240-word lyric plans about four minutes)
Weirdness (cfg) / score_temperature / music_temperature 1.0 / 0.7 / 1.0
repetition_penalty 1.2

Prompting

Style prompt

Start with the trigger, then one descriptive sentence in the order the training captions used: language, era and genre (always including "quiet storm"), the voice, instruments, mood, production, BPM. The four prompts in prompts/ are good starting points:

  • smoky_slowjam: 1990s slow jam, smoky contralto, pulsing bass, electric piano, pads, slow drums, 70 BPM.
  • blend_sax: 1980s sophisti-pop, smoky contralto with airy stacked harmonies, tenor sax, fretless bass, congas, 96 BPM.
  • breathy_slowjam: breathy soprano, whispery delivery, electric piano, synth bass, 66 BPM.
  • husky_ballad: husky contralto with melismatic runs, Rhodes, slow drum machine, 64 BPM.

Steering the voice

Every training caption described the lead with words from one fixed vocabulary, so these words carry weight:

axis words
register low contralto, deep contralto, warm mezzo, light soprano
texture smoky, breathy, husky, silky
character sweet, sultry, soulful, intimate
delivery cool restrained, minimal vibrato, whispery, melismatic runs, strong full-voiced, airy stacked harmonies

The four voice types in the data, as the captions wrote them:

  • Smoky contralto: low smoky contralto female vocal, sweet and sultry, cool restrained delivery with minimal vibrato
  • Breathy soprano: light breathy soprano female vocal, sweet and intimate, soft whispery delivery with airy stacked harmonies
  • Husky contralto: deep husky contralto female vocal, sultry and silky, smooth delivery with melismatic runs
  • Warm mezzo: warm husky mezzo female vocal, soulful and sultry, strong full-voiced delivery

Words shared between voices (contralto, husky, sweet, sultry) are the handles for blending: "low smoky contralto with airy stacked harmonies" mixes the first two. The smoky contralto has the most training data and is the most reliable.

Lyrics

Plain words, three real verses, a chorus whose second half changes, a bridge: about 180–270 words plans a three-and-a-half to four-and-a-half minute song. Very short lyrics (under ~150 words) make the planner pad the song and can come out garbled; very long ones plan past five minutes. Use YuE2 section tags on their own lines: [Verse 1], [Pre-Chorus], [Chorus], [Bridge], [Outro].

Which knob does what

  • File first: v2 (Velvet, Midnight) for the cleanest, most consistent results; Afterhours for the strongest voice character; Candlelight for more variety.
  • strength_model (decoder) changes the sound; strength_clip (planner) changes the writing.
  • Seeds decide the form as much as the lyric. If a render comes out garbled or ghostly, it usually wrote an over-long score (look at the .abc sidecar: good takes here have ~95–125 score lines); change the seed.

Training

Lossless recordings of 1980s–1990s quiet storm R&B, sophisti-soul and UK street soul by several female singers, as whole songs or halves cut at a pause between sung phrases. Captions were written by an audio LLM in YuE2's native descriptive style, then rewritten by hand so the voice description uses the fixed vocabulary above and every caption names the genre "quiet storm"; lyrics were transcribed the same way and checked.

Both generations: AI Toolkit YuE2 trainer, rank 32, LR 5e-5 with the planner at 0.4× (ar_lr_multiplier 0.4), planner KL 0.2, transcribed ABC scores (cot full) with 50 % score dropout.

  • v1 (Candlelight, Nocturne, Afterhours): 24 items, 800 steps (33 passes per item). The voice got strongest late, but by step 600+ the thinner voices in the set occasionally ran away and long high notes could grit.
  • v2 (Velvet, Midnight): 40 items, 850 steps (21 passes per item). More data at fewer passes: the same voice character with a planner that stayed a third closer to the base model.

Known limitations

  • English, female lead only. Every training song is a love song; very dark or aggressive lyrics are outside what it learned.
  • A score can run away (one section repeated or padded until the cap), which sounds garbled or ghostly. Change the seed.
  • Afterhours (v1 step 800) can grit on long sustained high notes; lower strength_model to ~0.8 or use a v2 file.
  • The breathy-soprano prompt is the least stable voice on the late v1 files, especially with dark lyrics.
  • Belted, raspy rock-soul power vocals are not in the data and do not work.

Files

qtstrm_candlelight.safetensors  118 MB   v1, step 500, planner + decoder LoRA (bf16)
qtstrm_nocturne.safetensors     118 MB   v1, step 700, same layout
qtstrm_afterhours.safetensors   118 MB   v1, step 800, same layout
qtstrm_velvet.safetensors       118 MB   v2, step 700, same layout
qtstrm_midnight.safetensors     118 MB   v2, step 800, same layout
demos/                          mp3 renders (192 kbps from the FLAC masters)
prompts/                        style prompts
lyrics/                         the demo lyrics

Support

These LoRAs are trained on my own GPU and released free. If they're useful to you and you'd like to chip in for compute, there's a Ko-fi: ko-fi.com/becausereasons <3

License and credits

Weights are released under CC BY-NC 4.0, inherited from the YuE2-3B base model. Non-commercial use only; attribute "QTSTRM LoRAs by becausereasons".

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for becausereasons/yue2-qtstrm-quiet-storm

Finetuned
Comfy-Org/YuE2
Adapter
(9)
this model