Instructions to use mlx-community/FlashVSR-v1.1-fp32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/FlashVSR-v1.1-fp32 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/FlashVSR-v1.1-fp32 --local-dir FlashVSR-v1.1-fp32
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/FlashVSR-v1.1-fp32
FlashVSR v1.1 (JunhaoZhuang/FlashVSR-v1.1,
OpenImagingLab/FlashVSR, arXiv 2510.12747) β
one-step, streaming, diffusion-based Γ4 video super-resolution β converted to MLX (fp32, the parity / reference lane) for the Swift/MLX port
xocialize/mlx-flashvsr-swift.
β οΈ Recommended for live-action footage β not anime, cartoons or graphics
FlashVSR is generative: it invents plausible detail at the target resolution rather than reconstructing it. On live action that is its strength β resolution-limited faces come back with real photographic detail where faithful upscalers give a blur. On drawn content (anime, cartoons, motion graphics, UI, text overlays) it renders flat colour and clean line art as photographic texture and pushes the drawing toward realism; use a faithful upscaler (e.g. Real-ESRGAN anime models) there instead. It is also aggressive with defocus: bokeh and soft backgrounds can come back as crisp invented texture.
Files
| file | component | params | notes |
|---|---|---|---|
dit_fp32.safetensors |
Wan2.1-1.3B-shaped DiT, DMD-distilled to one step at t = 1000 | 1,418,996,800 | patch_embedding Conv3d (O,D,H,W,I) |
lq_proj_fp32.safetensors |
Causal_LQ4x_Proj β the LQ conditioning path |
287,845,888 | Conv3d (O,D,H,W,I); RMS gammas (C) |
tcdecoder_fp32.safetensors |
TCDecoder β the LQ-conditioned tiny decoder (TAEHV-wide) | 45,338,371 | Conv2d (O,H,W,I) |
prompt_fp32.safetensors |
the fixed prompt context (1, 512, 4096) | β | no text encoder at runtime |
config.json |
architecture, pipeline defaults, key contract, provenance (source sha256s) |
Keys are upstream's verbatim; only conv layouts change (channels-last). Produced by the port's
oracle/convert_weights.py --lane fp32 from JunhaoZhuang/FlashVSR-v1.1 @ 27561b18. No Wan VAE and no umT5
are needed: the tiny pipeline decodes with the TCDecoder and conditions on the fixed prompt tensor.
Lanes. mlx-community/FlashVSR-v1.1-bf16 is the production lane (3.5 GB): every tensor rounded to bf16 β
bit-identical to casting the fp32 lane to bf16 at load (verified over all 899 parameters).
mlx-community/FlashVSR-v1.1-fp32 is the parity / reference lane (7.0 GB): the DiT and decoder as released, the
LQ projector and prompt (released in bf16) upcast exactly.
Parity with upstream
Gated against upstream's own code executed verbatim (CPU, fp32; 512Γ384 Γ33 frames, two goldens β upstream's default sparsity and a strongly sparse setting with empty query blocks):
- every component at isolated relative error β€ 1e-5 (DiT blocks fed the golden block inputs, LQ projector, decoder);
- end to end at upstream's own sensitivity floor: FlashVSR's locality-constrained sparse attention selects 128Γ128 blocks by a hard top-k, so a 1e-6 change to the input flips near-tied blocks β upstream itself moves to 52.9 dB under such a nudge, and so does this port (identical max-abs error). The port's E2E is 100β113 dB where upstream's floor is high.
The block-sparse attention (upstream: mit-han-lab's CUDA Block-Sparse-Attention) is implemented as a Metal kernel
that computes only the selected blocks β the LCSA path upstream's card asks third-party ports not to drop.
Quality
Γ4 (320Γ192 β 1280Γ768), full-reference against the native 1280Γ768 source, against upstream on PyTorch-MPS:
| clip | upstream (torch) SSIMULACRA2 | this port, bf16 | PSNR torch / port |
|---|---|---|---|
| live action | β17.0 | β16.3 | 27.49 / 27.59 |
| anime pan | β45.7 | β44.1 | 25.40 / 25.48 |
bf16 sits inside the fp32 seed-to-seed spread on every metric. (The anime row is why the recommendation above exists: the score reflects invented texture, with an image gradient twice the reference's.)
Memory
Streaming keeps memory independent of clip length; it scales with output pixels per frame. Measured through
MLXEngine on an M5 Max (process peak phys_footprint, bf16): 640Γ384 out 9.2 GB Β· 1280Γ768 out 19.1 GB Β·
1920Γ1152 out 33.7 GB. fp32 at 1280Γ768: 34.4 GB.
Use with mlx-flashvsr-swift
import MLXServeCore, MLXToolKit, MLXFlashVSR
let engine = MLXServeEngine()
let id = try await engine.register(FlashVSRUpscalePackage.registration,
configuration: FlashVSRConfiguration(precision: .fp32))
let out = try await engine.run(VideoUpscaleRequest(video: video, scale: 4), package: id) // Γ4 or Γ2
The engine downloads this repo into its model store on first use. Or drive the core directly:
import FlashVSRMLX
let pipe = try FlashVSRPipeline.load(directory: laneDir, dtypes: .parity)
let stream = FlashVSRStream(pipeline: pipe, height: 768, width: 1280) // Γ4 bicubic LQ size, multiples of 128
for frame in lqFrames { for out in try stream.push(frame) { /* β¦ */ } }
for out in try stream.finish() { /* one output per input frame */ }
Licences and provenance
Apache-2.0: FlashVSR code and v1.1 weights (OpenImagingLab / Junhao Zhuang et al.); the DiT is the Wan2.1 architecture (Apache-2.0); the TCDecoder derives from TAEHV (MIT). The training set (VSR-120K) is described by its authors but not released; this re-host takes the declared weight licence as governing.
@article{zhuang2025flashvsr,
title={FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution},
author={Zhuang, Junhao and Guo, Shi and Cai, Xin and Li, Xiaohui and Liu, Yihao and Yuan, Chun and Xue, Tianfan},
journal={arXiv preprint arXiv:2510.12747},
year={2025}
}
- Downloads last month
- 32
Quantized
Model tree for mlx-community/FlashVSR-v1.1-fp32
Base model
JunhaoZhuang/FlashVSR-v1.1