YuNet-Face-LiteRT / README.md
mlboydaisuke's picture
Card: base_model_relation: quantized (list under the base model's Quantizations)
1ff051b verified
|
Raw
History Blame Contribute Delete
7.62 kB
---
license: bsd-3-clause
library_name: LiteRT
pipeline_tag: object-detection
tags: [litert, tflite, on-device, android, gpu, face-detection, yunet, libfacedetection, landmarks]
base_model: ShiqiYu/libfacedetection
base_model_relation: quantized
---
# YuNet — LiteRT (on-device face detection, fully-GPU)
[YuNet](https://github.com/ShiqiYu/libfacedetection) (ShiqiYu/libfacedetection), a tiny fast face detector
(faces + 5 landmarks), converted to **LiteRT** and running **fully on the `CompiledModel` GPU** (ML Drift) on
Android. **0.076 M params / 0.3 MB fp16.**
![YuNet — face boxes + 5 landmarks on-device LiteRT GPU](samples/sample.png)
## On-device (Pixel 8a, Tensor G3 — verified)
| | |
|---|---|
| nodes on GPU | **146 / 146** LITERT_CL (full residency) |
| inference | **~4 ms** (640×640) |
| size | **0.3 MB** (fp16) |
| accuracy | device-vs-PyTorch corr **0.9999** (all 12 outputs) |
```
image[1,3,640,640] (BGR, 0-255) →[GPU: YuNet]→ 12 outputs: cls/obj/bbox/kps × strides {8,16,32}
```
## How it converts (litert-torch) — clean, no re-authoring
Pure CNN (depthwise-separable `ConvDPUnit`) + a **nearest-upsample** neck (`F.interpolate(mode="nearest")` →
`RESIZE_NEAREST_NEIGHBOR`, no transposed conv) + non-padded `MaxPool2d` (no `PADV2`). The head's per-stride
`permute/reshape/sigmoid` is baked in → 12 decode-ready outputs. Banned ops NONE, ≤4D, tflite-vs-torch corr
**1.0**, device-vs-torch corr **0.9999**.
## Decode (host-side) & preprocessing
**Preprocessing**: letterbox to 640×640, **BGR, 0-255, no normalization**. Anchor-free priors
(`px=col·s, py=row·s`, offset 0): score=`cls·obj`, box=center+`exp(wh)·s`, 5 landmarks `kps·s+prior`, then NMS.
## Minimal usage
**Android (Kotlin, CompiledModel GPU)**
```kotlin
val model = CompiledModel.create(context.assets, "yunet_fp16.tflite",
CompiledModel.Options(Accelerator.GPU), null)
val inputs = model.createInputBuffers(); val outputs = model.createOutputBuffers()
inputs[0].writeFloat(bgr) // [1,3,640,640] NCHW BGR, 0-255 (no normalization)
model.run(inputs, outputs)
// 12 outputs in order: cls x3 [1,N,1], obj x3 [1,N,1], bbox x3 [1,N,4], kps x3 [1,N,10]
// for strides {8,16,32}, N = (640/stride)^2 = 6400/1600/400. Decode = Python below.
val cls8 = outputs[0].readFloat()
```
**Python (desktop verification)**
```python
import math, numpy as np
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
im = Image.open("faces.jpg").convert("RGB").resize((640, 640))
bgr = np.asarray(im, np.float32)[:, :, ::-1] # BGR, 0-255
x = bgr.transpose(2, 0, 1)[None].copy() # [1,3,640,640]
it = Interpreter(model_path="yunet_fp16.tflite"); it.allocate_tensors()
it.set_tensor(it.get_input_details()[0]["index"], x); it.invoke()
o = [it.get_tensor(d["index"])[0] for d in it.get_output_details()]
# output order: cls x3, obj x3, bbox x3, kps x3 (strides 8, 16, 32)
dets = []
for li, s in enumerate([8, 16, 32]):
cls, obj, bb, kp = o[li][:, 0], o[3 + li][:, 0], o[6 + li], o[9 + li]
fw = 640 // s
for i in np.where(cls * obj > 0.6)[0]: # score threshold
px, py = (i % fw) * s, (i // fw) * s
cx, cy = bb[i, 0] * s + px, bb[i, 1] * s + py
w, h = math.exp(bb[i, 2]) * s, math.exp(bb[i, 3]) * s
lm = [(kp[i, 2 * j] * s + px, kp[i, 2 * j + 1] * s + py) for j in range(5)]
dets.append(([cx - w/2, cy - h/2, cx + w/2, cy + h/2], float(cls[i] * obj[i]), lm))
def iou(a, b): # greedy NMS, IoU 0.45
ix = max(0, min(a[2], b[2]) - max(a[0], b[0])); iy = max(0, min(a[3], b[3]) - max(a[1], b[1]))
u = (a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - ix*iy
return ix * iy / u if u > 0 else 0
dets.sort(key=lambda d: -d[1]); faces = []
for d in dets:
if all(iou(d[0], f[0]) < 0.45 for f in faces): faces.append(d)
for box, score, lm in faces: print(f"face {score:.2f}", np.round(box, 1), "landmarks", np.round(lm, 1))
```
## Performance
Measured on a **Pixel 8a** (Tensor G3, Android 16) with the standard TFLite [`benchmark_model`](https://ai.google.dev/edge/litert/models/measurement) tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.
| Runtime | Backend | Graph on GPU | Latency |
|---|---|---|---|
| LiteRT `CompiledModel` (`LITERT_CL`) | GPU | 146 / 146 | ~4 ms |
| TFLite `benchmark_model` (`TfLiteGpuDelegateV2`) | GPU (OpenCL) | 146 / 146 | 21.1 ms |
| TFLite `benchmark_model` | CPU (XNNPACK, 4 threads) | — | XNNPACK declined the graph |
**The two GPU rows are different runtimes, not a contradiction.** The `LITERT_CL` figure is the one recorded when this model shipped, taken through LiteRT's own `CompiledModel` accelerator — the path the Kotlin sample app and the LiteRT API use. The `TfLiteGpuDelegateV2` figure is the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. They agree on how much of the graph the GPU takes; they disagree on speed, and the classic delegate is the slower of the two here. Read the `TfLiteGpuDelegateV2` row as a reproducible floor, not as this model's speed on LiteRT.
XNNPACK declines these fp16 graphs — it reports `failed to delegate DEPTHWISE_CONV_2D` and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.
## Snapdragon NPU (Hexagon)
The NPU is **1.79x faster** than the GPU (1.85 ms against 3.31 ms) and loads 5.75x faster (100 ms against 577 ms).
| backend | compiled | inference (median / min) | load |
|---|---|---:|---:|
| NPU (Hexagon v81) | on-device JIT | 1.85 ms / 1.80 ms | 100 ms |
| GPU (Adreno) | — | 3.31 ms / 2.20 ms | 577 ms |
Measured on a **Samsung Galaxy S26** (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT `CompiledModel` 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status `NONE` throughout. Headroom 0.53, where 1.0 is the throttling threshold.
**The NPU rows ran the published file unchanged.** LiteRT compiled it for the Hexagon on the device at first load. That first compile took 979 ms here. The `load` column above is the cached load every later run pays. Recipe and the runtime libraries it needs: [NPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-npu.md).
GPU wiring: [GPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-gpu.md).
## Raspberry Pi 5 (CPU)
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT [`benchmark_model`](https://ai.google.dev/edge/litert/models/measurement) tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (`vcgencmd get_throttled` stayed `0x0`).
| File | Inference (median) | Spread (min–max) | Runs | Peak memory |
|---|---:|---:|---:|---:|
| `yunet_fp16.tflite` | 33.7 ms | 33.4–35.2 ms | 150 | 138 MB |
## License
[BSD-3-Clause](https://github.com/ShiqiYu/libfacedetection/blob/master/LICENSE). Upstream:
[ShiqiYu/libfacedetection](https://github.com/ShiqiYu/libfacedetection).