--- license: bsd-3-clause library_name: LiteRT pipeline_tag: object-detection tags: [litert, tflite, on-device, android, gpu, face-detection, yunet, libfacedetection, landmarks] base_model: ShiqiYu/libfacedetection base_model_relation: quantized --- # YuNet — LiteRT (on-device face detection, fully-GPU) [YuNet](https://github.com/ShiqiYu/libfacedetection) (ShiqiYu/libfacedetection), a tiny fast face detector (faces + 5 landmarks), converted to **LiteRT** and running **fully on the `CompiledModel` GPU** (ML Drift) on Android. **0.076 M params / 0.3 MB fp16.** ![YuNet — face boxes + 5 landmarks on-device LiteRT GPU](samples/sample.png) ## On-device (Pixel 8a, Tensor G3 — verified) | | | |---|---| | nodes on GPU | **146 / 146** LITERT_CL (full residency) | | inference | **~4 ms** (640×640) | | size | **0.3 MB** (fp16) | | accuracy | device-vs-PyTorch corr **0.9999** (all 12 outputs) | ``` image[1,3,640,640] (BGR, 0-255) →[GPU: YuNet]→ 12 outputs: cls/obj/bbox/kps × strides {8,16,32} ``` ## How it converts (litert-torch) — clean, no re-authoring Pure CNN (depthwise-separable `ConvDPUnit`) + a **nearest-upsample** neck (`F.interpolate(mode="nearest")` → `RESIZE_NEAREST_NEIGHBOR`, no transposed conv) + non-padded `MaxPool2d` (no `PADV2`). The head's per-stride `permute/reshape/sigmoid` is baked in → 12 decode-ready outputs. Banned ops NONE, ≤4D, tflite-vs-torch corr **1.0**, device-vs-torch corr **0.9999**. ## Decode (host-side) & preprocessing **Preprocessing**: letterbox to 640×640, **BGR, 0-255, no normalization**. Anchor-free priors (`px=col·s, py=row·s`, offset 0): score=`cls·obj`, box=center+`exp(wh)·s`, 5 landmarks `kps·s+prior`, then NMS. ## Minimal usage **Android (Kotlin, CompiledModel GPU)** ```kotlin val model = CompiledModel.create(context.assets, "yunet_fp16.tflite", CompiledModel.Options(Accelerator.GPU), null) val inputs = model.createInputBuffers(); val outputs = model.createOutputBuffers() inputs[0].writeFloat(bgr) // [1,3,640,640] NCHW BGR, 0-255 (no normalization) model.run(inputs, outputs) // 12 outputs in order: cls x3 [1,N,1], obj x3 [1,N,1], bbox x3 [1,N,4], kps x3 [1,N,10] // for strides {8,16,32}, N = (640/stride)^2 = 6400/1600/400. Decode = Python below. val cls8 = outputs[0].readFloat() ``` **Python (desktop verification)** ```python import math, numpy as np from PIL import Image from ai_edge_litert.interpreter import Interpreter im = Image.open("faces.jpg").convert("RGB").resize((640, 640)) bgr = np.asarray(im, np.float32)[:, :, ::-1] # BGR, 0-255 x = bgr.transpose(2, 0, 1)[None].copy() # [1,3,640,640] it = Interpreter(model_path="yunet_fp16.tflite"); it.allocate_tensors() it.set_tensor(it.get_input_details()[0]["index"], x); it.invoke() o = [it.get_tensor(d["index"])[0] for d in it.get_output_details()] # output order: cls x3, obj x3, bbox x3, kps x3 (strides 8, 16, 32) dets = [] for li, s in enumerate([8, 16, 32]): cls, obj, bb, kp = o[li][:, 0], o[3 + li][:, 0], o[6 + li], o[9 + li] fw = 640 // s for i in np.where(cls * obj > 0.6)[0]: # score threshold px, py = (i % fw) * s, (i // fw) * s cx, cy = bb[i, 0] * s + px, bb[i, 1] * s + py w, h = math.exp(bb[i, 2]) * s, math.exp(bb[i, 3]) * s lm = [(kp[i, 2 * j] * s + px, kp[i, 2 * j + 1] * s + py) for j in range(5)] dets.append(([cx - w/2, cy - h/2, cx + w/2, cy + h/2], float(cls[i] * obj[i]), lm)) def iou(a, b): # greedy NMS, IoU 0.45 ix = max(0, min(a[2], b[2]) - max(a[0], b[0])); iy = max(0, min(a[3], b[3]) - max(a[1], b[1])) u = (a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - ix*iy return ix * iy / u if u > 0 else 0 dets.sort(key=lambda d: -d[1]); faces = [] for d in dets: if all(iou(d[0], f[0]) < 0.45 for f in faces): faces.append(d) for box, score, lm in faces: print(f"face {score:.2f}", np.round(box, 1), "landmarks", np.round(lm, 1)) ``` ## Performance Measured on a **Pixel 8a** (Tensor G3, Android 16) with the standard TFLite [`benchmark_model`](https://ai.google.dev/edge/litert/models/measurement) tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean. | Runtime | Backend | Graph on GPU | Latency | |---|---|---|---| | LiteRT `CompiledModel` (`LITERT_CL`) | GPU | 146 / 146 | ~4 ms | | TFLite `benchmark_model` (`TfLiteGpuDelegateV2`) | GPU (OpenCL) | 146 / 146 | 21.1 ms | | TFLite `benchmark_model` | CPU (XNNPACK, 4 threads) | — | XNNPACK declined the graph | **The two GPU rows are different runtimes, not a contradiction.** The `LITERT_CL` figure is the one recorded when this model shipped, taken through LiteRT's own `CompiledModel` accelerator — the path the Kotlin sample app and the LiteRT API use. The `TfLiteGpuDelegateV2` figure is the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. They agree on how much of the graph the GPU takes; they disagree on speed, and the classic delegate is the slower of the two here. Read the `TfLiteGpuDelegateV2` row as a reproducible floor, not as this model's speed on LiteRT. XNNPACK declines these fp16 graphs — it reports `failed to delegate DEPTHWISE_CONV_2D` and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship. ## Snapdragon NPU (Hexagon) The NPU is **1.79x faster** than the GPU (1.85 ms against 3.31 ms) and loads 5.75x faster (100 ms against 577 ms). | backend | compiled | inference (median / min) | load | |---|---|---:|---:| | NPU (Hexagon v81) | on-device JIT | 1.85 ms / 1.80 ms | 100 ms | | GPU (Adreno) | — | 3.31 ms / 2.20 ms | 577 ms | Measured on a **Samsung Galaxy S26** (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT `CompiledModel` 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status `NONE` throughout. Headroom 0.53, where 1.0 is the throttling threshold. **The NPU rows ran the published file unchanged.** LiteRT compiled it for the Hexagon on the device at first load. That first compile took 979 ms here. The `load` column above is the cached load every later run pays. Recipe and the runtime libraries it needs: [NPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-npu.md). GPU wiring: [GPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-gpu.md). ## Raspberry Pi 5 (CPU) Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT [`benchmark_model`](https://ai.google.dev/edge/litert/models/measurement) tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (`vcgencmd get_throttled` stayed `0x0`). | File | Inference (median) | Spread (min–max) | Runs | Peak memory | |---|---:|---:|---:|---:| | `yunet_fp16.tflite` | 33.7 ms | 33.4–35.2 ms | 150 | 138 MB | ## License [BSD-3-Clause](https://github.com/ShiqiYu/libfacedetection/blob/master/LICENSE). Upstream: [ShiqiYu/libfacedetection](https://github.com/ShiqiYu/libfacedetection).