Spaces:
Running on Zero
A newer version of the Gradio SDK is available: 6.28.0
title: SAM 3 / 3.1 + Sapiens Pose ZeroGPU
emoji: π―
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 5.42.0
app_file: app.py
pinned: false
license: apache-2.0
hardware: zero-a10g
sam3-zerogpu
HuggingFace ZeroGPU Space that hosts Meta's SAM 3 / SAM 3.1 models for Cadayn β plus a bundled Sapiens-0.6B pose estimator (see "Bundled Sapiens pose" section below).
Endpoints in this Space
SAM 3 / SAM 3.1 (primary purpose):
api_track_objectβ SAM 3.1 multiplex video object trackingapi_annotate_imageβ SAM 3 per-image text-prompted detection + optional masksapi_annotate_images_batchβ batched variant, up to 8 items per GPU slot
Bundled Sapiens pose (added because of the 10 ZeroGPU Space cap):
api_pose_imageβ single image β 308 keypoints per detected personapi_pose_video_framesβ video URL + timestamps β per-frame multi-person poses with stable IDs across frames (ByteTrack)
Bundled NeuFlow v2 optical flow (Phase 2 β same cap reasoning):
api_optical_flowβ video URL + anchor timestamps; for each anchor computes flow between (t, t+pair_gap_s) and returns aggregated motion stats (mean / p95 / max magnitude px, dominant direction deg, direction consistency)
Bundled Video Depth Anything (Phase 4 β same cap reasoning):
api_video_depthβ video URL + anchor timestamps + optional per-frame bboxes; returns RELATIVE depth in [0, 1] (never metric), per-bbox p10/p50/p90 percentiles, and a nearβfar ordering of the bboxes per frame. Setis_relative: trueflag is always present in the response so downstream code never promises metric distances.
The backend client lives in
cadayn-backend/app/services/training/sam3p1_client.py. Both halves are
versioned independently; the contract below documents what they must agree
on.
API contract (v2 β Phase 3a W-3)
api_annotate_image inputs
| Position | Name | Type | Notes |
|---|---|---|---|
| 0 | image_url |
str |
HTTPS URL; signed URLs ok |
| 1 | text_prompts |
str |
comma-separated class names |
| 2 | confidence_threshold |
float |
0.0β1.0, default 0.3 |
| 3 | requested_masks |
bool |
v2 addition. True = extract masks; False = boxes only |
api_annotate_images_batch has the same v2 flag in position 2
(after items_json, confidence_threshold).
detection.mask_rle.format
Every returned detection includes a mask_rle dict with a format
field that callers must branch on:
uncompressed_rle_b64β real SAM 3 segmentation fromprocessor.segmentor theset_text_promptoutput dict. Decode withpycocotools.mask.decodeafter base64-decodingcounts.bbox_rle_b64β rectangular fallback built from the detection's bbox. Used when the installedsam3release does not expose masks. Treat these as bbox-equivalent; do not use them for fine-grained overlays.
The mask section of the detection is absent when requested_masks=False.
Deploy ordering with the backend
The Space and the backend are deployed from separate repos; the v2
contract takes 4 args on api_annotate_image. Mismatched versions break
differently depending on direction:
| Backend | Space | Result |
|---|---|---|
| v1 (3 args) | v1 (3 args) | OK β pre-Phase-3a baseline |
| v1 (3 args) | v2 (4 args) | OK β Gradio supplies requested_masks=True default |
| v2 (4 args) | v1 (3 args) | BREAKS β ValueError: Expected 3 inputs, received 4. Client retries once with 3 args and logs a warning |
| v2 (4 args) | v2 (4 args) | OK β steady state |
Recommended ordering when shipping a v2 change:
- Deploy the Space first (
./deploy-phase-2-spaces.shor manual push). - Wait for the Space to reach
RUNNING. - Deploy the backend β which now passes the 4th arg.
The client's one-shot retry is a belt-and-braces for the deploy window; it is not a substitute for ordering.
Env vars
SAM3_VIDEO_HF_SPACE_URLβapi_track_objectSAM3P1_IMAGE_HF_SPACE_URLβapi_annotate_image+ batchSAPIENS_HF_SPACE_URLβapi_pose_image+api_pose_video_frames(point at this same Space β the bundled endpoints share the container)NEUFLOW_HF_SPACE_URLβapi_optical_flow(same Space; Phase 2)VDA_HF_SPACE_URLβapi_video_depth(same Space; Phase 4)
All five env vars can point at the same Space URL β the bundled Space registers all seven Gradio routes.
Bundled Sapiens pose
Sapiens-0.6B + YOLOv8n person detection + supervision.ByteTrack share this container because the magboola HF account is at the 10 ZeroGPU Space cap and SAM was the most natural visual co-tenant.
The standalone copy lives at huggingface-spaces/sapiens-zerogpu/ (kept in
the tree but not deployed) for future extraction once the cap unblocks.
If/when extracted, sync app.py between the two locations and split the
pose-related entries out of requirements.txt.
Pose output schema
{
"ok": true,
"frames": [
{
"timestamp_s": 2.0,
"width": 1920,
"height": 1080,
"persons": [
{
"id": "p_3",
"bbox": {"x1": 412, "y1": 88, "x2": 712, "y2": 920, "conf": 0.91},
"keypoints": [
{"name": "nose", "x": 562.3, "y": 185.7, "conf": 0.94, "occluded": false}
],
"n_visible_keypoints": 287
}
]
}
],
"model": "facebook/sapiens-pose-0.6b-torchscript",
"max_persons_capped": false,
"elapsed_s": 4.12
}
Pose limits: 32 frames per video call, 12 persons per frame.
goliath_keypoints.txt ships with the Space β populated from
facebookresearch/sapiens to give keypoints semantic names; falls back
to kp_0..kp_307 if the file is empty.