--- title: SAM 3 / 3.1 + Sapiens Pose ZeroGPU emoji: 🎯 colorFrom: indigo colorTo: purple sdk: gradio sdk_version: 5.42.0 app_file: app.py pinned: false license: apache-2.0 hardware: zero-a10g --- # sam3-zerogpu HuggingFace ZeroGPU Space that hosts Meta's **SAM 3 / SAM 3.1** models for Cadayn β€” **plus a bundled Sapiens-0.6B pose estimator** (see "Bundled Sapiens pose" section below). ## Endpoints in this Space **SAM 3 / SAM 3.1 (primary purpose):** - `api_track_object` β€” SAM 3.1 multiplex video object tracking - `api_annotate_image` β€” SAM 3 per-image text-prompted detection + optional masks - `api_annotate_images_batch` β€” batched variant, up to 8 items per GPU slot **Bundled Sapiens pose (added because of the 10 ZeroGPU Space cap):** - `api_pose_image` β€” single image β†’ 308 keypoints per detected person - `api_pose_video_frames` β€” video URL + timestamps β†’ per-frame multi-person poses with stable IDs across frames (ByteTrack) **Bundled NeuFlow v2 optical flow (Phase 2 β€” same cap reasoning):** - `api_optical_flow` β€” video URL + anchor timestamps; for each anchor computes flow between (t, t+pair_gap_s) and returns aggregated motion stats (mean / p95 / max magnitude px, dominant direction deg, direction consistency) **Bundled Video Depth Anything (Phase 4 β€” same cap reasoning):** - `api_video_depth` β€” video URL + anchor timestamps + optional per-frame bboxes; returns RELATIVE depth in [0, 1] (never metric), per-bbox p10/p50/p90 percentiles, and a nearβ†’far ordering of the bboxes per frame. Set `is_relative: true` flag is always present in the response so downstream code never promises metric distances. The backend client lives in `cadayn-backend/app/services/training/sam3p1_client.py`. Both halves are versioned independently; the contract below documents what they must agree on. ## API contract (v2 β€” Phase 3a W-3) ### `api_annotate_image` inputs | Position | Name | Type | Notes | |---|---|---|---| | 0 | `image_url` | `str` | HTTPS URL; signed URLs ok | | 1 | `text_prompts` | `str` | comma-separated class names | | 2 | `confidence_threshold` | `float` | `0.0`–`1.0`, default `0.3` | | 3 | `requested_masks` | `bool` | **v2 addition.** `True` = extract masks; `False` = boxes only | `api_annotate_images_batch` has the same v2 flag in position 2 (after `items_json`, `confidence_threshold`). ### `detection.mask_rle.format` Every returned detection includes a `mask_rle` dict with a `format` field that callers must branch on: - `uncompressed_rle_b64` β€” real SAM 3 segmentation from `processor.segment` or the `set_text_prompt` output dict. Decode with `pycocotools.mask.decode` after base64-decoding `counts`. - `bbox_rle_b64` β€” rectangular fallback built from the detection's bbox. Used when the installed `sam3` release does not expose masks. Treat these as bbox-equivalent; do not use them for fine-grained overlays. The mask section of the detection is absent when `requested_masks=False`. ## Deploy ordering with the backend The Space and the backend are deployed from separate repos; the v2 contract takes 4 args on `api_annotate_image`. Mismatched versions break differently depending on direction: | Backend | Space | Result | |---|---|---| | v1 (3 args) | v1 (3 args) | OK β€” pre-Phase-3a baseline | | v1 (3 args) | v2 (4 args) | OK β€” Gradio supplies `requested_masks=True` default | | v2 (4 args) | v1 (3 args) | **BREAKS** β€” `ValueError: Expected 3 inputs, received 4`. Client retries once with 3 args and logs a warning | | v2 (4 args) | v2 (4 args) | OK β€” steady state | **Recommended ordering** when shipping a v2 change: 1. Deploy the Space first (`./deploy-phase-2-spaces.sh` or manual push). 2. Wait for the Space to reach `RUNNING`. 3. Deploy the backend β€” which now passes the 4th arg. The client's one-shot retry is a belt-and-braces for the deploy window; it is not a substitute for ordering. ## Env vars - `SAM3_VIDEO_HF_SPACE_URL` β€” `api_track_object` - `SAM3P1_IMAGE_HF_SPACE_URL` β€” `api_annotate_image` + batch - `SAPIENS_HF_SPACE_URL` β€” `api_pose_image` + `api_pose_video_frames` (point at this same Space β€” the bundled endpoints share the container) - `NEUFLOW_HF_SPACE_URL` β€” `api_optical_flow` (same Space; Phase 2) - `VDA_HF_SPACE_URL` β€” `api_video_depth` (same Space; Phase 4) All five env vars can point at the same Space URL β€” the bundled Space registers all seven Gradio routes. ## Bundled Sapiens pose Sapiens-0.6B + YOLOv8n person detection + supervision.ByteTrack share this container because the magboola HF account is at the 10 ZeroGPU Space cap and SAM was the most natural visual co-tenant. The standalone copy lives at `huggingface-spaces/sapiens-zerogpu/` (kept in the tree but **not deployed**) for future extraction once the cap unblocks. If/when extracted, sync `app.py` between the two locations and split the pose-related entries out of `requirements.txt`. ### Pose output schema ```json { "ok": true, "frames": [ { "timestamp_s": 2.0, "width": 1920, "height": 1080, "persons": [ { "id": "p_3", "bbox": {"x1": 412, "y1": 88, "x2": 712, "y2": 920, "conf": 0.91}, "keypoints": [ {"name": "nose", "x": 562.3, "y": 185.7, "conf": 0.94, "occluded": false} ], "n_visible_keypoints": 287 } ] } ], "model": "facebook/sapiens-pose-0.6b-torchscript", "max_persons_capped": false, "elapsed_s": 4.12 } ``` Pose limits: 32 frames per video call, 12 persons per frame. `goliath_keypoints.txt` ships with the Space β€” populated from `facebookresearch/sapiens` to give keypoints semantic names; falls back to `kp_0..kp_307` if the file is empty.