sam3-zerogpu / README.md
magboola's picture
deploy sam3-zerogpu
c495d1a verified
|
Raw
History Blame Contribute Delete
5.76 kB

A newer version of the Gradio SDK is available: 6.28.0

Upgrade
metadata
title: SAM 3 / 3.1 + Sapiens Pose ZeroGPU
emoji: 🎯
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 5.42.0
app_file: app.py
pinned: false
license: apache-2.0
hardware: zero-a10g

sam3-zerogpu

HuggingFace ZeroGPU Space that hosts Meta's SAM 3 / SAM 3.1 models for Cadayn β€” plus a bundled Sapiens-0.6B pose estimator (see "Bundled Sapiens pose" section below).

Endpoints in this Space

SAM 3 / SAM 3.1 (primary purpose):

  • api_track_object β€” SAM 3.1 multiplex video object tracking
  • api_annotate_image β€” SAM 3 per-image text-prompted detection + optional masks
  • api_annotate_images_batch β€” batched variant, up to 8 items per GPU slot

Bundled Sapiens pose (added because of the 10 ZeroGPU Space cap):

  • api_pose_image β€” single image β†’ 308 keypoints per detected person
  • api_pose_video_frames β€” video URL + timestamps β†’ per-frame multi-person poses with stable IDs across frames (ByteTrack)

Bundled NeuFlow v2 optical flow (Phase 2 β€” same cap reasoning):

  • api_optical_flow β€” video URL + anchor timestamps; for each anchor computes flow between (t, t+pair_gap_s) and returns aggregated motion stats (mean / p95 / max magnitude px, dominant direction deg, direction consistency)

Bundled Video Depth Anything (Phase 4 β€” same cap reasoning):

  • api_video_depth β€” video URL + anchor timestamps + optional per-frame bboxes; returns RELATIVE depth in [0, 1] (never metric), per-bbox p10/p50/p90 percentiles, and a nearβ†’far ordering of the bboxes per frame. Set is_relative: true flag is always present in the response so downstream code never promises metric distances.

The backend client lives in cadayn-backend/app/services/training/sam3p1_client.py. Both halves are versioned independently; the contract below documents what they must agree on.

API contract (v2 β€” Phase 3a W-3)

api_annotate_image inputs

Position Name Type Notes
0 image_url str HTTPS URL; signed URLs ok
1 text_prompts str comma-separated class names
2 confidence_threshold float 0.0–1.0, default 0.3
3 requested_masks bool v2 addition. True = extract masks; False = boxes only

api_annotate_images_batch has the same v2 flag in position 2 (after items_json, confidence_threshold).

detection.mask_rle.format

Every returned detection includes a mask_rle dict with a format field that callers must branch on:

  • uncompressed_rle_b64 β€” real SAM 3 segmentation from processor.segment or the set_text_prompt output dict. Decode with pycocotools.mask.decode after base64-decoding counts.
  • bbox_rle_b64 β€” rectangular fallback built from the detection's bbox. Used when the installed sam3 release does not expose masks. Treat these as bbox-equivalent; do not use them for fine-grained overlays.

The mask section of the detection is absent when requested_masks=False.

Deploy ordering with the backend

The Space and the backend are deployed from separate repos; the v2 contract takes 4 args on api_annotate_image. Mismatched versions break differently depending on direction:

Backend Space Result
v1 (3 args) v1 (3 args) OK β€” pre-Phase-3a baseline
v1 (3 args) v2 (4 args) OK β€” Gradio supplies requested_masks=True default
v2 (4 args) v1 (3 args) BREAKS β€” ValueError: Expected 3 inputs, received 4. Client retries once with 3 args and logs a warning
v2 (4 args) v2 (4 args) OK β€” steady state

Recommended ordering when shipping a v2 change:

  1. Deploy the Space first (./deploy-phase-2-spaces.sh or manual push).
  2. Wait for the Space to reach RUNNING.
  3. Deploy the backend β€” which now passes the 4th arg.

The client's one-shot retry is a belt-and-braces for the deploy window; it is not a substitute for ordering.

Env vars

  • SAM3_VIDEO_HF_SPACE_URL β€” api_track_object
  • SAM3P1_IMAGE_HF_SPACE_URL β€” api_annotate_image + batch
  • SAPIENS_HF_SPACE_URL β€” api_pose_image + api_pose_video_frames (point at this same Space β€” the bundled endpoints share the container)
  • NEUFLOW_HF_SPACE_URL β€” api_optical_flow (same Space; Phase 2)
  • VDA_HF_SPACE_URL β€” api_video_depth (same Space; Phase 4)

All five env vars can point at the same Space URL β€” the bundled Space registers all seven Gradio routes.

Bundled Sapiens pose

Sapiens-0.6B + YOLOv8n person detection + supervision.ByteTrack share this container because the magboola HF account is at the 10 ZeroGPU Space cap and SAM was the most natural visual co-tenant.

The standalone copy lives at huggingface-spaces/sapiens-zerogpu/ (kept in the tree but not deployed) for future extraction once the cap unblocks. If/when extracted, sync app.py between the two locations and split the pose-related entries out of requirements.txt.

Pose output schema

{
  "ok": true,
  "frames": [
    {
      "timestamp_s": 2.0,
      "width": 1920,
      "height": 1080,
      "persons": [
        {
          "id": "p_3",
          "bbox": {"x1": 412, "y1": 88, "x2": 712, "y2": 920, "conf": 0.91},
          "keypoints": [
            {"name": "nose", "x": 562.3, "y": 185.7, "conf": 0.94, "occluded": false}
          ],
          "n_visible_keypoints": 287
        }
      ]
    }
  ],
  "model": "facebook/sapiens-pose-0.6b-torchscript",
  "max_persons_capped": false,
  "elapsed_s": 4.12
}

Pose limits: 32 frames per video call, 12 persons per frame.

goliath_keypoints.txt ships with the Space β€” populated from facebookresearch/sapiens to give keypoints semantic names; falls back to kp_0..kp_307 if the file is empty.