sam3-zerogpu / README.md
magboola's picture
deploy sam3-zerogpu
c495d1a verified
|
Raw History Blame Contribute Delete
5.76 kB
---
title: SAM 3 / 3.1 + Sapiens Pose ZeroGPU
emoji: 🎯
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 5.42.0
app_file: app.py
pinned: false
license: apache-2.0
hardware: zero-a10g
---
# sam3-zerogpu
HuggingFace ZeroGPU Space that hosts Meta's **SAM 3 / SAM 3.1** models
for Cadayn β€” **plus a bundled Sapiens-0.6B pose estimator** (see "Bundled
Sapiens pose" section below).
## Endpoints in this Space
**SAM 3 / SAM 3.1 (primary purpose):**
- `api_track_object` β€” SAM 3.1 multiplex video object tracking
- `api_annotate_image` β€” SAM 3 per-image text-prompted detection + optional masks
- `api_annotate_images_batch` β€” batched variant, up to 8 items per GPU slot
**Bundled Sapiens pose (added because of the 10 ZeroGPU Space cap):**
- `api_pose_image` β€” single image β†’ 308 keypoints per detected person
- `api_pose_video_frames` β€” video URL + timestamps β†’ per-frame multi-person
poses with stable IDs across frames (ByteTrack)
**Bundled NeuFlow v2 optical flow (Phase 2 β€” same cap reasoning):**
- `api_optical_flow` β€” video URL + anchor timestamps; for each anchor
computes flow between (t, t+pair_gap_s) and returns aggregated motion
stats (mean / p95 / max magnitude px, dominant direction deg, direction
consistency)
**Bundled Video Depth Anything (Phase 4 β€” same cap reasoning):**
- `api_video_depth` β€” video URL + anchor timestamps + optional per-frame
bboxes; returns RELATIVE depth in [0, 1] (never metric), per-bbox
p10/p50/p90 percentiles, and a near→far ordering of the bboxes per
frame. Set `is_relative: true` flag is always present in the response
so downstream code never promises metric distances.
The backend client lives in
`cadayn-backend/app/services/training/sam3p1_client.py`. Both halves are
versioned independently; the contract below documents what they must agree
on.
## API contract (v2 β€” Phase 3a W-3)
### `api_annotate_image` inputs
| Position | Name | Type | Notes |
|---|---|---|---|
| 0 | `image_url` | `str` | HTTPS URL; signed URLs ok |
| 1 | `text_prompts` | `str` | comma-separated class names |
| 2 | `confidence_threshold` | `float` | `0.0`–`1.0`, default `0.3` |
| 3 | `requested_masks` | `bool` | **v2 addition.** `True` = extract masks; `False` = boxes only |
`api_annotate_images_batch` has the same v2 flag in position 2
(after `items_json`, `confidence_threshold`).
### `detection.mask_rle.format`
Every returned detection includes a `mask_rle` dict with a `format`
field that callers must branch on:
- `uncompressed_rle_b64` β€” real SAM 3 segmentation from `processor.segment`
or the `set_text_prompt` output dict. Decode with
`pycocotools.mask.decode` after base64-decoding `counts`.
- `bbox_rle_b64` β€” rectangular fallback built from the detection's bbox.
Used when the installed `sam3` release does not expose masks. Treat
these as bbox-equivalent; do not use them for fine-grained overlays.
The mask section of the detection is absent when `requested_masks=False`.
## Deploy ordering with the backend
The Space and the backend are deployed from separate repos; the v2
contract takes 4 args on `api_annotate_image`. Mismatched versions break
differently depending on direction:
| Backend | Space | Result |
|---|---|---|
| v1 (3 args) | v1 (3 args) | OK β€” pre-Phase-3a baseline |
| v1 (3 args) | v2 (4 args) | OK β€” Gradio supplies `requested_masks=True` default |
| v2 (4 args) | v1 (3 args) | **BREAKS** β€” `ValueError: Expected 3 inputs, received 4`. Client retries once with 3 args and logs a warning |
| v2 (4 args) | v2 (4 args) | OK β€” steady state |
**Recommended ordering** when shipping a v2 change:
1. Deploy the Space first (`./deploy-phase-2-spaces.sh` or manual push).
2. Wait for the Space to reach `RUNNING`.
3. Deploy the backend β€” which now passes the 4th arg.
The client's one-shot retry is a belt-and-braces for the deploy window;
it is not a substitute for ordering.
## Env vars
- `SAM3_VIDEO_HF_SPACE_URL` β€” `api_track_object`
- `SAM3P1_IMAGE_HF_SPACE_URL` β€” `api_annotate_image` + batch
- `SAPIENS_HF_SPACE_URL` β€” `api_pose_image` + `api_pose_video_frames`
(point at this same Space β€” the bundled endpoints share the container)
- `NEUFLOW_HF_SPACE_URL` β€” `api_optical_flow` (same Space; Phase 2)
- `VDA_HF_SPACE_URL` β€” `api_video_depth` (same Space; Phase 4)
All five env vars can point at the same Space URL β€” the bundled Space
registers all seven Gradio routes.
## Bundled Sapiens pose
Sapiens-0.6B + YOLOv8n person detection + supervision.ByteTrack share
this container because the magboola HF account is at the 10 ZeroGPU Space
cap and SAM was the most natural visual co-tenant.
The standalone copy lives at `huggingface-spaces/sapiens-zerogpu/` (kept in
the tree but **not deployed**) for future extraction once the cap unblocks.
If/when extracted, sync `app.py` between the two locations and split the
pose-related entries out of `requirements.txt`.
### Pose output schema
```json
{
"ok": true,
"frames": [
{
"timestamp_s": 2.0,
"width": 1920,
"height": 1080,
"persons": [
{
"id": "p_3",
"bbox": {"x1": 412, "y1": 88, "x2": 712, "y2": 920, "conf": 0.91},
"keypoints": [
{"name": "nose", "x": 562.3, "y": 185.7, "conf": 0.94, "occluded": false}
],
"n_visible_keypoints": 287
}
]
}
],
"model": "facebook/sapiens-pose-0.6b-torchscript",
"max_persons_capped": false,
"elapsed_s": 4.12
}
```
Pose limits: 32 frames per video call, 12 persons per frame.
`goliath_keypoints.txt` ships with the Space β€” populated from
`facebookresearch/sapiens` to give keypoints semantic names; falls back
to `kp_0..kp_307` if the file is empty.