NVIDIA imagegroundingbounding-boxesauto-labelvlm2d-grounding
Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, …
The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.
Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, …Fetch this skill’s definition over the open API — no key required.
curl -s /v1/skills/community__axehub:NVIDIA__NVIDIA__tao-generate-image-grounding