AXe Skills HubSearch /

← All skills

tao-generate-image-grounding

NV NVIDIA imagegroundingbounding-boxesauto-labelvlm2d-grounding

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, …

Definition

The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, …

Metadata

Category
Vision AI
Tier
community
Version
1.0.0

Use with an agent

Fetch this skill’s definition over the open API — no key required.

curl -s /v1/skills/community__axehub:NVIDIA__NVIDIA__tao-generate-image-grounding