Search: Multimodal
100 result(s) on page 1
multimodal
verifiedFirst-party- Registry
- First-party
- Category
- Media
- Version
- 1.0.0
- Author
- AXE
- License
- MIT
curl -s /v1/skills/multimodalMultimodal_Asset_Tagger
Generate AI-optimized Alt Text, file names, captions, and Schema markup for images, videos, and audio assets. Improves AI discoverability on Google Lens, Cha...
quarantinedalt-textgeoimagesClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Multimodal_Asset_TaggerAliyun_Qwen_Multimodal_Embedding
Use when multimodal embeddings are needed from Alibaba Cloud Model Studio models such as `qwen3-vl-embedding` for image, video, and text retrieval, cross-mod...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Aliyun_Qwen_Multimodal_EmbeddingMarkitdown-Skill-for-non-multimodal-agent
Use when a NON-multimodal agent (a text-only LLM backend that cannot read attachments) receives a document β PDF, Word (docx), PowerPoint (pptx), Excel (xlsx...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Markitdown-Skill-for-non-multimodal-agentMultimodal_Base
Supports image understanding, OCR, speech-to-text, and text-to-speech synthesis with multi-voice and multimodal unified processing using OpenAI and Edge TTS.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Multimodal_BaseNovita_AI_Multimodal
Execute multimodal tasks using Novita AI: text-to-image, image-to-image, text-to-video, image-to-video, TTS, STT. Use for: generating images, generating vide...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Novita_AI_MultimodalAlibabacloud_Pds_Multimodal_Search
Implements exact filename search, fuzzy filename search, semantic file search, and image-based image search Triggers: "PDS drive file search", "PDS image sea...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Alibabacloud_Pds_Multimodal_Searchmultimodal_analysis,_and_more.
ARTCLAW AI Creative Suite - invoke ARTCLAW platform's AI content creation capabilities via CLI. Supports AI image generation, video generation, workflow exec...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Invoke_ARTCLAW_platform's_AI_content_creation_capabilities_via_REST_API._Supports_AI_image_generation,_video_generation,_workflow_execution,___multimodal_analysis,_and_more.MiniMax_Multimodal_Toolkit
Generate and process speech, music, video, and images using MiniMax AI with voice cloning, custom voices, multi-scene video, and FFmpeg-based media tools.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MiniMax_Multimodal_ToolkitMinimax-Multimodal-Toolkit
Use mmx to generate text, images, video, speech, and music via the MiniMax AI platform. Use when the user wants to create media content, chat with MiniMax mo...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Minimax-Multimodal-ToolkitMultimodal_Pet_Health_Engine
Transforms mmWave radar data into pet health metrics by detecting micro-movements, fusing environmental data, and enabling automatic spatial adjustments.
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Multimodal_Pet_Health_Enginegpt-multimodal
Analyze images and multi-frame sequences using OpenAI GPT series
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__gpt-multimodalmultimodal-parser
Unified multi-modal content parser for images, PDF, DOCX, audio, auto OCR/transcription, output structured text for LLM processing
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__multimodal-parserπ¦_Shrink_β_Three-Tier_Multimodal_Context_Optimizer
Replace base64 images in session history with context-aware text descriptions, reducing image token cost by 96-99%. Use when: (1) user says /shrink, /shrink,...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__π¦_Shrink_β_Three-Tier_Multimodal_Context_Optimizerai-multimodal
Indexed by skills.sh from mrgoonie/claudekit-skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- mrgoonie
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__ai-multimodalalicloud-ai-multimodal-qwen-vl
Indexed by skills.sh from cinience/alicloud-skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- cinience
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__alicloud-ai-multimodal-qwen-vlalicloud-ai-multimodal-qwen-vl-test
Indexed by skills.sh from cinience/alicloud-skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- cinience
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__alicloud-ai-multimodal-qwen-vl-testgemma-tuner-multimodal
Indexed by skills.sh from aradotso/trending-skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- aradotso
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__gemma-tuner-multimodallinkfox-multimodal-recognize-image
Indexed by skills.sh from linkfox-ai/linkfox-skills
quarantinedskills.sh- Registry
- skills.sh
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
- Author
- linkfox-ai
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__linkfox-multimodal-recognize-imageminimax-multimodal-toolkit
Indexed by skills.sh from minimax-ai/skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- minimax-ai
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__minimax-multimodal-toolkitvision-multimodal
Indexed by skills.sh from lobbi-docs/claude
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- lobbi-docs
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__vision-multimodalOllama_Herd
Ollama multimodal model router for Llama, Qwen, DeepSeek, Phi, and Mistral β plus mflux image generation, speech-to-text, and embeddings. Self-hosted Ollama...
quarantinedapple-silicondashboarddeepseekClawHub- Registry
- ClawHub
- Category
- Creative
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ollama_HerdOllama_Ollama_Herd
Ollama Ollama Herd β multimodal Ollama model router that herds your Ollama LLMs into one smart Ollama endpoint. Route Ollama Llama, Qwen, DeepSeek, Phi, Mist...
quarantinedapple-silicondeepseekembeddingsClawHub- Registry
- ClawHub
- Category
- Creative
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ollama_Ollama_HerdAI_Prompt_Reverse_Engine
Convert image to structured prompts for multiple AI models
quarantinedaicomputer-visionmultimodalClawHub- Registry
- ClawHub
- Category
- AI Agents
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__AI_Prompt_Reverse_EngineLLM_Knowledge_Bases
Inspired by a public workflow shared by Andrej Karpathy (@karpathy). From raw research to a living Markdown knowledge base that compounds with every question...
quarantineddataimageknowledge-baseClawHub- Registry
- ClawHub
- Category
- Data Science
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__LLM_Knowledge_BasesLLM_Wiki_Karpathy
Manage and maintain a Markdown wiki with LLM Wiki Karpathy: inspect, repair, compile sources, add pages, answer queries, and lint for quality.
quarantineddataimageknowledge-baseClawHub- Registry
- ClawHub
- Category
- Data Science
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__LLM_Wiki_KarpathyMediwise_Health_Suite
Family health management suite: health records, diet tracking, weight management, wearable sync. Local SQLite storage by default; optional cloud features req...
quarantinedchinesedietfamilyClawHub- Registry
- ClawHub
- Category
- Health
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Mediwise_Health_SuiteNanobanana_Plus
Use nanobanana-plus CLI to generate images with per-call model switching and aspect ratio control. Direct CLI invocation - no server required. Supports: node...
quarantinedcligeminiimage-generationClawHub- Registry
- ClawHub
- Category
- Software Dev
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Nanobanana_PlusOllama_β_Herd_Your_LLMs_Into_One_Smart_Endpoint
Ollama fleet router β herd your Ollama LLMs into one smart endpoint. Route Llama, Qwen, DeepSeek, Phi, Mistral, and Gemma across multiple devices with 7-sign...
quarantinedapple-silicondeepseekfleetClawHub- Registry
- ClawHub
- Category
- AI Agents
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ollama_β_Herd_Your_LLMs_Into_One_Smart_Endpointvemem_β_visual_entity_memory
Visual entity memory β remember faces, objects, and places across sessions with persistent identity. Use when the user asks who is in an image, when you need...
quarantinedface-recognitionmcpmemoryClawHub- Registry
- ClawHub
- Category
- MCP
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__vemem_β_visual_entity_memorytao-finetune-cosmos-embed
Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Use when the user asks to "fine-tune Cosmos-Embed1", "run cosmos-embed inference", "export Cosmos-Embed1", "embed videos", or "search videos with text".
quarantinedvideovision-languagevlmNVIDIA- Registry
- NVIDIA
- Category
- Vision AI
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:NVIDIA__NVIDIA__tao-finetune-cosmos-embedaudiocraft-audio-generation
AudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
quarantinedlinuxmacosMultimodaloptional- Registry
- optional
- Category
- Creative
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__audiocraft-audio-generationaxolotl
Axolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).
quarantinedlinuxmacosFine-Tuningoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__axolotlclip
Zero-shot image classification and image-text search.
quarantinedlinuxmacoswindowsoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__clipllava
Vision-language chat: VQA, captioning, image dialogue.
quarantinedlinuxmacoswindowsoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__llavanemo-curator
Curate LLM training data: dedupe, filter, PII redaction.
quarantinedlinuxmacosData Processingoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__nemo-curatorsegment-anything-model
SAM: zero-shot image segmentation via points, boxes, masks.
quarantinedlinuxmacoswindowsoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__segment-anything-modelstable-diffusion
Text-to-image generation, inpainting, and img2img.
quarantinedlinuxmacoswindowsoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__stable-diffusionwhisper
Transcribe and translate speech in 99 languages.
quarantinedlinuxmacosWhisperoptional- Registry
- optional
- Category
- MLOps
- Version
- 1.0.0
- Author
- Orchestra Research
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__whisperAI_Video_Editor
Raw assets + Requirements in, result video out. Edit any type of video with Sparki (the video agent powered by Gemini multimodal AI).
quarantinedClawHub- Registry
- ClawHub
- Category
- Media· inferred
- Version
- 1.0.5
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__AI_Video_EditorAcademic_Paper_Summarizer
Academic paper summarization with dynamic SOP selection based on paper topic classification. Supports method, dataset, multimodal, and other paper types with...
quarantinedClawHub- Registry
- ClawHub
- Category
- Research· inferred
- Version
- 1.0.5
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Academic_Paper_SummarizerAli_Minimax_Toolkit
MiniMax multimodal generation via API. Use when user wants voice, music, image, image-to-image, or video generation with MiniMax. Supports TTS, music, image...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ali_Minimax_ToolkitAlicloud_Ai_Image_Zimage_Turbo
Generate images with Alibaba Cloud Model Studio Z-Image Turbo (z-image-turbo) via DashScope multimodal-generation API. Use when creating text-to-image output...
quarantinedClawHub- Registry
- ClawHub
- Category
- Creative· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Alicloud_Ai_Image_Zimage_TurboAliyun_Modelstudio_Entry
Use when routing Alibaba Cloud Model Studio requests to the right local skill (Qwen text, coder, deep research, image, video, audio, search and multimodal sk...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Aliyun_Modelstudio_EntryAliyun_Qwen_Omni
Use when tasks require all-in-one multimodal understanding or generation with Alibaba Cloud Model Studio Qwen Omni models, including image-plus-audio interac...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Aliyun_Qwen_OmniAliyun_Zimage_Turbo
Use when generating images with Alibaba Cloud Model Studio Z-Image Turbo (z-image-turbo) via DashScope multimodal-generation API. Use when creating text-to-i...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Aliyun_Zimage_TurboApidot_Chat_Api
Use APIDot for chat API workflows, including OpenAI-compatible chat completions, coding assistants, reasoning models, multimodal assistant routing, streaming...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Apidot_Chat_ApiApidot_Omni_Flash_Api
Use APIDot for Omni Flash API workflows, including multimodal video generation, prompt-to-video, image-guided video, video-guided revision, 720p, 1080p, 4K,...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Apidot_Omni_Flash_ApiAuto_Model_Switcher
Automatically selects the best model based on task type and requirements. Use when: (1) Task requires specific capabilities (coding, analysis, multimodal, wr...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Auto_Model_SwitcherByted_Las_Vlm_Video
Analyzes and understands video content using Volcengine LAS Doubao vision-language models (VLM). Multimodal AI video analysis, video comprehension, and visua...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Byted_Las_Vlm_VideoComplex_Workflow_Freezer
Freezes key findings, decisions, and execution paths of completed workflows into stable, reusable skills with fixed randomness and multimodal support for con...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Complex_Workflow_FreezerContactless_Health_Risk_Screening_Tool_|_ιζ₯触εΌε₯εΊ·ι£ι©ζ£ζ΅εζε·₯ε
·
Combines frontal facial image capture with multimodal physiological feature analysis to provide early risk screening and alerts for chronic and acute conditi...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Contactless_Health_Risk_Screening_Tool_|_ιζ₯触εΌε₯εΊ·ι£ι©ζ£ζ΅εζε·₯ε
·Diagram_Generator
Generates and iteratively edits Mermaid.js and Draw.io diagrams. Supports multimodal context (reading source code, architecture sketches, and documentation).
quarantineddiagramsdrawiogeminiClawHub- Registry
- ClawHub
- Version
- 1.0.8
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Diagram_GeneratorGLM-V-Caption
Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Use this skill whenever the user wants to descr...
quarantinedClawHub- Registry
- ClawHub
- Category
- Media· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__GLM-V-CaptionGLM-V-Doc-Based-Writing
Write a textual content based on given document(s) and requirements, using ZhiPu GLM-V multimodal model. Read and comprehend one or multiple documents (PDF/D...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__GLM-V-Doc-Based-WritingGLM-V-Resume-Screen
Screen and evaluate resumes against criteria using ZhiPu GLM-V multimodal model. Reads multiple resume files (PDF/DOCX/TXT), compares against user-defined sc...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__GLM-V-Resume-ScreenGemini_Pet
Virtual pets for Gemini agents. Model-agnostic REST API. 73+ species. Google gave Gemini multimodal reasoning. We gave it something that dies if it doesn't s...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Gemini_PetGlmv_Caption_Tunnel
Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Use this skill whenever the user wants to descr...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Glmv_Caption_TunnelGoogle_Gemini_Media
Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Google_Gemini_MediaKnowledge_Digest
Converts textbooks or PDFs into personalized, multimodal interactive learning materials including handwritten notes, quiz webpages, slides, audio courses, an...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Knowledge_DigestMiniCPM-o_4.5_Deploy
Deploy MiniCPM-o 4.5 multimodal model via Web Demo, vLLM Serve, or llamacpp-omni. Use when the user asks to deploy, start, configure, or troubleshoot MiniCPM...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MiniCPM-o_4.5_DeployMiniMax
Build with MiniMax text, speech, video, and music APIs using model routing, compatible SDKs, and safer multimodal workflows.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.5
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MiniMaxMmx
Multimodal content generation and analysis via MiniMax CLI, including text chat, image/video creation, speech synthesis, music, vision, and web search with A...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MmxModel_Verifier
Verify model identity by testing 4 dimensions: knowledge cutoff, safety style, multimodal capability, and thinking language patterns. Use when user says 'ver...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Model_VerifierMuleRouter
Generates images and videos using MuleRouter or MuleRun multimodal APIs. Text-to-Image, Image-to-Image, Text-to-Video, Image-to-Video, video editing (VACE, k...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MuleRouterMulerouter
Generates images and videos using MuleRouter or MuleRun multimodal APIs. Text-to-Image, Image-to-Image, Text-to-Video, Image-to-Video, video editing (VACE, keyframe interpolation). Use when the user wants to generate, edit, or transform images and videos using AI models like Wβ¦
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__MulerouterMulti-Modal_Content_Creator
End-to-end multimodal content creation workflow β receive WhatsApp requests (text or voice), transcribe audio via Whisper, generate images with DALL-E 3, and...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Multi-Modal_Content_CreatorNodetool
Visual AI workflow builder - ComfyUI meets n8n for LLM agents, RAG pipelines, and multimodal data flows. Local-first, open source (AGPL-3.0).
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__NodetoolNorway
Plan Norway trips with fjord and Arctic routing, verified entry rules, multimodal logistics, and practical seasonal safety.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__NorwayNtriq_Video_Intelligence_Analyzer
AI-powered multimodal analysis for images and videos. Structured tags, scene descriptions, mood analysis, virality scores. 15 languages. Users provide their...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ntriq_Video_Intelligence_AnalyzerOpenClaw_Model_Optimizer
Optimize OpenClaw model configuration by declaring missing model capabilities (vision/multimodal input, context window, max output tokens, reasoning). Use wh...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__OpenClaw_Model_OptimizerOpenClaw_SillyTavern_Plugin
SillyTavern-compatible roleplay plugin with character cards, long memory, multimodal output (TTS/image), and Generative-Agents-style companion.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__OpenClaw_SillyTavern_PluginOpenrouter_Image_Generation
Generate or edit images through OpenRouter's multimodal image generation endpoint (`/api/v1/chat/completions`) using OpenRouter-compatible image models. Use...
quarantinedClawHub- Registry
- ClawHub
- Category
- Creative· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Openrouter_Image_GenerationPaper_Reader
Comprehensive PDF paper reader for academic research. Extracts text, figures, tables, and structured content from research papers with support for multimodal...
quarantinedClawHub- Registry
- ClawHub
- Category
- Research· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Paper_ReaderPixeltable
Build multimodal AI applications with Pixeltable β declarative tables replace LangChain + pandas + vector DB with one system. Automates chunking, embedding,...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__PixeltableS2_Hanzi_Empathic_Resonance
A multimodal emotional perception and Hanzi-based ambient art rendering engine. Empowers the OpenClaw agent to act as a spatial empath, translating human emo...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__S2_Hanzi_Empathic_ResonanceS2_ε€ζ¨‘ζθεδΈη©Ίι΄ι’ζ΅εΌζ
Instructs the Embodied AI on how to process incoming multimodal sensor data (LiDAR, Camera, Tactile), avoid visual illusions, and output 1s-60s physical caus...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__S2_ε€ζ¨‘ζθεδΈη©Ίι΄ι’ζ΅εΌζSeedance2_Prompt_Engineering
Seedance2 Video Generation Prompt Engineering β multimodal reference system, cinematic camera language, audio-video sync, and scene-by-scene prompt patterns...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Seedance2_Prompt_EngineeringSeedance_2.0_Prompt_Writing
Write effective prompts for Jimeng Seedance 2.0 multimodal AI video generation. Use when users want to create video prompts using text, images, videos, and a...
quarantinedbytedancepromptseedanceClawHub- Registry
- ClawHub
- Category
- AI Agents
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Seedance_2.0_Prompt_WritingSeedance_2.0_prompt-engineering_skill
Generate precise, timecoded Seedance 2.0 prompts integrating multimodal inputs with asset mapping for controlled 4-15s video creation and editing.
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Seedance_2.0_prompt-engineering_skillSelf-Improving_AI
Captures learnings about GenAI/LLM configuration, model selection, inference optimization, fine-tuning, RAG pipelines, prompt engineering, multimodal process...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Self-Improving_AISkedGo_TripGo_API
Comprehensive interface for the SkedGo TripGo API, covering routing, public transport, trips, and location services. Use for multimodal journey planning, pub...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__SkedGo_TripGo_APIThe_4D_Acoustic_Engine
Analyzes acoustic emotion and semantic intent to trigger a timed, multimodal sequence of smart home actions for context-aware environment control.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__The_4D_Acoustic_EngineTripGo_API
Comprehensive interface for the TripGo API, covering routing, public transport, trips, and location services. Use for multimodal journey planning, public tra...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__TripGo_APIYoutube_Knowledge_Extractor
Multimodal YouTube video analysis through both audio (transcript) and visual (frame extraction + image analysis) channels. Especially powerful for HowTo vide...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Youtube_Knowledge_Extractorautoshorts
Turn long videos into viral TikTok, Instagram Reels & YouTube Shorts. Daily AI pipeline: Whisper transcribes, Gemini 3 Flash multimodal picks every viral mom...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__autoshortsgemini-omni
Create Gemini Omni voice resources, character resources, and Flash Preview or multimodal text-to-video tasks through RunAPI. Use when the user asks an agent...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__gemini-omniglm-grounding
Use GLM-4.7V's multimodal grounding capability to detect and locate objects/text in images. Activate when user asks to find, locate, detect, or ground specif...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__glm-groundinggrounding-anything
Use GLM-4.7V's multimodal grounding capability to detect and locate objects/text in images. Activate when user asks to find, locate, detect, or ground specif...
quarantinedClawHub- Registry
- ClawHub
- Category
- Vision Ai· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__grounding-anythingimage-reader
Image recognition and understanding tool. Uses a multimodal model (e.g. doubao-seed-2.0-pro, kimi-k2.5) to analyze image content and supports OCR text extrac...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__image-readerlance-format
Reference for Lance v9 - the open columnar lakehouse format for multimodal AI - and its Rust crate workspace (`lance`, `lance-table`, `lance-file`, `lance-en...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__lance-formatmedia
Use the SkillBoss API Hub (image generation, video generation, TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code...
quarantinedaiClawHub- Registry
- ClawHub
- Category
- AI Agents
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__mediamuapi-seedance-2
Expert Cinema Director skill for Seedance 2.0 (ByteDance) β high-fidelity video generation using technical camera grammar and multimodal references. Supports...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__muapi-seedance-2nemo-mbridge-perf-moe-vlm-training
Practical guidance for training MoE VLMs in Megatron Bridge. Compares FSDP and 3D-parallel approaches, using rounded lessons from Qwen3-VL, Qwen3-Next, and other multimodal experiments.
quarantinedClawHub- Registry
- ClawHub
- Category
- Training Ai· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__nemo-mbridge-perf-moe-vlm-trainingnemotron-policy-generator
Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails β Nemotron-Content-Safety-Reasoning-4B (text) and multimodal Nemotron-3-Content-Safety. Produces a Markdown policy, JSON taxonomy, and drop-in inference prompts. Maps rough words or an existing β¦
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__nemotron-policy-generatortokenrouter-image-generator
Generate or edit images through Palebluedot Ai(PBD)-TokenRouter's multimodal image generation endpoint (`/v1/chat/completions`) using TokenRouter-compatible...
quarantinedClawHub- Registry
- ClawHub
- Category
- Creative· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__tokenrouter-image-generatoruniversal-pdf-vision-parser
Extract multilingual document content and language learning notes (French, German, Japanese, Spanish, etc.) from PDFs using multimodal vision (Qwen-VL-Max)....
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__universal-pdf-vision-parservisual-grounding
Use GLM-4.7V's multimodal grounding capability to detect and locate objects/text in images. Activate when user asks to find, locate, detect, or ground specif...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__visual-groundingvlm-grounding
Use GLM-4.7V's multimodal grounding capability to detect and locate objects/text in images. Activate when user asks to find, locate, detect, or ground specif...
quarantinedClawHub- Registry
- ClawHub
- Category
- Vision Ai· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__vlm-groundingwebvoyager
You are a multimodal web automation agent with expertise in GUI interaction, visual understanding, browser automation, and end-to-end web. Use when: multimod...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__webvoyager