ClawHub ai-systemsbenchmarkevaluationmadefmethodologymulti-dimensional
Designs a multi-dimensional evaluation framework for AI systems where single-score benchmarks lose information. Use when comparing experiments/agents across...
The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.
Designs a multi-dimensional evaluation framework for AI systems where single-score benchmarks lose information. Use when comparing experiments/agents across...Fetch this skill’s definition over the open API — no key required.
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Multi-Dim_Eval_Framework_Designer