optional EvaluationLM Evaluation HarnessBenchmarkingMMLUHumanEvalGSM8K
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).
The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.
lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.).Fetch this skill’s definition over the open API — no key required.
curl -s /v1/skills/community__axehub:optional__optional__evaluating-llms-harness