ClawHub
How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-r…
The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.
How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-r…Fetch this skill’s definition over the open API — no key required.
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__mcore-run-on-slurm