AXe Skills HubSearch /

← All skills

trl-fine-tuning

AXe optional Post-TrainingTRLReinforcement LearningFine-TuningSFTDPO

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

Definition

The catalogue holds this skill’s description, not a full SKILL.md — no upstream address was recorded at ingestion, so the body cannot be fetched.

TRL: SFT, DPO, GRPO, RLOO reward modeling for LLM RLHF.

Metadata

Category
MLOps
Tier
optional
Version
1.0.0
License
MIT

Use with an agent

Fetch this skill’s definition over the open API — no key required.

curl -s /v1/skills/community__axehub:optional__optional__trl-fine-tuning

View source ↗