GRPO & RLVR Training This skill assumes finetuning-method-selection already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations ( lora-qlora-recipes ) or preference pairs ( preference-optimization ). What follows is when RL is the right tool, the reference recipe, the mandatory reward-inspection gate, and how to pick a GRPO variant when the base recipe misbehaves. Show more Installs 949 Repository wshobson/agents GitHub Stars 39.1K First Seen Jul 14, 2026 Security Audits Gen Agent Trust Hub Warn Socket Pass Snyk Pass
grpo-rlvr-training
安装
npx skills add https://github.com/wshobson/agents --skill grpo-rlvr-training