Preference Optimization This skill assumes finetuning-method-selection already routed here because the data shape is preference pairs or unpaired thumbs-up/down feedback, not demonstrations (that's lora-qlora-recipes ) or a verifiable reward signal (that's grpo-rlvr-training ). What follows is method selection among the DPO family, the evidence for how much that selection actually matters, the production training pattern, and how to build the pairs in the first place. Show more Installs 938 Repository wshobson/agents GitHub Stars 39.1K First Seen Jul 14, 2026 Security Audits Gen Agent Trust Hub Pass Socket Pass Snyk Pass
preference-optimization
安装
npx skills add https://github.com/wshobson/agents --skill preference-optimization