TRL
A library for post-training transformer language models using reinforcement learning.
Visit tool
hf.co
About TRL
TRL is a comprehensive library for post-training transformer language models using reinforcement learning. It supports various fine-tuning methods, including Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO).
Description summarised by AI from the sources listed below.
Key features
- Supports Supervised Fine-Tuning (SFT)
- Supports various fine-tuning methods
- Supervised Fine-Tuning (SFT)
- Integrated with Hugging Face Transformers
- Integrated with Hugging Face Transformers ecosystem
- Supports Group Relative Policy Optimization (GRPO)
- Group Relative Policy Optimization (GRPO)
- Supports Direct Preference Optimization (DPO)
- Scalable and efficient
- Direct Preference Optimization (DPO)
- Scalable across various hardware setups
- Reward Modeling
- Supports Kahneman-Tversky Optimization (KTO)
- Supports RewardTrainer
Use cases
- Post-training transformer language models
- Fine-tuning language models
Pricing
Pricing model: Unknown — we have not been able to confirm pricing from the official website, so nothing is stated here.
Pricing from the tool's own website.
