TRL

A library for post-training transformer language models using reinforcement learning.

AI AgentsResearchUnknownOpen source

About TRL

TRL is a comprehensive library for post-training transformer language models using reinforcement learning. It supports various fine-tuning methods, including Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), and Direct Preference Optimization (DPO).

Description summarised by AI from the sources listed below.

Key features

  • Supports Supervised Fine-Tuning (SFT)
  • Supports various fine-tuning methods
  • Supervised Fine-Tuning (SFT)
  • Integrated with Hugging Face Transformers
  • Integrated with Hugging Face Transformers ecosystem
  • Supports Group Relative Policy Optimization (GRPO)
  • Group Relative Policy Optimization (GRPO)
  • Supports Direct Preference Optimization (DPO)
  • Scalable and efficient
  • Direct Preference Optimization (DPO)
  • Scalable across various hardware setups
  • Reward Modeling
  • Supports Kahneman-Tversky Optimization (KTO)
  • Supports RewardTrainer

Use cases

  • Post-training transformer language models
  • Fine-tuning language models

Pricing

Pricing model: Unknown — we have not been able to confirm pricing from the official website, so nothing is stated here.

Pricing from the tool's own website.