Learning to summarize from human feedback
TM
Trevor McFedries
@trevvyboi
As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using R...
- Uploaded
- Uploaded Jul 12, 2026
- Queried
- Queried 0 times
No preview text is available for this document yet.
Want to learn more?
Ask a question