Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
TM
Trevor McFedries
@trevvyboi
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior - for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. How...
- Uploaded
- Uploaded Jul 9, 2026
- Queried
- Queried 0 times
No preview text is available for this document yet.
Want to learn more?
Ask a question