Sparrow --- Improving Alignment of Dialogue Agents via Targeted Human Judgements
A dialogue agent that uses rule-based reward models and targeted human judgements to be more helpful, correct, and harmless, introducing a scalable approach to alignment through decomposed evaluation criteria.