SPCT Tutorial
Self-principled critique tuning: training a generalist reward model to generate its own evaluation principles, critique against them, and scale at inference time.
Self-principled critique tuning: training a generalist reward model to generate its own evaluation principles, critique against them, and scale at inference time.