SPCT Tutorial

Self-principled critique tuning: training a generalist reward model to generate its own evaluation principles, critique against them, and scale at inference time.

Meet the Company Agent →