On-Policy Distillation — Dense Token-Level RL with Teacher Logprobs
Lightbulb Partners implements on-policy distillation using dense token-level feedback from teacher model logprobs during RL training, accelerating policy learning with richer supervision signals.