Skip to yearly menu bar Skip to main content


Probes as Training Signals: Iterative Retraining Overcomes Evasion and Reverses Reward Hacking

Lily Shi ⋅ Peter Hase

Abstract

Chat is not available.