ยท 4 days agoยท AI Alignment Forum
Evaluating Cooperation Holds Promise for Taming Evaluative AI Gaming
Behavioral evaluations may become worthless, which we think would be a disaster. Smart misaligned models may realize they are being evaluated ("eval awareness") and then act to look good to us so we don't realize they're misaligned ("eval gaming"). We think increasing eval cooperativeness might be a
#ai