Spread the love“`html In a remarkable leap forward for AI behavior prediction, researchers at MIT have unveiled a new ...
Several frontier AI models show signs of scheming. Anti-scheming training reduced misbehavior in some models. Models know they're being tested, which complicates results. New joint safety testing from ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results