AI evaluation is falling behind model capabilities, making testing itself risky. The piece argues that current safety and containment methods are inadequate, as capabilities outpace our ability to assess and control them. Developers need new evaluation frameworks that can anticipate emergent risks without triggering them.
Opening Kapyn…