Google DeepMind is testing double-blind, cryptographically secured AI benchmark evaluations. The pilot with Singapore's AI Safety Institute uses Confidential Space to hide test questions from Google and model weights from evaluators. It could set a new standard for tamper-proof model evaluation.
Opening Kapyn…