A systematic reproduction of 2,200 ICML papers reveals how few ML findings hold up beyond original settings. The effort highlights recurring pitfalls in evaluation, baselines, and hyperparameter reporting. For AI developers, it underscores the need to verify benchmark claims before building on them.
Opening Kapyn…