kapynResearch

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

DeepAmbigQA is a new benchmark and generation pipeline evaluating LLM answer completeness on ambiguous multi-hop questions. The benchmark tests models on both disambiguation and multi-step evidence integration using complex queries with multiple valid answers. This dataset helps developers better measure how retrieval-augmented models handle ambiguity and complex reasoning in open-domain search tasks.

Apple ML Research·Aug 6, 2026

Opening Kapyn…