LVSum is a new benchmark for evaluating timestamp-aware long video summarization in MLLMs. The human-annotated dataset includes 72 diverse videos across 13 domains, averaging 16 minutes in length with fine-grained temporal references. Developers working on multimodal video understanding can use this benchmark to test how well their models maintain temporal fidelity and semantic grounding over extended durations.
Opening Kapyn…