LLM & Text Generation2 min reading time

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

Apple Research Blog
Read full post
Researchers introduced DeepAmbigQA, a dataset of 3,600 multi-hop questions with half involving name ambiguity, to benchmark large language models' ability to provide complete answers. Tests show GPT-5 struggles with answer completeness, scoring low on exact matches for ambiguous and non-ambiguous questions. This highlights the challenge of developing QA systems that effectively resolve ambiguity and integrate multi-step reasoning.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

Harvey raises $550M more to develop AI tools for legal teams

SiliconANGLE