RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

Zhong, Victor; Shi, Weijia; Yih, Wen-tau; Zettlemoyer, Luke

Computer Science > Computation and Language

arXiv:2210.14353v1 (cs)

[Submitted on 25 Oct 2022 (this version), latest version 15 Nov 2022 (v2)]

Title:RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

Authors:Victor Zhong, Weijia Shi, Wen-tau Yih, Luke Zettlemoyer

View PDF

Abstract:We introduce RoMQA, the first benchmark for robust, multi-evidence, multi-answer question answering (QA). RoMQA contains clusters of questions that are derived from related constraints mined from the Wikidata knowledge graph. RoMQA evaluates robustness of QA models to varying constraints by measuring worst-case performance within each question cluster. Compared to prior QA datasets, RoMQA has more human-written questions that require reasoning over more evidence text and have, on average, many more correct answers. In addition, human annotators rate RoMQA questions as more natural or likely to be asked by people. We evaluate state-of-the-art large language models in zero-shot, few-shot, and fine-tuning settings, and find that RoMQA is challenging: zero-shot and few-shot models perform similarly to naive baselines, while supervised retrieval methods perform well below gold evidence upper bounds. Moreover, existing models are not robust to variations in question constraints, but can be made more robust by tuning on clusters of related questions. Our results show that RoMQA is a challenging benchmark for large language models, and provides a quantifiable test to build more robust QA methods.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2210.14353 [cs.CL]
	(or arXiv:2210.14353v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2210.14353

Submission history

From: Victor Zhong [view email]
[v1] Tue, 25 Oct 2022 21:39:36 UTC (4,789 KB)
[v2] Tue, 15 Nov 2022 17:30:07 UTC (4,789 KB)

Computer Science > Computation and Language

Title:RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators