FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

Liu, Yuanxin; Li, Lei; Ren, Shuhuai; Gao, Rundong; Li, Shicheng; Chen, Sishuo; Sun, Xu; Hou, Lu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2311.01813 (cs)

[Submitted on 3 Nov 2023 (v1), last revised 26 Dec 2023 (this version, v3)]

Title:FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

Authors:Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, Lu Hou

View PDF HTML (experimental)

Abstract:Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the automatic evaluation metrics are consistent with human standards. To address these problems, we propose FETV, a benchmark for Fine-grained Evaluation of Text-to-Video generation. FETV is multi-aspect, categorizing the prompts based on three orthogonal aspects: the major content, the attributes to control and the prompt complexity. FETV is also temporal-aware, which introduces several temporal categories tailored for video generation. Based on FETV, we conduct comprehensive manual evaluations of four representative T2V models, revealing their pros and cons on different categories of prompts from different aspects. We also extend FETV as a testbed to evaluate the reliability of automatic T2V metrics. The multi-aspect categorization of FETV enables fine-grained analysis of the metrics' reliability in different scenarios. We find that existing automatic metrics (e.g., CLIPScore and FVD) correlate poorly with human evaluation. To address this problem, we explore several solutions to improve CLIPScore and FVD, and develop two automatic metrics that exhibit significant higher correlation with humans than existing metrics. Benchmark page: this https URL.

Comments:	NeurIPS 2023 Datasets and Benchmarks Track
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2311.01813 [cs.CV]
	(or arXiv:2311.01813v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2311.01813

Submission history

From: Yuanxin Liu [view email]
[v1] Fri, 3 Nov 2023 09:46:05 UTC (10,378 KB)
[v2] Wed, 8 Nov 2023 11:53:01 UTC (10,252 KB)
[v3] Tue, 26 Dec 2023 05:27:46 UTC (10,252 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators