Multi-object event graph representation learning for Video Question Answering

Wang, Yanan; Haruta, Shuichiro; Zeng, Donghuo; Vizcarra, Julio; Kurokawa, Mori

Computer Science > Computer Vision and Pattern Recognition

arXiv:2409.07747 (cs)

[Submitted on 12 Sep 2024]

Title:Multi-object event graph representation learning for Video Question Answering

Authors:Yanan Wang, Shuichiro Haruta, Donghuo Zeng, Julio Vizcarra, Mori Kurokawa

View PDF HTML (experimental)

Abstract:Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships among objects extracted from videos to perform causal and temporal reasoning. While prior works have focused on modeling individual object movements using transformer-based methods, they falter when capturing complex scenarios involving multiple objects (e.g., "a boy is throwing a ball in a hoop"). We propose a contrastive language event graph representation learning method called CLanG to address this limitation. Aiming to capture event representations associated with multiple objects, our method employs a multi-layer GNN-cluster module for adversarial graph representation learning, enabling contrastive learning between the question text and its relevant multi-object event graph. Our method outperforms a strong baseline, achieving up to 2.2% higher accuracy on two challenging VideoQA datasets, NExT-QA and TGIF-QA-R. In particular, it is 2.8% better than baselines in handling causal and temporal questions, highlighting its strength in reasoning multiple object-based events.

Comments:	presented at MIRU2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2409.07747 [cs.CV]
	(or arXiv:2409.07747v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2409.07747

Submission history

From: Yanan Wang [view email]
[v1] Thu, 12 Sep 2024 04:42:51 UTC (7,344 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-object event graph representation learning for Video Question Answering

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Multi-object event graph representation learning for Video Question Answering

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators