Viewpoint Equivariance for Multi-View 3D Object Detection

Chen, Dian; Li, Jie; Guizilini, Vitor; Ambrus, Rares; Gaidon, Adrien

Computer Science > Computer Vision and Pattern Recognition

arXiv:2303.14548 (cs)

[Submitted on 25 Mar 2023 (v1), last revised 7 Apr 2023 (this version, v2)]

Title:Viewpoint Equivariance for Multi-View 3D Object Detection

Authors:Dian Chen, Jie Li, Vitor Guizilini, Rares Ambrus, Adrien Gaidon

View PDF

Abstract:3D object detection from visual sensors is a cornerstone capability of robotic systems. State-of-the-art methods focus on reasoning and decoding object bounding boxes from multi-view camera input. In this work we gain intuition from the integral role of multi-view consistency in 3D scene understanding and geometric learning. To this end, we introduce VEDet, a novel 3D object detection framework that exploits 3D multi-view geometry to improve localization through viewpoint awareness and equivariance. VEDet leverages a query-based transformer architecture and encodes the 3D scene by augmenting image features with positional encodings from their 3D perspective geometry. We design view-conditioned queries at the output level, which enables the generation of multiple virtual frames during training to learn viewpoint equivariance by enforcing multi-view consistency. The multi-view geometry injected at the input level as positional encodings and regularized at the loss level provides rich geometric cues for 3D object detection, leading to state-of-the-art performance on the nuScenes benchmark. The code and model are made available at this https URL.

Comments:	11 pages, 4 figures; accepted to CVPR 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as:	arXiv:2303.14548 [cs.CV]
	(or arXiv:2303.14548v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2303.14548

Submission history

From: Dian Chen [view email]
[v1] Sat, 25 Mar 2023 19:56:41 UTC (1,370 KB)
[v2] Fri, 7 Apr 2023 04:59:08 UTC (1,370 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Viewpoint Equivariance for Multi-View 3D Object Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Viewpoint Equivariance for Multi-View 3D Object Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators