Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Duan, Yinglin; Shi, Tianyang; Zou, Zhengxia; Qin, Jia; Zhao, Yifei; Yuan, Yi; Hou, Jie; Wen, Xiang; Fan, Changjie

Computer Science > Computer Vision and Pattern Recognition

arXiv:2009.12763 (cs)

[Submitted on 27 Sep 2020]

Title:Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Authors:Yinglin Duan (1), Tianyang Shi (1), Zhengxia Zou (2), Jia Qin (1 and 3), Yifei Zhao (1), Yi Yuan (1), Jie Hou (1), Xiang Wen (1 and 3), Changjie Fan (1) ((1) NetEase Fuxi AI Lab, (2) University of Michigan, Ann Arbor, (3) Zhejiang University)

View PDF

Abstract:Music-to-dance translation is a brand-new and powerful feature in recent role-playing games. Players can now let their characters dance along with specified music clips and even generate fan-made dance videos. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods suffer from limited training data pairs and the degradation of movements. This paper provides a new perspective for this task where we re-formulate the translation problem as a piece-wise dance phrase retrieval problem based on the choreography theory. With such a design, players are allowed to further edit the dance movements on top of our generation while other regression based methods ignore such user interactivity. Considering that the dance motion capture is an expensive and time-consuming procedure which requires the assistance of professional dancers, we train our method under a semi-supervised learning framework with a large unlabeled dataset (20x than labeled data) collected. A co-ascent mechanism is introduced to improve the robustness of our network. Using this unlabeled dataset, we also introduce self-supervised pre-training so that the translator can understand the melody, rhythm, and other components of music phrases. We show that the pre-training significantly improves the translation accuracy than that of training from scratch. Experimental results suggest that our method not only generalizes well over various styles of music but also succeeds in expert-level choreography for game players.

Comments:	14 pages, 8 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:	arXiv:2009.12763 [cs.CV]
	(or arXiv:2009.12763v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2009.12763

Submission history

From: Tianyang Shi [view email]
[v1] Sun, 27 Sep 2020 07:08:04 UTC (1,811 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators