Key-point Guided Deformable Image Manipulation Using Diffusion Model

Oh, Seok-Hwan; Jung, Guil; Kim, Myeong-Gee; Kim, Sang-Yun; Kim, Young-Min; Lee, Hyeon-Jik; Kwon, Hyuk-Sool; Bae, Hyeon-Min

Computer Science > Computer Vision and Pattern Recognition

arXiv:2401.08178 (cs)

[Submitted on 16 Jan 2024 (v1), last revised 19 Mar 2024 (this version, v3)]

Title:Key-point Guided Deformable Image Manipulation Using Diffusion Model

Authors:Seok-Hwan Oh, Guil Jung, Myeong-Gee Kim, Sang-Yun Kim, Young-Min Kim, Hyeon-Jik Lee, Hyuk-Sool Kwon, Hyeon-Min Bae

View PDF HTML (experimental)

Abstract:In this paper, we introduce a Key-point-guided Diffusion probabilistic Model (KDM) that gains precise control over images by manipulating the object's key-point. We propose a two-stage generative model incorporating an optical flow map as an intermediate output. By doing so, a dense pixel-wise understanding of the semantic relation between the image and sparse key point is configured, leading to more realistic image generation. Additionally, the integration of optical flow helps regulate the inter-frame variance of sequential images, demonstrating an authentic sequential image generation. The KDM is evaluated with diverse key-point conditioned image synthesis tasks, including facial image generation, human pose synthesis, and echocardiography video prediction, demonstrating the KDM is proving consistency enhanced and photo-realistic images compared with state-of-the-art models.

Comments:	24 pages
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2401.08178 [cs.CV]
	(or arXiv:2401.08178v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2401.08178

Submission history

From: Guil Jung [view email]
[v1] Tue, 16 Jan 2024 07:51:00 UTC (39,448 KB)
[v2] Mon, 18 Mar 2024 03:15:39 UTC (1 KB) (withdrawn)
[v3] Tue, 19 Mar 2024 03:47:39 UTC (20,722 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Key-point Guided Deformable Image Manipulation Using Diffusion Model

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Key-point Guided Deformable Image Manipulation Using Diffusion Model

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators