상세 보기
A photo cartoonization method based on text-to-image diffusion model
- Jeon, Hwyjoon;
- Shim, Jonghwa;
- Kim, Hyeonwoo;
- Hwang, Eenjun
WEB OF SCIENCE
6SCOPUS
9초록
In modern animation, background scenery such as complex buildings or elaborate structures requires a lot of time and effort to achieve a sense of realism and visual immersion. Various deep learning-based image cartoonization methods have been proposed to reduce these costs by transforming real images into high-quality animated scenes. However, these methods tend to oversimplify high-frequency patterns, producing flat and unrealistic images. Recent diffusion models trained using animated image-text datasets have shown good performance in generating high-quality animated images, but they cannot directly generate animated images from real images. In this paper, we propose a novel diffusion-based image-to-image cartoonization method using three lightweight adapters: cartoon-style adapter, color-structure adapter and semantic adapter. The cartoon-style adapter allows a single model to generate images in a variety of artistic styles. The color-structure adapter ensures that the overall shapes and color tones of the input image are preserved in the cartoonized image, while the semantic adapter ensures that semantic information of the input image for texture and detailed features are preserved. Through extensive experiments on animation backgrounds and real landscape datasets, we show that the proposed method can improve the FID score by up to 38 % and the CLIP-I score by up to 42 % compared to existing cartoonization methods.
키워드
- 제목
- A photo cartoonization method based on text-to-image diffusion model
- 저자
- Jeon, Hwyjoon; Shim, Jonghwa; Kim, Hyeonwoo; Hwang, Eenjun
- 발행일
- 2025-03-01
- 유형
- Article
- 저널명
- Neurocomputing
- 권
- 620