DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

Authors

  • Zhuowei Chen University of Science and Technology of China ByteDance
  • Shancheng Fang University of Science and Technology of China
  • Wei Liu ByteDance
  • Qian He ByteDance
  • Mengqi Huang University of Science and Technology of China
  • Zhendong Mao University of Science and Technology of China

DOI:

https://doi.org/10.1609/aaai.v38i2.27891

Keywords:

CV: Computational Photography, Image & Video Synthesis, CV: Language and Vision

Abstract

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity and follow the text prompts simultaneously for conditioned input face images and texts. Despite existing encoder-based methods achieving high efficiency and decent face similarity, the generated image often fails to follow the textual prompts. To ease this editability issue, we present DreamIdentity, to learn edit-friendly and accurate face-identity representations in the word embedding space. Specifically, we propose self-augmented editability learning to enhance the editability for projected embedding, which is achieved by constructing paired generated celebrity's face and edited celebrity images for training, aiming at transferring mature editability of off-the-shelf text-to-image models in celebrity to unseen identities. Furthermore, we design a novel dedicated face-identity encoder to learn an accurate representation of human faces, which applies multi-scale ID-aware features followed by a multi-embedding projector to generate the pseudo words in the text embedding space directly. Extensive experiments show that our method can generate more text-coherent and ID-preserved images with negligible time overhead compared to the standard text-to-image generation process.

Downloads

Published

2024-03-24

How to Cite

Chen, Z., Fang, S., Liu, W., He, Q., Huang, M., & Mao, Z. (2024). DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation. Proceedings of the AAAI Conference on Artificial Intelligence, 38(2), 1281-1289. https://doi.org/10.1609/aaai.v38i2.27891

Issue

Section

AAAI Technical Track on Computer Vision I