Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Chen, Muxi; Liu, Yi; Yi, Jian; Xu, Changran; Lai, Qiuxia; Wang, Hongliang; Ho, Tsung-Yi; Xu, Qiang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.05125 (cs)

[Submitted on 8 Mar 2024 (v1), last revised 28 Oct 2024 (this version, v2)]

Title:Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Authors:Muxi Chen, Yi Liu, Jian Yi, Changran Xu, Qiuxia Lai, Hongliang Wang, Tsung-Yi Ho, Qiang Xu

View PDF HTML (experimental)

Abstract:In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first, focusing on image qualities such as aesthetics and realism, and second, examining text conditions through concept coverage and fairness. We introduce an innovative aesthetic score prediction model that assesses the visual appeal of generated images and unveils the first dataset marked with low-quality regions in generated human images to facilitate automatic defect detection. Our exploration into concept coverage probes the model's effectiveness in interpreting and rendering text-based concepts accurately, while our analysis of fairness reveals biases in model outputs, with an emphasis on gender, race, and age. While our study is grounded in human imagery, this dual-faceted approach is designed with the flexibility to be applicable to other forms of image generation, enhancing our understanding of generative models and paving the way to the next generation of more sophisticated, contextually aware, and ethically attuned generative models. Code and data, including the dataset annotated with defective areas, are available at \href{this https URL}{this https URL}.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2403.05125 [cs.CV]
	(or arXiv:2403.05125v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.05125

Submission history

From: Muxi Chen [view email]
[v1] Fri, 8 Mar 2024 07:41:47 UTC (24,359 KB)
[v2] Mon, 28 Oct 2024 09:53:39 UTC (24,359 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators