AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception

Read original: arXiv:2404.09624 - Published 7/25/2024 by Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan, Zhichao Duan, Pengfei Chen, Leida Li, Weisi Lin, Guangming Shi

AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception

Overview

Proposes a multi-modality foundation model called "AesExpert" for image aesthetics perception
Combines visual and textual information to provide natural language feedback and aesthetic critiques
Utilizes instruction tuning to adapt the model for specific tasks related to image aesthetics assessment

Plain English Explanation

The researchers present a new AI model called "AesExpert" that is designed to perceive and analyze the aesthetics of images. Instead of just looking at the visual elements of an image, AesExpert also incorporates textual information to provide more nuanced and natural language-based feedback and critiques.

This multi-modality approach allows the model to better understand the artistic and subjective aspects of aesthetics, rather than relying solely on objective visual features. AesExpert can then use this understanding to generate detailed feedback and suggestions for improving the aesthetic qualities of an image.

The researchers also use a technique called "instruction tuning" to further adapt the model for specific tasks related to assessing image aesthetics. This allows AesExpert to be customized for different applications, such as providing feedback to photographers or evaluating the aesthetics of product designs.

Overall, the goal of this research is to develop AI systems that can engage with and appreciate the nuances of visual art and design, rather than just analyzing them through rigid technical measures. By incorporating both visual and textual understanding, AesExpert represents a step towards more holistic and human-like perception of aesthetic qualities.

Technical Explanation

The researchers propose a multi-modality foundation model called "AesExpert" for image aesthetics perception. The model combines visual and textual information to provide natural language feedback and aesthetic critiques.

The visual component of AesExpert is a convolutional neural network (CNN) trained on large-scale image datasets. The textual component is a transformer-based language model, which allows the system to understand and generate natural language related to aesthetic assessment.

By jointly learning from visual and textual data, AesExpert develops a more comprehensive understanding of what makes an image aesthetically pleasing or not. The researchers use instruction tuning to further adapt the model for specific tasks, such as providing feedback on photographic composition or evaluating the aesthetics of product designs.

Through experiments on benchmark datasets, the researchers demonstrate that AesExpert outperforms existing state-of-the-art models for image aesthetics assessment. The model's ability to generate natural language feedback is also shown to be valuable for applications like design critique and photography education.

Critical Analysis

The researchers acknowledge several limitations and areas for future work with AesExpert. First, the model's performance is still dependent on the quality and quantity of the training data, particularly for the textual component. Expanding the breadth of aesthetic domains and cultural perspectives represented in the training data could further improve the model's versatility.

Additionally, while the instruction tuning approach allows AesExpert to be customized for specific tasks, the researchers note that this process can be labor-intensive. Developing more efficient methods for adapting the model to new applications would enhance its real-world practicality.

Another potential concern is the model's ability to truly capture the subjective and contextual nature of aesthetic judgments. While AesExpert demonstrates promising results, assessing the nuances of human aesthetic perception remains a significant challenge for AI systems.

Overall, the AesExpert research represents an important step towards more holistic and human-like approaches to image aesthetics evaluation. By combining visual and textual understanding, the model moves beyond simplistic technical measures and begins to engage with the richer, more contextual aspects of artistic and design appreciation.

Conclusion

The AesExpert research proposes a novel multi-modality foundation model for image aesthetics perception. By integrating visual and textual information, the model can provide natural language feedback and aesthetic critiques that go beyond just analyzing technical visual features.

This work demonstrates the potential for AI systems to engage more deeply with the subjective and contextual nature of aesthetic judgments. While there are still limitations and areas for further development, AesExpert represents an important step towards building AI systems that can appreciate the nuances of visual art and design in a more human-like way.

As the field of artificial intelligence continues to advance, models like AesExpert may play a crucial role in enhancing our ability to understand, evaluate, and create aesthetically pleasing images and designs. This research contributes to the broader effort to develop AI that can truly engage with and appreciate the richness of human creativity and expression.

This summary was produced with help from an AI and may contain inaccuracies - check out the links to read the original source documents!

Follow @aimodelsfyi on 𝕏 →

Related Papers

AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception

Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan, Zhichao Duan, Pengfei Chen, Leida Li, Weisi Lin, Guangming Shi

The highly abstract nature of image aesthetics perception (IAP) poses significant challenge for current multimodal large language models (MLLMs). The lack of human-annotated multi-modality aesthetic data further exacerbates this dilemma, resulting in MLLMs falling short of aesthetics perception capabilities. To address the above challenge, we first introduce a comprehensively annotated Aesthetic Multi-Modality Instruction Tuning (AesMMIT) dataset, which serves as the footstone for building multi-modality aesthetics foundation models. Specifically, to align MLLMs with human aesthetics perception, we construct a corpus-rich aesthetic critique database with 21,904 diverse-sourced images and 88K human natural language feedbacks, which are collected via progressive questions, ranging from coarse-grained aesthetic grades to fine-grained aesthetic descriptions. To ensure that MLLMs can handle diverse queries, we further prompt GPT to refine the aesthetic critiques and assemble the large-scale aesthetic instruction tuning dataset, i.e. AesMMIT, which consists of 409K multi-typed instructions to activate stronger aesthetic capabilities. Based on the AesMMIT database, we fine-tune the open-sourced general foundation models, achieving multi-modality Aesthetic Expert models, dubbed AesExpert. Extensive experiments demonstrate that the proposed AesExpert models deliver significantly better aesthetic perception performances than the state-of-the-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision. Project homepage: https://yipoh.github.io/aes-expert/.

7/25/2024

Multi-modal Learnable Queries for Image Aesthetics Assessment

Zhiwei Xiong, Yunfan Zhang, Zhiqi Shen, Peiran Ren, Han Yu

Image aesthetics assessment (IAA) is attracting wide interest with the prevalence of social media. The problem is challenging due to its subjective and ambiguous nature. Instead of directly extracting aesthetic features solely from the image, user comments associated with an image could potentially provide complementary knowledge that is useful for IAA. With existing large-scale pre-trained models demonstrating strong capabilities in extracting high-quality transferable visual and textual features, learnable queries are shown to be effective in extracting useful features from the pre-trained visual features. Therefore, in this paper, we propose MMLQ, which utilizes multi-modal learnable queries to extract aesthetics-related features from multi-modal pre-trained features. Extensive experimental results demonstrate that MMLQ achieves new state-of-the-art performance on multi-modal IAA, beating previous methods by 7.7% and 8.3% in terms of SRCC and PLCC, respectively.

5/3/2024

UNIAA: A Unified Multi-modal Image Aesthetic Assessment Baseline and Benchmark

Zhaokun Zhou, Qiulin Wang, Bin Lin, Yiwei Su, Rui Chen, Xin Tao, Amin Zheng, Li Yuan, Pengfei Wan, Di Zhang

As an alternative to expensive expert evaluation, Image Aesthetic Assessment (IAA) stands out as a crucial task in computer vision. However, traditional IAA methods are typically constrained to a single data source or task, restricting the universality and broader application. In this work, to better align with human aesthetics, we propose a Unified Multi-modal Image Aesthetic Assessment (UNIAA) framework, including a Multi-modal Large Language Model (MLLM) named UNIAA-LLaVA and a comprehensive benchmark named UNIAA-Bench. We choose MLLMs with both visual perception and language ability for IAA and establish a low-cost paradigm for transforming the existing datasets into unified and high-quality visual instruction tuning data, from which the UNIAA-LLaVA is trained. To further evaluate the IAA capability of MLLMs, we construct the UNIAA-Bench, which consists of three aesthetic levels: Perception, Description, and Assessment. Extensive experiments validate the effectiveness and rationality of UNIAA. UNIAA-LLaVA achieves competitive performance on all levels of UNIAA-Bench, compared with existing MLLMs. Specifically, our model performs better than GPT-4V in aesthetic perception and even approaches the junior-level human. We find MLLMs have great potential in IAA, yet there remains plenty of room for further improvement. The UNIAA-LLaVA and UNIAA-Bench will be released.

4/16/2024

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

Tianhe Wu, Kede Ma, Jie Liang, Yujiu Yang, Lei Zhang

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality Assessment (IQA) remains largely unexplored. In this paper, we conduct a comprehensive and systematic study of prompting MLLMs for IQA. We first investigate nine prompting systems for MLLMs as the combinations of three standardized testing procedures in psychophysics (i.e., the single-stimulus, double-stimulus, and multiple-stimulus methods) and three popular prompting strategies in natural language processing (i.e., the standard, in-context, and chain-of-thought prompting). We then present a difficult sample selection procedure, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems. We assess three open-source and one closed-source MLLMs on several visual attributes of image quality (e.g., structural and textural distortions, geometric transformations, and color differences) in both full-reference and no-reference scenarios. Experimental results show that only the closed-source GPT-4V provides a reasonable account for human perception of image quality, but is weak at discriminating fine-grained quality variations (e.g., color differences) and at comparing visual quality of multiple images, tasks humans can perform effortlessly.

7/12/2024