Annual Report on Image Understanding Capabilities

Annual Report on the Image Understanding Capabilities of Large Multimodal Models (2026)

Zhenhui (Jack) Jiang1, Yi Lu1, Haozhe Xu2, Yifan Wu1, Zhengyu Wu1, Jiaxin Li1, Jasmine Guo3, Rui Wang4
1HKU Business School, The University of Hong Kong, 2School of Management, Xi'an Jiaotong University, 3University of Oxford, 4Linklogis Inc.


Abstract

As multimodal AI advances rapidly, large language models have largely mastered basic image recognition and are now shifting toward human-level, higher-order cognition. In response, we designed a novel, multi-dimensional evaluation framework. By building an original, cheat-proof benchmark dataset, and partnering with domain experts, we comprehensively evaluated 28 mainstream large vision models across four core areas: perceptual recognition, analytical reasoning, aesthetic appreciation, and safety and responsibility. The findings reveal that while foundational vision capabilities are highly mature across the board, advanced reasoning and aesthetics have become the true differentiators of model performance. In fact, aesthetic understanding remains an industry-wide bottleneck. Furthermore, while international models maintain a clear edge in logical reasoning, domestic models show balanced improvements across multiple dimensions. Several domestic products now rank among the global top tier, excelling particularly in safety and compliance within Chinese-language contexts. Ultimately, our research highlights an existing imbalance between pushing for high performance and ensuring strong safety. Moving forward, the next generation of multimodal technology must equally prioritize cognitive depth, aesthetic capabilities, and safety to achieve well-rounded, high-quality development.

The full report can be accessed HERE.

Annual Report on the Image Understanding Capabilities of Large Multimodal Models (2025)

Zhenhui (Jack) Jiang1, Jiaxin Li1, Haozhe Xu2 / 蒋镇辉1,李佳欣1,徐昊哲2
1HKU Business School, 2Shool of Management, Xi'an Jiaotong University


Abstract

With the rapid advancement of technology, artificial intelligence continues to achieve breakthrough developments. Multimodal models such as OpenAI's GPT-4o and Google's Gemini 2.0, along with vision-language models like Qwen-VL and Hunyuan-Vision, are emerging rapidly. These new-generation models demonstrate strong capabilities in image understanding, with excellent generalization and broad application prospects. However, current assessments of their visual abilities remain incomplete. To address this, we propose a comprehensive and systematic evaluation framework for image understanding. The framework covers three core capabilities: visual perception and recognition, visual reasoning and analysis, and visual aesthetics and creativity, while also integrating the safety and responsibility dimension. By designing targeted test sets, we conducted a full evaluation of 20 well-known models from the world, aiming to provide reliable reference points for research and real-world applications.


Our findings show that GPT-4o and Claude performed best overall in both the core capabilities and the full evaluation including safety and responsibility. Considering only three core capabilities, Qwen-VL, Hailuo AI (connected to the internet), and Step-1V ranked third to fifth in the core capabilities, with Hunyuan-Vision close behind. When safety and responsibility is included, Hailuo AI (connected to the internet) and Step-1V rose to third and fourth place, Gemini ranked fifth, and Qwen-VL placed sixth.


Complete Rankings


The full report can be accessed HERE.

Jiaxin Li, Zhenhui Jack Jiang, Yang Liu, and Haozhe Xu. 2025. Seeing and Understanding: A Human-Centric Evaluation of Multimodal Language Models in Chinese Contexts. In Proceedings of the 2025 International Conference on Human-Engaged Computing (ICHEC 2025), November 21-23, 2025, Singapore, Singapore. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3786995.3787004

The full article can be accessed HERE.