paper-with-me

Papers

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

2026-02-27 · Tianxiang Du, Hulingxiao He, Yuxin Peng arxiv

The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We define this capability as aesthetic guidance (AG) -- an essential but largely underexplored domain in computational aesthetics. Existing multimodal large language models (MLLMs) primarily offer overly positive feedback, failing to identify issues or provide actionable guidance. Without AG capability, they cannot effectively identify distracting regions or optimize compositional balance, thus also struggling in aesthetic cropping, which aims to refine photo composition through reframing after capture. To address this, we introduce AesGuide, the first large-scale AG dataset and benchmark with 10,748 photos annotated with aesthetic scores, analyses, and guidance. Building upon it, we propose Venus, a two-stage framework that first empowers MLLMs with AG capability through progressively complex aesthetic questions and then activates their aesthetic cropping power via CoT-based rationales. Extensive experiments show that Venus substantially improves AG capability and achieves state-of-the-art (SOTA) performance in aesthetic cropping, enabling interpretable and interactive aesthetic refinement across both stages of photo creation. Code is available at https://github.com/PKU-ICST-MIPL/Venus_CVPR2026.

📄 PDF Abstract BibTeX arXiv:2602.23980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning

2025-03-19 · Yang Tan, Chen Liu, Jingyuan Gao, Banghao Wu 외

Natural language processing (NLP) has significantly influenced scientific domains beyond human language, including protein engineering, where pre-trained protein language models (PLMs) have demonstrated remarkable succes…

BenchmarkingLanguage ModelingLanguage ModellingRetrieval

Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding

2025-12-08 · Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng 외 arxiv

Vision-language models (VLMs) have demonstrated impressive multimodal comprehension capabilities and are being deployed in an increasing number of online video understanding applications. While recent efforts extensively…

Scene Segmentation

UI-Venus Technical Report: Building High-performance UI Agents with RFT

2025-08-14 · Zhangxuan Gu, Zhengwen Zeng, Zhenyu Xu, Xingran Zhou 외 arxiv

We present UI-Venus, a native UI agent that takes only screenshots as input based on a multimodal large language model. UI-Venus achieves SOTA performance on both UI grounding and navigation tasks using only several hund…

VENUS: Visual Editing with Noise Inversion Using Scene Graphs

2026-01-12 · Thanh-Nhan Vo, Trong-Thuan Nguyen, Tam V. Nguyen, Minh-Triet Tran arxiv

State-of-the-art text-based image editing models often struggle to balance background preservation with semantic consistency, frequently resulting either in the synthesis of entirely new images or in outputs that fail to…

Text-based Image Editing

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

2026-02-06 · Yichen Gong, Zhuohan Cai, Sunhao Dai, Yuqi Zhou 외 arxiv

Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of real-world mobile usage. To this end, we introduce VenusBench-Mobile, …