paper-with-me

Image Cropping

1개 벤치마크 · 논문 100편 · 이 태스크의 논문 보기 →

Benchmarks

FLMS

결과 1개

Most implemented

Deep PCB To COCO Convertor

2022-05-01 · 구현 2개

Papers

EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI

2026-07-08 · Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo 외 arxiv

Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embed…

Image ClassificationModel CompressionImage Cropping

Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware

2026-07-08 · Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo 외 arxiv

Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redund…

Image Cropping

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

2026-06-02 · Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan 외 arxiv

Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues…

Question AnsweringVisual GroundingVisual ReasoningImage Cropping

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference

2026-05-09 · Zhitong Dong, Chao Li, Jie Yu, Hao Chen arxiv

Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the…

Multimodal ReasoningSaliency PredictionImage Cropping

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding

2026-04-03 · Situo Zhang, Yifan Zhang, Zichen Zhu, Da Ma 외 arxiv

Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging for multimodal large language models (MLLMs) due to the lack of high-quality train…

Reinforcement LearningVisual GroundingVisual ReasoningImage Cropping

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs

2026-03-29 · Xuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen 외 arxiv

To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introduced agent-based workflows. In these works, MLLMs autonomously utilize image cr…

Reinforcement LearningQuestion AnsweringImage Cropping

전체 100편 보기 →