Image Cropping
1개 벤치마크 · 논문 100편 · 이 태스크의 논문 보기 →
Benchmarks
FLMS
Most implemented
Kornia: an Open Source Differentiable Computer Vision Library for PyTorch
See Better Before Looking Closer: Weakly Supervised Data Augmentation Network for Fine-Grained Visual Classification
A2-RL: Aesthetics Aware Reinforcement Learning for Image Cropping
FiT: Flexible Vision Transformer for Diffusion Model
Resolution Enhancement Processing on Low Quality Images Using Swin Transformer Based on Interval Dense Connection Strategy
Deep PCB To COCO Convertor
Papers
EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI
Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embed…
Image ClassificationModel CompressionImage CroppingSmart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware
Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. However, as an image usually contains much spatial redund…
Image CroppingVistaHop: Benchmarking Long-Horizon Visual DeepSearch
Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues…
Question AnsweringVisual GroundingVisual ReasoningImage CroppingCROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference
Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the…
Multimodal ReasoningSaliency PredictionImage CroppingCharTool: Tool-Integrated Visual Reasoning for Chart Understanding
Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging for multimodal large language models (MLLMs) due to the lack of high-quality train…
Reinforcement LearningVisual GroundingVisual ReasoningImage CroppingLearning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introduced agent-based workflows. In these works, MLLMs autonomously utilize image cr…
Reinforcement LearningQuestion AnsweringImage Cropping