A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment
Deep convolutional neural networks (CNN) have recently been shown to generate promising results for aesthetics assessment. However, the performance of these deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be transformed via cropping, warping, or padding, which often alter image composition, reduce image resolution, or cause image distortion. Thus the aesthetics of the original images is impaired because of potential loss of fine grained details and holistic image layout. However, such fine grained details and holistic image layout is critical for evaluating an image's aesthetics. In this paper, we present an Adaptive Layout-Aware Multi-Patch Convolutional Neural Network (A-Lamp CNN) architecture for photo aesthetic assessment. This novel scheme is able to accept arbitrary sized images, and learn from both fined grained details and holistic image layout simultaneously. To enable training on these hybrid inputs, we extend the method by developing a dedicated double-subnet neural network structure, i.e. a Multi-Patch subnet and a Layout-Aware subnet. We further construct an aggregation layer to effectively combine the hybrid features from these two subnets. Extensive experiments on the large-scale aesthetics assessment benchmark (AVA) demonstrate significant performance improvement over the state-of-the-art in photo aesthetic assessment.
Code (0)
등록된 구현이 없습니다.
Tasks
Aesthetics Quality AssessmentSimilar Papers 제목 키워드 기반
LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
Document layout comprises both structural and visual (eg. font-sizes) information that is vital but often ignored by machine learning models. The few existing models which do use layout information only consider textual …
document understandingCLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
We present CLAMP-ViT, a data-free post-training quantization method for vision transformers (ViTs). We identify the limitations of recent techniques, notably their inability to leverage meaningful inter-patch relationshi…
Contrastive Learningobject-detectionObject DetectionQuantizationLaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
Large language models (LLMs) experience significant performance degradation when the input exceeds the pretraining context window, primarily due to the out-of-distribution (OOD) behavior of Rotary Position Embedding (RoP…
"MS-Patch-Clamp" or the Possibility of Mass Spectrometry Hybridization with Patch-Clamp Setups for Single Cell Metabolomics and Channelomics
In this projecting work we propose a mass spectrometric patch-clamp equipment with the capillary performing both a local potential registration at the cell membrane and the analyte suction simultaneously. This paper prov…
InformativenessLayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control…
Image GenerationLayout-to-Image GenerationObject