paper-with-me

홈 › Papers

CAPTAIN: Comprehensive Composition Assistance for Photo Taking

2018-11-10 · Farshid Farhat, Mohammad Mahdi Kamani, James Z. Wang

Many people are interested in taking astonishing photos and sharing with others. Emerging hightech hardware and software facilitate ubiquitousness and functionality of digital photography. Because composition matters in photography, researchers have leveraged some common composition techniques to assess the aesthetic quality of photos computationally. However, composition techniques developed by professionals are far more diverse than well-documented techniques can cover. We leverage the vast underexplored innovations in photography for computational composition assistance. We propose a comprehensive framework, named CAPTAIN (Composition Assistance for Photo Taking), containing integrated deep-learned semantic detectors, sub-genre categorization, artistic pose clustering, personalized aesthetics-based image retrieval, and style set matching. The framework is backed by a large dataset crawled from a photo-sharing Website with mostly photography enthusiasts and professionals. The work proposes a sequence of steps that have not been explored in the past by researchers. The work addresses personal preferences for composition through presenting a ranked-list of photographs to the user based on user-specified weights in the similarity measure. The matching algorithm recognizes the best shot among a sequence of shots with respect to the user's preferred style set. We have conducted a number of experiments on the newly proposed components and reported findings. A user study demonstrates that the work is useful to those taking photos.

📄 PDF Abstract BibTeX arXiv:1811.04184

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringImage RetrievalRetrievalset matching

Similar Papers 제목 키워드 기반

Solving the Perspective-2-Point Problem for Flying-Camera Photo Composition

2018-06-01 · CVPR 2018 6 · Ziquan Lan, David Hsu, Gim Hee Lee

Drone-mounted flying cameras will revolutionize photo-taking. The user, instead of holding a camera in hand and manually searching for a viewpoint, will interact directly with image contents in the viewfinder through …

Form

PhotoFramer: Multi-modal Image Composition Instruction

2025-11-30 · Zhiyuan You, Ke Wang, He Zhang, Xin Cai 외 arxiv

Composition matters during the photo-taking process, yet many casual users struggle to frame well-composed images. To provide composition guidance, we introduce PhotoFramer, a multi-modal composition instruction framewor…

Adaptive In-conversation Team Building for Language Model Agents

2024-05-29 · Linxin Song, Jiale Liu, Jieyu Zhang, Shaokun Zhang 외

Leveraging multiple large language model (LLM) agents has shown to be a promising approach for tackling complex tasks, while the effective design of multiple agents for a particular application remains an art. It is thus…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

Learning to Compose with Professional Photographs on the Web

2017-02-01 · Yi-Ling Chen, Jan Klopp, Min Sun, Shao-Yi Chien 외

Photo composition is an important factor affecting the aesthetics in photography. However, it is a highly challenging task to model the aesthetic properties of good compositions due to the lack of globally applicable rul…

Image Cropping

CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models

2025-12-11 · Tong Zhang, Carlos Hinojosa, Bernard Ghanem arxiv

Diffusion models can unintentionally reproduce training examples, raising privacy and copyright concerns as these systems are increasingly deployed at scale. Existing inference-time mitigation methods typically manipulat…