paper-with-me

Papers Descriptive

“Descriptive” 태그가 달린 논문 1,477편 · 필터 해제

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

2025-07-17 · Huakang Chen, Yuepeng Jiang, Guobin Ma, Chunbo Hao 외

Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current syste…

Descriptive

Assay2Mol: large language model-based drug design using BioAssay context

2025-07-16 · Yifan Deng, Spencer S. Ericksen, Anthony Gitter

Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate the functional responses of candidate molecules against disease targets. Un…

DescriptiveDrug DesignDrug DiscoveryIn-Context Learning+3

Describe Anything Model for Visual Question Answering on Text-rich Images

2025-07-16 · Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong, Anh-Khoi Nguyen 외

Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed descriptions of any specific image areas…

DescriptiveLanguage ModelingLanguage ModellingQuestion Answering+2

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

2025-07-09 · Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui 외

Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the …

DescriptiveText GenerationVideo Generation

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor

2025-07-04 · Ethan Lin, Linxi Zhao, Atharva Sehgal, Jennifer J. Sun

Text-based visual descriptors--ranging from simple class names to more descriptive phrases--are widely used in visual concept discovery and image classification with vision-language models (VLMs). Their effectiveness, ho…

Descriptiveimage-classificationImage Classification

Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization

2025-07-03 · De Cheng, Zhipeng Xu, Xinyang Jiang, Dongsheng Li 외

Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains. Notably, recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have demons…

DescriptiveDisentanglementDomain GeneralizationLarge Language Model+1

Dataset Distillation via Vision-Language Category Prototype

2025-06-30 · Yawen Zou, Guang Li, Duo Su, Zi Wang 외

Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consump…

Dataset DistillationDescriptiveLarge Language Model

Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization

2025-06-25 · Zhiwang Zhang, Dong Xu, Wanli Ouyang, Chuanqi Tan

In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of s…

Dense Video CaptioningDescriptiveSentenceSentence Summarization+1

Experiential marketing strategy and tourism demand in the contribution of the positioning of the floating islands Los Uros, Puno

2025-06-22 · Guina Flores Montalico

Experiential focused on creating memorable and meaningful experiences for consumers, has emerged as a key strategy in promoting tourist destinations. particularly in destinations seeking to highlight their unique cultura…

DescriptiveMarketing

DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

2025-06-21 · Mihir Godbole, Xiangbo Gao, Zhengzhong Tu

Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision…

Autonomous DrivingDescriptiveLarge Language Model

A Simple Contrastive Framework Of Item Tokenization For Generative Recommendation

2025-06-20 · Penglong Zhai, Yifang Yuan, Fanyi Di, Jie Li 외

Generative retrieval-based recommendation has emerged as a promising paradigm aiming at directly generating the identifiers of the target candidates. However, in large-scale recommendation systems, this approach becomes …

Contrastive LearningDescriptiveQuantizationRecommendation Systems+1

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

2025-06-19 · Kexin Huang, Qian Tu, Liwei Fan, Chenchen Yang 외

In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech …

BenchmarkingDescriptiveInstruction FollowingSpeech Synthesis+2

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

2025-06-18 · Anuradha Chopra, Abhinaba Roy, Dorien Herremans

Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, tha…

Caption GenerationDescriptiveKey DetectionLarge Language Model+2

Uncovering Intention through LLM-Driven Code Snippet Description Generation

2025-06-18 · Yusuf Sulistyo Nugroho, Farah Danisha Salam, Brittany Reid, Raula Gaikovina Kula 외

Documenting code snippets is essential to pinpoint key areas where both developers and users should pay attention. Examples include usage examples and other Application Programming Interfaces (APIs), which are especially…

Descriptive

Evolvable Conditional Diffusion

2025-06-16 · Zhao Wei, Chin Chun Ooi, Abhishek Gupta, Jian Cheng Wong 외

This paper presents an evolvable conditional diffusion method such that black-box, non-differentiable multi-physics models, as are common in domains like computational fluid dynamics and electromagnetics, can be effectiv…

DenoisingDescriptivescientific discovery

A Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation

2025-06-16 · Xiaoyang Wei, Camille Kurtz, Florence Cloppet

Performance evaluation for Content-Based Image Retrieval (CBIR) remains a crucial but unsolved problem today especially in the medical domain. Various evaluation metrics have been discussed in the literature to solve thi…

Content-Based Image RetrievalDescriptiveImage RetrievalKnowledge Graphs+2

Rethinking Optimization: A Systems-Based Approach to Social Externalities

2025-06-15 · Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

Optimization is widely used for decision making across various domains, valued for its ability to improve efficiency. However, poor implementation practices can lead to unintended consequences, particularly in socioecono…

Descriptive

Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables

2025-06-13 · Yitong Zhou, Mingyue Cheng, Qingyang Mao, Yucong Luo 외

Chemical tables encode complex experimental knowledge through symbolic expressions, structured variables, and embedded molecular graphics. Existing benchmarks largely overlook this multimodal and domain-specific complexi…

BenchmarkingDescriptiveQuestion AnsweringTable Recognition

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

2025-06-11 · Aaron Foss, Chloe Evans, Sasha Mitts, Koustuv Sinha 외

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to…

counterfactualDescriptiveQuestion AnsweringVideo Question Answering+1

ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model

2025-06-11 · Jialong Zuo, Yongtai Deng, Mengdan Tan, Rui Jin 외

In real-word scenarios, person re-identification (ReID) expects to identify a person-of-interest via the descriptive query, regardless of whether the query is a single modality or a combination of multiple modalities. Ho…

cross-modal alignmentDescriptivePerson Re-Identification
1–20 / 1,477 다음 →