Papers Descriptive
“Descriptive” 태그가 달린 논문 1,477편 · 필터 해제
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current syste…
DescriptiveAssay2Mol: large language model-based drug design using BioAssay context
Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate the functional responses of candidate molecules against disease targets. Un…
DescriptiveDrug DesignDrug DiscoveryIn-Context Learning+3Describe Anything Model for Visual Question Answering on Text-rich Images
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed descriptions of any specific image areas…
DescriptiveLanguage ModelingLanguage ModellingQuestion Answering+2FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the …
DescriptiveText GenerationVideo GenerationBeyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor
Text-based visual descriptors--ranging from simple class names to more descriptive phrases--are widely used in visual concept discovery and image classification with vision-language models (VLMs). Their effectiveness, ho…
Descriptiveimage-classificationImage ClassificationPrompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains. Notably, recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have demons…
DescriptiveDisentanglementDomain GeneralizationLarge Language Model+1Dataset Distillation via Vision-Language Category Prototype
Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consump…
Dataset DistillationDescriptiveLarge Language ModelShow, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of s…
Dense Video CaptioningDescriptiveSentenceSentence Summarization+1Experiential marketing strategy and tourism demand in the contribution of the positioning of the floating islands Los Uros, Puno
Experiential focused on creating memorable and meaningful experiences for consumers, has emerged as a key strategy in promoting tourist destinations. particularly in destinations seeking to highlight their unique cultura…
DescriptiveMarketingDRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
Understanding the short-term motion of vulnerable road users (VRUs) like pedestrians and cyclists is critical for safe autonomous driving, especially in urban scenarios with ambiguous or high-risk behaviors. While vision…
Autonomous DrivingDescriptiveLarge Language ModelA Simple Contrastive Framework Of Item Tokenization For Generative Recommendation
Generative retrieval-based recommendation has emerged as a promising paradigm aiming at directly generating the identifiers of the target candidates. However, in large-scale recommendation systems, this approach becomes …
Contrastive LearningDescriptiveQuantizationRecommendation Systems+1InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech …
BenchmarkingDescriptiveInstruction FollowingSpeech Synthesis+2SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, tha…
Caption GenerationDescriptiveKey DetectionLarge Language Model+2Uncovering Intention through LLM-Driven Code Snippet Description Generation
Documenting code snippets is essential to pinpoint key areas where both developers and users should pay attention. Examples include usage examples and other Application Programming Interfaces (APIs), which are especially…
DescriptiveEvolvable Conditional Diffusion
This paper presents an evolvable conditional diffusion method such that black-box, non-differentiable multi-physics models, as are common in domains like computational fluid dynamics and electromagnetics, can be effectiv…
DenoisingDescriptivescientific discoveryA Semantically-Aware Relevance Measure for Content-Based Medical Image Retrieval Evaluation
Performance evaluation for Content-Based Image Retrieval (CBIR) remains a crucial but unsolved problem today especially in the medical domain. Various evaluation metrics have been discussed in the literature to solve thi…
Content-Based Image RetrievalDescriptiveImage RetrievalKnowledge Graphs+2Rethinking Optimization: A Systems-Based Approach to Social Externalities
Optimization is widely used for decision making across various domains, valued for its ability to improve efficiency. However, poor implementation practices can lead to unintended consequences, particularly in socioecono…
DescriptiveBenchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
Chemical tables encode complex experimental knowledge through symbolic expressions, structured variables, and embedded molecular graphics. Existing benchmarks largely overlook this multimodal and domain-specific complexi…
BenchmarkingDescriptiveQuestion AnsweringTable RecognitionCausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to…
counterfactualDescriptiveQuestion AnsweringVideo Question Answering+1ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single Model
In real-word scenarios, person re-identification (ReID) expects to identify a person-of-interest via the descriptive query, regardless of whether the query is a single modality or a combination of multiple modalities. Ho…
cross-modal alignmentDescriptivePerson Re-Identification