Papers Concept Alignment
“Concept Alignment” 태그가 달린 논문 36편 · 필터 해제
FinTagging: An LLM-ready Benchmark for Extracting and Structuring Financial Information
We introduce FinTagging, the first full-scope, table-aware XBRL benchmark designed to evaluate the structured information extraction and semantic alignment capabilities of large language models (LLMs) in the context of X…
Concept AlignmentMulti-class ClassificationRoboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to genera…
Concept Alignmentobject-detectionObject DetectionReplace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image
Text-to-Image (T2I) has been prevalent in recent years, with most common condition tasks having been optimized nicely. Besides, counterfactual Text-to-Image is obstructing us from a more versatile AIGC experience. For th…
Concept AlignmentcounterfactualAn Explanation of Intrinsic Self-Correction via Linear Representations and Latent Concepts
We provide an explanation for the performance gains of intrinsic self-correction, a process where a language model iteratively refines its outputs without external feedback. More precisely, we investigate how prompting i…
Concept AlignmentLanguage ModelingLanguage ModellingHandling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
Adapting vision-language models (VLMs) to downstream tasks with pseudolabels has gained increasing attention. A major obstacle is that the pseudolabels generated by VLMs tend to be imbalanced, leading to inferior perform…
Concept AlignmentTraining-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis
Conditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented with a narrow scope, handling a restricted…
Concept AlignmentImage GenerationEnhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
Retrieval-Augmented Generation (RAG) systems face significant performance gaps when applied to technical domains requiring precise information extraction from complex documents. Current evaluation methodologies relying o…
Concept AlignmentRAGRetrievalRetrieval-augmented Generation+1Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
In the contemporary era of intelligent connectivity, Affective Computing (AC), which enables systems to recognize, interpret, and respond to human behavior states, has become an integrated part of many AI systems. As one…
Concept AlignmentEmotion RecognitionFacial Expression RecognitionConceptCLIP: Towards Trustworthy Medical AI via Concept-Enhanced Contrastive Langauge-Image Pre-training
Trustworthiness is essential for the precise and interpretable application of artificial intelligence (AI) in medical imaging. Traditionally, precision and interpretability have been addressed as separate tasks, namely m…
ArticlesConcept AlignmentMedical Image AnalysisRadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinical workflow. Current approaches either fo…
Concept AlignmentImage CaptioningRetrieval-augmented GenerationText-Video Retrieval with Global-Local Semantic Consistent Learning
Adapting large-scale image-text pre-training models, e.g., CLIP, to the video domain represents the current state-of-the-art for text-video retrieval. The primary approaches involve transferring text-video pairs to a com…
Concept AlignmentRetrievalVideo RetrievalAnchor and Broadcast: An Efficient Concept Alignment Approach for Evaluation of Semantic Graphs
In this paper, we present AnCast, an intuitive and efficient tool for evaluating graph-based meaning representations (MR). AnCast implements evaluation metrics that are well understood in the NLP community, and they incl…
AMR Graph SimilarityConcept AlignmentRelationImproving Concept Alignment in Vision-Language Concept Bottleneck Models
Concept Bottleneck Models (CBM) map images to human-interpretable concepts before making class predictions. Recent approaches automate CBM construction by prompting Large Language Models (LLMs) to generate text concepts …
ClassificationConcept AlignmentA Self-explaining Neural Architecture for Generalizable Concept Learning
With the wide proliferation of Deep Neural Networks in high-stake applications, there is a growing demand for explainability behind their decision-making process. Concept learning models attempt to learn high-level 'conc…
Concept AlignmentContrastive LearningDecision MakingDomain AdaptationConcept-Attention Whitening for Interpretable Skin Lesion Diagnosis
The black-box nature of deep learning models has raised concerns about their interpretability for successful deployment in real-world clinical applications. To address the concerns, eXplainable Artificial Intelligence (X…
Concept AlignmentDiagnosticExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the per…
Concept AlignmentInstruction FollowingLanguage ModellingVisual Question Answering (VQA)SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
Misinformation is a prevalent societal issue due to its potential high risks. Out-of-context (OOC) misinformation, where authentic images are repurposed with false text, is one of the easiest and most effective ways to m…
Concept AlignmentExplanation GenerationLanguage ModelingLanguage Modelling+4Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples
Current multimodal models leveraging contrastive learning often face limitations in developing fine-grained conceptual understanding. This is due to random negative samples during pretraining, causing almost exclusively …
Concept AlignmentContrastive LearningImage-text Retrieval$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
Despite the recent advances in personalized text-to-image (P-T2I) generative models, it remains challenging to perform finetuning-free multi-subject-driven T2I in a resource-efficient manner. Predominantly, contemporary …
Concept AlignmentGPUPhilosophyMICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment
Black-box deep learning approaches have showcased significant potential in the realm of medical image analysis. However, the stringent trustworthiness requirements intrinsic to the medical field have catalyzed research i…
Concept AlignmentExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Medical Image Analysis