Papers Multimodal Large Language Model
“Multimodal Large Language Model” 태그가 달린 논문 347편 · 필터 해제
LRMR: LLM-Driven Relational Multi-node Ranking for Lymph Node Metastasis Assessment in Rectal Cancer
Accurate preoperative assessment of lymph node (LN) metastasis in rectal cancer guides treatment decisions, yet conventional MRI evaluation based on morphological criteria shows limited diagnostic performance. While some…
DiagnosticLarge Language ModelMultimodal Large Language ModelMFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection
Smoke is the first visible indicator of a wildfire.With the advancement of deep learning, image-based smoke detection has become a crucial method for detecting and preventing forest fires. However, the scarcity of smoke …
Fire DetectionImage GenerationLarge Language ModelMultimodal Large Language ModelKptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle with capturing fine-grained semantic inform…
Keypoint DetectionLanguage ModelingLanguage ModellingLarge Language Model+1Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
AI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if c…
Large Language ModelMultimodal Large Language ModelVideo UnderstandingTalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previou…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1BlueLM-2.5-3B Technical Report
We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our k…
Large Language ModelMultimodal Large Language ModelCoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal spatial control, as their generation pr…
DenoisingLarge Language ModelMultimodal Large Language ModelMask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
Text-to-image retrieval (TIR) aims to find relevant images based on a textual query, but existing approaches are primarily based on whole-image captions and lack interpretability. Meanwhile, referring expression segmenta…
Cross-Modal RetrievalImage CaptioningImage RetrievalLarge Language Model+9ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries,…
Audio GenerationLarge Language ModelMultimodal Large Language ModelOracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography
As one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of approximately 4,500 OBS characters, onl…
DeciphermentLarge Language ModelMultimodal Large Language ModelVisual LocalizationMedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data,…
DiagnosticLarge Language ModelMultimodal Large Language ModelShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabili…
GPUImage GenerationLanguage ModelingLanguage Modelling+4DreamJourney: Perpetual View Generation with Video Diffusion Models
Perpetual view generation aims to synthesize a long-term video corresponding to an arbitrary camera trajectory solely from a single input image. Recent methods commonly utilize a pre-trained text-to-image diffusion model…
Image to 3DLarge Language ModelMultimodal Large Language ModelPerpetual View GenerationThe Condition Number as a Scale-Invariant Proxy for Information Encoding in Neural Units
This paper explores the relationship between the condition number of a neural network's weight tensor and the extent of information encoded by the associated processing unit, viewed through the lens of information theory…
Large Language ModelMultimodal Large Language ModelASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
Multimodal Large Language Model (MLLM) often suffer from hallucinations. They over-rely on partial cues and generate incorrect responses. Recently, methods like Visual Contrastive Decoding (VCD) and Instruction Contrasti…
HallucinationLanguage ModelingLanguage ModellingLarge Language Model+2VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
Data visualization generation using Large Language Models (LLMs) has shown promising results but often produces suboptimal visualizations that require human intervention for improvement. In this work, we introduce VIS-Sh…
Data VisualizationLanguage ModelingLanguage ModellingLarge Language Model+1CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model
Multimodal Large Language Models (MLLMs) have rapidly evolved with the growth of Large Language Models (LLMs) and are now applied in various fields. In finance, the integration of diverse modalities such as text, charts,…
Decision MakingFinancial AnalysisLanguage ModelingLanguage Modelling+2VGR: Visual Grounded Reasoning
In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or scie…
Large Language ModelMathMultimodal Large Language ModelVisual ReasoningPHRASED: Phrase Dictionary Biasing for Speech Translation
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in speech translation tasks. In this paper, we…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
The rapid progress of Multimodal Large Language Models(MLLMs) has transformed the AI landscape. These models combine pre-trained LLMs with various modality encoders. This integration requires a systematic understanding o…
cross-modal alignmentLarge Language ModelMultimodal Large Language ModelRepresentation Learning