paper-with-me

Papers Multimodal Large Language Model

“Multimodal Large Language Model” 태그가 달린 논문 347편 · 필터 해제

LRMR: LLM-Driven Relational Multi-node Ranking for Lymph Node Metastasis Assessment in Rectal Cancer

2025-07-15 · Yaoxian Dong, Yifan Gao, Haoyue Li, Yanfen Cui 외

Accurate preoperative assessment of lymph node (LN) metastasis in rectal cancer guides treatment decisions, yet conventional MRI evaluation based on morphological criteria shows limited diagnostic performance. While some…

DiagnosticLarge Language ModelMultimodal Large Language Model

MFGDiffusion: Mask-Guided Smoke Synthesis for Enhanced Forest Fire Detection

2025-07-15 · Guanghao Wu, Chen Xu, Hai Song, Chong Wang 외

Smoke is the first visible indicator of a wildfire.With the advancement of deep learning, image-based smoke detection has become a crucial method for detecting and preventing forest fires. However, the scarcity of smoke …

Fire DetectionImage GenerationLarge Language ModelMultimodal Large Language Model

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

2025-07-15 · Jie Yang, Wang Zeng, Sheng Jin, Lumin Xu 외

The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle with capturing fine-grained semantic inform…

Keypoint DetectionLanguage ModelingLanguage ModellingLarge Language Model+1

Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI

2025-07-14 · Jiangkai Wu, Zhiyuan Ren, LiMing Liu, Xinggong Zhang

AI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if c…

Large Language ModelMultimodal Large Language ModelVideo Understanding

TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

2025-07-08 · Yujie Hu, Xuanyu Zhang, Weiqi Li, Jian Zhang

Virtual try-on has made significant progress in recent years. This paper addresses how to achieve multifunctional virtual try-on guided solely by text instructions, including full outfit change and local editing. Previou…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

BlueLM-2.5-3B Technical Report

2025-07-08 · Baojiao Xiong, Boheng Chen, Chengzhi Wang, Daxiong Luo 외

We present BlueLM-2.5-3B, a compact and unified dense Multimodal Large Language Model (MLLM) designed for efficient edge-device deployment, offering strong general-purpose and reasoning capabilities. To the best of our k…

Large Language ModelMultimodal Large Language Model

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step

2025-07-06 · Zheyuan Liu, Munan Ning, Qihui Zhang, Shuo Yang 외

Current text-to-image (T2I) generation models struggle to align spatial composition with the input text, especially in complex scenes. Even layout-based approaches yield suboptimal spatial control, as their generation pr…

DenoisingLarge Language ModelMultimodal Large Language Model

Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval

2025-06-28 · Li-Cheng Shen, Jih-Kang Hsieh, Wei-Hua Li, Chu-Song Chen

Text-to-image retrieval (TIR) aims to find relevant images based on a textual query, but existing approaches are primarily based on whole-image captions and lack interpretability. Meanwhile, referring expression segmenta…

Cross-Modal RetrievalImage CaptioningImage RetrievalLarge Language Model+9

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

2025-06-26 · Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang 외

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries,…

Audio GenerationLarge Language ModelMultimodal Large Language Model

OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography

2025-06-26 · Caoshuo Li, Zengmao Ding, Xiaobin Hu, Bang Li 외

As one of the earliest ancient languages, Oracle Bone Script (OBS) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of approximately 4,500 OBS characters, onl…

DeciphermentLarge Language ModelMultimodal Large Language ModelVisual Localization

MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

2025-06-23 · Yuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue 외

Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data,…

DiagnosticLarge Language ModelMultimodal Large Language Model

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

2025-06-22 · Junying Chen, Zhenyang Cai, Pengcheng Chen, Shunian Chen 외

Recent advances in multimodal generative models have unlocked photorealistic, instruction-aligned image generation, yet leading systems like GPT-4o-Image remain proprietary and inaccessible. To democratize these capabili…

GPUImage GenerationLanguage ModelingLanguage Modelling+4

DreamJourney: Perpetual View Generation with Video Diffusion Models

2025-06-21 · Bo Pan, Yang Chen, Yingwei Pan, Ting Yao 외

Perpetual view generation aims to synthesize a long-term video corresponding to an arbitrary camera trajectory solely from a single input image. Recent methods commonly utilize a pre-trained text-to-image diffusion model…

Image to 3DLarge Language ModelMultimodal Large Language ModelPerpetual View Generation

The Condition Number as a Scale-Invariant Proxy for Information Encoding in Neural Units

2025-06-19 · Oswaldo Ludwig

This paper explores the relationship between the condition number of a neural network's weight tensor and the extent of information encoded by the associated processing unit, viewed through the lens of information theory…

Large Language ModelMultimodal Large Language Model

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

2025-06-17 · Yujun Wang, Jinhe Bi, Yunpu Ma, Soeren Pirk

Multimodal Large Language Model (MLLM) often suffer from hallucinations. They over-rely on partial cues and generate incorrect responses. Recently, methods like Visual Contrastive Decoding (VCD) and Instruction Contrasti…

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+2

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation

2025-06-16 · Bo Pan, Yixiao Fu, Ke Wang, Junyu Lu 외

Data visualization generation using Large Language Models (LLMs) has shown promising results but often produces suboptimal visualizations that require human intervention for improvement. In this work, we introduce VIS-Sh…

Data VisualizationLanguage ModelingLanguage ModellingLarge Language Model+1

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

2025-06-16 · Jiangtong Li, Yiyun Zhu, Dawei Cheng, Zhijun Ding 외

Multimodal Large Language Models (MLLMs) have rapidly evolved with the growth of Large Language Models (LLMs) and are now applied in various fields. In finance, the integration of diverse modalities such as text, charts,…

Decision MakingFinancial AnalysisLanguage ModelingLanguage Modelling+2

VGR: Visual Grounded Reasoning

2025-06-13 · Jiacong Wang, Zijian Kang, Haochen Wang, Haiyong Jiang 외

In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or scie…

Large Language ModelMathMultimodal Large Language ModelVisual Reasoning

PHRASED: Phrase Dictionary Biasing for Speech Translation

2025-06-10 · Peidong Wang, Jian Xue, Rui Zhao, Junkun Chen 외

Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in speech translation tasks. In this paper, we…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

2025-06-05 · Jisu An, Junseok Lee, Jeoungeun Lee, Yongseok Son

The rapid progress of Multimodal Large Language Models(MLLMs) has transformed the AI landscape. These models combine pre-trained LLMs with various modality encoders. This integration requires a systematic understanding o…

cross-modal alignmentLarge Language ModelMultimodal Large Language ModelRepresentation Learning
1–20 / 347 다음 →