paper-with-me

Papers

MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection

2025-07-06 · Liman Wang, Hanyang Zhong, Tianyuan Wang, Shan Luo, Jihong Zhu arxiv

Choosing appropriate fabrics is critical for meeting functional and quality demands in robotic textile manufacturing, apparel production, and smart retail. We propose MLLM-Fabric, a robotic framework leveraging multimodal large language models (MLLMs) for fabric sorting and selection. Built on a multimodal robotic platform, the system is trained through supervised fine-tuning and explanation-guided distillation to rank fabric properties. We also release a dataset of 220 diverse fabrics, each with RGB images and synchronized visuotactile and pressure data. Experiments show that our Fabric-Llama-90B consistently outperforms pretrained vision-language baselines in both attribute ranking and selection reliability. Code and dataset are publicly available at https://github.com/limanwang/MLLM-Fabric.

📄 PDF Abstract BibTeX arXiv:2507.04351

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

2025-08-30 · Guangzong Si, Hao Yin, Xianfei Li, Qing Ding 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved impressive advances, yet object hallucination remains a persistent challenge. Existing methods, based on the flawed assumption that omission and fabrication hallucin…

A Survey of Multimodal Large Language Model from A Data-centric Perspective

2024-05-26 · Tianyi Bai, Hao Liang, Binwang Wan, Yanran Xu 외

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. …

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

2025-05-30 · Bowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang 외

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to ade…

HallucinationMultimodal ReasoningVisual Reasoning

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

2026-05-25 · Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li, Yu-Jhe Li 외 arxiv

Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. T…

Instruction FollowingImage Generation

Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

2026-01-10 · Jen-tse Huang, Chang Chen, Shiyang Lai, Wenxuan Wang 외 arxiv

Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impress…

Logical Fallacies