paper-with-me

Papers

MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model

2025-11-14 · Manyu Li, Ruian He, Chenxi Ma, Weimin Tan, Bo Yan arxiv

Multimodal Large Language Models are increasingly applied to biomedical imaging, yet scientific reasoning for microscopy remains limited by the scarcity of large-scale, high-quality training data. We introduce MicroVQA++, a three-stage, large-scale and high-quality microscopy VQA corpus derived from the BIOMEDICA archive. Stage one bootstraps supervision from expert-validated figure-caption pairs sourced from peer-reviewed articles. Stage two applies HiCQA-Graph, a novel heterogeneous graph over images, captions, and QAs that fuses NLI-based textual entailment, CLIP-based vision-language alignment, and agent signals to identify and filter inconsistent samples. Stage three uses a MultiModal Large Language Model (MLLM) agent to generate multiple-choice questions (MCQ) followed by human screening. The resulting release comprises a large training split and a human-checked test split whose Bloom's level hard-sample distribution exceeds the MicroVQA benchmark. Our work delivers (i) a quality-controlled dataset that couples expert literature with graph-based filtering and human refinement; (ii) HiCQA-Graph, the first graph that jointly models (image, caption, QA) for cross-modal consistency filtering; (iii) evidence that careful data construction enables 4B-scale MLLMs to reach competitive microscopy reasoning performance (e.g., GPT-5) and achieve state-of-the-art performance among open-source MLLMs. Code and dataset will be released after the review process concludes.

📄 PDF Abstract BibTeX arXiv:2511.11407

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research

2025-03-17 · CVPR 2025 1 · James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano 외

Scientific research demands sophisticated reasoning over multimodal data, a challenge especially prevalent in biology. Despite recent advances in multimodal large language models (MLLMs) for AI-assisted research, existin…

ArticlesBenchmarkingMultimodal ReasoningMultiple-choice+4

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

2026-05-11 · Manyu Li, Ruian He, Chenxi Ma, Weimin Tan 외 arxiv

Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as microscopy remains limited by the scarcity of domain-specific training dat…

FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy

2026-02-27 · Hyejin Park, Jiwon Yoon, Sumin Park, Suree Kim 외 arxiv

Accurate focus quality assessment (FQA) in fluorescence microscopy is challenging due to stain-dependent optical variations that induce heterogeneous focus behavior across images. Existing methods, however, treat focus q…

LUCYD: A Feature-Driven Richardson-Lucy Deconvolution Network

2023-07-16 · Tomáš Chobola, Gesine Müller, Veit Dausmann, Anton Theileis 외

The process of acquiring microscopic images in life sciences often results in image degradation and corruption, characterised by the presence of noise and blur, which poses significant challenges in accurately analysing …

Image Restoration

A Poisson-Gaussian Denoising Dataset with Real Fluorescence Microscopy Images

2018-12-26 · CVPR 2019 6 · Yide Zhang, Yinhao Zhu, Evan Nichols, Qingfei Wang 외

Fluorescence microscopy has enabled a dramatic development in modern biology. Due to its inherently weak signal, fluorescence microscopy is not only much noisier than photography, but also presented with Poisson-Gaussian…

Denoising