paper-with-me

홈 › Papers

Fine-tuning a vision-language model for fracture-surface morphology recognition

2026-05-08 · Quanliang Liu, Jungtaek Kim, Kangwook Lee, Hyunseok Oh arxiv

Vision-language models (VLMs) have shown strong potential for scientific image understanding, but general-purpose models often lack the domain-specific visual knowledge required for reliable materials characterization. In this work, we fine-tuned an open-source VLM (Qwen3-VL-32B-Instruct) for fracture-surface image analysis using a curated dataset of 13,168 open-source, literature-mined fracture-surface images. Morphology annotations were generated by GPT-5.2-Reasoning (high) from both the images and relevant excerpts of their source papers, and the dataset was further enriched with targeted manual collection and rotation-based augmentation. The resulting specialist model outperforms flagship proprietary multimodal models on a benchmark of 100 manually annotated images. It achieves a precision of 0.92, compared to 0.35 for the base Qwen3-VL-32B-Instruct, 0.58 for GPT-5.5-Reasoning (high), and 0.78 for Gemini 3.1 Pro-Reasoning (high). Dataset ablations show that manual collection of rare-feature images and augmentation via image rotation are both beneficial to improve recognition of less common fracture morphology features. We further discuss integrated use of the fine-tuned model with proprietary models to combine fracture-specific visual accuracy with broader multimodal reasoning for autonomous fractography. Although focused on fracture-surface images, this work demonstrates how VLMs can be adapted through targeted collection and fine-tuning on novel feature images to recognize those features and support downstream decision-making in autonomous microscopy workflows.

📄 PDF Abstract BibTeX arXiv:2605.07145

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Towards Cross-Scale Attention and Surface Supervision for Fractured Bone Segmentation in CT

2024-05-02 · Yu Zhou, Xiahao Zou, Yi Wang

Bone segmentation is an essential step for the preoperative planning of fracture trauma surgery. The automated segmentation of fractured bone from computed tomography (CT) scans remains challenging, due to the large diff…

Computed Tomography (CT)Segmentation

Quantitative Matching of Forensic Evidence Fragments Utilizing 3D Microscopy Analysis of Fracture Surface Replicas

2021-09-24 · Bishoy Dawood, Carlos Llosa-Vite, Geoffrey Z. Thompson, Barbara K. Lograsso 외

Fractured surfaces carry unique details that can provide an accurate quantitative comparison to support comparative forensic analysis of those fractured surfaces. In this study, a statistical analysis comparison protocol…

Articles

Vision-based inspection system employing computer vision & neural networks for detection of fractures in manufactured components

2019-01-25 · Sarthak J. Shetty

We are proceeding towards the age of automation and robotic integration of our production lines [5]. Effective quality-control systems have to be put in place to maintain the quality of manufactured components. Among dif…

BIG-bench Machine Learning

Digital Image-Based Stress–Permeability Relationships of Rough Fractures Using Numerical Contact Mechanics and Stokes Equation

2022-01-10 · Transport in Porous Media 2022 1 · Amanzhol Kubeyev, Nathaniel Forbes Inskip, Tomos Phillips, Yihuai Zhang 외

Flow in fractures is sensitive to their geometrical surface characteristics. The surface can undergo deformation if there is a change in stress. Natural fractures have complex geometries and rough surfaces which complica…

Contact mechanics

Learning Generalizable Features for Tibial Plateau Fracture Segmentation Using Masked Autoencoder and Limited Annotations

2025-02-05 · Peiyan Yue, Die Cai, Chu Guo, Mengxing Liu 외

Accurate automated segmentation of tibial plateau fractures (TPF) from computed tomography (CT) requires large amounts of annotated data to train deep learning models, but obtaining such annotations presents unique chall…

Computed Tomography (CT)