paper-with-me

홈 › Papers

Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs

2026-07-12 · Shubham Rao arxiv

Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly adapting small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM receives each training image along with its correct label and generates a justified visual explanation. A 3B-parameter model is then fine-tuned on these reasoning-augmented examples via LoRA. By conditioning on correct answers, we ensure all training reasoning is directed toward the correct conclusion, which is critical because frontier models score as low as 24.1% on our hardest task. We validate on four industrial classification tasks spanning three image modalities using only 18 to 30 labeled images per task. Across 4 seeds per task (32 training runs), our method outperforms direct fine-tuning on all 16 seed-task combinations, with mean improvements of +1.7 to +4.4 percentage points. A controlled equal-budget experiment confirms the improvement comes from reasoning quality, not additional training steps. An unconditioned baseline demonstrates that with out answer-conditioning, wrong reasoning degrades performance by 17.8 percentage points. On weld radiograph classification, the fine-tuned 3B model outperforms GPT-4.1 by 10.0pp using just 24 training images.

📄 PDF Abstract BibTeX arXiv:2607.10666

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

2026-07-16 · Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee 외 arxiv

A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When s…

Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation

2026-02-27 · Gianluigi Silvestri, Edoardo Cetin arxiv

Reasoning-oriented language models achieve strong performance by generating long chain-of-thought traces at inference time. However, this capability comes with substantial and often excessive computational cost, which ca…

Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step

2023-06-24 · Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 외

Chain-of-thought prompting (e.g., "Let's think step-by-step") primes large language models to verbalize rationalization for their predictions. While chain-of-thought can lead to dramatic performance gains, benefits appea…

Diversity

Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts

2025-10-21 · Seungjun Yu, Junsung Park, Youngsun Lim, Hyunjung Shim arxiv

We present a two-phase vision-language QA system for autonomous driving that answers high-level perception, prediction, and planning questions. In Phase-1, a large multimodal LLM (Qwen2.5-VL-32B) is conditioned on six-ca…

Autonomous Driving

Reducing Hallucinations: Enhancing VQA for Flood Disaster Damage Assessment with Visual Contexts

2023-12-21 · Yimin Sun, Chao Wang, Yan Peng

The zero-shot performance of visual question answering (VQA) models relies heavily on prompts. For example, a zero-shot VQA for disaster scenarios could leverage well-designed Chain of Thought (CoT) prompts to stimulate …

HallucinationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)