paper-with-me

Papers

Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning

2024-02-15 · Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Jiuxiang Gu, Tianyi Zhou

Instruction tuning is critical to large language models (LLMs) for achieving better instruction following and task adaptation capabilities but its success heavily relies on the training data quality. Many recent methods focus on improving the data quality but often overlook the compatibility of the data with the student model being finetuned. This paper introduces Selective Reflection-Tuning, a novel paradigm that synergizes a teacher LLM's reflection and introspection for improving existing data quality with the data selection capability of the student LLM, to automatically refine existing instruction-tuning data. This teacher-student collaboration produces high-quality and student-compatible instruction-response pairs, resulting in sample-efficient instruction tuning and LLMs of superior performance. Selective Reflection-Tuning is a data augmentation and synthesis that generally improves LLM finetuning and self-improvement without collecting brand-new data. We apply our method to Alpaca and WizardLM data and achieve much stronger and top-tier 7B and 13B LLMs.

📄 PDF Abstract BibTeX arXiv:2402.10110

Code (2)

tianyi-lab/reflection_tuning 공식 구현
mingliiii/reflection_tuning

Tasks

Data AugmentationInstruction Following

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models

2025-08-08 · Lingyuan Liu, Mengxiang Zhang arxiv

Knowledge Distillation (KD) is a fundamental technique for compressing large language models (LLMs) into compact, efficient student models. However, existing white-box KD methods mainly focus on balancing ground truth an…

Knowledge Distillation

Symmetry-Aware Likelihood-Orbit Aggregation for Selective Left-Right Claim Verification

2026-09-15 · Zhouzhi Xiong, Chuxi Zhang, Weizhen He, Yi Chen 외 arxiv

Frozen vision-language models (VLMs) remain unreliable on fine-grained left-right claims, and raw claim likelihoods need not reliably rank verification errors. After a horizontal-reflection intervention is fixed, how sho…

To Use or to Refuse? Re-Centering Student Agency with Generative AI in Engineering Design Education

2025-10-22 · Thijs Willems, Sumbul Khan, Qian Huang, Bradley Camburn 외 arxiv

This pilot study traces students' reflections on the use of AI in a 13-week foundational design course enrolling over 500 first-year engineering and architecture students at the Singapore University of Technology and Des…

MolReFlect: Towards Fine-grained In-Context Alignment between Molecules and Texts

2024-11-22 · arXiv preprint 2024 11 · Jiatong Li, Yunqing Liu, Wei Liu, Jingdi Lei 외

Molecule discovery is a pivotal research field, impacting everything from the medicines we take to the materials we use. Recently, Large Language Models (LLMs) have been widely adopted in molecule understanding and gener…

DescriptiveMolecule CaptioningText-based de novo Molecule Generation

MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts

2024-11-22 · Jiatong Li, Yunqing Liu, Wei Liu, Jingdi Le 외

Molecule discovery is a pivotal research field, impacting everything from the medicines we take to the materials we use. Recently, Large Language Models (LLMs) have been widely adopted in molecule understanding and gener…

Descriptive