paper-with-me

홈 › Papers

Constructing Synthetic Instruction Datasets for Improving Reasoning in Domain-Specific LLMs: A Case Study in the Japanese Financial Domain

2026-03-02 · Yuma Okochi, Fabio Milentiansen Sim, Tomoyasu Okada arxiv

In adapting LLMs to specific domains, achieving both domain expertise and reasoning ability remains an urgent challenge. This study proposes a general method for constructing high-quality synthetic instruction data for any domain, starting from domain-specific vocabulary. As a demonstration, we applied this method to the financial domain and constructed a large-scale instruction dataset totaling approximately 9.5 billion tokens with Chain-of-Thought reasoning traces. Evaluation results confirmed performance improvements over baseline models on financial benchmarks, demonstrating the effectiveness of our approach. We also report findings on the impact of reasoning trace length on performance and its limitations. Lastly, we open-source our models and datasets on https://huggingface.co/nri-ai .

📄 PDF Abstract BibTeX arXiv:2603.01353

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KORMo: Korean Open Reasoning Model for Everyone

2025-10-10 · Minjun Kim, Hyeonseok Lim, Hangyeol Yoo, Inho Won 외 arxiv

This work presents the first large-scale investigation into constructing a fully open bilingual large language model (LLM) for a non-English language, specifically Korean, trained predominantly on synthetic data. We intr…

FRIDA to the Rescue! Analyzing Synthetic Data Effectiveness in Object-Based Common Sense Reasoning for Disaster Response

2025-02-25 · Mollie Shichman, Claire Bonial, Austin Blodgett, Taylor Hudson 외

Large Language Models (LLMs) have the potential for substantial common sense reasoning. However, these capabilities are often emergent in larger models. This means smaller models that can be run locally are less helpful …

Common Sense ReasoningDisaster ResponseInformation Retrieval

SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification

2025-11-12 · Mohamed Elaraby, Jyoti Prakash Maheswari arxiv

Large Language Models (LLMs) with extended context windows promise direct reasoning over long documents, reducing the need for chunking or retrieval. Constructing annotated resources for training and evaluation, however,…

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions

2026-04-09 · Ashima Suvarna, Kendrick Phan, Mehrab Beikzadeh, Hritik Bansal 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has substantially improved reasoning in formal domains such as mathematics and code, but extending these gains beyond STEM remains challenging. Extending RLVR beyond …

Reinforcement Learning

LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain

2024-04-02 · Joel Niklaus, Lucia Zheng, Arya D. McCarthy, Christopher Hahn 외

Instruction tuning is an important step in making language models useful for direct user interaction. However, the legal domain is underrepresented in typical instruction datasets (e.g., only 10 out of 1600+ tasks in Sup…

Argument MiningDecision MakingLanguage ModelingLanguage Modelling+2