paper-with-me

Papers

Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?

2025-01-29 · Pouya Pezeshkpour, Estevam Hruschka

Large Language Models (LLMs) have demonstrated remarkable performance on various tasks, yet their ability to extract and internalize deeper insights from domain-specific datasets remains underexplored. In this study, we investigate how continual pre-training can enhance LLMs' capacity for insight learning across three distinct forms: declarative, statistical, and probabilistic insights. Focusing on two critical domains: medicine and finance, we employ LoRA to train LLMs on two existing datasets. To evaluate each insight type, we create benchmarks to measure how well continual pre-training helps models go beyond surface-level knowledge. We also assess the impact of document modification on capturing insights. The results show that, while continual pre-training on original documents has a marginal effect, modifying documents to retain only essential information significantly enhances the insight-learning capabilities of LLMs.

📄 PDF Abstract BibTeX arXiv:2501.17840

Code (1)

megagonlabs/insight_miner 공식 구현

Similar Papers 제목 키워드 기반

Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning

2025-11-17 · Xinlan Wu, Bin Zhu, Feng Han, Pengkun Jiao 외 arxiv

Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LMMs) in food analysis suffer from catastr…

Semantic SimilarityContinual Learning

CLoRA: Parameter-Efficient Continual Learning with Low-Rank Adaptation

2025-07-26 · Shishir Muralidhara, Didier Stricker, René Schuster arxiv

In the past, continual learning (CL) was mostly concerned with the problem of catastrophic forgetting in neural networks, that arises when incrementally learning a sequence of tasks. Current CL methods function within th…

parameter-efficient fine-tuningSemantic SegmentationContinual Learning

SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs

2025-02-05 · Dinithi Jayasuriya, Sina Tayebati, Davide Ettori, Ranganath Krishnan 외

We propose SPARC, a lightweight continual learning framework for large language models (LLMs) that enables efficient task adaptation through prompt tuning in a lower-dimensional space. By leveraging principal component a…

Continual Learning

Continual Cross-Dataset Adaptation in Road Surface Classification

2023-09-05 · Paolo Cudrano, Matteo Bellusci, Giuseppe Macino, Matteo Matteucci

Accurate road surface classification is crucial for autonomous vehicles (AVs) to optimize driving conditions, enhance safety, and enable advanced road mapping. However, deep learning models for road surface classificatio…

Autonomous VehiclesClassificationContinual Learning

Continual-NExT: A Unified Comprehension And Generation Continual Learning Framework

2026-02-20 · Jingyang Qiao, Zhizhong Zhang, Xin Tan, Jingyu Gong 외 arxiv

Dual-to-Dual MLLMs refer to Multimodal Large Language Models, which can enable unified multimodal comprehension and generation through text and image modalities. Although exhibiting strong instantaneous learning and gene…

Continual Learning