paper-with-me

홈 › Papers

Penrose Tiled Low-Rank Compression and Section-Wise Q&A Fine-Tuning: A General Framework for Domain-Specific Large Language Model Adaptation

2025-03-28 · Chuan-Wei Kuo, Siyu Chen, Chenqi Yan, Yu Yang Fredrik Liu

Large language models (LLMs) hold great promise for specialized scientific domains such as materials science, yet adapting them efficiently and accurately to domain-specific knowledge remains challenging due to limited data and high knowledge density. We propose a two-stage framework that combines structured model compression with a scientific fine-tuning regimen to address this challenge. In the compression stage, we decompose the LLM's weight matrices into local low-rank "rank blocks" and arrange these blocks in a Penrose-like non-periodic tiling pattern. Each block is then compacted via spectral transformations (e.g., discrete cosine or Fourier transforms), and a Kullback-Leibler (KL) divergence-based alignment loss preserves the distributional similarity between the compressed model's representations and those of the original full model. In the adaptation stage, the compressed model is further tuned using a human-like scientific reading protocol: it processes technical materials science documents section by section, engaging in a structured question-and-answer routine for each section. This section-wise Q&A fine-tuning strategy extracts explicit reasoning traces and gradually injects domain knowledge, while minimizing catastrophic forgetting of the model's general language capabilities. By balancing efficient compression with targeted adaptation, our two-stage approach enables precise specialization of LLMs to high-value domains under data-scarce conditions. We present this principled yet exploratory pipeline and outline its potential for advancing materials science knowledge integration, laying the groundwork for comprehensive empirical evaluation in future work.

📄 PDF Abstract BibTeX arXiv:2503.22074

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelLow-rank compressionModel Compression

Similar Papers 제목 키워드 기반

Fast Computation of Moore-Penrose Inverse Matrices

2008-04-30 · Pierre Courrieu

Many neural learning algorithms require to solve large least square systems in order to obtain synaptic weights. Moore-Penrose inverse matrices allow for solving such systems, even with rank deficiency, and they provide …

Spatially adaptive image compression using a tiled deep network

2018-02-07 · David Minnen, George Toderici, Michele Covell, Troy Chinen 외

Deep neural networks represent a powerful class of function approximators that can learn to compress and reconstruct images. Existing image compression algorithms based on neural networks learn quantized representations …

Image Compression

Barwise Compression Schemes for Audio-Based Music Structure Analysis

2022-02-10 · Axel Marmoret, Jérémy E. Cohen, Frédéric Bimbot

Music Structure Analysis (MSA) consists in segmenting a music piece in several distinct sections. We approach MSA within a compression framework, under the hypothesis that the structure is more easily revealed by a simpl…

Building Cross-Sectional Systematic Strategies By Learning to Rank

2020-12-13 · Daniel Poh, Bryan Lim, Stefan Zohren, Stephen Roberts

The success of a cross-sectional systematic strategy depends critically on accurately ranking assets prior to portfolio construction. Contemporary techniques perform this ranking step either with simple heuristics or by …

Information RetrievalLearning-To-RankRetrieval

Differentiable SVD based on Moore-Penrose Pseudoinverse for Inverse Imaging Problems

2024-11-21 · Yinghao Zhang, Yue Hu

Low-rank regularization-based deep unrolling networks have achieved remarkable success in various inverse imaging problems (IIPs). However, the singular value decomposition (SVD) is non-differentiable when duplicated sin…

compressed sensingImage Compressed SensingMRI Reconstruction