Papers Small Language Model
“Small Language Model” 태그가 달린 논문 109편 · 필터 해제
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
Recent advancements in reasoning-focused language models such as OpenAI's O1 and DeepSeek-R1 have shown that scaling test-time computation-through chain-of-thought reasoning and iterative exploration-can yield substantia…
Code GenerationMathreinforcement-learningReinforcement Learning+2Domain-Adaptive Small Language Models for Structured Tax Code Prediction
Every day, multinational firms process thousands of transactions, each of which must adhere to tax regulations that vary by jurisdiction and are often nuanced. The determination of product and service tax codes, such as …
DecoderSmall Language ModelTowards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models
Large Language Models (LLMs) have showcased remarkable generalizability in language comprehension and hold significant potential to revolutionize human-computer interaction in smart homes. Existing LLM-based smart home a…
Privacy PreservingSmall Language ModelCounterfactual Influence as a Distributional Quantity
Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metric to study memorization, quantifying how…
counterfactualimage-classificationImage ClassificationMemorization+1Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragraphs from PubMed scientific articles, as…
ArticlesContinual PretrainingLanguage ModelingLanguage Modelling+4Distilling On-device Language Models for Robot Planning with Minimal Human Intervention
Large language models (LLMs) provide robots with powerful contextual reasoning abilities and a natural human interface. Yet, current LLM-enabled robots typically depend on cloud-hosted models, limiting their usability in…
Small Language ModelLightweight Relevance Grader in RAG
Retrieval-Augmented Generation (RAG) addresses limitations of large language models (LLMs) by leveraging a vector database to provide more accurate and up-to-date information. When a user submits a query, RAG executes a …
Language ModelingLanguage ModellingRAGRetrieval-augmented Generation+1HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
Large Language models have demonstrated promising performance in research ideation across scientific domains. Hypothesis development, the process of generating a highly specific declarative statement connecting a researc…
Language ModelingLanguage ModellingSmall Language ModelvalidTowards a Small Language Model Lifecycle Framework
Background: The growing demand for efficient and deployable language models has led to increased interest in Small Language Models (SLMs). However, existing research remains fragmented, lacking a unified lifecycle perspe…
Language ModelingLanguage ModellingmodelSmall Language ModelWhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
Mean Opinion Score (MOS) prediction for text to music systems requires evaluating both overall musical quality and text prompt alignment. This paper introduces WhisQ, a multimodal architecture that addresses this dual-as…
cross-modal alignmentLanguage ModelingLanguage ModellingRepresentation Learning+1Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
Recently, Large Language Models (LLMs) have demonstrated significant potential for data annotation, markedly reducing the labor costs associated with downstream applications. However, existing methods mostly adopt an agg…
Small Language Modeltext-classificationText ClassificationAdaptive Task Vectors for Large Language Models
In-Context Learning (ICL) enables Large Language Models (LLMs) to perform tasks without parameter updates by conditioning on a few demonstrations provided in the prompt. Despite its success, ICL suffers from several limi…
In-Context LearningSmall Language ModelZero-Shot Vision Encoder Grafting via LLM Surrogates
Vision language models (VLMs) typically pair a modestly sized vision encoder with a large language model (LLM), e.g., Llama-70B, making the decoder the primary computational burden during training. To reduce costs, a pot…
DecoderLanguage ModelingLanguage ModellingLarge Language Model+1A Lightweight Multi-Expert Generative Language Model System for Engineering Information and Knowledge Extraction
Despite recent advancements in domain adaptation techniques for large language models, these methods remain computationally intensive, and the resulting models can still exhibit hallucination issues. Most existing adapta…
Domain AdaptationHallucinationLanguage ModelingLanguage Modelling+1Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster
Chain-of-thought (CoT) distillation allows a large language model (LLM) to guide a small language model (SLM) in reasoning tasks. Existing methods train the SLM to learn the long rationale in one iteration, resulting in …
Heuristic SearchLanguage ModelingLanguage ModellingLarge Language Model+1Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model
The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages li…
Language ModelingLanguage ModellingMedical Question AnsweringPatient QA+2TinyRS-R1: Compact Multimodal Language Model for Remote Sensing
Remote-sensing applications often run on edge hardware that cannot host today's 7B-parameter multimodal language models. This paper introduces TinyRS, the first 2B-parameter multimodal small language model (MSLM) optimiz…
Language ModelingLanguage ModellingOpen-Ended Question AnsweringQuestion Answering+4Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
To support emerging language-based applications using dispersed and heterogeneous computing resources, the hybrid language model (HLM) offers a promising architecture, where an on-device small language model (SLM) genera…
Language ModelingLanguage ModellingLarge Language ModelSmall Language ModelMilChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Remote Sensing
Remarkable capabilities in understanding and generating text-image content have been demonstrated by recent advancements in multimodal large language models (MLLMs). However, their effectiveness in specialized domains-pa…
Language ModelingLanguage ModellingSmall Language ModelSadeed: Advancing Arabic Diacritization Through Small Language Model
Arabic text diacritization remains a persistent challenge in natural language processing due to the language's morphological richness. In this paper, we introduce Sadeed, a novel approach based on a fine-tuned decoder-on…
Arabic Text DiacritizationBenchmarkingDecoderLanguage Modeling+6