paper-with-me

Continual Pretraining

3개 벤치마크 · 논문 147편 · 이 태스크의 논문 보기 →

Benchmarks

ACL-ARC

결과 1개

AG News

결과 1개

SciERC

결과 1개

Most implemented

Rho-1: Not All Tokens Are What You Need

2024-04-11 · 구현 3개

Continual Pre-training of Language Models

2023-02-07 · 구현 2개

Papers

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

2026-09-08 · Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen 외 hf

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how vi…

Continual Pretraining

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

2026-08-06 · Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li arxiv

Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted thr…

Continual Pretraining

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

2026-07-23 · Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li 외 arxiv

Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale because they rely on specialized hardware, centralized operators, or …

Continual PretrainingRobot Manipulation

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

2026-07-22 · Abigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun 외 arxiv

Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computatio…

Continual Pretraining

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

2026-07-08 · Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia 외 arxiv

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We in…

Continual Pretraining

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

2026-07-02 · Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach 외 arxiv

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-th…

Continual Pretraining

전체 147편 보기 →