paper-with-me

Language Modelling

59개 벤치마크 · 논문 17,661편 · 이 태스크의 논문 보기 →

Benchmarks

WikiText-103

결과 90개

enwik8

결과 42개

The Pile

결과 39개

WikiText-2

결과 38개

LAMBADA

결과 37개

One Billion Word

결과 27개

Text8

결과 24개

Hutter Prize

결과 18개

OpenWebText

결과 12개

SALMon

결과 10개

C4

결과 9개

BIG-bench-lite

결과 3개

Wiki-40B

결과 3개

CLUE (AFQMC)

결과 2개

CLUE (C3)

결과 2개

CLUE (CMNLI)

결과 2개

CLUE (CMRC2018)

결과 2개

CLUE (DRCD)

결과 2개

CLUE (OCNLI_50K)

결과 2개

CLUE (WSC1.1)

결과 2개

FewCLUE (BUSTM)

결과 2개

FewCLUE (CHID-FC)

결과 2개

FewCLUE (CLUEWSC-FC)

결과 2개

FewCLUE (EPRSTMT)

결과 2개

FewCLUE (OCNLI-FC)

결과 2개

VietMed

결과 2개

Bookcorpus2

결과 1개

Gutenberg PG-19

결과 1개

OpenSubtitles

결과 1개

PubMed Central

결과 1개

StackExchange

결과 1개

USPTO Backgrounds

결과 1개

Ubuntu IRC

결과 1개

2000 HUB5 English

결과 1개

A1

결과 1개

Books3

결과 1개

Curation Corpus

결과 1개

DM Mathematics

결과 1개

FreeLaw

결과 1개

GitHub

결과 1개

HackerNews

결과 1개

NIH ExPorter

결과 1개

OpenWebtext2

결과 1개

PhilPapers

결과 1개

Pile CC

결과 1개

Text8 dev

결과 1개

enwik8 dev

결과 1개

enwiki8

결과 1개

Most implemented

A Neural Algorithm of Artistic Style

2015-08-26 · 구현 284개

Semi-supervised Sequence Learning

2015-11-04 · 구현 161개

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

Papers

MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

2026-08-26 · Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay arxiv

Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model t…

Language Modelling

FourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention

2026-07-08 · Athanasios Zeris arxiv

FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fixed random spectral filter achieves val=…

Language Modelling

Abstract representational geometry supports inference in large language models

2026-06-22 · Yunan Zeng, Yuwang Wang arxiv

A defining feature of human intelligence is the ability to adapt to changing environments by inferring latent task structure from sparse observations. Neuroscientific research indicates that this capability relies on the…

Language Modelling

Fine-Tuning Large Language Models for Quantum Reasoning

2026-06-20 · Katherine Ip, Casey R. Myers, Udaya Parampalli, James Quach 외 arxiv

Large language models (LLMs) exhibit abilities beyond natural language modelling and text generation. Recent advances in their reasoning capabilities have spurred interest in applying LLMs to complex scientific tasks req…

Language ModellingText Generation

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

2026-06-09 · Xuanchen Li, Tianrui Wang, Yuheng Lu, Zikang Huang 외 arxiv

Speech-to-text (S2T) systems for recognition (ASR) and translation (S2TT) typically generate discrete text tokens. In contrast, continuous-target language modelling performs generation in a continuous space, yet its pote…

Speech RecognitionLanguage Modelling

Runtime-Certified Bounded-Error Quantized Attention

2026-05-20 · Dean Calver arxiv

KV cache quantization reduces the memory cost of long-context LLM inference, but introduces approximation error that is typically validated only empirically. Existing systems rely on average-case robustness, with no mech…

Language Modelling

전체 17,661편 보기 →