Language Modelling
59개 벤치마크 · 논문 17,661편 · 이 태스크의 논문 보기 →
Benchmarks
Artificial Analysis Math Index
Artificial Analysis Coding Index
WikiText-103
Penn Treebank (Word Level)
enwik8
The Pile
WikiText-2
LAMBADA
One Billion Word
Text8
Hutter Prize
OpenWebText
SALMon
C4
BIG-bench-lite
Wiki-40B
CLUE (AFQMC)
CLUE (C3)
CLUE (CMNLI)
CLUE (CMRC2018)
CLUE (DRCD)
CLUE (OCNLI_50K)
CLUE (WSC1.1)
FewCLUE (BUSTM)
FewCLUE (CHID-FC)
FewCLUE (CLUEWSC-FC)
FewCLUE (EPRSTMT)
FewCLUE (OCNLI-FC)
VietMed
Bookcorpus2
Gutenberg PG-19
OpenSubtitles
PubMed Central
StackExchange
USPTO Backgrounds
Ubuntu IRC
2000 HUB5 English
A1
Books3
Curation Corpus
DM Mathematics
FreeLaw
GitHub
HackerNews
NIH ExPorter
OpenWebtext2
PhilPapers
Pile CC
Text8 dev
enwik8 dev
enwiki8
Most implemented
A Neural Algorithm of Artistic Style
Semi-supervised Sequence Learning
LoRA: Low-Rank Adaptation of Large Language Models
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Universal Language Model Fine-tuning for Text Classification
Papers
MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum
Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model t…
Language ModellingFourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention
FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fixed random spectral filter achieves val=…
Language ModellingAbstract representational geometry supports inference in large language models
A defining feature of human intelligence is the ability to adapt to changing environments by inferring latent task structure from sparse observations. Neuroscientific research indicates that this capability relies on the…
Language ModellingFine-Tuning Large Language Models for Quantum Reasoning
Large language models (LLMs) exhibit abilities beyond natural language modelling and text generation. Recent advances in their reasoning capabilities have spurred interest in applying LLMs to complex scientific tasks req…
Language ModellingText GenerationSpeech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation
Speech-to-text (S2T) systems for recognition (ASR) and translation (S2TT) typically generate discrete text tokens. In contrast, continuous-target language modelling performs generation in a continuous space, yet its pote…
Speech RecognitionLanguage ModellingRuntime-Certified Bounded-Error Quantized Attention
KV cache quantization reduces the memory cost of long-context LLM inference, but introduces approximation error that is typically validated only empirically. Existing systems rely on average-case robustness, with no mech…
Language Modelling