Document Classification
21개 벤치마크 · 논문 670편 · 이 태스크의 논문 보기 →
Benchmarks
Reuters-21578
Cora
HOC
BBCSport
Amazon
AAPD
Classic
IMDb-M
Recipe
SciDocs (MAG)
SciDocs (MeSH)
WOS-5736
LUN
MPQA
Reuters De-En
Reuters En-De
WOS-11967
WOS-46985
Yelp-14
Most implemented
Graph Attention Networks
Semi-Supervised Classification with Graph Convolutional Networks
Revisiting Semi-Supervised Learning with Graph Embeddings
On Calibration of Modern Neural Networks
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Papers
UCSC NLP at SemEval-2026 Task 10: Boundary-Aware Span Extraction and RoBERTa Classification for Conspiracy Detection
We present our systems for SemEval-2026 Task 10 (PsyCoMark), addressing conspiracy marker extraction (Subtask 1) and document-level conspiracy detection (Subtask 2). For marker extraction, we formulate the task as multi-…
Document ClassificationRevising RVL-CDIP: Quantifying Errors and Test-Train Overlap
RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overlap, both of which may impact model perfor…
Document ClassificationmoBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT
Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of the ModernBERT-base checkpoint on 60 bi…
Natural Language UnderstandingDocument ClassificationInformation RetrievalEnhancing BiGRU with a KAN Block for Legal Document Classification and Summarization
This study introduces a novel architecture of KAN-based BiGRU model for the task of classification and summarization of legal documents in a low-resource multilingual setup. In order to tackle problems associated with do…
Document ClassificationSecurity Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System
Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructure, while rule-based tools often miss threats that depend on context. T…
Document ClassificationMulti-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear lit…
Document Classification