paper-with-me

Document Classification

21개 벤치마크 · 논문 670편 · 이 태스크의 논문 보기 →

Benchmarks

Reuters-21578

결과 24개

Cora

결과 18개

HOC

결과 15개

BBCSport

결과 12개

Amazon

결과 9개

Twitter

결과 9개

AAPD

결과 6개

Classic

결과 6개

IMDb-M

결과 6개

Recipe

결과 6개

SciDocs (MAG)

결과 6개

SciDocs (MeSH)

결과 6개

WOS-5736

결과 6개

LUN

결과 3개

MPQA

결과 3개

Reuters De-En

결과 3개

Reuters En-De

결과 3개

WOS-11967

결과 3개

WOS-46985

결과 3개

Yelp-14

결과 3개

Most implemented

Graph Attention Networks

2017-10-30 · 구현 93개

On Calibration of Modern Neural Networks

2017-06-14 · 구현 15개

Papers

UCSC NLP at SemEval-2026 Task 10: Boundary-Aware Span Extraction and RoBERTa Classification for Conspiracy Detection

2026-07-06 · Dom Marhoefer, Milos Suvakovic, Glenn Grant-Richards, Aidan Pinero 외 arxiv

We present our systems for SemEval-2026 Task 10 (PsyCoMark), addressing conspiracy marker extraction (Subtask 1) and document-level conspiracy detection (Subtask 2). For marker extraction, we formulate the task as multi-…

Document Classification

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

2026-06-30 · Stefan Larson, Attila Nagy, Sam Desai, Cyrus Desai 외 arxiv

RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overlap, both of which may impact model perfor…

Document Classification

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT

2026-06-21 · Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos, Giovana Kerche Bonás arxiv

Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of the ModernBERT-base checkpoint on 60 bi…

Natural Language UnderstandingDocument ClassificationInformation Retrieval

Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization

2026-05-27 · Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Most. Aysha Siddika Sumona, Shahnewaz Siddique 외 arxiv

This study introduces a novel architecture of KAN-based BiGRU model for the task of classification and summarization of legal documents in a low-resource multilingual setup. In order to tackle problems associated with do…

Document Classification

Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System

2026-05-19 · Ivan Dobrovolskyi arxiv

Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructure, while rule-based tools often miss threats that depend on context. T…

Document Classification

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

2026-05-11 · Denghao Ma, Qing Liu, Zulong Chen, Chuanfei Xu 외 arxiv

Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear lit…

Document Classification

전체 670편 보기 →