paper-with-me

Papers

SChuBERT: Scholarly Document Chunks with BERT-encoding boost Citation Count Prediction.

2020-11-01 · EMNLP (sdp) 2020 11 · Thomas van Dongen, Gideon Maillette de Buy Wenniger, Lambert Schomaker

Predicting the number of citations of scholarly documents is an upcoming task in scholarly document processing. Besides the intrinsic merit of this information, it also has a wider use as an imperfect proxy for quality which has the advantage of being cheaply available for large volumes of scholarly documents. Previous work has dealt with number of citations prediction with relatively small training data sets, or larger datasets but with short, incomplete input text. In this work we leverage the open access ACL Anthology collection in combination with the Semantic Scholar bibliometric database to create a large corpus of scholarly documents with associated citation information and we propose a new citation prediction model called SChuBERT. In our experiments we compare SChuBERT with several state-of-the-art citation prediction models and show that it outperforms previous methods by a large margin. We also show the merit of using more training data and longer input for number of citations prediction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Citation PredictionPrediction

Similar Papers 제목 키워드 기반

SChuBERT: Scholarly Document Chunks with BERT-encoding boost Citation Count Prediction

2020-12-21 · Thomas van Dongen, Gideon Maillette de Buy Wenniger, Lambert Schomaker

Predicting the number of citations of scholarly documents is an upcoming task in scholarly document processing. Besides the intrinsic merit of this information, it also has a wider use as an imperfect proxy for quality w…

Citation PredictionPrediction

MultiSChuBERT: Effective Multimodal Fusion for Scholarly Document Quality Prediction

2023-08-15 · Gideon Maillette de Buy Wenniger, Thomas van Dongen, Lambert Schomaker

Automatic assessment of the quality of scholarly documents is a difficult task with high potential impact. Multimodality, in particular the addition of visual information next to text, has been shown to improve the perfo…

Chunking

schuBERT: Optimizing Elements of BERT

2020-05-09 · ACL 2020 6 · Ashish Khetan, Zohar Karnin

Transformers \citep{vaswani2017attention} have gradually become a key component for many state-of-the-art natural language representation models. A recent Transformer based model- BERT \citep{devlin2018bert} achieved sta…

RoR: Read-over-Read for Long Document Machine Reading Comprehension

2021-09-10 · Findings (EMNLP) 2021 11 · Jing Zhao, Junwei Bao, Yifan Wang, Yongwei Zhou 외

Transformer-based pre-trained models, such as BERT, have achieved remarkable results on machine reading comprehension. However, due to the constraint of encoding length (e.g., 512 WordPiece tokens), a long document is us…

Machine Reading ComprehensionReading ComprehensionTriviaQA

A Granular Grassmannian Clustering Framework via the Schubert Variety of Best Fit

2025-12-29 · Karim Salta, Michael Kirby, Chris Peterson arxiv

In many classification and clustering tasks, it is useful to compute a geometric representative for a dataset or a cluster, such as a mean or median. When datasets are represented by subspaces, these representatives beco…