paper-with-me

홈 › Papers

Learning to Rank Scientific Documents from the Crowd

2016-11-04 · Jesse M Lingeman, Hong Yu

Finding related published articles is an important task in any science, but with the explosion of new work in the biomedical domain it has become especially challenging. Most existing methodologies use text similarity metrics to identify whether two articles are related or not. However biomedical knowledge discovery is hypothesis-driven. The most related articles may not be ones with the highest text similarities. In this study, we first develop an innovative crowd-sourcing approach to build an expert-annotated document-ranking corpus. Using this corpus as the gold standard, we then evaluate the approaches of using text similarity to rank the relatedness of articles. Finally, we develop and evaluate a new supervised model to automatically rank related scientific articles. Our results show that authors' ranking differ significantly from rankings by text-similarity-based models. By training a learning-to-rank model on a subset of the annotated corpus, we found the best supervised learning-to-rank model (SVM-Rank) significantly surpassed state-of-the-art baseline systems.

📄 PDF Abstract BibTeX arXiv:1611.01400

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDocument RankingLearning-To-Ranktext similarity

Similar Papers 제목 키워드 기반

LLM-Based Compact Reranking with Document Features for Scientific Retrieval

2025-05-19 · Runchu Tian, Xueqiang Xu, Bowen Jin, SeongKu Kang 외

Scientific retrieval is essential for advancing academic discovery. Within this process, document reranking plays a critical role by refining first-stage retrieval results. However, large language model (LLM) listwise re…

Large Language ModelRerankingRetrieval

Key2Vec: Automatic Ranked Keyphrase Extraction from Scientific Articles using Phrase Embeddings

2018-06-01 · NAACL 2018 6 · Debanjan Mahata, John Kuriakose, Rajiv Ratn Shah, Roger Zimmermann

Keyphrase extraction is a fundamental task in natural language processing that facilitates mapping of documents to a set of representative phrases. In this paper, we present an unsupervised technique (Key2Vec) that lever…

ArticlesChunkingKeyphrase ExtractionNamed Entity Recognition (NER)+3

+VeriRel: Verification Feedback to Enhance Document Retrieval for Scientific Fact Checking

2025-08-14 · Xingyu Deng, Xi Wang, Mark Stevenson arxiv

Identification of appropriate supporting evidence is critical to the success of scientific fact checking. However, existing approaches rely on off-the-shelf Information Retrieval algorithms that rank documents based on r…

Information RetrievalDocument RankingFact Checking

CNLP-NITS @ LongSumm 2021: TextRank Variant for Generating Long Summaries

2021-06-01 · NAACL (sdp) 2021 6 · Darsh Kaushik, Abdullah Faiz Ur Rahman Khilji, Utkarsh Sinha, Partha Pakray

The huge influx of published papers in the field of machine learning makes the task of summarization of scholarly documents vital, not just to eliminate the redundancy but also to provide a complete and satisfying crux o…

Extractive Summarization

Translation Using JAPIO Patent Corpora: JAPIO at WAT2016

2016-12-01 · WS 2016 12 · Satoshi Kinoshita, Tadaaki Oshio, Tomoharu Mitsuhashi, Terumasa Ehara

We participate in scientific paper subtask (ASPEC-EJ/CJ) and patent subtask (JPC-EJ/CJ/KJ) with phrase-based SMT systems which are trained with its own patent corpora. Using larger corpora than those prepared by the work…

Information RetrievalMachine TranslationTranslation