paper-with-me

홈 › Papers

LT Expertfinder: An Evaluation Framework for Expert Finding Methods

2019-06-01 · NAACL 2019 6 · Tim Fischer, Steffen Remus, Chris Biemann

Expert finding is the task of ranking persons for a predefined topic or search query. Finding experts for a specified area is an important task and has attracted much attention in the information retrieval community. Most approaches for this task are evaluated in a supervised fashion, which depend on predefined topics of interest as well as gold standard expert rankings. Famous representatives of such datasets are enriched versions of DBLP provided by the ArnetMiner projet or the W3C Corpus of TREC. However, manually ranking experts can be considered highly subjective and detailed rankings are hardly distinguishable. Evaluating these datasets does not necessarily guarantee a good or bad performance of the system. Particularly for dynamic systems, where topics are not predefined but formulated as a search query, we believe a more informative approach is to perform user studies for directly comparing different methods in the same view. In order to accomplish this in a user-friendly way, we present the LT Expert Finder web-application, which is equipped with various query-based expert finding methods that can be easily extended, a detailed expert profile view, detailed evidence in form of relevant documents and statistics, and an evaluation component that allows the qualitative comparison between different rankings.

📄 PDF Abstract BibTeX

Code (1)

uhh-lt/lt-expertfinder 공식 구현

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

On the Biased Assessment of Expert Finding Systems

2024-10-07 · Jens-Joris Decorte, Jeroen Van Hautte, Chris Develder, Thomas Demeester

In large organisations, identifying experts on a given topic is crucial in leveraging the internal knowledge spread across teams and departments. So-called enterprise expert retrieval systems automatically discover and s…

Retrieval

Integrated Framework for LLM Evaluation with Answer Generation

2025-09-24 · Sujeong Lee, Hayoung Lee, Seongsoo Heo, Wonik Choi arxiv

Reliable evaluation of large language models is essential to ensure their applicability in practical scenarios. Traditional benchmark-based evaluation methods often rely on fixed reference answers, limiting their ability…

Answer Generation

Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models

2025-04-17 · Sudesh Ramesh Bhagat, Ibne Farabi Shihab, Anuj Sharma

This study investigates the relationship between deep learning (DL) model accuracy and expert agreement in classifying crash narratives. We evaluate five DL models -- including BERT variants, USE, and a zero-shot classif…

Sentence

Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks

2026-03-06 · Jena D. Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad 외 arxiv

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, along with meta-evaluation frameworks that s…

Human-Centered Evaluation of an LLM-Based Process Modeling Copilot: A Mixed-Methods Study with Domain Experts

2026-03-13 · Chantale Lauer, Peter Pfeiffer, Nijat Mehdiyev arxiv

Integrating Large Language Models (LLMs) into business process management tools promises to democratize Business Process Model and Notation (BPMN) modeling for non-experts. While automated frameworks assess syntactic and…