paper-with-me

홈 › Papers

Best Practices for Text Annotation with Large Language Models

2024-02-05 · Petter Törnberg

Large Language Models (LLMs) have ushered in a new era of text annotation, as their ease-of-use, high accuracy, and relatively low costs have meant that their use has exploded in recent months. However, the rapid growth of the field has meant that LLM-based annotation has become something of an academic Wild West: the lack of established practices and standards has led to concerns about the quality and validity of research. Researchers have warned that the ostensible simplicity of LLMs can be misleading, as they are prone to bias, misunderstandings, and unreliable results. Recognizing the transformative potential of LLMs, this paper proposes a comprehensive set of standards and best practices for their reliable, reproducible, and ethical use. These guidelines span critical areas such as model selection, prompt engineering, structured prompting, prompt stability analysis, rigorous model validation, and the consideration of ethical and legal implications. The paper emphasizes the need for a structured, directed, and formalized approach to using LLMs, aiming to ensure the integrity and robustness of text annotation practices, and advocates for a nuanced and critical engagement with LLMs in social scientific research.

📄 PDF Abstract BibTeX arXiv:2402.05129

Code (0)

등록된 구현이 없습니다.

Tasks

Model SelectionPrompt Engineeringtext annotation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MetaLint: Easy-to-Hard Generalization for Code Linting

2025-07-15 · Atharva Naik, Lawanya Baghel, Dhakshin Govindarajan, Darsh Agrawal 외 arxiv

Large language models excel at code generation but struggle with code linting, particularly in generalizing to unseen or evolving best practices beyond those observed during training. We introduce MetaLint, a meta-learni…

Code Generation

Semantic annotation for computational pathology: Multidisciplinary experience and best practice recommendations

2021-06-25 · Noorul Wahab, Islam M Miligy, Katherine Dodd, Harvir Sahota 외

Recent advances in whole slide imaging (WSI) technology have led to the development of a myriad of computer vision and artificial intelligence (AI) based diagnostic, prognostic, and predictive algorithms. Computational P…

Diagnostic

Reference and coreference in situated dialogue

2021-06-01 · NAACL (ALVR) 2021 6 · Sharid Loáiciga, Simon Dobnik, David Schlangen

In recent years several corpora have been developed for vision and language tasks. We argue that there is still significant room for corpora that increase the complexity of both visual and linguistic domains and which ca…

Prompt Refinement or Fine-tuning? Best Practices for using LLMs in Computational Social Science Tasks

2024-08-02

Large Language Models are expressive tools that enable complex tasks of text understanding within Computational Social Science. Their versatility, while beneficial, poses a barrier for establishing standardized best prac…

Magic Words or Methodical Work? Challenging Conventional Wisdom in LLM-Based Political Text Annotation

2026-03-27 · Lorcan McLaren, James Cross, Zuzanna Krakowska, Robin Rauner 외 arxiv

Political scientists are rapidly adopting large language models (LLMs) for text annotation, yet the sensitivity of annotation results to implementation choices remains poorly understood. Most evaluations test a single mo…

Prompt Engineering