paper-with-me

Papers

LLMbench: A Comparative Close Reading Workbench for Large Language Models

2026-04-16 · David M. Berry arxiv

LLMbench is a browser-based workbench for the comparative close reading of large language model (LLM) outputs. Where existing tools for LLM comparison, such as Google PAIR's LLM Comparator are engineered for quantitative evaluation and user-rating metrics, LLMbench is oriented towards the hermeneutic practices of the digital humanities. Two model responses to the same prompt are side by side in annotatable panels with four analytical overlays (Probabilities for token-level log-probability inspection, Differences for word-level diff across the two panels, Tone for Hyland-style metadiscourse analysis, and Structure for sentence-level parsing with discourse connective highlighting), alongside five analytical modes, Stochastic Variation, Temperature Gradient, Prompt Sensitivity, Token Probabilities, and Cross-Model Divergence, that make the probabilistic structure of generated text legible at the token level. The tool treats the generated text as a research object in its own right from a probability distribution, a text that could have been otherwise, and provides visualisations including continuous heatmaps, entropy sparklines, pixel maps, and three-dimensional probability terrains, that show the counterfactual history from which each word emerged. This paper describes the tool's architecture, its six modes, and its design rationale, and argues that log-probability data, currently underused in humanistic and social-scientific readings of AI, is an important resource for a critical studies of generative AI models.

📄 PDF Abstract BibTeX arXiv:2604.15508

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

2025-09-19 · Nithin Somasekharan, Ling Yue, Yadi Cao, Weichao Li 외 arxiv

Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a critical and labor-intensive component -- …

SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding

2025-11-26 · Tae-Min Choi, Tae Kyeong Jeong, Garam Kim, Jaemin Lee 외 arxiv

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA)…

Visual Question AnsweringScene Understanding

WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting

2024-05-01 · Olly Styles, Sam Miller, Patricio Cerda-Mardini, Tanaya Guha 외

We introduce WorkBench: a benchmark dataset for evaluating agents' ability to execute tasks in a workplace setting. WorkBench contains a sandbox environment with five databases, 26 tools, and 690 tasks. These tasks repre…

Scheduling

A Workbench for Autograding Retrieve/Generate Systems

2024-05-21 · Laura Dietz

This resource paper addresses the challenge of evaluating Information Retrieval (IR) systems in the era of autoregressive Large Language Models (LLMs). Traditional methods relying on passage-level judgments are no longer…

DiversityInformation RetrievalRetrieval

NLP Workbench: Efficient and Extensible Integration of State-of-the-art Text Mining Tools

2023-03-02 · Peiran Yao, Matej Kosmajac, Abeer Waheed, Kostyantyn Guzhva 외

NLP Workbench is a web-based platform for text mining that allows non-expert users to obtain semantic understanding of large-scale corpora using state-of-the-art text mining models. The platform is built upon latest pre-…

Entity LinkingRelation ExtractionSemantic ParsingSentiment Analysis