paper-with-me

홈 › Papers

SentSpace: Large-Scale Benchmarking and Evaluation of Text using Cognitively Motivated Lexical, Syntactic, and Semantic Features

2022-07-01 · NAACL (ACL) 2022 7 · Greta Tuckute, Aalok Sathe, Mingye Wang, Harley Yoder, Cory Shain, Evelina Fedorenko

SentSpace is a modular framework for streamlined evaluation of text. SentSpacecharacterizes textual input using diverse lexical, syntactic, and semantic features derivedfrom corpora and psycholinguistic experiments. Core sentence features fall into three primaryfeature spaces: 1) Lexical, 2) Contextual, and 3) Embeddings. To aid in the analysis of computed features, SentSpace provides a web interface for interactive visualization and comparison with text from large corpora. The modular design of SentSpace allows researchersto easily integrate their own feature computation into the pipeline while benefiting from acommon framework for evaluation and visualization. In this manuscript we will describe thedesign of SentSpace, its core feature spaces, and demonstrate an example use case by comparing human-written and machine-generated (GPT2-XL) sentences to each other. We findthat while GPT2-XL-generated text appears fluent at the surface level, psycholinguistic normsand measures of syntactic processing reveal key differences between text produced by humansand machines. Thus, SentSpace provides a broad set of cognitively motivated linguisticfeatures for evaluation of text within natural language processing, cognitive science, as wellas the social sciences.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSentence

Similar Papers 제목 키워드 기반

Personalized Benchmarking with the Ludwig Benchmarking Toolkit

2021-11-08 · Avanika Narayan, Piero Molino, Karan Goel, Willie Neiswanger 외

The rapid proliferation of machine learning models across domains and deployment settings has given rise to various communities (e.g. industry practitioners) which seek to benchmark models across tasks and objectives of …

BenchmarkingHyperparameter Optimizationtext-classificationText Classification

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation

2025-05-17 · Jiarui Wang, Huiyu Duan, Ziheng Jia, Yu Zhao 외

Recent advancements in large multimodal models (LMMs) have driven substantial progress in both text-to-video (T2V) generation and video-to-text (V2T) interpretation tasks. However, current AI-generated videos (AIGVs) sti…

BenchmarkingQuestion AnsweringText-to-Video GenerationVideo Alignment+1

NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking

2024-06-21 · Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng 외

Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possi…

Autonomous DrivingBenchmarkingNavSim

FedScale: Benchmarking Model and System Performance of Federated Learning at Scale

2021-05-24 · Fan Lai, Yinwei Dai, Sanjay S. Singapuram, Jiachen Liu 외

We present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging …

BenchmarkingFederated Learningimage-classificationImage Classification+6

MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization

2026-03-26 · Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang, Zheng Hui 외 arxiv

Recent advancements in Large Language Models (LLMs) have expanded context windows to million-token scales, yet benchmarks for evaluating memory remain limited to short-session synthetic dialogues. We introduce \textsc{Me…