paper-with-me

Topic coverage

3개 벤치마크 · 논문 30편 · 이 태스크의 논문 보기 →

Benchmarks

Most implemented

Papers

X+Slides: Benchmarking Audience-Conditioned Slide Generation

2026-06-17 · Haodong Chen, Xuanhe Zhou, Wei Zhou, Xinyue Shao 외 arxiv

Automatically generating slide decks from source documents is an important application of large language models (LLMs). Existing benchmarks primarily assess slide completeness and technical depth, while overlooking the t…

Topic coverage

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

2026-06-06 · Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma 외 arxiv

Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence. Here, overconfidence is defined b…

Topic coverage

Agent-Orchestrated Adaptive RAG: A Comparative Study on Structured and Multi-Hop Retrieval

2026-06-04 · Anuj Maharjan, Devinder Kaur, Richard Molyet arxiv

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipelines rely on static, single-step retrieval that limits performance on c…

Topic coverage

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

2026-06-02 · Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia, Abbas Goher Khan 외 arxiv

High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, we…

Topic coverage

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

2026-05-11 · Shu Wang, Shansong Zhou, Xinyang Wang, Shiwei Wang 외 arxiv

Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or across multiple documents into coherent answers. However, this settin…

Question AnsweringTopic coverage

From Similarity to Structure: Training-free LLM Context Compression with Hybrid Graph Priors

2026-04-25 · Yitian Zhou, Chaoning Zhang, Jiaquan Zhang, Zhenzhen Huang 외 arxiv

Long-context large language models remain computationally expensive to run and often fail to reliably process very long inputs, which makes context compression an important component of many systems. Existing compression…

Topic coverage

전체 30편 보기 →