paper-with-me

홈 › Papers

Minerva: A Programmable Memory Test Benchmark for Language Models

2025-02-05 · Menglin Xia, Victor Ruehle, Saravan Rajmohan, Reza Shokri

How effectively can LLM-based AI assistants utilize their memory (context) to perform various tasks? Traditional data benchmarks, which are often manually crafted, suffer from several limitations: they are static, susceptible to overfitting, difficult to interpret, and lack actionable insights--failing to pinpoint the specific capabilities a model lacks when it does not pass a test. In this paper, we present a framework for automatically generating a comprehensive set of tests to evaluate models' abilities to use their memory effectively. Our framework extends the range of capability tests beyond the commonly explored (passkey, key-value, needle in the haystack) search, a dominant focus in the literature. Specifically, we evaluate models on atomic tasks such as searching, recalling, editing, matching, comparing information in context memory, performing basic operations when inputs are structured into distinct blocks, and maintaining state while operating on memory, simulating real-world data. Additionally, we design composite tests to investigate the models' ability to perform more complex, integrated tasks. Our benchmark enables an interpretable, detailed assessment of memory capabilities of LLMs.

📄 PDF Abstract BibTeX arXiv:2502.03358

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning

2026-01-15 · Darshan Singh, Arsha Nagrani, Kawshik Manikantan, Harman Singh 외 arxiv

Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, i…

Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs

2026-01-31 · Md Tanvirul Alam, Aritran Piplai, Ionut Cardei, Nidhi Rastogi 외 arxiv

Cyber threat intelligence (CTI) analysts routinely convert noisy, unstructured security artifacts into standardized, automation-ready representations. Although large language models (LLMs) show promise for this task, exi…

Reinforcement Learning

MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis

2021-07-13 · Haocheng Ren, Hao Zhang, Jia Zheng, Jiaxiang Zheng 외

With the rapid development of data-driven techniques, data has played an essential role in various computer vision tasks. Many realistic and synthetic datasets have been proposed to address different problems. However, t…

2D Semantic SegmentationDepth EstimationRoom Layout Estimation

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

2026-05-14 · Arsha Nagrani, Jasper Uijilings, Shyamal Buch, Tobias Weyand 외 arxiv

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation o…

Visual Reasoning

Minerva: A File-Based Ransomware Detector

2023-01-26 · Dorjan Hitaj, Giulio Pagnotta, Fabio De Gaspari, Lorenzo De Carli 외

Ransomware attacks have caused billions of dollars in damages in recent years, and are expected to cause billions more in the future. Consequently, significant effort has been devoted to ransomware detection and mitigati…

feature selection