paper-with-me

홈 › Papers

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

2026-04-17 · Jiacheng Shen, Masato Hagiwara, Milad Alizadeh, Ellen Gilsenan-McMahon, Marius Miron, David Robinson, Emmanuel Chemla, Sara Keen, Gagan Narula, Mathieu Laurière, Matthieu Geist, Olivier Pietquin arxiv

Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.

📄 PDF Abstract BibTeX arXiv:2604.16241

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BAGEL: Bootstrapping Agents by Guiding Exploration with Language

2024-03-12 · Shikhar Murty, Christopher Manning, Peter Shaw, Mandar Joshi 외

Following natural language instructions by executing actions in digital environments (e.g. web-browsers and REST APIs) is a challenging task for language model (LM) agents. Unfortunately, LM agents often fail to generali…

In-Context LearningLanguage ModelingLanguage Modelling

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

2026-07-03 · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue 외 arxiv

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vi…

multimodal generation

Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation

2026-02-08 · Shayan Ali Hassan, Tao Ni, Zafar Ayyub Qazi, Marco Canini arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation. However, these systems remain susceptible to malicious prompts that induce unsafe or po…

Natural Language Understanding

BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation

2026-02-10 · Yucheng Hu, Jianke Zhang, Yuanfei Luo, Yanjiang Guo 외 arxiv

Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recent Vision-Language-Action (VLA) models ha…

Emerging Properties in Unified Multimodal Pretraining

2025-05-20 · Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou 외

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open0source foundational model that natively supports multimoda…

Image EditingImage GenerationImage Manipulation+2