paper-with-me

홈 › Papers

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

2026-04-18 · Isabella Luong, Joyee Chen, Arturs Kanepajs, Jasmine Brazilek, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu arxiv

Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt to three adversarial pressure rounds drawn from a five-type taxonomy: Social, Cultural, Economic, Pragmatic, and Epistemic. We score conversations on two dimensions: Animal Welfare Value Stability (AWVS, primary) and Animal Welfare Moral Sensitivity (AWMS, diagnostic). We evaluate seven frontier models: Claude Opus 4.7, GPT-5.5, DeepSeek V4, Llama 3.3 70B, Mistral Small, Grok 4.3, and Gemini 3.1 Flash Lite. Multi-turn evaluation captures behavior single-turn benchmarks miss: 4 of 7 models change rank relative to Turn 1 scores, including Gemini Flash Lite, which drops from fifth on AWMS to last on AWVS. AWMS and AWVS are positively but imperfectly correlated, suggesting moral-recognition tests capture a stable but incomplete component of model behavior under pressure. MANTA also enables a species-by-pressure interaction matrix unavailable to prior benchmarks, showing welfare robustness depends jointly on the animal and pressure applied; companion animals score above wild animals, which score above farmed animals and invertebrates. We release the dataset, scripted pressure plans, judge prompts, and analysis code.

📄 PDF Abstract BibTeX arXiv:2605.16301

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

2022-12-14 · Nathan Godey, Roman Castagné, Éric de la Clergerie, Benoît Sagot

Static subword tokenization algorithms have been an essential component of recent works on language modeling. However, their static nature results in important flaws that degrade the models' downstream performance and ro…

Language ModelingLanguage Modelling

Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence

2024-12-10 · Wenbo Huang, Jinghui Zhang, Guang Li, Lei Zhang 외

In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their applic…

Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action Recognition+1

MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems

2026-07-30 · Mao-xun Huang, Jerry Wang, Yi-Cheng Lai, Zhengxin Zhang 외 arxiv

Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent specialization, information exchange, and intermediate validation. However, existing systems typically trea…

Mathematical Reasoning

On a Well-behaved Relational Generalisation of Rough Set Approximations

2016-12-05 · Alexa Gopaulsingh

We examine non-dual relational extensions of rough set approximations and find an extension which satisfies surprisingly many of the usual rough set properties. We then use this definition to give an explanation for an o…

Survey

MANTA -- Model Adapter Native generations that's Affordable

2024-09-22 · Ansh Chaurasia

The presiding model generation algorithms rely on simple, inflexible adapter selection to provide personalized results. We propose the model-adapter composition problem as a generalized problem to past work factoring in …

DiversitymodelSynthetic Data Generation