paper-with-me

홈 › Papers

ANTONIO: Towards a Systematic Method of Generating NLP Benchmarks for Verification

2023-05-06 · Marco Casadio, Luca Arnaboldi, Matthew L. Daggitt, Omri Isac, Tanvi Dinkar, Daniel Kienitz, Verena Rieser, Ekaterina Komendantskaya

Verification of machine learning models used in Natural Language Processing (NLP) is known to be a hard problem. In particular, many known neural network verification methods that work for computer vision and other numeric datasets do not work for NLP. Here, we study technical reasons that underlie this problem. Based on this analysis, we propose practical methods and heuristics for preparing NLP datasets and models in a way that renders them amenable to known verification methods based on abstract interpretation. We implement these methods as a Python library called ANTONIO that links to the neural network verifiers ERAN and Marabou. We perform evaluation of the tool using an NLP dataset R-U-A-Robot suggested as a benchmark for verifying legally critical NLP applications. We hope that, thanks to its general applicability, this work will open novel possibilities for including NLP verification problems into neural network verification competitions, and will popularise NLP problems within this community.

📄 PDF Abstract BibTeX arXiv:2305.04003

Code (1)

antonionlp/antonio 공식 구현 tf

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

Systematic Generation of Diverse Benchmarks for DNN Verification

2020-07-14 · Dong Xu, David Shriver, Matthew B. Dwyer, Sebastian Elbaum

The field of verification has advanced due to the interplay of theoretical development and empirical evaluation. Benchmarks play an important role in this by supporting the assessment of the state-of-the-art and comparis…

Stress-Testing Neural Network Verifiers with Provably Robust Instances

2026-05-16 · David Troxell, Yulia Alexandr, Sofia Hunt, Stephanie Lei 외 arxiv

Neural network verifiers aim to provide formal guarantees on model behavior, but existing verification benchmarks are fundamentally limited by their lack of ground-truth labels. As a result, verifier evaluation relies on…

Exploring the Potential of Large Language Models in Public Transportation: San Antonio Case Study

2025-01-07 · Ramya Jonnala, Gongbo Liang, Jeong Yang, Izzat Alsmadi

The integration of large language models (LLMs) into public transit systems presents a transformative opportunity to enhance urban mobility. This study explores the potential of LLMs to revolutionize public transportatio…

Decision Making

Flawed Waveform Design of Augusto Aubry, Antonio DeMaio et al

2018-02-26

arXiv admin note: This submission has been withdrawn by arXiv administrators due to unprofessional personal attack.

Formal Evidence Generation for Assurance Cases for Robotic Software Models

2026-02-03 · Fang Yan, Simon Foster, Ana Cavalcanti, Ibrahim Habli 외 arxiv

Robotics and Autonomous Systems are increasingly deployed in safety-critical domains, so that demonstrating their safety is essential. Assurance Cases (ACs) provide structured arguments supported by evidence, but generat…