paper-with-me

Papers

RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource Languages

2025-07-08 · Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Large language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks. This paper introduces RabakBench, a new multilingual safety benchmark localized to Singapore's unique linguistic context, covering Singlish, Chinese, Malay, and Tamil. RabakBench is constructed through a scalable three-stage pipeline: (i) Generate - adversarial example generation by augmenting real Singlish web content with LLM-driven red teaming; (ii) Label - semi-automated multi-label safety annotation using majority-voted LLM labelers aligned with human judgments; and (iii) Translate - high-fidelity translation preserving linguistic nuance and toxicity across languages. The final dataset comprises over 5,000 safety-labeled examples across four languages and six fine-grained safety categories with severity levels. Evaluations of 11 popular open-source and closed-source guardrail classifiers reveal significant performance degradation. RabakBench not only enables robust safety evaluation in Southeast Asian multilingual settings but also offers a reproducible framework for building localized safety datasets in low-resource environments. The benchmark dataset, including the human-verified translations, and evaluation code are publicly available.

📄 PDF Abstract BibTeX arXiv:2507.05980

Code (1)

govtech-responsibleai/rabakbench 공식 구현

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

Connecting Vision and Language with Video Localized Narratives

2023-02-22 · CVPR 2023 1 · Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut 외

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, th…

Question AnsweringVideo Narrative GroundingVideo Question Answering

NymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data

2026-03-19 · Daniel DeTone, Federica Bogo, Eric-Tuan Le, Duncan Frost 외 arxiv

The Nymeria Dataset, released in 2024, is a large-scale collection of in-the-wild human activities captured with multiple egocentric wearable devices that are spatially localized and temporally synchronized. It provides …

Point Clouds

Generating Labels for Regression of Subjective Constructs using Triplet Embeddings

2019-04-02 · Karel Mundnich, Brandon M. Booth, Benjamin Girault, Shrikanth Narayanan

Human annotations serve an important role in computational models where the target constructs under study are hidden, such as dimensions of affect. This is especially relevant in machine learning, where subjective labels…

regressionTriplet

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

2025-02-04 · Hongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen 외

User interface understanding with vision-language models has received much attention due to its potential for enabling next-generation software automation. However, existing UI datasets either only provide large-scale co…

On The Continuous Steering of the Scale of Tight Wavelet Frames

2015-12-07 · Zsuzsanna Püspöki, John Paul Ward, Daniel Sage, Michael Unser

In analogy with steerable wavelets, we present a general construction of adaptable tight wavelet frames, with an emphasis on scaling operations. In particular, the derived wavelets can be "dilated" by a procedure compara…