paper-with-me

Papers

PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

2026-03-15 · Mohammad Javad Ranjbar Kalahroodi, Mohammad Amini, Parmis Bathayan, Heshaam Faili, Azadeh Shakery arxiv

Persian poses unique audio understanding challenges through its classical poetry, traditional music, and pervasive code-switching - none captured by existing benchmarks. We introduce PARSA-Bench (Persian Audio Reasoning and Speech Assessment Benchmark), the first benchmark for evaluating large audio-language models on Persian language and culture, comprising 16 tasks and over 8,000 samples across speech understanding, paralinguistic analysis, and cultural audio understanding. Ten tasks are newly introduced, including poetry meter and style detection, traditional Persian music understanding, and code-switching detection. Text-only baselines consistently outperform audio counterparts, suggesting models may not leverage audio-specific information beyond what transcription alone provides. Culturally-grounded tasks expose a qualitatively distinct failure mode: all models perform near random chance on vazn detection regardless of scale, suggesting prosodic perception remains beyond the reach of current models. The dataset is publicly available at https://huggingface.co/datasets/MohammadJRanjbar/PARSA-Bench

📄 PDF Abstract BibTeX arXiv:2603.14456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ELAB: Extensive LLM Alignment Benchmark in Persian Language

2025-04-17 · Zahra Pourbahman, Fatemeh Rajabi, Mohammadhossein Sadeghi, Omid Ghahroodi 외

This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing…

FairnessRed Teaming

A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset

2023-01-21 · Javad Peymanfard, Samin Heydarian, Ali Lashini, Hossein Zeinali 외

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multi…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Reading+4

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems

2025-05-27 · Nima Sedghiyeh, Sara Sadeghi, Reza Khodadadi, Farzin Kashani 외

Although Automatic Speech Recognition (ASR) systems have become an integral part of modern technology, their evaluation remains challenging, particularly for low-resource languages such as Persian. This paper introduces …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

FaMTEB: Massive Text Embedding Benchmark in Persian Language

2025-02-17 · Erfan Zinvandi, Morteza Alikhani, Mehran Sarmadi, Zahra Pourbahman 외

In this paper, we introduce a comprehensive benchmark for Persian (Farsi) text embeddings, built upon the Massive Text Embedding Benchmark (MTEB). Our benchmark includes 63 datasets spanning seven different tasks: classi…

ChatbotMTEB BenchmarkRerankingRetrieval+2

PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian

2026-02-01 · Jamshid Mozafari, Seyed Parsa Mousavinasab, Adam Jatowt arxiv

Reasoning-focused Question Answering (QA) has advanced rapidly with Large Language Models (LLMs), yet high-quality benchmarks for low-resource languages remain scarce. Persian, spoken by roughly 130 million people, lacks…

Question Answering