paper-with-me

홈 › Papers

Sanvaad: A Multimodal Accessibility Framework for ISL Recognition and Voice-Based Interaction

2025-12-06 · Kush Revankar, Shreyas Deshpande, Araham Sayeed, Ansh Tandale, Sarika Bobde arxiv

Communication between deaf users, visually im paired users, and the general hearing population often relies on tools that support only one direction of interaction. To address this limitation, this work presents Sanvaad, a lightweight multimodal accessibility framework designed to support real time, two-way communication. For deaf users, Sanvaad includes an ISL recognition module built on MediaPipe landmarks. MediaPipe is chosen primarily for its efficiency and low computational load, enabling the system to run smoothly on edge devices without requiring dedicated hardware. Spoken input from a phone can also be translated into sign representations through a voice-to-sign component that maps detected speech to predefined phrases and produces corresponding GIFs or alphabet-based visualizations. For visually impaired users, the framework provides a screen free voice interface that integrates multilingual speech recognition, text summarization, and text-to-speech generation. These components work together through a Streamlit-based interface, making the system usable on both desktop and mobile environments. Overall, Sanvaad aims to offer a practical and accessible pathway for inclusive communication by combining lightweight computer vision and speech processing tools within a unified framework.

📄 PDF Abstract BibTeX arXiv:2512.06485

Code (0)

등록된 구현이 없습니다.

Tasks

Text SummarizationSpeech Recognition

Similar Papers 제목 키워드 기반

Sign Language Recognition Analysis using Multimodal Data

2019-09-24 · Al Amin Hosain, Panneer Selvam Santhalingam, Parth Pathak, Jana Kosecka 외

Voice-controlled personal and home assistants (such as the Amazon Echo and Apple Siri) are becoming increasingly popular for a variety of applications. However, the benefits of these technologies are not readily accessib…

Activity RecognitionHuman Activity RecognitionSign Language Recognition

GenAI Voice Mode in Programming Education

2025-09-12 · Sven Jacobs, Natalie Kiesler arxiv

Real-time voice interfaces using multimodal Generative AI (GenAI) can potentially address the accessibility needs of novice programmers with disabilities (e.g., related to vision). Yet, little is known about how novices …

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…

Speaker Recognition

Development of an Inclusive Educational Platform Using Open Technologies and Machine Learning: A Case Study on Accessibility Enhancement

2025-01-22 · Jimi Togni

This study addresses the pressing challenge of educational inclusion for students with special needs by proposing and developing an inclusive educational platform. Integrating machine learning, natural language processin…

Object Recognitionspeech-recognitionSpeech RecognitionText Generation+2

Greater accessibility can amplify discrimination in generative AI

2026-03-23 · Carolin Holtermann, Minh Duc Bui, Kaitlyn Zhou, Valentin Hofmann 외 arxiv

Hundreds of millions of people rely on large language models (LLMs) for education, work, and even healthcare. Yet these models are known to reproduce and amplify social biases present in their training data. Moreover, te…