paper-with-me

홈 › Papers

MAviS: A Multimodal Conversational Assistant For Avian Species

2026-03-07 · Yevheniia Kryklyvets, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jinxing Zhou, Fahad Shabzan Khan, Rao Anwer, Salman Khan, Hisham Cholakkal arxiv

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models face challenges when it comes to specialized topics like avian species, making it harder to provide accurate and contextually relevant information in these areas. To address this limitation, we introduce the MAviS-Dataset, a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species, comprising both pretraining and instruction-tuning subsets enriched with structured question-answer pairs. Building on the MAviS-Dataset, we introduce MAviS-Chat, a multimodal LLM that supports audio, vision, and text and is designed for fine-grained species understanding, multimodal question answering, and scene-specific description generation. Finally, for quantitative evaluation, we present MAviS-Bench, a benchmark of over 25,000 QA pairs designed to assess avian species-specific perceptual and reasoning abilities across modalities. Experimental results show that MAviS-Chat outperforms the baseline MiniCPM-o-2.6 by a large margin, achieving state-of-the-art open-source results and demonstrating the effectiveness of our instruction-tuned MAviS-Dataset. Our findings highlight the necessity of domain-adaptive multimodal LLMs for ecological applications.

📄 PDF Abstract BibTeX arXiv:2603.07294

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

2026-01-10 · Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak 외 arxiv

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent S…

Speech Recognition

Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data

2024-06-19 · Michael Doell, Dominik Kuehn, Vanessa Suessle, Matthew J. Burnett 외

Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings is time-consuming and challenged by the presence of background noise in recordings. Existing models for sound event detection (SE…

Event DetectionSound Event Detection

Investigating Target Class Influence on Neural Network Compressibility for Energy-Autonomous Avian Monitoring

2026-02-19 · Nina Brolich, Simon Geis, Maximilian Kasper, Alexander Barnhill 외 arxiv

Biodiversity loss poses a significant threat to humanity, making wildlife monitoring essential for assessing ecosystem health. Avian species are ideal subjects for this due to their popularity and the ease of identifying…

MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling

2025-08-11 · Qian Wang, Ziqi Huang, Ruoxi Jia, Paul Debevec 외 arxiv

Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expressiveness. To mitigate these limitation…

Visual StorytellingAudio GenerationVideo Generation

Implications of the Trivers-Willard Sex Ratio Hypothesis for Avian Species and Poultry Production, And a Summary of the Historic Context of this Research

2017-08-30

At a theoretical level, the Trivers-Willard Sex Ratio Hypothesis applies to both avian species and mammals. This article, however, conjectures that at the statistical level, sex ratio effects are likely to produce sharpe…