paper-with-me

Papers

Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring

2025-07-30 · Sinh Trong Vu, Hieu Trung Pham, Dung Manh Nguyen, Hieu Minh Hoang, Nhu Hoang Le, Thu Ha Pham, Tai Tan Mai arxiv

Classroom behavior monitoring is a critical aspect of educational research, with significant implications for student engagement and learning outcomes. Recent advancements in Visual Question Answering (VQA) models offer promising tools for automatically analyzing complex classroom interactions from video recordings. In this paper, we investigate the applicability of several state-of-the-art open-source VQA models, including LLaMA2, LLaMA3, QWEN3, and NVILA, in the context of classroom behavior analysis. To facilitate rigorous evaluation, we introduce our BAV-Classroom-VQA dataset derived from real-world classroom video recordings at the Banking Academy of Vietnam. We present the methodology for data collection, annotation, and benchmark the performance of the selected VQA models on this dataset. Our initial experimental results demonstrate that all four models achieve promising performance levels in answering behavior-related visual questions, showcasing their potential in future classroom analytics and intervention systems.

📄 PDF Abstract BibTeX arXiv:2507.22369

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering

2019-10-01 · ICCV 2019 10 · Haonan Luo, Guosheng Lin, Zichuan Liu, Fayao Liu 외

Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to …

Embodied Question AnsweringQuestion AnsweringSegmentationVideo Segmentation+3

Learning Convolutional Text Representations for Visual Question Answering

2017-05-18 · Zhengyang Wang, Shuiwang Ji

Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural netwo…

General Classificationimage-classificationtext-classificationVisual Question Answering+1

Exploring Question Decomposition for Zero-Shot VQA

2023-10-25 · NeurIPS 2023 11

Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies. We explore a question decom…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

EDUQA: Educational Domain Question Answering System using Conceptual Network Mapping

2019-11-12 · Abhishek Agarwal, Nikhil Sachdeva, Raj Kamal Yadav, Vishaal Udandarao 외

Most of the existing question answering models can be largely compiled into two categories: i) open domain question answering models that answer generic questions and use large-scale knowledge base along with the targete…

Answer GenerationOpen-Domain Question AnsweringQuestion AnsweringRetrieval

What Would a Teacher Do? Predicting Future Talk Moves

2021-06-09 · Findings (ACL) 2021 8 · Ananya Ganesh, Martha Palmer, Katharina Kann

Recent advances in natural language processing (NLP) have the ability to transform how classroom learning takes place. Combined with the increasing integration of technology in today's classrooms, NLP systems leveraging …

Question Answering