paper-with-me

홈 › Papers

QuesGenie: Intelligent Multimodal Question Generation

2025-08-27 · Ahmed Mubarak, Amna Ahmed, Amira Nasser, Aya Mohamed, Fares El-Sadek, Mohammed Ahmed, Ahmed Salah, Youssef Sobhy arxiv

In today's information-rich era, learners have access to abundant educational resources, but the lack of practice materials tailored to these resources presents a significant challenge. This project addresses that gap by developing a multi-modal question generation system that can automatically generate diverse question types from various content formats. The system features four major components: multi-modal input handling, question generation, reinforcement learning from human feedback (RLHF), and an end-to-end interactive interface. This project lays the foundation for automated, scalable, and intelligent question generation, carefully balancing resource efficiency, robust functionality and a smooth user experience.

📄 PDF Abstract BibTeX arXiv:2509.03535

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Generation

Similar Papers 제목 키워드 기반

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

2026-02-27 · Xiang Deng, Feng Gao, Yong Zhang, Youxin Pang 외 arxiv

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or s…

Instruction FollowingQuestion Answering

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

2025-10-23 · Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran 외 arxiv

This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimod…

Question Answering

MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance

2025-09-10 · Kaikai Zhao, Zhaoxiang Liu, Peng Wang, Xin Wang 외 arxiv

General-domain large multimodal models (LMMs) have achieved significant advances in various image-text tasks. However, their performance in the Intelligent Traffic Surveillance (ITS) domain remains limited due to the abs…

Object LocalizationObject Counting

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

2024-02-20 · Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai 외

Asking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to…

Question GenerationQuestion-Generation

Efficient Multimodal Planning Agent for Visual Question-Answering

2026-01-28 · Zhuo Chen, Xinyu Geng, Xinyu Wang, Yong Jiang 외 arxiv

Visual Question-Answering (VQA) is a challenging multimodal task that requires integrating visual and textual information to generate accurate responses. While multimodal Retrieval-Augmented Generation (mRAG) has shown p…