paper-with-me

Papers

Beyond the Textual: Generating Coherent Visual Options for MCQs

2025-08-26 · Wanqiang Wang, Longzhu He, Wei Zheng arxiv

Multiple-choice questions (MCQs) play a crucial role in fostering deep thinking and knowledge integration in education. However, previous research has primarily focused on generating MCQs with textual options, but it largely overlooks the visual options. Moreover, generating high-quality distractors remains a major challenge due to the high cost and limited scalability of manual authoring. To tackle these problems, we propose a Cross-modal Options Synthesis (CmOS), a novel framework for generating educational MCQs with visual options. Our framework integrates Multimodal Chain-of-Thought (MCoT) reasoning process and Retrieval-Augmented Generation (RAG) to produce semantically plausible and visually similar answer and distractors. It also includes a discrimination module to identify content suitable for visual options. Experimental results on test tasks demonstrate the superiority of CmOS in content discrimination, question generation and visual option generation over existing methods across various subjects and educational levels.

📄 PDF Abstract BibTeX arXiv:2508.18772

Code (0)

등록된 구현이 없습니다.

Tasks

Question Generation

Similar Papers 제목 키워드 기반

ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context

2024-07-13 · Sixiao Zheng, Yanwei Fu

Visual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes. Existing autoregressive methods, which rely on previous frame-sentenc…

Image GenerationStory ContinuationStory VisualizationText-to-Image Generation+1

Latent Beam Diffusion Models for Decoding Image Sequences

2025-03-26 · Guilherme Fernandes, Vasco Ramos, Regev Cohen, Idan Szpektor 외

While diffusion models excel at generating high-quality images from text prompts, they struggle with visual consistency in image sequences. Existing methods generate each image independently, leading to disjointed narrat…

CI-VID: A Coherent Interleaved Text-Video Dataset

2025-07-02 · Yiming Ju, Jijin Hu, Zhengxiong Luo, Haoge Deng 외 arxiv

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets ar…

Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

2024-05-16 · João Bordalo, Vasco Ramos, Rodrigo Valério, Diogo Glória-Silva 외

Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language Models (LLMs) have become adept at genera…

Beyond Words: Multimodal LLM Knows When to Speak

2025-05-20 · Zikai Liao, Yi Ouyang, Yi-Lun Lee, Chen-Ping Yu 외

While large language model (LLM)-based chatbots have demonstrated strong capabilities in generating coherent and contextually relevant responses, they often struggle with understanding when to speak, particularly in deli…

Large Language Model