paper-with-me

홈 › Papers

Enhancing Textbooks with Visuals from the Web for Improved Learning

2023-04-18 · Janvijay Singh, Vilém Zouhar, Mrinmaya Sachan

Textbooks are one of the main mediums for delivering high-quality education to students. In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge. However, many textbooks lack these interesting visuals to support student learning. In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web. We collect a dataset of e-textbooks in the math, science, social science and business domains. We then set up a text-image matching task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a matching optimization problem. Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the precise formulation of the optimization problem matters. We release the dataset of textbooks with an associated image bank to inspire further research in this intersectional area of computer vision and NLP for education.

📄 PDF Abstract BibTeX arXiv:2304.08931

Code (1)

eth-nlped/textbook-enrichment 공식 구현

Tasks

Math

Similar Papers 제목 키워드 기반

OmniCaptioner: One Captioner to Rule Them All

2025-04-09 · Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao 외

We propose OmniCaptioner, a versatile visual captioning framework for generating fine-grained textual descriptions across a wide variety of visual domains. Unlike prior methods limited to specific image types (e.g., natu…

AllImage CaptioningImage GenerationText to Image Generation+2

Enhancing Programming eTextbooks with ChatGPT Generated Counterfactual-Thinking-Inspired Questions

2023-06-01 · Arun Balajiee Lekshmi Narayanan, Rully Agus Hendrawan, Venktesh V

Digital textbooks have become an integral part of everyday learning tasks. In this work, we consider the use of digital textbooks for programming classes. Generally, students struggle with utilizing textbooks on programm…

counterfactual

VisualSpeech: Enhance Prosody with Visual Context in TTS

2025-01-31 · Shumin Que, Anton Ragni

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic informatio…

Prosody Predictiontext-to-speechText to Speech

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL

2025-05-29 · Yichen Feng, Zhangchen Xu, Fengqing Jiang, Yuetai Li 외

Vision language models (VLMs) are expected to perform effective multimodal reasoning and make logically coherent decisions, which is critical to tasks such as diagram understanding and spatial problem solving. However, c…

Arithmetic ReasoningImage GenerationLogical ReasoningMultimodal Reasoning

OpenStaxQA: A multilingual dataset based on open-source college textbooks

2025-10-03 · Pranav Gupta arxiv

We present OpenStaxQA, an evaluation benchmark specific to college-level educational applications based on 43 open-source college textbooks in English, Spanish, and Polish, available under a permissive Creative Commons l…