paper-with-me

Papers

Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation

2026-04-20 · Nathikan Yodthapa, Thanapong Intharah, Sahan Bulathwela arxiv

Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augmentation strategy for training small language models (sLMs) to perform topic-controlled summarisation. We propose a pairwise data augmentation method that combines contexts from different documents to create contrastive training examples, enabling models to learn the relationship between topics and summaries more effectively. Using the SciTLDR dataset enriched with Wikipedia-derived topics, we systematically evaluate how augmentation scale affects model performance. Results show consistent improvements in win rate and semantic alignment as the augmentation scale increases, while the amount of real training data remains fixed. Consequently, a T5-base model trained with our augmentation approach achieves competitive performance relative to larger models, despite using significantly fewer parameters and substantially fewer real training examples.

📄 PDF Abstract BibTeX arXiv:2604.18087

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

A Novel Approach to Scalable and Automatic Topic-Controlled Question Generation in Education

2025-01-09 · Ziqing Li, Mutlu Cukurova, Sahan Bulathwela

The development of Automatic Question Generation (QG) models has the potential to significantly improve educational practices by reducing the teacher workload associated with creating educational content. This paper intr…

Data AugmentationQuestion GenerationQuestion-GenerationSpecificity

ODE-free Neural Flow Matching for One-Step Generative Modeling

2026-04-07 · Xiao Shou arxiv

Diffusion and flow matching models generate samples by learning time-dependent vector fields whose integration transports noise to data, requiring tens to hundreds of network evaluations at inference. We instead learn th…

Image Generation

Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

2025-08-12 · Francesco Fabbri, Gustavo Penha, Edoardo D'Amico, Alice Wang 외 arxiv

Evaluating personalized recommendations remains a central challenge, especially in long-form audio domains like podcasts, where traditional offline metrics suffer from exposure bias and online methods such as A/B testing…

Automated Feature-Topic Pairing: Aligning Semantic and Embedding Spaces in Spatial Representation Learning

2021-09-22 · Dongjie Wang, Kunpeng Liu, David Mohaisen, Pengyang Wang 외

Automated characterization of spatial data is a kind of critical geographical intelligence. As an emerging technique for characterization, Spatial Representation Learning (SRL) uses deep neural networks (DNNs) to learn n…

Representation Learning

DINO-RotateMatch: A Rotation-Aware Deep Framework for Robust Image Matching in Large-Scale 3D Reconstruction

2025-12-03 · Kaichen Zhang, Tianxiang Sheng, Xuanming Shi arxiv

This paper presents DINO-RotateMatch, a deep-learning framework designed to address the chal lenges of image matching in large-scale 3D reconstruction from unstructured Internet images. The method integrates a dataset-ad…

3D ReconstructionImage Matching