paper-with-me

홈 › Papers

Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

2024-10-28 · Michael Pieler, Marco Bellagente, Hannah Teufel, Duy Phung, Nathan Cooper, Jonathan Tow, Paulo Rocha, Reshinth Adithyan, Zaid Alyafeai, Nikhil Pinnaparaju, Maksym Zhuravinskyi, Carlos Riquelme

Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data. We build upon previous work by replicating existing results on C4 and extending them with our optimized rephrasing pipeline to the English, German, Italian, and Spanish Oscar subsets of CulturaX. Our pipeline leads to increased performance on standard evaluation benchmarks in both the mono- and multilingual setup. In addition, we provide a detailed study of our pipeline, investigating the choice of the base dataset and LLM for the rephrasing, as well as the relationship between the model size and the performance after pre-training. By exploring data with different perceived quality levels, we show that gains decrease with higher quality. Furthermore, we find the difference in performance between model families to be bigger than between different model sizes. This highlights the necessity for detailed tests before choosing an LLM to rephrase large amounts of data. Moreover, we investigate the effect of pre-training with synthetic data on supervised fine-tuning. Here, we find increasing but inconclusive results that highly depend on the used benchmark. These results (again) highlight the need for better benchmarking setups. In summary, we show that rephrasing multilingual and low-quality data is a very promising direction to extend LLM pre-training data.

📄 PDF Abstract BibTeX arXiv:2410.20796

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

OSCAR OSCAR is a new learning method that uses object tags detected in images as anchor points to ease the learning of image-text alignment. The model take a triple as input…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Sound Natural: Content Rephrasing in Dialog Systems

2020-11-03 · EMNLP 2020 11 · Arash Einolghozati, Anchit Gupta, Keith Diedrick, Sonal Gupta

We introduce a new task of rephrasing for a more natural virtual assistant. Currently, virtual assistants work in the paradigm of intent slot tagging and the slot values are directly passed as-is to the execution engine.…

Language ModellingParaphrase Generation

MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs

2025-07-13 · Shulin Huang, Linyi Yang, Yue Zhang arxiv

Large language models exhibit cultural biases and limited cross-cultural understanding capabilities, particularly when serving diverse global user populations. We propose MCEval, a novel multilingual evaluation framework…

RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation

2025-11-06 · Xiangjun Zhang, Litong Gong, Yinglin Zheng, Yansong Liu 외 arxiv

Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompts rather than well-designed ones. The pr…

Text-to-Video Generation

Can Pre-training help VQA with Lexical Variations?

2020-11-01 · Findings of the Association for Computational Linguistics 2020 · Shailza Jolly, Shubham Kapoor

Rephrasings or paraphrases are sentences with similar meanings expressed in different ways. Visual Question Answering (VQA) models are closing the gap with the oracle performance for datasets like VQA2.0. However, these …

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Rephrasing Profanity in Chinese Text

2017-08-01 · WS 2017 8 · Hui-Po Su, Zhen-Jie Huang, Hao-Tsung Chang, Chuan-Jie Lin

This paper proposes a system that can detect and rephrase profanity in Chinese text. Rather than just masking detected profanity, we want to revise the input sentence by using inoffensive words while keeping their origin…

Sentence