paper-with-me

홈 › Papers

Synthetic Books

2022-01-24 · Varvara Guljajeva

The article explores new ways of written language aided by AI technologies, like GPT-2 and GPT-3. The question that is stated in the paper is not about whether these novel technologies will eventually replace authored books, but how to relate to and contextualize such publications and what kind of new tools, processes, and ideas are behind them. For that purpose, a new concept of synthetic books is introduced in the article. It stands for the publications created by deploying AI technology, more precisely autoregressive language models that are able to generate human-like text. Supported by the case studies, the value and reasoning of the synthetic books are discussed. The paper emphasizes that artistic quality is an issue when it comes to AI-generated content. The article introduces projects that demonstrate an interactive input by an artist and/or audience combined with the deep-learning-based language models. In the end, the paper focuses on understanding the neural aesthetics of written language in the art context.

📄 PDF Abstract BibTeX arXiv:2201.09518

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Image quality prediction using synthetic and natural codebooks: comparative results

2022-12-20 · Maxim Koroteev, Kirill Aistov, Valeriy Berezovskiy, Pavel Frolov

We investigate a model for image/video quality assessment based on building a set of codevectors representing in a sense some basic properties of images, similar to well-known CORNIA model. We analyze the codebook buildi…

CPUVideo Quality Assessment

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

2026-07-30 · Jiawen Tao, Miao Peng, Yaoming Li, Xiaokun Yuan 외 arxiv

Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a different factor: whether related conten…

Improving Text Relationship Modeling with Artificial Data

2020-10-27 · Peter Organisciak, Maggie Ryan

Data augmentation uses artificially-created examples to support supervised machine learning, adding robustness to the resulting models and helping to account for limited availability of labelled data. We apply and evalua…

BIG-bench Machine LearningClassificationData AugmentationGeneral Classification

Using Audio Books for Training a Text-to-Speech System

2014-05-01 · LREC 2014 5 · Chalam, Aimilios aris, Pirros Tsiakoulis, Sotiris Karabetsos 외

Creating new voices for a TTS system often requires a costly procedure of designing and recording an audio corpus, a time consuming and effort intensive task. Using publicly available audiobooks as the raw material of a …

DiversitySpeech Synthesistext-to-speechText to Speech

One Thousand and One Pairs: A "novel" challenge for long-context language models

2024-06-24 · Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal 외

Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-leng…

RetrievalSentence