paper-with-me

Papers

Decoding, Fast and Slow: A Case Study on Balancing Trade-Offs in Incremental, Character-level Pragmatic Reasoning

2021-08-01 · INLG (ACL) 2021 8 · Sina Zarrieß, Hendrik Buschmeier, Ting Han, Simeon Schüz

Recent work has adopted models of pragmatic reasoning for the generation of informative language in, e.g., image captioning. We propose a simple but highly effective relaxation of fully rational decoding, based on an existing incremental and character-level approach to pragmatically informative neural image captioning. We implement a mixed, ‘fast’ and ‘slow’, speaker that applies pragmatic reasoning occasionally (only word-initially), while unrolling the language model. In our evaluation, we find that increased informativeness through pragmatic decoding generally lowers quality and, somewhat counter-intuitively, increases repetitiveness in captions. Our mixed speaker, however, achieves a good balance between quality and informativeness.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningInformativenessLanguage ModelingLanguage ModellingRolling Shutter Correction

Similar Papers 제목 키워드 기반

Slow-Fast Inference: Training-Free Inference Acceleration via Within-Sentence Support Stability

2026-03-12 · Xingyu Xie, Zhaochen Yu, Yue Liao, Tao Wang 외 arxiv

Long-context autoregressive decoding remains expensive because each decoding step must repeatedly process a growing history. We observe a consistent pattern during decoding: within a sentence, and more generally within a…

Cross-Balancing for Data-Informed Design and Efficient Analysis of Observational Studies

2025-11-19 · Ying Jin, José Zubizarreta arxiv

Causal inference starts with a simple idea: compare groups that differ by treatment, not much else. Traditionally, similar groups are constructed using only observed covariates; however, it remains a long-standing challe…

Causal Inference

Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding

2024-06-18 · Kaiyan Zhang, Jianyu Wang, Ning Ding, Biqing Qi 외

Large Language Models (LLMs) exhibit impressive capabilities across various applications but encounter substantial challenges such as high inference latency, considerable training costs, and the generation of hallucinati…

Hallucination

Sparse is Enough in Scaling Transformers

2021-11-24 · NeurIPS 2021 12 · Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser 외

Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging…

Text Summarization

Regularization Implies balancedness in the deep linear network

2025-11-03 · Kathryn Lindsey, Govind Menon arxiv

We use geometric invariant theory (GIT) to study the deep linear network (DLN). The Kempf-Ness theorem is used to establish that the $L^2$ regularizer is minimized on the balanced manifold. We introduce related balancing…