paper-with-me

홈 › Papers

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

2025-08-07 · Nikita Dragunov, Temurbek Rahmatullaev, Elizaveta Goncharova, Nikita Kurdiukov, Aysel Mirzoeva, Anna Borisiuk, Andrey Kuznetsov, Anton Razzhigaev arxiv

The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-squared error or diffusion objectives. We present SONAR-LLM, a decoder-only transformer that "thinks" in the same continuous SONAR embedding space, yet is supervised through token-level cross-entropy propagated via the frozen SONAR decoder. This hybrid objective retains the semantic abstraction of LCM while eliminating its diffusion sampler and restoring a likelihood-based training signal. Across model sizes from 39M to 1.3B parameters, SONAR-LLM attains competitive generation quality. We report scaling trends, ablations, benchmark results, and release the complete training code and all pretrained checkpoints to foster reproducibility and future research.

📄 PDF Abstract BibTeX arXiv:2508.05305

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SONAR: Sentence-Level Multimodal and Language-Agnostic Representations

2023-08-22 · Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot

We introduce SONAR, a new multilingual and multimodal fixed-size sentence embedding space. Our single text encoder, covering 200 languages, substantially outperforms existing sentence embeddings such as LASER3 and LabSE …

DecoderMachine TranslationSentenceSentence Embedding+5

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

2026-03-17 · Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot, Ioannis Tsiamas 외 arxiv

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSONAR, a new family of omnilingual, cross-…

FLiP: Towards understanding and interpreting multimodal multilingual sentence embeddings

2026-04-20 · Santosh Kesiraju, Bolaji Yusuf, Šimon Sedláček, Oldřich Plchot 외 arxiv

This paper presents factorized linear projection (FLiP) models for understanding pretrained sentence embedding spaces. We train FLiP models to recover the lexical content from multilingual (LaBSE), multimodal (SONAR) and…

Unified Vision-Language Modeling via Concept Space Alignment

2026-03-01 · Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk arxiv

We introduce V-SONAR, a vision-language embedding space extended from the text-only embedding space SONAR (Omnilingual Embeddings Team et al., 2026), which supports 1500 text languages and 177 speech languages. To constr…

Question AnsweringVideo CaptioningVideo Retrieval

FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation

2025-11-19 · Tingrui Shen, Yiheng Zhang, Chen Tang, Chuan Ping 외 arxiv

Autoregressive models can generate high-quality 3D meshes by sequentially producing vertices and faces, but their token-by-token decoding results in slow inference, limiting practical use in interactive and large-scale a…