paper-with-me

Papers

Transmuting prompts into weights

2025-10-09 · Hanna Mazzawi, Benoit Dherin, Michael Munn, Adrian Goldwaser, Michael Wunder, Javier Gonzalvo arxiv

A growing body of research has demonstrated that the behavior of large language models can be effectively controlled at inference time by directly modifying their internal states, either through vector additions to their activations or through updates to their weight matrices. These techniques, while powerful, are often guided by empirical heuristics, such as deriving ``steering vectors'' from the average activations of contrastive prompts. Building on the foundational work of Dherin et al. (2025), who discovered that a prompt's influence mathematically maps to token-dependent implicit weight updates and introduced the initial concept of a static thought patch for prompt compression, we elevate this framework into a robust algorithm for direct model editing. We derive a principled method for condensing this transient information into token-independent thought vectors and thought matrices. These constructs provide a theoretical explanation for existing vector-and-matrix-based model editing techniques and offer a direct, computationally-grounded method for transmuting textual input into reusable weight updates for complex architectures and new knowledge injection.

📄 PDF Abstract BibTeX arXiv:2510.08734

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CoRA: Collaborative Information Perception by Large Language Model's Weights for Recommendation

2024-08-20 · YuTing Liu, Jinghao Zhang, Yizhou Dang, Yuliang Liang 외

Involving collaborative information in Large Language Models (LLMs) is a promising technique for adapting LLMs for recommendation. Existing methods achieve this by concatenating collaborative features with text tokens in…

Collaborative FilteringGeneral KnowledgeWorld Knowledge

MAiVAR-T: Multimodal Audio-image and Video Action Recognizer using Transformers

2023-08-01 · Muhammad Bilal Shaikh, Douglas Chai, Syed Mohammed Shamsul Islam, Naveed Akhtar

In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Au…

Action RecognitionTemporal Action Localization

Dynamic Prompt Optimizing for Text-to-Image Generation

2024-04-05 · CVPR 2024 1 · Wenyi Mo, Tianyu Zhang, Yalong Bai, Bing Su 외

Text-to-image generative models, specifically those based on diffusion models like Imagen and Stable Diffusion, have made substantial advancements. Recently, there has been a surge of interest in the delicate refinement …

Image GenerationText to Image GenerationText-to-Image Generation

Continuous Prompt Generation from Linear Combination of Discrete Prompt Embeddings

2023-12-16 · Pascal Passigan, Kidus Yohannes, Joshua Pereira

The wayward quality of continuous prompts stresses the importance of their interpretability as unexpected and unpredictable behaviors appear following training, especially in the context of large language models automati…

Natural Language Understanding

OmniDataComposer: A Unified Data Structure for Multimodal Data Fusion and Infinite Data Generation

2023-08-08 · Dongyang Yu, Shihao Wang, Yuan Fang, Wangpeng An

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object TrackingOptical Character Recognition+6