paper-with-me

Papers

Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models

2026-07-28 · Yongyi Cui, Yue Li, Tianbao Jiang, Xin Yi arxiv

Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerprints, i.e., trigger--target pairs embedded in model behavior, offer a practical, black-box-verifiable ownership signal, but existing methods decouple the two stages of the fingerprint life cycle: how a fingerprint is constructed and how it is injected. Existing fingerprinting frameworks suffer from two limitations. Natural-language fingerprints are prone to accidental activation, and garbled fingerprints are easily filtered by perplexity-based detection. Furthermore, decoupling construction from injection leaves the latter unaware of the trigger's linguistic structure, missing the opportunity for targeted optimization. We argue that fingerprint construction should drive injection, and present a unified fingerprinting framework that jointly optimizes both stages. First, LCF constructs code-mixing fingerprints by combining low-resource languages under a semantic-density substitution rule and grammar-biased mixing, yielding triggers whose perplexity sits far below garbled baselines while avoiding the accidental-activation failures of natural-language triggers. Second, LCFEdit injects each fingerprint with a null-space projection derived from high-resource multilingual representations that preserves knowledge, augmented by a cross-lingual alignment step that steers the weight update toward the fingerprint language's representation subspace. This construction-aware injection ensures that the update is linguistically informed and therefore more stable. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate persistent ownership verification with negligible impact on utility.

📄 PDF Abstract BibTeX arXiv:2607.25633

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing

2026-03-01 · Maomao Li, Yunfei Liu, Yu Li arxiv

Image-driven video editing aims to propagate edit contents from the modified first frame to the remaining frames. Existing methods usually invert the source video to noise using a pre-trained image-to-video (I2V) model a…

Unsupervised Protoform Reconstruction through Parsimonious Rule-guided Heuristics and Evolutionary Search

2025-06-12 · Promise Dodzi Kpoglu

We propose an unsupervised method for the reconstruction of protoforms i.e., ancestral word forms from which modern language forms are derived. While prior work has primarily relied on probabilistic models of phonologica…

From Construction to Injection: Edit-Based Fingerprints for Large Language Models

2025-09-03 · Yue Li, Xin Yi, Dongsheng Shi, Yongyi Cui 외 arxiv

Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial misuse. In black-box deployment, verification is hindered by defensive filtering of…

DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German

2014-05-01 · LREC 2014 5 · Beno{\^\i}t Sagot

We introduce DeLex, a freely-avaible, large-scale and linguistically grounded morphological lexicon for German developed within the Alexina framework. We extracted lexical information from the German wiktionary and devel…

Morphological AnalysisMorphological Inflection

MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction

2026-06-12 · Ruijie Xu, Xinnan Zhu, Jiayu Ying, Daoguo Dong 외 arxiv

Text-driven 3D scene generation is a promising technique for digital content creation, embodied AI simulation, and interactive design, yet practical workflows often require refining, extending, or correcting existing sce…

Scene Generation