paper-with-me

홈 › Papers

Citation-Driven Multi-View Training for Patent Embeddings: QaECTER and Sophia-Bench

2026-04-24 · Younes Djemmal, You Zuo, Kim Gerdes, Kirian Guiller arxiv

Patent retrieval underpins critical decisions in innovation, examination, and IP strategy, yet progress has been hampered by the absence of benchmarks that reflect the diversity of real world search scenarios. We address this gap with two contributions. First, we introduce Sophiabench, a large-scale patent retrieval benchmark comprising 10,000 queries and 75,000 corpus documents stratified across ten years, eight IPC technology sections, and twelve filing jurisdictions. Unlike prior benchmarks, Sophia-bench tests retrieval using 12 different query types-from structured patent fields to AI-generated summaries-and evaluates results against citation-based ground truth enhanced with a novel domain-relevance metric (InScope). Together, these enable systematic measurement of how well models perform across query types, technology domains, and jurisdictions. Second, we introduce QaECTER, a 344M-parameter embedding model trained on patent citation graphs and multi-view self-alignment. Despite its compact size, QaECTER establishes a new state of the art for patent retrieval. It outperforms the \#1 model on the English retrieval text embedding benchmark (RTEB), a model 23x larger, as well as all existing patent specific models across every query type, IPC section, and jurisdiction on Sophia-bench, with gains of up to 7.2% average NDCG@10 over the next-best model. These results are confirmed on an independent external benchmark, where QaECTER surpasses all prior models without requiring task-specific instruction prompts. Both the benchmark and the model are designed for practical deployment in large-scale patent search systems.

📄 PDF Abstract BibTeX arXiv:2604.22897

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Patent Representation Learning via Self-supervision

2025-11-03 · You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie, Benoît Sagot arxiv

We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding the same text twice under independent dropout masks, but applying this recipe to l…

Representation Learning

Patent Citation Dynamics Modeling via Multi-Attention Recurrent Networks

2019-05-22 · Taoran Ji, Zhiqian Chen, Nathan Self, Kaiqun Fu 외

Modeling and forecasting forward citations to a patent is a central task for the discovery of emerging technologies and for measuring the pulse of inventive progress. Conventional methods for forecasting these forward ci…

Citation PredictionPoint Processes

PatSTEG: Modeling Formation Dynamics of Patent Citation Networks via The Semantic-Topological Evolutionary Graph

2024-02-03 · Ran Miao, Xueyu Chen, Liang Hu, Zhifei Zhang 외

Patent documents in the patent database (PatDB) are crucial for research, development, and innovation as they contain valuable technical information. However, PatDB presents a multifaceted challenge compared to publicly …

Graph Learning

Early identification of important patents through network centrality

2017-10-25 · Mariani Manuel Sebastian, Medo Matus, Lafond François

One of the most challenging problems in technological forecasting is to identify as early as possible those technologies that have the potential to lead to radical changes in our society. In this paper, we use the US pat…

From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

2025-09-11 · Grazia Sveva Ascione, Nicolò Tamagnone arxiv

Classifying patents by their relevance to the UN Sustainable Development Goals (SDGs) is crucial for tracking how innovation addresses global challenges. However, the absence of a large, labeled dataset limits the use of…

Transfer Learning