PatFig: Generating Short and Long Captions for Patent Figures
This paper introduces Qatent PatFig, a novel large-scale patent figure dataset comprising 30,000+ patent figures from over 11,000 European patent applications. For each figure, this dataset provides short and long captions, reference numerals, their corresponding terms, and the minimal claim set that describes the interactions between the components of the image. To assess the usability of the dataset, we finetune an LVLM model on Qatent PatFig to generate short and long descriptions, and we investigate the effects of incorporating various text-based cues at the prediction stage of the patent figure captioning process.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Patent Figure Classification using Large Vision-language Models
Patent figure classification facilitates faceted search in patent retrieval systems, enabling efficient prior art search. Existing approaches have explored patent figure classification for only a single aspect and for as…
ClassificationFew-Shot LearningMultiple-choiceQuestion Answering+2IMPACT: A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents
In this paper, we introduce IMPACT (Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents), a large-scale multimodal patent dataset with detailed captions for design patent figures. Our dataset in…
Cross-Modal RetrievalImage ClassificationImage RetrievalPatent classification+5DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
In the field of design patent analysis, traditional tasks such as patent classification and patent image retrieval heavily depend on the image data. However, patent images -- typically consisting of sketches with abstrac…
Contrastive LearningImage RetrievalCLID: Controlled-Length Image Descriptions with Limited Data
Controllable image captioning models generate human-like image descriptions, enabling some kind of control over the generated captions. This paper focuses on controlling the caption length, i.e. a short and concise descr…
controllable image captioningImage CaptioningPatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
Writing comprehensive and accurate descriptions of technical drawings in patent documents is crucial to effective knowledge sharing and enabling the replication and protection of intellectual property. However, automatio…
Large Language ModelMultimodal Large Language ModelPatent Figure Description Generation