Comparison of SMT and NMT trained with large Patent Corpora: Japio at WAT2017
Japio participates in patent subtasks (JPC-EJ/JE/CJ/KJ) with phrase-based statistical machine translation (SMT) and neural machine translation (NMT) systems which are trained with its own patent corpora in addition to the subtask corpora provided by organizers of WAT2017. In EJ and CJ subtasks, SMT and NMT systems whose sizes of training corpora are about 50 million and 10 million sentence pairs respectively achieved comparable scores for automatic evaluations, but NMT systems were superior to SMT systems for both official and in-house human evaluations.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMachine TranslationNMTSentenceTranslationSimilar Papers 제목 키워드 기반
Translation Using JAPIO Patent Corpora: JAPIO at WAT2016
We participate in scientific paper subtask (ASPEC-EJ/CJ) and patent subtask (JPC-EJ/CJ/KJ) with phrase-based SMT systems which are trained with its own patent corpora. Using larger corpora than those prepared by the work…
Information RetrievalMachine TranslationTranslationImproving Chemical Named Entity Recognition in Patents with Contextualized Word Embeddings
Chemical patents are an important resource for chemical information. However, few chemical Named Entity Recognition (NER) systems have been evaluated on patent documents, due in part to their structural and linguistic co…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Engineering Knowledge Graph from Patent Database
We propose a large, scalable engineering knowledge graph, comprising sets of (entity, relationship, entity) triples that are real-world engineering facts found in the patent database. We apply a set of rules based on the…
Knowledge GraphsPatent Figure Classification using Large Vision-language Models
Patent figure classification facilitates faceted search in patent retrieval systems, enabling efficient prior art search. Existing approaches have explored patent figure classification for only a single aspect and for as…
ClassificationFew-Shot LearningMultiple-choiceQuestion Answering+2Patent Language Model Pretraining with ModernBERT
Transformer-based language models such as BERT have become foundational in NLP, yet their performance degrades in specialized domains like patents, which contain long, technical, and legally structured text. Prior approa…