CN-Probase: A Data-driven Approach for Large-scale Chinese Taxonomy Construction
Taxonomies play an important role in machine intelligence. However, most well-known taxonomies are in English, and non-English taxonomies, especially Chinese ones, are still very rare. In this paper, we focus on automatic Chinese taxonomy construction and propose an effective generation and verification framework to build a large-scale and high-quality Chinese taxonomy. In the generation module, we extract isA relations from multiple sources of Chinese encyclopedia, which ensures the coverage. To further improve the precision of taxonomy, we apply three heuristic approaches in verification module. As a result, we construct the largest Chinese taxonomy with high precision about 95% called CN-Probase. Our taxonomy has been deployed on Aliyun, with over 82 million API calls in six months.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Utilizing Probase in Open Directory Project-based Text Classification
Open Directory Project (ODP) has been successfully utilized in text classification due to its representation ability of various categories. However, ODP includes a limited number of entities, which play an important role…
ClassificationGeneral Classificationtext-classificationText ClassificationLarge-scale Multi-granular Concept Extraction Based on Machine Reading Comprehension
The concepts in knowledge graphs (KGs) enable machines to understand natural language, and thus play an indispensable role in many applications. However, existing KGs have the poor coverage of concepts, especially fine-g…
DescriptiveKnowledge GraphsMachine Reading ComprehensionReading ComprehensionBeyond Word Embeddings: Learning Entity and Concept Representations from Large Scale Knowledge Bases
Text representations using neural word embeddings have proven effective in many NLP applications. Recent researches adapt the traditional word embedding models to learn vectors of multiword expressions (concepts/entities…
Semantic ParsingWord EmbeddingsBuilding Large Chinese Corpus for Spoken Dialogue Research in Specific Domains
Corpus is a valuable resource for information retrieval and data-driven natural language processing systems,especially for spoken dialogue research in specific domains. However,there is little non-English corpora, partic…
Information RetrievalRetrievalSentenceBBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark
To advance Chinese financial natural language processing (NLP), we introduce BBT-FinT5, a new Chinese financial pre-training language model based on the T5 model. To support this effort, we have built BBT-FinCorpus, a la…
Language ModelingLanguage Modelling