paper-with-me

Papers

Exploring Representations from Unlabeled Data with Co-training for Chinese Word Segmentation

2013-10-01 · EMNLP 2013 10 · Longkai Zhang, Houfeng Wang, Xu sun, Mairgup Mansur
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Chinese Word SegmentationFeature Engineering

Similar Papers 제목 키워드 기반

Neural Chinese Word Segmentation with Lexicon and Unlabeled Data via Posterior Regularization

2019-04-26 · Junxin Liu, Fangzhao Wu, Chuhan Wu, Yongfeng Huang 외

Existing methods for CWS usually rely on a large number of labeled sentences to train word segmentation models, which are expensive and time-consuming to annotate. Luckily, the unlabeled data is usually easy to collect a…

Chinese Word SegmentationSegmentation

ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information

2021-06-30 · ACL 2021 5 · Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng 외

Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we …

Language ModelingLanguage ModellingMachine Reading ComprehensionNamed Entity Recognition+5

EBMs vs. CL: Exploring Self-Supervised Visual Pretraining for Visual Question Answering

2022-06-29 · Violetta Shevchenko, Ehsan Abbasnejad, Anthony Dick, Anton Van Den Hengel 외

The availability of clean and diverse labeled data is a major roadblock for training models on complex tasks such as visual question answering (VQA). The extensive work on large vision-and-language models has shown that …

Contrastive LearningOut of Distribution (OOD) DetectionQuestion AnsweringSelf-Supervised Learning+3

To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging

2020-10-27 · EMNLP 2020 11 · Kasturi Bhattacharjee, Miguel Ballesteros, Rishita Anubhai, Smaranda Muresan 외

Leveraging large amounts of unlabeled data using Transformer-like architectures, like BERT, has gained popularity in recent times owing to their effectiveness in learning general representations that can then be further …

C-Pack: Packed Resources For General Chinese Embeddings

2023-09-14 · Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff 외

We introduce C-Pack, a package of resources that significantly advance the field of general Chinese embeddings. C-Pack includes three critical resources. 1) C-MTEB is a comprehensive benchmark for Chinese text embeddings…

MTEB Benchmark