paper-with-me

홈 › Papers

M6: A Chinese Multimodal Pretrainer

2021-03-01 · Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, Jie Zhang, Jianwei Zhang, Xu Zou, Zhikang Li, Xiaodong Deng, Jie Liu, Jinbao Xue, Huiling Zhou, Jianxin Ma, Jin Yu, Yong Li, Wei Lin, Jingren Zhou, Jie Tang, Hongxia Yang

In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We propose a cross-modal pretraining method called M6, referring to Multi-Modality to Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. We scale the model size up to 10 billion and 100 billion parameters, and build the largest pretrained model in Chinese. We apply the model to a series of downstream applications, and demonstrate its outstanding performance in comparison with strong baselines. Furthermore, we specifically design a downstream task of text-guided image generation, and show that the finetuned M6 can create high-quality images with high resolution and abundant details.

📄 PDF Abstract BibTeX arXiv:2103.00823

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Picturized and Recited with Dialects: A Multimodal Chinese Representation Framework for Sentiment Analysis of Classical Chinese Poetry

2025-05-19 · Xiaocong Du, Haoyu Pei, Haipeng Zhang

Classical Chinese poetry is a vital and enduring part of Chinese literature, conveying profound emotional resonance. Existing studies analyze sentiment based on textual meanings, overlooking the unique rhythmic and visua…

Representation LearningSentenceSentiment Analysis

CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

2024-06-28 · Zhong-Zhi Li, Ming-Liang Zhang, Fei Yin, Zhi-Long Ji 외

Due to the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Despite the datasets like MathVista proposed benchmarks for as…

DiversityMath

Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings

2025-05-30 · Shujian Yang, Shiyao Cui, Chuanrui Hu, Haicheng Wang 외

Detecting toxic content using language models is important but challenging. While large language models (LLMs) have demonstrated strong performance in understanding Chinese, recent studies show that simple character subs…

In-Context Learning

CMNER: A Chinese Multimodal NER Dataset based on Social Media

2024-02-21 · Yuanze Ji, Bobo Li, Jun Zhou, Fei Li 외

Multimodal Named Entity Recognition (MNER) is a pivotal task designed to extract named entities from text with the support of pertinent images. Nonetheless, a notable paucity of data for Chinese MNER has considerably imp…

Miscellaneousnamed-entity-recognitionNamed Entity RecognitionNER

A Large-Scale Chinese Multimodal NER Dataset with Speech Clues

2021-08-01 · ACL 2021 5 · Dianbo Sui, Zhengkun Tian, Yubo Chen, Kang Liu 외

In this paper, we aim to explore an uncharted territory, which is Chinese multimodal named entity recognition (NER) with both textual and acoustic contents. To achieve this, we construct a large-scale human-annotated Chi…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1