paper-with-me

Papers

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

2024-03-26 · Yuelin Bai, Xinrun Du, Yiming Liang, Yonggang Jin, Junting Zhou, Ziqiang Liu, Feiteng Fang, Mingshan Chang, Tianyu Zheng, Xincheng Zhang, Nuo Ma, Zekun Wang, Ruibin Yuan, Haihong Wu, Hongquan Lin, Wenhao Huang, Jiajun Zhang, Chenghua Lin, Jie Fu, Min Yang, Shiwen Ni, Ge Zhang

Remarkable progress on English instruction tuning has facilitated the efficacy and reliability of large language models (LLMs). However, there remains a noticeable gap in instruction tuning for Chinese, where the complex linguistic features pose significant challenges. Existing datasets, generally distilled from English-centric LLMs, are not well-aligned with Chinese users' interaction patterns. To bridge this gap, we introduce COIG-CQIA, a new Chinese instruction tuning dataset derived from various real-world resources and undergoing rigorous human verification. We conduct extensive experiments on COIG-CQIA, and compare them with strong baseline models and datasets. The experimental results show that models trained on COIG-CQIA achieve highly competitive performance in diverse benchmarks. Additionally, our findings offer several insights for designing effective Chinese instruction-tuning datasets and data-mixing strategies. Our dataset are available at https://huggingface.co/datasets/m-a-p/COIG-CQIA.

📄 PDF Abstract BibTeX arXiv:2403.18058

Code (0)

등록된 구현이 없습니다.

Tasks

All

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Chinese Open Instruction Generalist: A Preliminary Release

2023-04-17 · Ge Zhang, Yemin Shi, Ruibo Liu, Ruibin Yuan 외

Instruction tuning is widely recognized as a key technique for building generalist language models, which has attracted the attention of researchers and the public with the release of InstructGPT~\citep{ouyang2022trainin…

COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes

2025-10-16 · Yunwen Li, Shuangshuang Ying, Xingwei Qu, Xin Li 외 arxiv

Large language models exhibit systematic deficiencies in creative writing, particularly in non-English contexts where training data is scarce and lacks process-level supervision. We present COIG-Writer, a novel Chinese c…

Cross-Lingual TransferMathematical Reasoning

Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation

2024-01-12 · Tianyu Zheng, Shuyue Guo, Xingwei Qu, Jiawei Guo 외

In this paper, we introduce Kun, a novel approach for creating high-quality instruction-tuning datasets for large language models (LLMs) without relying on manual annotations. Adapting a self-training algorithm based on …

Instruction FollowingTranslation

Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation

2025-12-09 · Young Kyung Kim, Oded Schlesinger, Yuzhou Zhao, J. Matias Di Martino 외 arxiv

While state-of-the-art image generation models achieve remarkable visual quality, their internal generative processes remain a "black box." This opacity limits human observation and intervention, and poses a barrier to e…

Text-to-Image Generation

Panda LLM: Training Data and Evaluation for Open-Sourced Chinese Instruction-Following Large Language Models

2023-05-04 · Fangkai Jiao, Bosheng Ding, Tianze Luo, Zhanfeng Mo

This project focuses on enhancing open-source large language models through instruction-tuning and providing comprehensive evaluations of their performance. We explore how various training data factors, such as quantity,…

Instruction Following