paper-with-me

홈 › Papers

Yi-Lightning Technical Report

2024-12-02 · Alan Wake, Bei Chen, C. X. Lv, Chao Li, Chengen Huang, Chenglin Cai, Chujie Zheng, Daniel Cooper, Fan Zhou, Feng Hu, Ge Zhang, Guoyin Wang, Heng Ji, Howard Qiu, Jiangcheng Zhu, Jun Tian, Katherine Su, Lihuan Zhang, Liying Li, Ming Song, Mou Li, Peng Liu, Qicheng Hu, Shawn Wang, Shijun Zhou, Shiming Yang, Shiyong Li, Tianhang Zhu, Wen Xie, Wenhao Huang, Xiang He, Xiaobo Chen, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Yanpeng Li, Yongke Zhao, Yongzhen Luo, Yuchi Xu, Yuxuan Sha, Zhaodong Yan, Zhiyuan Liu, Zirui Zhang, Zonghong Dai

This technical report presents Yi-Lightning, our latest flagship large language model (LLM). It achieves exceptional performance, ranking 6th overall on Chatbot Arena, with particularly strong results (2nd to 4th place) in specialized categories including Chinese, Math, Coding, and Hard Prompts. Yi-Lightning leverages an enhanced Mixture-of-Experts (MoE) architecture, featuring advanced expert segmentation and routing mechanisms coupled with optimized KV-caching techniques. Our development process encompasses comprehensive pre-training, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF), where we devise deliberate strategies for multi-stage training, synthetic data construction, and reward modeling. Furthermore, we implement RAISE (Responsible AI Safety Engine), a four-component framework to address safety issues across pre-training, post-training, and serving phases. Empowered by our scalable super-computing infrastructure, all these innovations substantially reduce training, deployment and inference costs while maintaining high-performance standards. With further evaluations on public academic benchmarks, Yi-Lightning demonstrates competitive performance against top-tier LLMs, while we observe a notable disparity between traditional, static benchmark results and real-world, dynamic human preferences. This observation prompts a critical reassessment of conventional benchmarks' utility in guiding the development of more intelligent and powerful AI systems for practical applications. Yi-Lightning is now available through our developer platform at https://platform.lingyiwanwu.com.

📄 PDF Abstract BibTeX arXiv:2412.01253

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotLarge Language ModelMathMixture-of-Experts

Similar Papers 제목 키워드 기반

Lightning UQ Box: A Comprehensive Framework for Uncertainty Quantification in Deep Learning

2024-10-04 · Nils Lehmann, Jakob Gawlikowski, Adam J. Stewart, Vytautas Jancauskas 외

Uncertainty quantification (UQ) is an essential tool for applying deep neural networks (DNNs) to real world tasks, as it attaches a degree of confidence to DNN outputs. However, despite its benefits, UQ is often left out…

BenchmarkingUncertainty Quantification

A deep learning network for cloud-to-ground lightning nowcasting with multisource data

2020-05-01 · journal 2020 5 · Kanghui Zhou, Yongguang Zheng, Wansheng Dong, and Tingbo Wang

Precise and timely lightning nowcasting is still a great challenge for meteorologists. In this study, a new semantic segmentation deep learning network for cloud-to-ground (CG) lightning nowcasting, named LightningNet, h…

Semantic Segmentation

Lightning Mapping: Techniques, Challenges, and Opportunities

2020-05-31 · Ammar Alammari, Ammar Ahmed Alkahtani, Mohd Riduan Ahmad, Fuad M. Noman 외

Despite the significant progress made in studying the lightning phenomenon, precise location and mapping of its occurrence remain a challenge. Lightning mapping can be determined by studying the electromagnetic radiation…

Kalman Filter and Wavelet Cross-correlation for VHF Broadband Interferometer Lightning Mapping

2020-04-25 · Ammar Alammari, Ammar Alkahtani, Mohd Riduan, Fuad Noman 외

A lightning mapping system based on perpendicular crossed baseline interferometer (ITF) technology has been developed rapidly in recent years. Several processing methods have been proposed to estimate the temporal locati…

Denoising

Reproducibility Report: Contrastive Learning of Socially-aware Motion Representations

2022-08-18 · Roopsa Sen, Sidharth Sinha, Parv Maheshwari, Animesh Jha 외

The following paper is a reproducibility report for "Social NCE: Contrastive Learning of Socially-aware Motion Representations" {\cite{liu2020snce}} published in ICCV 2021 as part of the ML Reproducibility Challenge 2021…

Contrastive Learning