paper-with-me

홈 › Papers

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models

2025-07-14 · Luolin Xiong, Haofen Wang, Xi Chen, Lu Sheng, Yun Xiong, Jingping Liu, Yanghua Xiao, Huajun Chen, Qing-Long Han, Yang Tang arxiv

DeepSeek, a Chinese Artificial Intelligence (AI) startup, has released their V3 and R1 series models, which attracted global attention due to their low cost, high performance, and open-source advantages. This paper begins by reviewing the evolution of large AI models focusing on paradigm shifts, the mainstream Large Language Model (LLM) paradigm, and the DeepSeek paradigm. Subsequently, the paper highlights novel algorithms introduced by DeepSeek, including Multi-head Latent Attention (MLA), Mixture-of-Experts (MoE), Multi-Token Prediction (MTP), and Group Relative Policy Optimization (GRPO). The paper then explores DeepSeek engineering breakthroughs in LLM scaling, training, inference, and system-level optimization architecture. Moreover, the impact of DeepSeek models on the competitive AI landscape is analyzed, comparing them to mainstream LLMs across various fields. Finally, the paper reflects on the insights gained from DeepSeek innovations and discusses future trends in the technical and engineering development of large AI models, particularly in data, training, and reasoning.

📄 PDF Abstract BibTeX arXiv:2507.09955

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems

2026-01-20 · Badri N. Patro, Vijay S. Agneeswaran arxiv

The field of artificial intelligence has undergone a revolution from foundational Transformer architectures to reasoning-capable systems approaching human-level performance. We present LLMOrbit, a comprehensive circular …

From ChatGPT to DeepSeek AI: A Comprehensive Analysis of Evolution, Deviation, and Future Implications in AI-Language Models

2025-04-04 · Simrandeep Singh, Shreya Bansal, Abdulmotaleb El Saddik, Mukesh Saini

The rapid advancement of artificial intelligence (AI) has reshaped the field of natural language processing (NLP), with models like OpenAI ChatGPT and DeepSeek AI. Although ChatGPT established a strong foundation for con…

Multiple-choice

Technically Love: The Evolution of Human-AI Romance Discourse on Reddit

2026-03-09 · Tyler Chang, Jina Huh-Yoo, Afsaneh Razi arxiv

Human-AI romantic relationships are increasingly common, yet little is understood about how public discourse around them emerges and shifts over time. Prior research has examined user experiences and ethical concerns, bu…

MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

2023-12-28 · Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang 외

In this work, we introduce a novel evaluation paradigm for Large Language Models (LLMs) that compels them to transition from a traditional question-answering role, akin to a student, to a solution-scoring role, akin to a…

GSM8KLanguage Model EvaluationLanguage ModelingLanguage Modelling+3

RoAD Benchmark: How LiDAR Models Fail under Coupled Domain Shifts and Label Evolution

2026-01-09 · Subeen Lee, Siyeong Lee, Namil Kim, Jaesik Choi arxiv

For 3D perception systems to operate reliably in real-world environments, they must remain robust to evolving sensor characteristics and changes in object taxonomies. However, existing adaptive learning paradigms struggl…

Self-Supervised LearningContinual LearningAutonomous Driving