DeepSeek plans to train models with up to 8 trillion parameters
According to Huxiu, DeepSeek is currently training a new model with approximately 2 trillion parameters, while founder Liang Wenfeng stated at a recent investor meeting that the company plans to scale future models up to 8 trillion parameters. This expansion represents a fivefold increase over the 1.6 trillion total parameters of DeepSeek's V4-Pro model, released in April.
Key facts
- 01DeepSeek's V4-Pro model, released in April, features 1.6 trillion total parameters and 49 billion active parameters.
- 02The planned 8-trillion-parameter model is five times the size of the V4-Pro model.
- 03Competitors such as Kimi K3 (2.8 trillion parameters) and ByteDance (up to 10 trillion parameters) are also pursuing larger model scales.
AI-generated from the sources below. Always check the originals.
How each country tells it
So far our sources show coverage from one country. Perspectives appear once we find reports from a second country.
Sources
Summaries are AI-generated from the linked sources and may contain errors; always check the originals. We summarise and link; we never republish articles. Photos come from openly licensed libraries, official publicity material and brand logos, credited to their sources. If you own an image and want it credited differently or removed, email info@coda.news and we will act promptly.
- 8万亿,梁文锋也开始卷“超大” · 虎嗅