Tech companies and AI startups increasingly adopt smaller, faster language models like DeepSeek V4.1 Flash and Gemini 3.5 Flash to handle high-concurrency AI agent tasks economically
According to 36Kr, a Shenzhen AI startup with about 10 employees shifted from GPT-5.6 to DeepSeek V4.1 Flash to run hundreds of agents concurrently, finding it faster and cheaper. The move coincides with Google I/O introducing Gemini 3.5 Flash, while token usage at Google reportedly rose from 500 billion in March to over 3 trillion by May.
Key facts
- 01The Shenzhen startup has about 10 people.
- 02They previously used GPT-5.6 and could only handle two to three tasks at once.
- 03They chose DeepSeek V4.1 Flash because it is faster and cheaper for their development tasks since September.
- 04Google I/O introduced Gemini 3.5 Flash and Google’s internal token usage rose from 500 billion in March to over 3 trillion by May.
AI-generated from the sources below. Always check the originals.
How each country tells it
So far our sources show coverage from one country. Perspectives appear once we find reports from a second country.
Sources
Summaries are AI-generated from the linked sources and may contain errors; always check the originals. We summarise and link; we never republish articles. Photos come from openly licensed libraries, official publicity material and brand logos, credited to their sources. If you own an image and want it credited differently or removed, email info@coda.news and we will act promptly.
- 大模型,开始集体变小 · 36氪