【文章标题】:A dark horse enters China’s AI race: StartLux 中国AI竞赛杀出一匹黑马:StartLux

【文章正文】: Something big has happened - a dark horse has emerged among China’s top domestic large model developers. 重大事件发生——中国头部大模型开发商阵营中杀出一匹黑马。

A new company, with its first model having only 27B parameters, took second place overall in the CAICT MCP specialized test. 这家新公司的首款模型仅270亿参数,却在中国信通院MCP专项测试中斩获综合排名第二。

Ranked ahead of it is the killer weapon Liang Wenfeng has kept under wraps for nearly a year: DeepSeek-V4-Pro, boasting a parameter count of 1.6 trillion. 排在它前面的,是梁文锋雪藏近一年的杀手锏:参数高达1.6万亿的DeepSeek-V4-Pro。

The difference between the two is a mere 1.3 percentage points. 二者差距仅有1.3个百分点。

This dark horse is the StartLux-V1.0-27B-Preview, from StartLux (formerly Yuandian Xinghui). 这匹黑马正是来自StartLux(原远点星辉)的StartLux-V1.0-27B-Preview。

Looks a bit unfamiliar, doesn’t it? Don’t worry, its founder and CEO is an old acquaintance: Chen Danyan. 看起来有些陌生?别急,其创始人兼CEO是位老熟人:陈大年。

Known as the “godfather” of programmers in the internet era, his achievements are a matter of public record: 这位被誉为互联网时代程序员”教父”的人物,其成就有目共睹:

Shanda Network’s co-founder and the head of Lian Shang Network, one of China’s earliest programmers to introduce the concept of “shareware” 盛大网络联合创始人、联商网掌舵人,中国最早引入”共享软件”概念的程序员之一

After a decade of retirement, he has made a comeback, this time betting on local models. 退休十年后重出江湖,这次他押注的是本地模型。

This has to do with Chen Dawei’s recent rare public appearance. 这与陈大年近期罕见的公开亮相有关。

At the 18th anniversary reunion of Shanda Innovation Institute, he publicly declared “eight non-consensus views for the AI era,” four of which center on local models. 在盛大创新院18周年聚会上,他公开宣示”AI时代的八个非共识观点”,其中四点围绕本地模型展开:

Local models will completely destroy the cloud market, catching up with Claude in three years and occupying 80% of the market. The model competition based on parameters is going to be outdated… 本地模型将彻底摧毁云市场,三年追平Claude并占据80%市场。基于参数的模型竞争即将过时…

StartLux is the best representation of his idea. StartLux正是其理念的最佳代言。

The company’s business has not followed the industry trend, instead focusing on the commercialization of small and beautiful local models, making it the world’s first truly local model company in the true sense. 该公司业务没有跟随行业风潮,而是专注小而美本地模型的商业化落地,成为全球首个真正意义上的纯本地模型公司。

As the first market-oriented scorecard, StartLux-V1.0-27B-Preview does not rely on the cloud and can run directly on consumer-grade PCs. 作为首份市场化成绩单,StartLux-V1.0-27B-Preview不依赖云端,可直接在消费级PC上运行。

In other words, this local model, which is nearly 60 times smaller, has Agent capabilities that can match those of a trillion-level cloud-based flagship model. 也就是说,这个体积小近60倍的本地模型,其Agent能力竟能与万亿级云端旗舰机型匹敌。

What justification is there for this? 凭何有此底气?

Small Model Achieves Big Results, 27B Outperforms 1.6T 小模型干大事,270亿胜过1.6万亿

Before the answer is revealed, let’s take a look at who the comparison is being made to. 揭晓答案前,先看对标的是谁。

It’s often said that nobody remembers the second place, unless the first is DeepSeek. 常言道没人记得第二名,除非第一是深度求索。

Moreover, the gap is minimal, making it well worth discussing. 更何况差距微乎其微,极具讨论价值。

The results come from the authoritative institution, China Academy of Information and Communications Technology’s trustworthy AI large model benchmark test MCP special item, which sets six types of tasks around real application scenarios: 成绩源自权威机构中国信通院的可信AI大模型基准测试MCP专项,围绕真实应用场景设置六类任务:

Location navigation, web search, browser automation, financial analysis, code repository management, 3D design. 位置导航、网页搜索、浏览器自动化、金融分析、代码仓库管理、3D设计。

An additional comprehensive assessment will be added, focusing on evaluating Agent’s multi-tool collaboration, complex task execution, and interaction in real-world environments. 另设综合评估项,重点考察Agent的多工具协作、复杂任务执行和现实环境交互能力。

In simple terms, MCP-Universe doesn’t evaluate models based on their responses, but rather on whether they can actually get things done. This is also the most fundamental aspect of judging an Agent’s quality. 简言之,MCP-Universe不考核模型怎么说,而是看它能否真正办成事。这也是判断Agent优劣最根本的维度。

The results showed that StartLux-V1.0-27B-Preview had a comprehensive score of 39.25%, ranking second. 结果显示StartLux-V1.0-27B-Preview综合得分39.25%,位列第二。

DeepSeek-V4-Flash-0731, with over 284 billion parameters, and Step-3.7-Flash, at 198 billion, trail DeepSeek-V4-Pro by just 1.3 percentage points. 2840亿参数的DeepSeek-V4-Flash-0731、1980亿的阶跃3.7-Flash,与DeepSeek-V4-Pro差距仅1.3个百分点。

With the same 27B parameter scale, StartLux also surpasses Qwen 3.6 by 5.34 percentage points. 同属270亿参数规模,StartLux还以5.34个百分点优势超越通义千问3.6。

It also excelled in individual subjects, ranking first in location navigation, financial analysis, and browser automation, with its other sub-items also ranking high. 单科成绩同样亮眼,位置导航、金融分析、浏览器自动化三项第一,其余子项也名列前茅。

Let’s take a look at two case studies, putting data aside. 暂搁数据,看两个实操案例。

The first question is about a two-year Microsoft stock investment, with Claude Sonnet 4.6 as the competing topic. 首题是微软股票两年期投资测算,竞品为Claude Sonnet 4.6。

Claude’s answer is: 47,254,收益率89.02%。

Video link: https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfg 视频链接:https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfg

StartLux gave: 47,499.09,收益率90.00%。

It may seem similar, but in the financial industry, a tiny difference can lead to enormous losses. 看似相差无几,但在金融领域,毫厘之差可能造成巨额损失。

Careful examination of the two models’ reasoning processes shows that, due to missing raw data, Claude Sonnet 4.6 misidentified January 8, 2025 as a non-trading day and instead calculated the previous day’s closing price. 细究两者推理过程发现:因原始数据缺失,Claude Sonnet 4.6将2025年1月8日误判为非交易日,转而计算前一日收盘价。

Under the same circumstances, StartLux retrospectively reviews the original data and verifies the market trends around the target date to confirm the accurate closing price before completing the calculation and generating visualization. 相同情况下,StartLux会回溯原始数据,核验目标日期前后市场走势确认准确收盘价,再完成计算并生成可视化。

Ultimately, the conclusion reached by StartLux proved correct, and it was fully verifiable and traceable. 最终证明StartLux得出的结论正确无误,且全程可验证、可追溯。

The second task was more straightforward: both models were asked to search for flight tickets in a browser at the same time. 第二项任务更直接:让两个模型同时在浏览器中查询机票。

I’m unable to open a browser or interact with live websites like Google Flights. I can only process text and don’t have real-time browsing capabilities. To find this flight yourself, here’s what you’d do: 1. Go to google.com/flights 2. Enter Singapore (SIN) → Beijing 3. Select one-way, se (译文延续)“我无法打开浏览器或与Google Flights等实时网站交互。我只能处理文本,不具备实时浏览能力。如需查询该航班,请按以下步骤操作:1.访问google.com/flights 2.输入新加坡(SIN)→北京 3.选择单程…”

🔗 知识库双向关联