Start reading free
The Alibaba group headquarters in Hangzhou

Thomas LOMBARD, designed by HASSELL (architects)[1] / CC BY-SA 3.0 · Wikimedia Commons

科技 · Tech

中国的 AI 模型,这个夏天追上来了

China's AI models caught up this summer

HSK 5 · 3 min read · By Gao Peng
Pinyin

8 yuè 3 āledàoqiánwéizhǐzuìderéngōngzhìnéngxíng Qwen3.8-Max,cānshùliàngwèiliǎngwànqiān亿péngshèbàodàozàiduōxiàngshìzhōngzhègexíngdebiǎoxiànjīngjiējìnměiguódelèichǎnpǐn

On 3 August Alibaba released its largest AI model to date, Qwen3.8-Max, with 2.4 trillion parameters. Bloomberg reported that on a range of tests the model already performs close to comparable American products.
3 августа Alibaba выпустила свою крупнейшую на сегодня модель искусственного интеллекта Qwen3.8-Max с 2,4 триллиона параметров. Как сообщает Bloomberg, по ряду тестов она уже близка к сопоставимым американским продуктам.

zhèshìsānzhōunèisānzhèyàngdezàizhīqiányuèzhīànmiàntuīchūle Kimi K3,DeepSeek gēngxīnledexíngzhèsānkuǎnxíngdōukāifàngquánzhòngdefāngshì

It was the third such release in three weeks. Before it, Moonshot AI had put out Kimi K3, and DeepSeek had also updated its own model. All three were released with open weights.
Это третий подобный релиз за три недели. До него Moonshot AI представила Kimi K3, а DeepSeek тоже обновила свою модель. Все три выпущены с открытыми весами.

cānshùjiěwèixíngnèitiáozhěngdexiǎodānyuánmendezhízàixùnliànguòchéngzhōngmànmànquèdìngcānshùyuèduōxíngnéngzhudeguījiùyuèduōxùnliànshǐ使yòngshíyàodesuànyuè

A parameter can be understood as one of many small adjustable units inside the model, whose value is settled gradually during training. The more parameters, the more patterns the model can remember — and the more computing power needed to train and use it.
Параметр можно понимать как один из множества маленьких настраиваемых элементов внутри модели, значение которого постепенно определяется при обучении. Чем больше параметров, тем больше закономерностей модель может запомнить и тем больше вычислительных мощностей нужно для обучения и работы.

zhēnzhèngyǐnguānzhùdeshishìfēnshùérshìdiàoyòngxíngdechéngběnduōjiāméibàodàozhèxiēxíngměidiàoyòngdejiàmíngxiǎnshuǐpíngchàbuduōdeměiguóchǎnpǐnyǒudechājiējìnshíbèi

What really drew attention was not the test scores but the cost of calling the model. Several outlets reported that the price per call of these models is markedly lower than American products at a similar level, in some cases by close to a factor of ten.
Внимание привлекли не баллы в тестах, а стоимость обращения к модели. По сообщениям ряда изданий, цена одного запроса к этим моделям заметно ниже, чем у американских продуктов сопоставимого уровня, местами почти в десять раз.

chéngběnchābiédefenyuányīnchūkǒuguǎnzhìyǒuguānjìnniánláiměiguóxiànzhìxiàngzhōngguóchūkǒuzuìxiānjìnderéngōngzhìnéngxīnpiànzhōngguódeshíyànshìhěnnándàoyàngqiángdesuànshìzhuǎnxiàngzàijiàgòuxiàoshàngxúnzhǎokōngjiān

Part of the cost difference has to do with export controls. In recent years the United States has restricted exports of its most advanced AI chips to China. Chinese labs struggle to obtain equally powerful computing capacity, and have turned instead to finding room in architectural efficiency.
Часть разницы в стоимости связана с экспортным контролем. В последние годы США ограничивают поставки самых передовых ИИ-чипов в Китай. Китайским лабораториям трудно получить столь же мощные вычислительные ресурсы, и они ищут резерв в эффективности архитектуры.

bèiguǎng广fàncǎiyòngdebànshìzhǒngjiàohùnzhuānjiādejiàgòuxíngdecānshùzǒngliàngsuīránhěndànzàichǔqǐngqiúshízhǐyòngdàozhōngfencānshùcānsuànzhèyàngzǒngliàngměishíyòngdàodeliàngjiùfēnkāileshǐ使yòngdechéngběnsuízhīxiàjiàng

The widely adopted answer is an architecture called mixture-of-experts. The total parameter count is large, but for any single request only a portion is used, while the rest takes no part in the computation. Total size and per-request usage are thereby separated, and running costs fall.
Широко применяемое решение — архитектура под названием «смесь экспертов». Общее число параметров велико, но при обработке одного запроса используется лишь часть, а остальные в вычислении не участвуют. Общий объём и фактически задействованный на запрос объём тем самым разделяются, а стоимость работы снижается.

kāifàngquánzhòngzhǐdeshìgōngxùnliànwánchéngdecānshùchūláizhèyàngshǐ使yòngzhějiùxiàzǎixiàlaizàidediànnǎoshàngshǐ使yòng

Open weights means the company publishes the finished, trained parameters. Users can then download them themselves and run the model on their own computers.
Открытые веса означают, что компания публикует готовые, обученные параметры. Пользователь может сам скачать их и запустить модель на своём компьютере.

zhèkāiyuánruǎnjiànbìngyàngchūláideshìxùnliàndejiēguǒshixùnliàndeguòchéngyònglexiēshùyòngleshénmexùnliànfāngbāndōuhuìshuōmíngsuǒnèibānshuōkāifàngquánzhòngérshuōkāiyuán

This is not the same as open-source software. What is published is the result of training, not the process: which data and which training methods were used are generally not disclosed. So the industry usually says open weights, not open source.
Это не то же самое, что открытый исходный код. Публикуется результат обучения, а не процесс: какие данные и какие методы обучения использовались, как правило, не раскрывается. Поэтому в отрасли обычно говорят «открытые веса», а не «открытый код».

guòduìshǐ使yòngzhěláishuōzhègebiéréngránhěnzhòngyàoquánzhòngchūláihòuxíngfàngzàigòudeshàngshǐ使yòngshùyòngkāiběngòuduìzhèdiǎnzuìguānxīndeshìliáojīnróngjiàogòu

For users, though, the distinction still matters. Once the weights are published, the model can be run on an organisation's own servers, and data need not leave the institution. The ones who care about this most are healthcare, finance and education bodies.
Однако для пользователей это различие по-прежнему важно. После публикации весов модель можно запускать на собственных серверах организации, и данные не покидают учреждение. Больше всего это волнует медицинские, финансовые и образовательные организации.

tiáokuǎnshìchuánkuàimàndelìngyuányīnzhèkuǎnxíngyòngdeshìjiàokuānsōngdekāiyuánshānggòuzàimendechǔshàngkāichǎnpǐnyòngfèi

Licence terms are another reason for how fast they spread. These models use relatively permissive open-source licences, so commercial organisations can build products on top of them without paying.
Условия лицензии — ещё одна причина скорости распространения. Эти модели используют сравнительно свободные открытые лицензии, поэтому коммерческие организации могут строить на их основе продукты бесплатно.

zàizhōngguóxiàngzhònggōngdeshēngchéngshìréngōngzhìnéngchǎnpǐnànguīdìngbèiàn。2023 niánkāishǐshíxíngdeshēngchéngshìréngōngzhìnéngguǎnzànxíngbànduìchūleyāoqiúzhèshìguónèichǎnpǐnshíjiānjiàozhōngdeyuányīnzhī

In China, generative AI products offered to the public must be filed with the authorities as required. The Interim Measures for the Administration of Generative AI Services, in force since 2023, set out that requirement. This is one reason domestic releases tend to cluster in time.
В Китае продукты генеративного ИИ, предоставляемые публике, обязаны пройти регистрацию согласно правилам. Такое требование установлено «Временными мерами по управлению услугами генеративного ИИ», действующими с 2023 года. Отчасти поэтому релизы внутри страны идут волнами.

zàiyánjiāoxuéfāngmiànzhèxíngchǔzhōngwéndenéngqiánmíngxiǎngāoleduànwéngǎixiědàomǒu HSK shuǐpínghuòzhězhǐchūxuézhězuòwénrándeshuōxiànzàidōujīngnéngzuòdàole

In language teaching, this batch of models handles Chinese noticeably better than before. Rewriting a passage to a given HSK level, or pointing out unnatural phrasing in a learner's composition, is now within reach.
В преподавании языка это поколение моделей заметно лучше работает с китайским, чем прежде. Переписать отрывок под заданный уровень HSK или указать на неестественные обороты в сочинении учащегося — теперь вполне посильные задачи.

háiyàoshuōmíngdeshìshìchéngshíshǐ使yòngdeyànbìngyàngshìfēnshùgāodexíngzàixiědàizhèlèirènwushàngnénghěnqiángdànzàiliáotiānduìhuàshíquèxiǎndejiàoyìnggòurán

It should also be said that test scores and the actual experience of using a model are not the same thing. A model with high scores may be strong at tasks like writing code, yet come across as stiff and unnatural in conversation.
Стоит также оговорить, что баллы в тестах и реальный опыт использования — не одно и то же. Модель с высокими баллами может отлично справляться с задачами вроде написания кода, но в разговоре звучать скованно и неестественно.

qiánháiqīngchuzhèlúnjiàxiàjiàngnéngnéngzhíchíxiàjiāgōngzàikāifàngquánzhòngshàngdezuòhuìhuìsuízheshānghuàdetuījìnérgǎibiàntóngyàngháiméiyǒuàn

For now it is unclear whether this round of price declines can keep going. Whether the companies' approach to open weights will change as commercialisation advances is likewise still unanswered.
Пока неясно, сможет ли нынешнее снижение цен продолжаться. Изменится ли подход компаний к открытым весам по мере коммерциализации — тоже пока без ответа.

Words from this story

人工智能
réngōng zhìnéng
artificial intelligence; AI
模型
móxíng
model
参数
cānshù
parameter
芯片
xīnpiàn
chip
成本
chéngběn
cost
权重
quánzhòng
weights (of a model)
混合
hùnhé
to mix; mixed
服务器
fúwùqì
server
许可
xǔkě
licence; permission
备案
bèi'àn
to file with the authorities
Sources: Alibaba Adds to China AI Breakthroughs With New Qwen Model (Bloomberg) · China turns up the heat with open model blitz (The Register) · Chinese AI has leveled up, and brought renewed focus on the open weight model shift (CNBC) · Cyberspace Administration of China

Discussion · 0

It’s quiet in here… suspiciously quiet. Say something 👀

NihaoCard