01.AI has released Yi-34B, a language model the Chinese startup says performs competitively with larger systems. The announcement puts the company, founded by computer scientist Kai-Fu Lee, into a global contest over model performance, access and the computing resources needed to build AI.
Benchmark claims and model access
In results published by the company, Yi-34B performs at or above the level of Meta's Llama2-70B and Falcon-180B. Those models have more than twice and five times as many parameters, respectively. 01.AI also says its smaller Yi-6B performs at the level of Llama2-34B.
These are company-reported benchmark results, which offer a snapshot of performance under the tests the company selected. They do not, by themselves, establish how the models will compare across every task or use case. Still, the results make model size a central part of the story: 01.AI is presenting a smaller system as competitive with substantially larger models.
The models are available to developers and researchers under the Apache 2.0 license on HuggingFace, ModelScope and Github. The company describes them as intended for users around the world, rather than only China.
There is a caveat for commercial users. Although the stated license permits free commercial use, 01.AI says users seeking that use should submit an application through its website. The article describes this requirement as puzzling in light of the license terms, leaving prospective users to reconcile the two before relying on the models commercially.
Training data and practical costs
According to 01.AI's website, Yi-34B was trained from scratch on a corpus of three trillion tokens, which the company calls high quality. Lee attributes the model's benchmark performance against larger competitors to that data quality.
The explanation points to a broader lesson about language model development: performance is shaped by the material used in training as well as by the number of parameters. The source also notes that other research has found data quality to have a critical effect on large language model training.
A smaller model can also be cheaper to run, according to the article. That could matter to developers weighing access against operating costs. The release therefore combines two claims: Yi models can perform competitively, and their smaller size may make them less expensive to use.
A young company with ambitious plans
01.AI reached a valuation of more than $1 billion in less than eight months. Alibaba Group Holding Ltd.'s cloud division participated in its latest funding round, according to the source.
Lee, who leads venture capital firm Sinovation Ventures, is also CEO of 01.AI. He began assembling its team in March 2023, and the company began operations in June. It employs more than 100 people, including experienced industry figures who have worked on Google Bard and TensorFlow.
Lee has a Ph.D. in computer science from Carnegie Mellon University and has held leadership positions at Apple, Microsoft and Google. His background and the company's rapid growth help explain the attention around Yi, but the model's long-term standing will depend on how it performs beyond the benchmarks announced by its developer.
Compute constraints and what comes next
Lee has said China is behind in the large language model race, while describing Yi-34B as competitive globally. He expects that scaling up training with high-quality data could produce dramatically better AI models as early as next year and has hinted at further releases.
He has also said 01.AI's next proprietary model should be able to compete with OpenAI's GPT-4. That remains a projection, rather than a result established by the Yi-34B release.
Access to computing hardware is another part of the company's plans. Lee said 01.AI had overdrawn its bank account and bought plenty of chips after the U.S. banned Nvidia from exporting to China, and that it would need those chips for the foreseeable future. The detail underlines how model ambitions depend not only on data and research, but also on the hardware available to train and operate the systems.
For developers, Yi-34B offers a new model to examine, with published performance claims and an access policy that deserves careful attention. For 01.AI, it is an early public step toward a broader lineup and a test of whether its data and compute strategy can sustain the claims made for its first release.