Who gets to build the next generation of open-source AI?

Open-source AI has widened access to powerful language models, giving researchers and developers room to adapt them and build new tools. But many projects depend on models and resources from a small number of well-funded organizations, which could limit future work if those organizations restrict access.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story weighs broader access to AI against reliance on a few organizations, without a clear dominant risk.

Who gets to build the next generation of open-source AI?

Free access to AI models has helped researchers and developers experiment, adapt existing systems and create new tools. But much of that activity depends on a small number of organizations willing and able to release models or the resources used to train them. If those organizations change course, the open-source AI boom could lose some of its foundation.

A wave of projects built on shared work

Open-source AI has quickly produced alternatives to chatbots and language models from large technology companies. Projects such as HuggingChat, StableLM and StableVicuna have given developers models they can study and modify. Other releases mentioned in the same wave include Alpaca, Dolly and Cerebras-GPT.

The projects differ, but many draw on a common base. HuggingChat and several other models build on Meta’s LLaMA, while some projects use models or datasets from EleutherAI. This means that a new release can make it easier for others to experiment without requiring every team to train a large model from the beginning.

That access can broaden who gets to take part. Developers can explore how models work, researchers can examine their behavior, and new applications can emerge from work by people outside the largest AI labs. The open-source approach can also make flaws easier to find, although access by itself does not ensure that a model is reliable or safe.

The cost of starting from scratch

Training a large language model from scratch remains difficult for most groups. It takes substantial computing resources, and the process can become more complicated as models grow. Training across multiple GPUs requires coordinating the hardware, while failures can force teams to restart work and add to the expense.

Stability AI CEO Emad Mostaque described the challenge in direct terms: “We melted a bunch of GPUs building StableLM.” He also cautioned that StableLM does not come close to matching GPT-4. The source article describes language models as harder to train than the image model Stable Diffusion, which could run on a good home computer and helped spur open-source development around image-making AI.

Stella Biderman of EleutherAI put the current ceiling for many groups roughly in the range of 6 to 10 billion parameters. GPT-3 has 175 billion parameters, while LLaMA has 65 billion. These figures help explain why adapting an existing model is more practical for many teams than building a comparable system themselves.

EleutherAI’s own history also shows how much a new project can depend on shared information and outside support. OpenAI did not release GPT-3, but it shared enough information for Biderman and colleagues to work out how to replicate it. The group assembled the Pile, a large text dataset, and used it to train an open-source model. A cloud computing company sponsored the largest model EleutherAI trained; Biderman said paying for it out of pocket would have cost about $400,000.

Openness depends on decisions by companies

Meta’s LLaMA became a starting point for many projects, making Meta AI’s choices important to the wider ecosystem. Joelle Pineau, Meta AI’s managing director, said releasing code to outsiders was the right approach at the time, while questioning whether the same strategy would continue for the next five years. Meta may decide to limit releases if it sees greater risks from misuse or liability.

OpenAI, meanwhile, is already reversing its previous open policy because of competition fears, according to the article. If major organizations restrict access to their models, later projects may have fewer strong systems to adapt. The likely effect would reach beyond any one model: it could narrow the set of people and groups able to develop new AI tools.

Access also involves trade-offs. Large language models can produce misinformation, prejudice and hate speech, and can be used to make propaganda or power malware factories. Meta AI may keep models trained on Facebook user data in-house because of the risk that private information could leak. For other releases, it may use a click-through license that limits use to research.

That approach has limits. LLaMA was released under a license, but someone posted the full model and instructions for running it on 4chan within days. Pineau said she still considered the decision the right trade-off for that model, while expressing disappointment that the release had been shared in this way.

Finding a balance between access and accountability

Other groups are also weighing how broadly to release models. Hugging Face introduced a process that requires people to request and receive approval before downloading many models on its platform. Margaret Mitchell, the company’s chief ethics scientist, supports what she calls “responsible democratization”: releasing models in a controlled way that takes account of the risk of harm or misuse.

The debate is therefore not simply whether open source is valuable. Broader access can spread experimentation and give more people visibility into AI systems. At the same time, unrestricted access can make harmful uses easier, and the effort needed to train foundational models means that many open-source projects remain tied to organizations with deep resources.

The future of open-source AI will depend partly on whether those organizations keep sharing models, data and technical information, and on how they judge the risks of doing so. If access narrows, the community may still adapt models already available. But the next wave of work could have fewer starting points and less room for independent groups to push the technology forward.