Share computing power to make large language models easier to run

Petals lets people pool computing resources to run Bloom, a large language model, through a distributed network. It could lower the cost of access, but its public network can expose input text, and results may be slower than ChatGPT.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

Petals broadens access to language models while raising privacy and speed trade-offs, with neither a strong control nor dependency concern dominating.

Share computing power to make large language models easier to run

Petals is an open-source project that lets people contribute computing power to run large text-generating AI models. Instead of requiring one user to own a powerful computer, the system splits work across servers connected over the internet. The approach could make models such as Bloom easier to access, though it brings trade-offs in speed, reliability and privacy.

Pooling hardware to run a large model

Petals was developed as a collaborative project involving researchers at Hugging Face, Yandex Research and the University of Washington. Its code was released publicly, and volunteers can provide hardware to handle part of a model’s workload. Together, participating servers can carry out tasks that would be difficult for many users to run on a local machine.

To use the network, people install an open-source library and follow instructions on a website to connect. They can then generate text using Bloom through Petals, or set up a server and contribute computing resources to the network. The project also gives researchers access to a flexible system they can adapt and study.

That flexibility matters because large language models are resource-intensive. The source article notes that running Bloom locally requires a GPU that can cost hundreds to thousands of dollars. A shared network offers another route: access depends on the collective capacity of connected machines, rather than each user buying the necessary hardware.

Lower cost comes with slower responses

Petals aims to reduce the cost of running text-generating AI, which has largely been available through well-funded companies and research labs. Its creators see distributed computing as a way to make these tools more broadly accessible and support collaborative work on machine learning models.

In the source article’s tests, response times varied with the prompt. A simple translation took a couple of seconds, while more complex requests took well over 20 seconds. One longer answer took close to three minutes. Those results were slower than ChatGPT, though using Petals was free at the time of the report.

The network’s performance depends on how much computing power people contribute and whether their servers stay available. If a server disconnects, Petals tries to find a replacement automatically. Servers also disconnect after around 1.5 seconds without activity to conserve resources; the system can resume sessions, though users may experience a slight delay.

At the time of the report, the project lead said that multiple users with GPUs of different capacity had joined since the network’s launch in early December. The project planned a rewards system in which contributors could receive “Bloom points” to use for higher priority or increased security guarantees, or potentially exchange for other rewards.

Privacy and model behavior remain concerns

Petals’ distributed design creates a privacy risk. Its project page warns that servers may be able to recover input text from information passed through the network. That could include private details such as names and phone numbers. A malicious server could also record or modify data, potentially changing generated code so that it does not work as intended.

Researchers involved in the project advised against sending sensitive information through the public network. One suggested option for groups that need to handle such data is to create a private swarm hosted by organizations and people they trust. That arrangement could let small startups or labs share computing resources while limiting access to their data.

Using Petals also does not remove problems found in text-generating systems themselves. The source article points to the tendency of leading models to produce toxic or biased text. Petals is intended for research and academic use, at least for the time being, and the project’s researchers expect it to face issues as it develops.

A research tool with a wider ambition

Petals offers a different way to access a large language model: distribute the computing work among volunteers instead of relying on a single machine or provider. If enough servers participate, the network could support chatbots and other interactive applications, including fine-tuning models, according to its lead developer.

That promise depends on solving practical challenges as well as attracting contributors. Users must weigh the lower cost against slower responses and the risk of exposing inputs on a public network. For research and experimentation, Petals points toward a more shared model of access; for private data, its own researchers recommend a trusted, private network.