Images captured inside homes by pre-production Roomba devices later appeared on Facebook and Discord. An investigation by MIT Technology Review traced the pictures to a training-data process involving employees, paid data collectors, and workers labeling images for a contractor. The leak drew attention to how personal information can move through an AI data supply chain.
From home recordings to online posts
The images came to the attention of investigative reporter Eileen Guo after someone flagged pictures circulating online. Their low camera angles showed rooms from near floor level, and some included people and animals who did not appear to know they were being recorded.
Guo spent months determining where the images came from. The investigation identified iRobot, the maker of the Roomba, as the source. One picture showed a woman sitting on a toilet, bringing the sensitivity of the material into sharp focus.
When contacted by MIT Technology Review, iRobot confirmed that the 15 images under discussion came from its devices. The company said they were pre-production machines, not products released to consumers.
The people in the images were not customers. They were iRobot employees and people the company called “paid data collectors,” who had agreed to take part in testing. But the investigation found that agreement did not necessarily mean they understood how the recordings could be handled, or that they might end up online.
Why people were reviewing the images
To train an AI system, companies often need examples that reflect real environments. A person may review each image, video, or other item and add labels and context so a computer can learn what it contains. This work is known as data annotation.
For a robot vacuum, labeled images can help the system recognize its surroundings. The investigation noted that the leaked camera views pointed diagonally upward, capturing walls and ceilings as well as the floor. That kind of information could help a device identify which room it is in, supporting broader plans to build more home robots.
At the time, the data labelers handling these images worked for Scale AI, a company hired by iRobot. Guo described the work as low-paid workers categorizing images to teach AI systems what they show. The investigation found that some labelers shared images on Facebook and Discord.
Guo said iRobot could have chosen other approaches, such as having people work in an office with more controlled processes or completing annotation in-house. After the company learned of the leak, it said it had begun an investigation, ended its contract with Scale AI, and would take steps to prevent a similar incident. It did not explain what those steps would be.
Consent can travel only so far
The investigation raised questions about what participants understood when they agreed to testing. They knew the vacuums would record video in their homes, Guo said, but might not have understood that people would label and view the recordings or that third parties outside the country could receive them. The possibility that images could reach social media was not understood by the people interviewed for the reporting.
There is also a gap between one person agreeing to a device’s use and everyone in a home being recorded. Albert Fox Cahn, executive director of the Surveillance Technology Oversight Project, argued that agreements may cover the product maker’s interests without giving meaningful protection to other people captured in a home. Visitors and other household members may never be asked for consent.
Cahn also questioned whether employees can freely agree to this kind of testing when their employer is involved. He said an employment relationship can make it difficult to refuse. That concern is distinct from the data leak itself, but it shows why a signed agreement may not settle the question of whether consent was meaningful.
A wider challenge for AI data
The Roomba case illustrates a less visible part of AI development: personal data may pass from a device to a company, a contractor, and individual annotators before it is used to train a system. Each step creates another point where access and handling matter.
In the United States, the episode said companies are not required to disclose this kind of data use, and privacy policies may allow information to be used to improve products and services, including AI training. People may therefore agree to broad terms through ordinary product use without knowing how their data will be reviewed or shared.
Synthetic data, which does not come directly from real people, is one possible alternative. The episode noted that companies like Dyson are starting to use it, while experts cited in the reporting said it could still be 10 years to multiple decades before it becomes the future of this work. For now, the investigation leaves a practical question for connected-device makers: how can they make training data useful while limiting who can see personal recordings and what those people can do with them?