Who Sees What a Robot Vacuum Records at Home?

Development Roombas captured private household images that reached contracted data workers, who shared screenshots in online groups. The incident shows how footage gathered to train robot vacuum systems can pass through several companies and people before it is used.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 1 ►

The story centers on private household footage captured by robots and exposed through the AI training data chain.

Who Sees What a Robot Vacuum Records at Home?

Robot vacuums with cameras can do more than map a room or steer around a cable. Images collected to improve their navigation may also be reviewed by people who label them for artificial intelligence training. A leak of private footage from development Roombas exposed how that data can travel beyond the home.

How private images reached online groups

In the fall of 2020, gig workers in Venezuela posted household images to online forums where they discussed their work. The images came from development versions of iRobot’s Roomba J7 series and included scenes captured from low to the ground, among them footage of a woman on a toilet and a boy lying in a hallway.

iRobot confirmed that the images were recorded by special development robots. The company said those devices had hardware and software modifications that were never present in consumer products. They were given to paid collectors and employees who signed written agreements to send video and other data back for training. iRobot said the robots carried a bright green “video recording in progress” sticker and that collectors were told to remove sensitive items, including children, from the robot’s operating space.

The images were sent to Scale AI, a company that contracts workers around the world to label audio, photo, and video data for artificial intelligence. MIT Technology Review obtained 15 screenshots that had been posted in closed social media groups. iRobot said the sharing violated a written non-disclosure agreement and said it was ending its relationship with the service provider that leaked the images.

Why training data involves human reviewers

Machine learning systems learn to recognize patterns from examples. For a robot vacuum, those examples can include images of rooms, furniture, cables, and objects on the floor. People often need to identify and label the items in those images so the system can learn what it is seeing. That work is called data annotation.

Real homes are useful training environments because each one is different. A cable, sock, or piece of furniture can appear in many forms, and a robot needs to respond to that variety as it moves through a room. But using realistic images also means that a training set can contain personal details from ordinary household life.

The screenshots were a small part of a larger process. iRobot has said it shared over 2 million images with Scale AI and an unknown quantity with other data annotation platforms. The path described in the report ran from homes in North America, Europe, and Asia to iRobot, then to Scale AI, and onward to contracted workers, including the workers in Venezuela who posted images in private groups.

Camera equipped vacuums raise different privacy questions

Robot vacuums have developed from machines that relied on bump and cliff sensors to models that use mapping, computer vision, or lidar to navigate. Computer vision uses algorithms trained on images and video to identify information in a scene. Cameras can help a vacuum avoid an obstacle or distinguish a couch from something else on the floor.

Those same capabilities make the devices more revealing than a basic cleaner. A camera-equipped vacuum can move through rooms while people are using them, and its view may include people, belongings, and the layout of a home. Dennis Giese, a researcher who studies security vulnerabilities in internet-connected devices, described robot vacuums as having powerful sensors and the ability to drive around inside the home.

Companies described different ways of assembling training data. iRobot said most of its image data came from real homes used by employees or volunteers recruited through third-party vendors, while some came from staged recordings. It also offers customers an option in its app to send selected obstacle images. Roborock said it creates images in its labs or works with vendors in China; Dyson described home trials and synthetic training data as sources.

Consent and the data supply chain

The development robots in the incident were not consumer devices, but the questions reach beyond that distinction. Consumers may agree to data collection through product use and privacy policies, while having little sense of who will handle the information or how many organizations it may pass through. A company can describe safeguards and still rely on outside providers whose workers access the material.

iRobot said its consent form made clear that service providers could process the footage, but it did not allow MIT Technology Review to review the agreements or make collectors available to discuss what they understood. The incident highlights a practical gap: people may expect a vacuum to record information for navigation, yet not picture a human reviewer looking at raw footage.

For manufacturers, training robots on diverse data can improve how they operate in unfamiliar homes. For households, the route from a camera to an annotation worker is part of the privacy picture. Understanding that route matters as robot vacuums become more capable and their recorded images become useful to more people in the systems that train them.