Why Cara’s AI scraping crisis is forcing a new defense

Cara, an image-sharing and portfolio app used by about 1.5 million artists, was hit by three major scrapes beginning on August 13. The first scraper, now identified as Heft, later apologized, deleted his dataset, and began working with Jingna Zhang on an open-source tool to help protect artists.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

Unauthorized mass scraping for AI training increases harm and loss of control for artists, though it is not about autonomous AI danger.

Why Cara’s AI scraping crisis is forcing a new defense

Cara was built around a promise that matters deeply to many artists: a place to show work without embracing the unauthorized use of that work for AI training. That promise is now under pressure after a series of major scraping incidents exposed both the limits of technical defenses and the emotional cost of putting art online.

Since early 2023, photographer Jingna Zhang and a small group of volunteers have maintained Cara as both a social media platform and a portfolio app. The site has attracted about 1.5 million artists, many of them drawn by its stance against AI exploitation and by tools meant to make scraping less useful.

What happened to Cara

Beginning on August 13, Cara faced three major scrapes. The first became public when a Reddit user, MandarinDawnPoppy994, posted a 12-terabyte archive containing 12 million works from the platform. The archive represented more or less Cara’s entire library of publicly available images.

The post appeared on r/DefendingAIArt, where the user described the scrape as a technical stunt and said the process cost him less than $10. Zhang told WIRED the Cara team learned about it when users tagged them, after the scraper was, in her words, “gloating and looking for other people to join him to do something with the dataset on Reddit.”

The incident immediately became more than a technical problem. It sparked arguments across AI-related forums about whether scraping a site built for artists resisting AI training was defensible, even if scrapers believe the law has not caught up with the issue.

For Cara, the scrape also created practical damage. The activity spiked server fees and alarmed creators who had moved to the platform from larger services such as Instagram, where the source article says all content is explicitly available to Meta as training data.

More scrapes followed

The first scrape did not end the problem. According to Zhang, other people were apparently encouraged by MandarinDawnPoppy994 and carried out what she viewed as “copycat” attacks.

A second scraper pulled about 8.5 million links from Cara, along with metadata including usernames, titles, and tags. Those materials were uploaded to Hugging Face by a user named “CaptiveDreamer.”

After Hugging Face received many takedown requests, it said it would notify the user to remove personal metadata. But the company said it would not remove the URLs because “no copies of the artworks are hosted here” and the links “point to the copies the artists published on Cara.” It also stated that “further copyright reports on the same basis will not change this outcome.”

Then, on August 22, a third scraper obtained 123,000 images from Cara. That scrape also included text posts and user bios containing personal information, and the material was shared on Academic Torrents.

In response, Zhang launched a GoFundMe for legal fees with a goal of $120,000. She said the funds would support work on possible cyber and copyright law strategies for defending Cara. As of Thursday, the campaign had raised more than $100,000, and Cara was looking for additional legal assistance.

The limits of an artist-safe platform

Cara already offers protective features, including Glaze, which is intended to mask an image’s style from scrapers in a way that can disrupt AI mimicry. The platform also filters out AI images. But the source article makes clear that preventing scraping itself is extremely difficult.

Zhang said the team has tried to take reasonable steps “within limits without making it horrible to use.” Temporary measures such as login gates can add friction, but she does not see them as a complete answer to a broader internet problem.

That tension is important. Artists want visibility, portfolios, community, and control. Scraping attacks exploit the fact that public visibility and total protection do not easily coexist online.

Some Cara users have responded by deleting portfolios or leaving the site. Zhang said she supports people who do that if it makes them feel better, but she also warned that moving elsewhere does not guarantee safety. In her view, bigger platforms are scraped more often, which makes the situation feel worse rather than easier.

From scraper to collaborator

The most unexpected part of the story is what happened next. The person behind the first scrape later regretted what he had done, apologized, deleted the dataset, and began helping Zhang think through defenses.

WIRED identifies him as “Heft,” a student in North America with a software background and an interest in digital preservation and archival projects. He asked to be referred to by one of his screen names because he said he had received doxing and death threats over the Cara incident.

Heft told WIRED that scraping Cara began as a technical project and that he had not intended to make the data public. He later said he “made a foolish decision to attempt to ragebait with the dataset on Reddit” and “was carried away by trolling in the comments.”

He said he knew the post would provoke artists, but did not expect the depth of distress it caused. He saw people describing panic attacks and deleting portfolios. After speaking directly with artists, he said he better understood how personal their work was and how strongly they valued ownership.

In his own assessment, the choice to target Cara and present the scrape as he did was “cruel and thoughtless.” He added, “I missed the consequences that this would have beyond causing a bit of anger.”

What comes next

Heft has since joined Cara’s Discord server as a troubleshooter. Zhang said he has been explaining which proposed fixes are unlikely to work and identifying weak points that allow scraping to continue.

His view is that no site can be made truly “unscrapable.” That belief has shaped the new work he and Zhang are now pursuing: an open-source tool intended to help protect artists after scraping has occurred.

The details of that countermeasure are not fully described in the source article. But the shift is still significant. Cara’s crisis shows that the debate over AI scraping is not only about datasets, servers, or platform rules. It is also about whether people who understand how scraping works can help build practical defenses for the artists most affected by it.