Court loss narrows Google’s DMCA fight over AI web scraping

Google plans to amend its complaint after a court dismissed its DMCA case against SerpApi at an early stage. The fight now turns on whether Google can show it was authorized to protect specific copyrighted material in search features such as “knowledge panels.”

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story mildly leans Terminator because it concerns automated scraping bypassing platform controls, but it is primarily a routine legal dispute.

Court loss narrows Google’s DMCA fight over AI web scraping

Google is not dropping its legal push against AI web scraping after a major setback in court. The company says it will amend its complaint against SerpApi, a web scraper accused of bypassing anti-scraping technology and selling access to scraped Google search results through an unauthorized “Google Search API” service.

The dispute has become a test of how far the Digital Millennium Copyright Act can stretch when platforms try to stop bots from collecting material visible in public search results. Reddit has pursued a similar theory, and its own case is now under pressure after Google’s loss.

Why Google turned to the DMCA

Google sued SerpApi last December, arguing that SerpApi circumvented anti-scraping technology. Google said that technology existed to protect copyrighted content appearing in search results, including content licensed for “knowledge panels” that appear for some well-known people or entities.

That made the case unusual from the start. The source article notes that Google search results can’t be copyrighted. Google’s position instead depends on the idea that some content within certain search features may be protected because it was licensed from rights holders.

Google also alleged that SerpApi’s activity violated its terms and made it impossible to profit from, or offset the cost of, “billions” of bot searches. In a blog, Google said the lawsuit was a “last resort” against “malicious scraping” that it claimed violated rights holders’ choices over who can access their content.

Reddit filed a similar lawsuit in October, accusing SerpApi and Perplexity of scraping Reddit content that appears in Google results. Reddit claimed SerpApi bypassed two levels of protection: Reddit’s own anti-scraping controls and Google’s controls over scraping Reddit content from search results.

The court’s early dismissal changed the fight

Last week, a court granted SerpApi’s motion to dismiss at an early stage in Google’s lawsuit. The judge found that Google did not have DMCA standing because it did not own the content in the search results and had not shown that it was acting on behalf of rights holders.

Meredith Rose, a senior policy counsel with expertise in the DMCA at Public Knowledge, told Ars that the early dismissal was notable. In her view, the problem was that Google had not adequately identified the copyrighted material it was protecting.

Rose described Google’s and Reddit’s use of the DMCA as “bizarre” and “not what the law had sort of contemplated as a use case,” while also saying it was not “surprising.” She said the DMCA has historically been used to quickly halt unwanted content uses and push licensing discussions, which may explain why companies reached for it as AI scraping increased over the past three years.

SerpApi framed the ruling as a defense of the open web. It told Ars that Google and Reddit appeared to be trying to use the DMCA to control content they did not author and do not own.

What Google must show next

The case is not finished. Google has 21 days to amend its complaint, and spokesperson José Castañeda told Ars that Google plans to do so. He said Google was “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” otherwise challenging Google’s standing.

Castañeda also said: “We look forward to filing an amended complaint, as the Court invited us to do, and we remain committed to protecting our services and partners from unauthorized access.”

According to the source article, Google’s narrower path is tied to “knowledge panels.” If Google can allege that rights holders directly authorized it to use anti-scraping technology to prevent unauthorized access to their content, it may be able to block a limited amount of SerpApi’s scraping.

That argument could be difficult to make cleanly. Rose warned that Google may have “talked themselves into a little bit of a corner here, both in this litigation and historically.” If Google argues too broadly that knowledge panels contain copyrighted material, it could raise uncomfortable questions about content in those panels that Google has not licensed.

Rose suggested Google needs to draw a narrow line: only the portions that are reproductions of copyrighted material and explicitly licensed should support the amended complaint. Otherwise, she said, Google could invite another fair use fight it likely would not want.

Reddit faces a related risk

Reddit did not respond to Ars’ requests for comment, but it said in a filing before last week’s hearing that it was prepared to discuss how Google’s court loss affected its own case. SerpApi said Reddit was not among the attendees in the courtroom.

The judge in Reddit’s hearing appeared focused on whether Reddit’s agreement with Google actually authorized Google to protect Reddit’s copyrighted content. That question matters because Rose told Ars the Google ruling does not look favorable for Reddit.

As Rose explained, the judge in Google’s case identified who can bring a DMCA claim in this context: the copyright owner, the exclusive licensee, or the party deploying and manufacturing the technological protection measure at issue. Rose’s view was blunt: “Reddit is none of those things.”

That does not decide Reddit’s case by itself, but it shows why the same legal theory may be difficult to sustain. Both cases now appear to center on authorization, ownership, licensing, and whether public search scraping can be treated as harm to rights holders under the DMCA.

The broader web scraping question

The fight reflects a practical problem for large platforms: AI bots can place pressure on services, licensing relationships, and control over how public web content is accessed. But this case shows that a tool built for copyright protection may not easily fit every dispute over automated access.

For SerpApi, the issue is whether Google and Reddit can wall off public search results through legal claims over content they do not own. For Google, the issue is whether it can protect services and partners from unauthorized access when some search features include licensed material.

The next version of Google’s complaint will matter because it must be more specific. If Google can connect its anti-scraping technology to rights-holder authorization and identifiable copyrighted content, the lawsuit may continue in a narrower form. If it cannot, SerpApi’s early win may become the defining turn in the case.