Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology11 min read

“Google and Reddit do not own the Internet," web scraper says after court win - Ars Technica

Google's and Reddit's use of DMCA to fight web scraper is bizarre, expert says. Discover insights about “google and reddit do not own the internet," web scraper

TechnologyInnovationBest PracticesGuideTutorial
“Google and Reddit do not own the Internet," web scraper says after court win - Ars Technica
Listen to Article
0:00
0:00
0:00

“Google and Reddit do not own the Internet," web scraper says after court win - Ars Technica

Overview

“Google and Reddit do not own the Internet,” web scraper says after court win

Google’s and Reddit’s use of DMCA to fight web scraper is bizarre, expert says.

Details

After a big court loss last week, Google has confirmed that it won’t give up its fight to block AI bots from scraping its search results. And Reddit is weirdly along for the ride.

Curiously invoking the Digital Millennium Copyright Act (DMCA), Google sued Serp Api last December. The search giant accused the web scraper of circumventing its anti-scraping technology and then selling content scraped from Google search results through an unauthorized “Google Search API” software service.

According to Google, the anti-scraping tech was in place to protect copyrighted content in search results. Allegedly, Serp Api’s circumvention threatened to disrupt Google’s relationships with rights holders, including some who license content to Google to appear in so-called “knowledge panels” that are displayed in some search results for well-known people or entities.

It was an odd use of the DMCA, since Google search results can’t be copyrighted. But Google was apparently emboldened to explore the legal theory after Reddit filed a very similar lawsuit in October, accusing Serp Api and Google-rival Perplexity of scraping Reddit content that appears in Google results.

In a blog, Google cited Reddit’s lawsuit when announcing its own challenge, which it said it filed as a “last resort” to block “malicious scraping” that violates rights holders’ choices over who can access their content.

Specifically, Google alleged that Serp Api’s circumvention violated its terms and made it impossible to profit from—or offset the cost of—“billions” of bot searches. And before it, Reddit claimed that Serp Api was evading two levels of security: Reddit’s own controls blocking scraping on its platform and Google controls blocking scraping of Reddit content in search results.

Meredith Rose, a senior policy counsel with expertise in the DMCA for a nonprofit public interest group called Public Knowledge, told Ars that Google and Reddit seem to be “sort of grasping at whatever tool is available” in the face of the sudden, continuous rise of AI scraping over the past three years. And while the way they’re using the DMCA is “bizarre”—and “not what the law had sort of contemplated as a use case”—she says it’s not “surprising.” Historically, the DMCA has been an effective tool to quickly stop disfavored content uses and force discussions around licensing, so turning to it may have been an obvious starting point, given Google’s goals.

But Google’s and Reddit’s unusual DMCA arguments don’t seem to be winning ones. Last week, a court took the somewhat rare step of granting Serp Api’s motion to dismiss very early on in Google’s lawsuit. In that case, the judge found that Google had no DMCA standing to sue Serp Api, since it didn’t own any of the content in the search results and has not shown that it’s acting on behalf of any rights holders.

“That does not happen terribly often,” Rose told Ars. “It really boiled down to Google didn’t allege enough about what it was protecting that was copyrighted.”

Likely the timing of that decision wasn’t great for Reddit, which faced a hearing on Serp Api’s motion to dismiss its lawsuit last Thursday. It’s unclear which way the court will rule in that case, but Rose told Ars that the Google ruling doesn’t bode well for Reddit since Reddit can’t claim that it is the content owner or exclusive licensee of content in search results.

“The judge in the Google case said, ‘Well, in order to have standing to bring a lawsuit under the DMCA, you can be the copyright owner or the exclusive licensee or the person who is deploying and manufacturing the technological protection measure at issue,’” Rose told Ars. “Reddit is none of those things.”

Serp Api is hoping that the fight will be over soon, telling Ars that the costly legal battle is worth sticking it out to defend the open web.

“The bottom line is that both Google and Reddit appear to be engaged in attempts to use the DMCA to wall off the open Internet by retroactively claiming control over content that they didn’t author and don’t own,” Serp Api told Ars.

Although Rose agreed with Serp Api that, in granting the motion to dismiss, the court gave Serp Api a big win, the fight is not over yet, as Google has a narrow path forward to keep its war against web scraping alive.

Google acknowledged that search results can’t be copyrighted but argued that “knowledge panels” sometimes include copyrighted content that Google licenses from rights holders. If Google can amend its complaint to argue that rights holders directly authorized Google to use its anti-scraping technology to prevent unauthorized access to content, then Google may be able to block a very limited amount of Serp Api’s scraping.

Google’s spokesperson, José Castañeda, told Ars that Google plans to amend the complaint and is “pleased to see that the Court rejected nearly all of Serp Api’s legal arguments” otherwise attempting to dispute Google’s standing.

“We look forward to filing an amended complaint, as the Court invited us to do, and we remain committed to protecting our services and partners from unauthorized access,” Castañeda said.

However, Rose told Ars that Google has somewhat “talked themselves into a little bit of a corner here, both in this litigation and historically.”

For Google, it could be “very dangerous” to argue that the knowledge panel is “chock full of copyrighted material,” Rose suggested. Since the search giant doesn’t license all the content in the knowledge box, Google could risk future lawsuits if the act of algorithmically creating the knowledge box without licenses suddenly becomes viewed as infringement, Rose said.

“They have to make an argument somehow that there are parts of that that are reproductions of copyrighted material, but the only parts that are reproductions of copyrighted material are the ones they’ve explicitly licensed,” Rose suggested. “Otherwise, they’re admitting that they have been reproducing stuff without licensing it, and that gets them into another fair use fight that they probably don’t want to have.”

Google was given 21 days to amend its complaint, at which point it will become clearer how it plans to thread the needle to keep its DMCA fight going.

Reddit did not respond to Ars’ requests to comment but claimed in a filing ahead of last week’s hearing that it was prepared to discuss how Google’s court loss impacted its case.

Perhaps notably, Serp Api said that Reddit was not among attendees in the courtroom. Serp Api did not comment much on Reddit’s arguments at the hearing but said that the judge appeared to be focused on the nuances of the legal questions. Most particularly, the judge seemed interested in whether Reddit’s agreement with Google actually authorized Google to protect its copyrighted content.

It would appear then that both courts have somewhat narrowed the fight to this key question, but Serp Api seems confident that neither Google nor Reddit can show evidence that scraping public search results harms rights holders.

Asked for comment on Google’s plan to amend its complaint, Serp Api told Ars that there may be little point in continuing to argue over “snippets of text that appear in Google’s Knowledge Panels” after Google launched its attack to supposedly defend “hundreds of thousands of publishers” that appear in search results.

“We hope Google drops this misguided attack on the open Internet, but if necessary, we are prepared to defend Serp Api, our customers, and our principles,” Serp Api said.

Serp Api told Ars that its business has continued to grow while it has fought the DMCA lawsuits but that its customers, including tech giants like Nvidia, Uber, and Adobe, have faced uncertainty as both cases have dragged on.

They “rely on Serp Api every day to provide structured access to search data, which is and has always been a lawful and legitimate business,” Serp Api said.

Beyond its own customers, people invested in the open web have also pondered what the impact of a Google win could be, Serp Api suggested.

“Of course, litigation is incredibly expensive and disruptive,” Serp Api said. “However, we feel strongly that platforms such as Google and Reddit do not own the Internet and that their attempts to weaponize the DMCA threaten all users of the open web.”

Rose told Ars that while Serp Api hasn’t necessarily “endeared itself to a lot of people,” she thinks that “they’re in the right from a policy perspective here.”

Ever since the summer of 2023—when a sort of “API apocalypse” began—publishers have been cutting off access to the open web, Rose said, by attempting to find ways to quickly stop high volumes of AI scraping like Serp Api does.

“This mass reaction to scraping” has caused a “re-enclosure of a lot of the web,” Rose said. And concerningly, that doesn’t just gum up AI training efforts, it also harms research, archiving, journalism, public health reporting, and other important work that depends on anonymous crawling and automated scraping at scale, Rose said.

Serp Api told Ars that it has gotten “overwhelming” feedback from its customers, researchers, and SEO providers in support of its efforts to defend against the Google and Reddit lawsuits. The web scraper is optimistic that the court will grant its motion to dismiss against Reddit and last week celebrated that Google’s lawsuit may only proceed on very limited claims.

In a win, Serp Api warned that Reddit planned to act as a “toll collector,” requiring payments for scraping when it doesn’t even own the content.

“Reddit seeks to consolidate its control over its users’ content so that it will be in a better position to tax that content. Reddit is not acting to protect its users; Reddit is clearing a path to exploit them,” Serp Api warned.

Similarly, Serp Api accused Google of asking the court to ignore that it’s the “largest scraper on the planet,” while agreeing to cut off other web scrapers from information that is “100 percent public.”

For Rose, there’s no winning for online users if, at the end of the fight, public access to data is cut off.

“It’s a very fraught time,” Rose said, acknowledging that it’s not just Reddit and Google but many publishers across the web that are currently tempted to use “whatever tool is available in the toolbox to tamper down” AI scraping.

Some publishers may be financially motivated, like Google and Reddit, while others may be protecting their infrastructure or taking a moral stance, Rose said. Whatever the motivation is, “it’s leading to this kind of broader ecosystem-wide consequence of re-enclosure,” she warned.

And although Serp Api expects to win, the problem with DMCA cases, she suggested, is that it can be difficult to predict the outcome.

“I’m very curious how this is going to play out with Reddit,” Rose said. “I’m always a little bit cynical because to some extent, whenever you get into the copyright realm, a lot of judges just decide it’s all vibes.”

Advance Publications, which owns Ars Technica parent Condé Nast, is the largest shareholder in Reddit.

  1.          First teaser for Apple TV's Neuromancer debuts at SDCC
    
  2.          Activist charged with felony after giving border agent "duress code" that wiped his phone
    
  3.          Artist sues AI meme generator for selling deeply personal comic as ad template
    
  4.          Space X eyes tower catch for next Starship after auspicious end to 13th flight
    
  5.          I wanted a clock that never needed setting. Things escalated.
    

Ars Technica has been separating the signal from the noise for over 25 years. With our unique combination of technical savvy and wide-ranging interest in the technological arts and sciences, Ars is the trusted source in a sea of information. After all, you don’t need to know everything, only what’s important.

Key Takeaways

  • “Google and Reddit do not own the Internet,” web scraper says after court win

  • Google’s and Reddit’s use of DMCA to fight web scraper is bizarre, expert says

  • After a big court loss last week, Google has confirmed that it won’t give up its fight to block AI bots from scraping its search results

  • Curiously invoking the Digital Millennium Copyright Act (DMCA), Google sued Serp Api last December

  • According to Google, the anti-scraping tech was in place to protect copyrighted content in search results

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.