Feature

Reddit’s scraping ruling is a reason to audit access and data suppliers—not expect an AI payout

What Reddit’s July 2026 ruling actually decided, why vendor use matters, and what small publishers should document before changing crawler policies.

Impetuous · · 4 Min Read

Blocking an AI crawler on your website does not address every route by which your content might reach an answer engine. Reddit’s lawsuit targets one such route: obtaining Reddit excerpts through Google search results using a third-party scraping tool.

The July 31, 2026 ruling allowed key claims against Perplexity and SerpApi to proceed; it did not establish liability or award publishers compensation. For a small publisher, the useful response is to document content rights, access controls and supplier behavior—not assume that every unwanted scrape now creates a winning lawsuit. The court’s order identifies exactly which claims survived.

What the court actually decided

U.S. District Judge Paul A. Engelmayer, in the Southern District of New York, considered motions to dismiss from SerpApi and Perplexity. At this stage, the court accepted the complaint’s factual allegations as true and drew reasonable inferences in Reddit’s favor. Those allegations still require proof.

Reddit alleged that SerpApi’s tools bypassed Google’s SearchGuard JavaScript challenges and CAPTCHAs, extracted Reddit snippets from search results, and enabled Perplexity to use that material in its answer system. The alleged route was through Google—not simply a crawler fetching Reddit’s origin server. The opinion’s factual background describes that mechanism.

The court preserved:

  • DMCA Section 1201(a)(1)(A) access-control circumvention claims against both companies.
  • A Section 1201(a)(2) circumvention-tool trafficking claim against SerpApi.
  • New York civil-conspiracy claims against both.

It dismissed the Section 1201(b) claim against SerpApi and the unjust-enrichment and unfair-competition claims against both defendants. The distinction matters: SearchGuard was plausibly an access control, but Reddit had not adequately alleged that it controlled copying after access was obtained. The opinion’s analysis and disposition therefore describe a partial, procedural win—not a blanket ruling against scraping. Perplexity denies the allegations, and SerpApi disputes Reddit’s characterization of accessing public search results. Reuters reported both responses.

“We used a vendor” was not the decisive issue

The ruling did not decide that merely purchasing already-scraped data creates secondary DMCA liability.

Instead, the court found it plausible that Perplexity directly participated in circumvention: Reddit alleged that Perplexity set operational parameters and conducted queries using SerpApi’s tool. Footnote 25 expressly leaves the secondary-liability question unresolved because direct circumvention was sufficiently alleged. See the opinion’s circumvention analysis.

That is the relevant distinction for publishers buying research feeds, search-results APIs or retrieval services. A vendor relationship does not describe what your system actually does. Receiving a licensed archive and configuring a service to evade access challenges are different workflows. Due diligence should examine the workflow, not just the invoice.

Public visibility and deletion promises matter differently

The court rejected the argument that SearchGuard could not control access because humans could view the same material. As alleged, it distinguished authorized human access from automated bulk access. Public readability therefore did not defeat Reddit’s access-control theory at the pleading stage.

Authorization mattered too. The court found it plausible that Reddit’s user agreement and its alleged licensing terms authorized protective measures, including Google’s SearchGuard. Installing a bot challenge alone does not establish the same facts for another publisher. The opinion examines access controls and copyright-owner authorization separately.

The court also accepted a specific reputational-injury theory: unauthorized copies could undermine Reddit’s efforts to honor user deletions and protect candid participation. This supported standing; it was not a separate, general privacy ruling requiring all AI services to delete any publisher’s content. The opinion’s standing analysis explains that harm.

Nor should operators read Google’s separate July 20 dismissal as proof that search-result scraping is always lawful. Engelmayer distinguished that complaint because it lacked sufficient allegations that copyright owners authorized SearchGuard. Reddit alleged more about its licensing restrictions and authorization of protective measures. Loeb & Loeb’s legal analysis explains the difference.

A proportionate workflow for a small publisher

These are proposed operating steps, not a recipe for establishing a legal claim.

1. Inventory rights before threatening enforcement. Separate staff-authored articles, commissioned work, licensed assets and reader submissions. Retain the agreements governing each, including authority to license content and authorize protective measures. Do not assume that hosting content means owning its copyright or that a contributor license authorizes every downstream use.

2. Separate crawler policy from enforcement. Record which access you permit, then test whether your controls enforce that policy without blocking intended reader and search access. The robots exclusion standard explicitly says its rules are not access authorization or a substitute for security measures. A robots.txt disallow is not equivalent to a CAPTCHA or authenticated endpoint. RFC 9309 makes that distinction; the practical companion is whether robots.txt stops AI browser agents.

3. Preserve evidence of the route, not just the output. Keep dated policy versions, rule changes, relevant request logs and examples of reproduced passages. An answer quoting your article does not, by itself, show which service fetched it or which control was bypassed. Preserve logs securely and avoid collecting unnecessary reader data.

4. Audit the services your publishing system buys. Ask suppliers where material comes from, what permits collection and reuse, whether access challenges are bypassed, and how corrections or deletions propagate. Record whether your integration merely receives data or initiates and configures collection. If the supplier cannot explain the route, pause that integration for review rather than treating “public data” as sufficient assurance.

Prioritize a bounded rights-and-access audit over a speculative enforcement campaign. Escalate a documented incident to counsel who can assess the works, authorization, bypass mechanism and injury. This ruling makes those details worth preserving; it does not remove the need to establish them.