Reddit announced that it will limit the Wayback Machine from archiving most of its site. The company said it found cases where AI firms used archived Reddit content for scraping. As a result, the archive will not be able to crawl post pages, comments, or user profiles. Only the Reddit homepage will remain available for archiving under the new limits.
Reddit will stop the Wayback Machine from indexing post detail pages, comment threads, and profile pages. That means snapshots of individual discussions will no longer be saved in the public web archive. The archive will still be able to record the site home page, which shows trending posts and headlines for a given date. Reddit said it told the Internet Archive about the limits before they began to take effect.

Why Reddit is taking this step
Reddit said it detected instances where companies scraped content from the Wayback Machine and then used that data in ways that violated Reddit policy. The company pointed to cases where archived pages provided a path for third parties to collect large sets of user-generated content. Reddit claimed the restriction of the archive access aids in user privacy and ensures that material removed by the user will not be republished via the archive.
The Wayback Machine makes an archive of what people on the public web can view and holds these copies in a way that people get an idea of how a page appeared on a certain date in the past. Those snapshots serve the interests of archivists and researchers so that they may analyze the historical development of link rot, news, and site transitions. The Internet Archive has served as a key asset in preserving a historical account of online material that stands to be erased.
What the Internet Archive says
The Internet Archive said it continues to talk with Reddit about the change. The archive said it aims to preserve the public web and that it wants to find solutions that protect users while keeping historical records. The two organizations have a history of working through web access and removal requests and they said dialogue remains open.
Limiting archive access cuts off a source that some researchers and journalists use to track conversations and recover deleted material. That will make it harder to study how stories and threads evolved over time. Meanwhile, Reddit argues that the action defends the privacy of users and minimizes the possibility of the deleted material being recycled towards mass AI training or other applications that are not permitted on the site.
The wider context about scraping and AI
Platforms have been tightening access after AI firms used public data to train models. Reddit has already made deals that let some companies access its data under license. It has also sued at least one AI firm for alleged scraping. The new step shows how platforms are trying to control which tools can access their content and how that content may be reused.
Legal and ethical questions
Archivists state that the web archive preserves cultural memory, which promotes accountability. Platforms claim that they have to defend their users and apply the terms of service. The conflict also throws up legal and ethical issues as to the ownership of public content, to what extent it is desirable that it should be permanently preserved or should be allowed to die in privacy, and how to prevent bad scraping without destroying history. There is a possibility of courts and policymakers having to spell out the guidelines in light of these controversies.

Practical steps for researchers and readers
If an important Reddit discussion vanishes from the Wayback Machine, users should save copies and cite original sources quickly. Researchers can request access or work with the archive and with site owners to preserve key materials. Journalists should document and archive sources in multiple places and explain any gaps so readers understand what the record shows and what is missing.
Reddit moved to block large-scale archiving of its posts and profiles to stop third parties from harvesting archived content for AI scraping. The change shifts how the public record of Reddit conversations will be kept. Readers and researchers will need to adapt if they want to preserve important material going forward.
