Reddit CEO Steve Huffman is taking a stand against major tech companies like Microsoft, demanding they pay for the right to scrape Reddit’s data. This move comes after Reddit successfully struck deals with Google and OpenAI, allowing these companies to use Reddit’s data under agreed-upon terms.
In a recent interview, Huffman emphasized the importance of these agreements. “Without these agreements, we don’t have any say or knowledge of how our data is displayed and what it’s used for, which has put us in a position now of blocking folks who haven’t been willing to come to terms with how we’d like our data to be used or not used,” he said. He specifically named Microsoft, Anthropic, and Perplexity as companies that have refused to negotiate, describing the situation as “a real pain in the ass to block these companies.”
In an effort to protect its data, Reddit has ramped up its fight against unauthorized web crawlers. At the beginning of July, the site’s robots.txt file was updated to block crawlers without agreements. This led to Reddit results becoming visible only in Google search results, where Reddit is paid for its data, and not in other search engines like Bing.
Huffman accused Microsoft of using Reddit’s data to train its AI and summarizing content in Bing results without informing Reddit. He also claimed that Reddit’s data has been sold through the Bing API to other search engines. In response to Microsoft AI CEO Mustafa Suleyman’s comment that public data on the internet is “freeware,” Huffman countered, “We’ve had Microsoft, Anthropic, and Perplexity act as though all of the content on the internet is free for them to use. That’s their real position.”
Following Reddit’s actions to block Bing from crawling its site, Microsoft’s head of search, Jordi Ribas, stated on X that “Reddit has blocked Bing from crawling their site for search, favoring another search engine and impacting competition from Bing and Bing-powered engines.” Microsoft spokesperson Caitlin Roulston echoed this, telling The Verge that “we honor the directions provided by websites that do not want content on their pages to be used with our generative AI models.”
Huffman pointed to OpenAI’s SearchGPT, which includes Reddit results thanks to a prior agreement, as the model he wants to replicate. According to Reddit spokesperson Tim Rathschmidt, none of the current content licensing deals Reddit has include exclusive use cases for its data.
By advocating for licensing deals, Reddit aligns itself with traditional media publishers who seek compensation for content used to feed generative AI. “I think the traditional value exchange from search engines has changed,” Huffman remarked. “Search and summarization and training are merging, and the value exchange of crawling in exchange for traffic back is becoming muddied.”
In response to Huffman’s comments, Anthropic spokesperson Jennifer Martinez stated, “Reddit has been on our block list for web crawling since mid-May and we haven’t added any URLs from Reddit to our crawler since then. We respect robots.txt, the industry accepted signal for blocking web crawling.”

Your First 10 AI Skills
10 practical AI skills, copy-paste prompts and a 7-day plan to start using AI with confidence.
Download the guide →
