BDRevise24 — English Edition

বিজ্ঞাপন

আপনার বিজ্ঞাপন এখানে — HEADER
Technology

New Filings Reveal Internal OpenAI and Microsoft Debates on Web Scraping

Newly unsealed filings in The New York Times copyright lawsuit against OpenAI and Microsoft have revealed internal discussions regarding AI scraping. The documents shed light on how automated systems interact with publisher content and impact web traffic.

BDRevise24 Desk
New Filings Reveal Internal OpenAI and Microsoft Debates on Web Scraping
Photo: TechCrunch

Newly unsealed filings in The New York Times copyright lawsuit against OpenAI and Microsoft have brought to light a series of internal discussions regarding artificial intelligence scraping and its direct impact on news publishers. The legal battle, which was originally brought by The New York Times three years ago, has continued to expose sensitive internal communications between executives and technical staff at the major artificial intelligence and technology corporations.

Among the newly revealed details from the September 17, 2026 filings are significant statistics concerning the performance of search and answer engines. Specifically, Microsoft's Copilot answer engine caused a staggering 93 percent drop in click-through rates for The New York Times domain when compared directly to traditional Bing search results. This dramatic decline highlights the stark economic differences between conventional web traffic referral models and modern generative artificial intelligence deployment.

The court documents also detail the massive scale of data ingestion utilized during model development. The filings reveal that OpenAI's mid-training datasets contained upwards of 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting. Furthermore, a Common Crawl-derived dataset was shown to include more than 2 million documents sourced directly from nytimes.com, while the Project Mango training dataset contained over 160,903 unique works from various news publishers.

Internal correspondence cited within the unsealed documents featured commentary from key figures across both companies, including individuals such as Rebecca Bellan, Justin Sullivan, Brent Hecht, Satya Nadella, Nick Turley, Greg Brockman, and Nick Ryder. Addressing the precarious nature of the industry's reliance on journalism, Brent Hecht noted in an internal communication, "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’"

Further commentary within the administrative and executive ranks touched upon the necessity of proper content acquisition frameworks. Satya Nadella remarked that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training." Meanwhile, other internal reactions captured in the filings included brief responses such as Greg Brockman stating, "ah nice," as the companies navigated the rapid development of their large language model systems alongside platforms like Replit.

বিজ্ঞাপন

আপনার বিজ্ঞাপন এখানে — IN_ARTICLE

More in Technology

See all

বিজ্ঞাপন

আপনার বিজ্ঞাপন এখানে — FOOTER