How Tech Giants Cut Corners to Harvest Data for A.I.

OpenAI, Google and Meta ignored corporate policies, altered their own rules and discussed skirting copyright law as they sought online information to train their newest artificial intelligence systems.

In late 2021, OpenAI faced a supply problem.

The artificial intelligence lab had exhausted every reservoir of reputable English-language text on the internet as it developed its latest A.I. system. It needed more data to train the next version of its technology — lots more.

So OpenAI researchers created a speech recognition tool called Whisper. It could transcribe the audio from YouTube videos, yielding new conversational text that would make an A.I. system smarter.

Some OpenAI employees discussed how such a move might go against YouTube’s rules, three people with knowledge of the conversations said. YouTube, which is owned by Google, prohibits use of its videos for applications that are “independent” of the video platform.

Ultimately, an OpenAI team transcribed more than one million hours of YouTube videos, the people said. The team included Greg Brockman, OpenAI’s president, who personally helped collect the videos, two of the people said. The texts were then fed into a system called GPT-4, which was widely considered one of the world’s most powerful A.I. models and was the basis of the latest version of the ChatGPT chatbot.

Source: nytimes.com

H2 Interactive sets Steam launch date for IGS Classic Arcade Collection

Juspay integrates Hyperswitch into Recurly to expand subscription payment options

FB Success Story +155% FTD, 135% ROI

FB Success Story +155% FTD, 135% ROI

Blueprint Gaming upgrades Super Graphics Upside Down for Lottomart exclusive

H2 Interactive sets Steam launch date for IGS Classic Arcade Collection

Juspay integrates Hyperswitch into Recurly to expand subscription payment options

FB Success Story +155% FTD, 135% ROI

FB Success Story +155% FTD, 135% ROI

Blueprint Gaming upgrades Super Graphics Upside Down for Lottomart exclusive

AppsFlyer: UK app activity spikes on every goal during England knockout games

Blocks & Headlines: Blockchain.com, OpenWorld, JPM Coin, Citi Token Services, GBBC, HTX, TRM Labs, NTT DOCOMO Global and XDC Network – July 22, 2026