Users have repeatedly expressed concern about the origin of the data that large companies use to train AI models. In this case, Apple boasts of defending the privacy of its users by not using their data to train Apple Intelligence. However, no one has said anything about the internet and YouTube videos.
According to Wired, big tech companies like Nvidia, Anthropic, and Apple have used material from thousands of YouTube videos to train their AIs. This has happened due to the use these companies have made of the material from Eleuther AI, a non-profit organization that aims to help more independent AI developers.
Eleuther AI downloaded subtitle files from over 170,000 YouTube videos, which were then compiled into a large dataset called Pile. This organization, dedicated to open-source artificial intelligence, has made most of this data available for anyone to use, including big tech companies. A notable example is Apple, which has claimed to use Pile to train OpenELM, an advanced AI model developed just a few weeks before the launch of Apple Intelligence.
Founded in July 2020, the main objective pursued by Eleuther AI is to assist in the development of open-source AI models. Despite primarily targeting small developers, the data it collects also captures the interest of tech giants like Apple.
However, it is worth noting that neither Apple nor the other companies had directly used the YouTube data. These had already been collected previously by Eleuther AI, so it is they who would have violated YouTube’s terms and conditions. Either way, this situation sheds light on a fairly common situation: indiscriminate data theft.