Oxford lets OpenAI train its AI models on Bodleian Library
TL;DR
The University of Oxford has allowed OpenAI to train its AI models on historical texts from the Bodleian Library, according to internal documents reported by The Guardian. Material digitised by OpenAI has been used to populate the company's training set, as tech firms search academic institutions for fresh data. University staff have raised concerns about the reputational risk of partnering with the company behind ChatGPT.
Nauti's Take
Digitising historical collections is a real opportunity: old texts become searchable, and models learn language and knowledge beyond today's internet. The risk lies in the terms, since the deal only surfaced through internal documents and a university is handing cultural heritage to a commercial provider.
Libraries and archives should only sign such partnerships with transparent contracts and their own rights to the digitised material.