An historic library reading room with towering dark wooden bookshelves, a scanner's glow illuminating an open centuries-old manuscript
AI & Singularity

Oxford Let OpenAI Train on Bodleian Texts — and Didn't Say So Out Loud

Meeting minutes obtained by the Guardian show the training use was discussed internally but never announced publicly. Oxford says the material is out-of-copyright and 'modest in scale'; the scans will be published openly within months.

OpenAItraining dataOxfordBodleian Librarycopyright

Oxford University has allowed OpenAI to use digitised texts from the Bodleian Library to train its AI models, according to internal documents obtained by the Guardian and seen by the student paper Cherwell. The digitised material has been used to “populate the OpenAI training set”, according to minutes of a 2024 meeting between the library and the company — a purpose that neither party’s public announcement of the partnership mentioned.

The public story, when Oxford and OpenAI launched their five-year partnership in March 2025, was digitisation: the BBC reported at the time that part of the Bodleian’s public collection would be digitised to make it “more widely available for students and researchers”. The documents now reported show the project had a second aim. “Populat[ing] the OpenAI training set with knowledge that uses under-represented data on the open web” is recorded as a key goal. An OpenAI spokesperson said the company was proud to ensure “the AI models of today preserve the world’s historical knowledge for the future”; a university spokesperson said the digitised text was out-of-copyright, “modest in scale”, and that the Bodleian retains rights to the scans and will publish them openly online within months.

Why a 900-year-old library is attractive to an AI lab is the more interesting story, and the Guardian’s reporting supplies the context: the open web is saturating with AI-generated text, which is worse training data, so developers are turning to physical, often historical, collections. Booksellers in the UK and Ireland have reported a spate of bulk orders for obscure titles — and rival Anthropic has spent tens of millions on “destructive scanning”, slicing spines off acquired books to scan them before pulping. The Next Web notes the Oxford deal is notable precisely because the Bodleian’s collections remain intact. By June 2025, 125,000 images from historical dissertations had been shared, alongside 10,000 16th-century “broadside ballads”.

The internal record is not a leak of wrongdoing — it is a record of ordinary institutional ambivalence. FOI’d meeting minutes show Bodleian governance committee staff raising reputational risk and the deal’s tension with the university’s environmental commitments, given the energy intensity of AI. That ambivalence got resolved in the direction of partnership, with the training-use aspect simply not put in the brochure. As our earlier coverage of Oxford’s own warmth-accuracy research showed, the university is one of the more clear-eyed institutions about what these models do and don’t do — which makes the quiet training use more interesting, not less: this is what a best-case library deal looks like, and it still involved the training purpose going unannounced.

The New Zealand angle is the precedent, not the parties. Australia’s copyright-and-datacentre accommodation with AI firms showed institutions trading access for infrastructure; the Oxford deal shows heritage institutions trading access for digitisation money and model access. GLAM institutions — galleries, libraries, museums — hold precisely the human-made, well-provenanced text that model builders now desperately want, and New Zealand’s own collections, from the Alexander Turnbull Library down, will face the same pitch. The choice Oxford made — public-private digitisation where the AI company gets training data and the public eventually gets open scans — is arguably a better bargain than the book-buying-and-pulping alternative operating elsewhere. But it was a bargain struck mostly in private, by an institution with 23 million items and no obligation to tell its students the training use until journalists FOI’d the minutes. That is the part worth watching here, where the government is still deciding what our own AI settings should be and national collections are being courted by the same companies.


Sources: The Guardian; Cherwell; The Next Web; BBC.

Sources: The Guardian, 'Oxford lets OpenAI train its AI models on Bodleian Library' (September 26, 2026), Cherwell, 'Oxford partnership allows OpenAI to use rare Bodleian texts to train ChatGPT' (September 27, 2026), The Next Web, 'Oxford let OpenAI train AI models on Bodleian texts, the Guardian reports' (September 26, 2026)