Duke Humfrey’s Library, the oldest reading room of the Bodleian Library in Oxford (illustrative). Image: Diliff / Wikimedia Commons, CC BY-SA 3.0, cropped

One of the world’s most famous libraries has become AI training data. The University of Oxford has let OpenAI use historical texts digitised from its Bodleian Library to train its AI models, the Guardian reports, citing internal documents that say the material has been used to “populate the OpenAI training set.”

What Oxford shared

Oxford announced a partnership with OpenAI in March 2025, using OpenAI’s software to digitise texts from the Bodleian so students and researchers could reach them more easily. That announcement didn’t say the scans would also be used to train OpenAI’s models.

According to university meeting minutes obtained by the Guardian through a freedom of information request:

  • By June 2025, 125,000 images from historical dissertations had been shared with OpenAI, including PhD theses from European and American universities written in the 19th and 20th centuries.
  • Other scans include a rare collection of 10,000 16th-century “broadside ballads”, song sheets once sold on Tudor street corners.
  • Staff have discussed digitising 18th-century Irish state papers, the private letters of the novelist Maria Edgeworth and Dorothy Hodgkin’s penicillin notebooks.
  • The contract raises the prospect of digitising much more of the Bodleian’s 23 million items, and the minutes mention an “Ask the Bod” chatbot.

The same minutes record concerns from staff, including members of the Bodleian’s governance committee, about the reputational risk of partnering with OpenAI and about the environmental impact of a deal with such an energy-hungry technology.

What Oxford and OpenAI say

Oxford says the project is “modest in scale,” covers only out-of-copyright material and isn’t exclusive to OpenAI. A spokesperson said the Bodleian keeps the rights to the scans and will start publishing them openly online in the next few months, and rejected any suggestion that the AI training was hidden from staff or students. Digitisation was the university’s main interest, the spokesperson said, but staff had been open that the project would also provide training data.

OpenAI said it was “proud” to help ensure “the AI models of today preserve the world’s historical knowledge for the future,” adding that with more than a billion people using the technology, “it’s important it reflects different cultures, histories and perspectives.”

Oxford is the only UK member of OpenAI’s NextGenAI programme, which has similar arrangements with US libraries and universities including Boston Public Library, Caltech, MIT and the University of Michigan.

Why AI companies want old books

The deal is part of a wider rush for fresh text. The open web is filling up with AI-generated writing, which makes it less useful for training new models, so developers are turning to physical and historical collections that have never been online.

Unlike some of those efforts, the Bodleian’s books stay intact. The Guardian notes that Anthropic has spent tens of millions of dollars buying books and slicing off their spines to scan them before they’re pulped, though it says it doesn’t do this to rare or antiquarian books. Secondhand booksellers have also reported a run of orders for obscure titles, from a guide to farm tools in 18th-century Africa to biographies of 1950s car drivers, which they suspect are being bought for AI training.

Why it matters

Libraries hold exactly what AI companies now want most: vast amounts of human writing that isn’t already on the internet. Oxford’s deal shows how that knowledge can be opened up to everyone and handed to a tech company in the same move, and it’s likely to put pressure on other universities to say plainly what their own AI partnerships involve.

Source: The Guardian.

Related