J
Harvard adds copyright-free fuel to the AI fire.
With funding from Microsoft and OpenAI, the university’s Institutional Data Initiative (IDI) is releasing a dataset for training AI that contains nearly one million public-domain books — around five times larger than the controversial Books3 dataset.
The aim is to “level the playing field” for smaller AI developers who don’t have access to the massive datasets used by tech giants, according to IDI executive director Greg Leppert.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
Loading comments
Getting the conversation ready...
Most Popular
Most Popular
- The Dell XPS 13 is the first real competitor to the MacBook Neo
- Apple releases iOS 27 with Siri AI overhaul
- The Steam Frame is made for irresponsible hardware nerds like me
- Apple Home’s new security camera features cost as much as $60 a month
- Microsoft issues emergency Windows 11 update to fix its record-breaking patch











