HomeNews ReportsMark Zuckerberg permitted Meta to use pirated books to train its AI model, say...

Mark Zuckerberg permitted Meta to use pirated books to train its AI model, say authors accusing Meta of copyright infringement

Meta employees reportedly used Library Genesis (LibGen), a well-known database of pirated books, and some of them expressed concern over torrenting on company laptops.

Meta founder and CEO Mark Zuckerberg reportedly allowed his company to use pirated copies of copyrighted books to train its artificial intelligence systems. Meta is currently facing lawsuits for copyright infringement from various authors and comedian Sarah Silverman who accused the company of misusing their works to train its large language model Llama.

As per reports, documents submitted in California federal court revealed that internal files of Meta show that the company was aware of books being pirated using torrents. Meta employees reportedly used Library Genesis (LibGen), a well-known database of pirated books. The complainants had sought permission from the court to submit an updated complaint on Wednesday.

The company’s internal communications showed employees expressing concern over downloading books from LibGen. One of the engineers reportedly expressed reservations about ‘torrenting from a corporate laptop’. It is the allegation of the authors that Zuckerberg approved the use of LibGen despite the concerns of its executive team.

Denying the allegations of copyright infringement, Meta took the defence of the ‘fair use’ doctrine. It argued that the plaintiffs were aware of the use of LibGen by Meta since July 2024 and had enough time to use this information in their complaints.

Meta’s redaction attempts were dismissed by Judge Vince Chhabria of the US District Court for the Northern District of California as ‘preposterous’ and intended to avoid negative publicity instead of protecting business interests. The judge warned Meta against making any requests for redaction in future stating that any unreasonable broad selling requests would result in all material being unsealed.

While Meta had earlier admitted to using Books3, a dataset of around 196,000 books for training its Llama language model, it did not publicly disclose the direct use of LibGen data.

Join OpIndia's official WhatsApp channel

  Support Us  

For likes of 'The Wire' who consider 'nationalism' a bad word, there is never paucity of funds. They have a well-oiled international ecosystem that keeps their business running. We need your support to fight them. Please contribute whatever you can afford

OpIndia Staff
OpIndia Staffhttps://www.opindia.com
Staff reporter at OpIndia

Related Articles

Trending now

Private hospital nurse Sirisha raped and murdered by bike mechanic Mohammed Sanaullah in Telangana: All you need to know

Sanaullah found out that Sirisha was working as a nurse at the nearby Prime Care Hospital and stayed in a hospital-provided room. Gradually, the police said, Sanaullah developed ‘sexual interest’ in Sirisha that drove him to raping her.

Russia’s big defence pitch: Why Putin is pushing S-500 and Su-57 on India

India's own defence industry is becoming a stronger competitor. Programmes such as Tejas, AMCA, Project Kusha and the indigenous ballistic missile defence system reflect New Delhi's long-term effort to build critical military capabilities domestically.
- Advertisement -