Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

  • kromem@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    1 year ago

    Well, here’s straight from one of the suits against them:

    “The OpenAI Books2 dataset can be estimated to contain about 294,000 titles. The only ‘internet-based books corpora’ that have ever offered that much material are notorious ‘shadow library’ websites like Library Genesis (aka LibGen), Z-Library (aka B-ok), Sci-Hub, and Bibliotik. The books aggregated by these websites have also been available in bulk via torrent systems.”

    I’m not even sure how they would have logistically gone about purchasing 294,000 books in bulk in digital form to be fed into training. Using the existing collections seems much more likely, but I suppose we’ll see what turns up in litigation.

    Also, the penalty for downloading copyrighted material if willful infringement is up to $250,000 per work. So it’s quite a bit more than the cost of one book on the line…