• Tyler_Zoro
    link
    fedilink
    English
    arrow-up
    1
    ·
    3 years ago

    I think you can make some reasonable arguments about how AI training is good or bad, but this is not that. This is just a terrible, terrible article that no one should take seriously.

    Describes GPT as:

    a glorified copy and paste machine

    Which is so far from what it actually is that I would expect even the layperson to understand that.

    I’m not a big fan of these conglomerates ingesting other people’s work and then reselling it

    Which… they aren’t. Also, how is GPT (or even the company that created it) a “conglomerate”? Does the author even know what that word means?

    I can 100% confirm that Brave lets you ingest copyrighted material through their Brave Search API

    So… you can do web searches through Brave’s API. Okay… and this was resolved in the courts over a decade ago…

    Brave offers numerous API products, some of which are specifically designed for AI. This one, Data for AI, lets you “Feed results to AI models for inference”, while their premium version of this same API lets you “Cache/store data to train AI models” not only with “regular” rights but also “storage rights”.

    This appears to be the author conflating the rights to the API search results with the content of source websites.

    You can go download content from any website you like. That’s not what they’re giving you rights to. They are giving you rights to THEIR API responses which summarize and categorize web data in a search index similar to those used by Google and Microsoft.

    As far as search engines go, they can get away with it because linking back to a Wikipedia article on the same page as the search results is considered attribution.

    That is not AT ALL why linking to a source page with duplicated text from that site is considered fair use! If that were the case, then Google could never include any search results for pages that were not clearly licensed (which is an awful lot of the web).