It is widely believed (and in some cases acknowledged) that a lot of models are trained on copyrighted data scraped from the web. In some cases, even scrapes of ebook piracy websites - google 'books3' to learn more.
Some companies (such as those working on AI) believe this is legal, others (such as the copyright holders to those books) believe it isn't.
In any case, IMHO it's unlikely any cutting edge models will be offering us their training data any time soon.
Some companies (such as those working on AI) believe this is legal, others (such as the copyright holders to those books) believe it isn't.
In any case, IMHO it's unlikely any cutting edge models will be offering us their training data any time soon.