On another site we discussed LLMs and copyright and this is what I wrote about it there, and I liked it so much I thought I'd bring it here.
The underlying social reason why it should be this way is that copyright is (allegedly) for the purpose of protecting human effort in order to induce additional creations. Only if you work from the standpoint of it being only for protecting revenues (whether profit, tax, or both) does it even make sense to give copyright to AI creations.
The reason why it arguably is a problem, though I think it is critical to say that this really hasn't been decided in court anywhere from what I can tell, is a combination of two things: first, the definition of a derivative work is loose and open to subjective interpretation, and second that's just how the law is written. An LLM is software and you aren't, even if your argument is that they work the same... which they don't.
That last point is relevant to the argument as well. That's the technical reason why they shouldn't be the same. LLMs don't have ideas or concepts. They have only tokens and statistical relationships between them. The model records ways in which a token representing a thing is related to tokens representing other things, but there's no mind there to have a concept of the thing. I'm deliberately avoiding providing any example here to make a point — any thing is equally inscrutable to the same LLM producing a pronouncement about it.
That is relevant to this question because while we are capable of learning things about things because we have the concept of things, the LLM does not have understanding of any concepts because understanding is not a thing which it does, in fact it has nothing to do understanding with. The LLM is producing a facsimile of the very least and last part of the thinking process, but it's doing it with a combination of statistics and random numbers and that's it. That's why it can be equally confident about two sets of diametrically opposed results produced from the same query — it's specifically because there's nothing behind the facade.
Therefore, since the AI is not translating the ideas represented by tokens into ideas and concepts (since those are fundamentally not things an LLM can apprehend) at no point is it turning input into learning. "Training" is asserted to be the best word we have for what we are doing, but that's really not a thing because the LLM is the weights and the code that produces the chains of output. There's nothing else there! We have a mind with consciousness and ideas actually mean things to us, and we react to those meanings and they affect our output — when this isn't true, we have a word for that, and it is "psychopath".
The other argument to be made about LLMs not washing away copyright is that even if they were considered to be thinking, that is arguably not enough to make what they do legal. Even for humans we have the concept of taint, for which we invented clean room engineering as a circumvention strategy. But the LLM doesn't have any separation, as the model is a direct product of the input data.