AI Training on Public Data: Theft or Fair Use?
With the explosion of generative AI, many models like ChatGPT and Midjourney require vast amounts of data for training, often sourced from text, images, and videos on the internet. Creators and rights holders accuse AI companies of using their works without permission, while AI companies argue it falls under 'fair use' or mere 'learning.' At the heart of the debate: what copyright rules should govern AI training? Should original rights holders be compensated?
⚡ Quick take — no essay needed
Pro
How is it theft if it's literally public? We all look at art and get inspired. AI is just learning patterns, it's not copy-pasting.
Honestly, it’s just learning from what’s public. Humans do the exact same thing when we study art for inspiration, so I don't see how AI is any different.
Honestly, it's just like how human artists learn by looking at other people's work online. It’s public data anyway, so calling it theft is a bit of a stretch.
Con
Honestly, how is this different from a human artist looking at inspiration online? It's public data, they're just learning patterns, not copy-pasting.
Honestly, how is scraping someone's portfolio to train a commercial model "fair use"? If you profit off my art, you should at least ask first and pay up.