Do androids dream of electric copyright?

On March 26, 2026, digital artist Austin Beaulier launched a class action lawsuit against Roblox, a video game developer for the use of ‘web-scraping’, in which AI models trawl the internet for data that can be used to train AI models.[1] This is an area of uncertainty about the rights of creators and the obligations of companies providing data for AI training: these class actions join a list of twenty-eight active web-scraping related lawsuits in the United States right now.[2]
The contention is that potentially millions of 3D models created by hobbyists and small companies were uploaded to digital repositories such as GitHub, under Creative Commons license terms that generally required attribution when used and / or restrict commercial use. Many of the relevant 3D models had an additional ‘No-AI’ tag indicating the creator’s legal non-consent.[3] The AI training dataset in question, Objaverse-XL, preserved precise links to the models downloaded for training purposes, however, there was no attempt to segregate works with a restrictive license or ‘No-AI’ tag when the dataset was subsequently used to train AI models. [4]
It is difficult to predict what the court here will decide – or in similar cases in the future. Such use is very much a legal contested grey area. A similar US case, Bartz v. Anthropic PBC, determined that it is permissible for AIs to be trained off of copyrighted data, on the basis the output itself is sufficiently detached from each piece of training data.[5] However, in Kadrey v Meta, although the court did not rule against Meta for using copyrighted materials in AI training, the court underlined that its decision did not determine that such use was actually lawful, and that future cases may decide to consider it a copyright violation when AI might threaten to dominate the market for creative works.[6] Similar cases are in progress in the EU, UK and other courts.
Whatever the courts might decide, it is clear that a final legal position on this use of copyrighted training data will have serious ramifications. Should the legal position ultimately favour the rights of creators to maintain control over derivative uses of their data, then AI training may come under threat. Creating effective safeguards will likely be a significant burden: in 2022 it was estimated that there are more than 2.5 billion creative commons licensed works online, that Objaverse-XL was a repository of approximately 10.2 million unique 3D models, and that there are six different creative commons license types, with varying use obligations.[7] As these trackers are automated, then human oversight will likely be required to ensure compliance over this complex nest of data. Given the complexity and quantity of data, how this would practically work is unclear – it may be that companies have to factor non-compliance as a possible financial and legal risk when creating training data sets.
Perhaps even more significantly for companies seeking to use AIs for creative generation, if a significant proportion of high-quality data is ‘locked away’ from models, the quality of AI output predictably suffers. A recent study investigating the performance of leading LLMs with and without access to copyrighted material – copyrighted suggesting it is both recent and high-quality – found that the LLMs with access had 23.4% better accuracy when answering reading comprehension questions.[8] Without unrestricted access to the dataset of Objaverse-XL, it is impossible to quantify the impact that restrictions will have on its outputs, especially considering the relative complexity of 3D models. However, should the trained output from the dataset cease to be of a consistently high quality, using AI to facilitate rapid creative design may become less worthwhile in contexts such as game design which demand high-quality visuals.
Should the legal position however be that using copyrighted material for training AIs is acceptable, then this will likely have a deleterious effect on the willingness of creators to generate and upload work. After DeviantArt, a popular digital hub for artists, allowed creators’ work to be used for training the site’s propriety AI, artists uploaded approximately 21% fewer works after the change over a year long period – even though artists could opt-out their creations.[9] It is difficult to measure what impact (but the impact would presumably be greater) this has on artists who have no ability to opt-out their work. A 2025 survey of 335 creative workers showed that 68% felt they had diminished job security as a result of AI.[10] Should artists (and those who might choose to develop artistic skills) increasingly find that it is uneconomical or undesirable to continue investing in their skillset and creating content, the AI models that are threatening them may themselves become undermined. Without novel, high-quality training data, then AI will be increasingly (and probably inadvertently) trained on AI generated work. A recent paper has described the effect of image generation AI being trained on an image dataset principally created by generative AI, under the creatively named Model Autophagy Disorder (MAD) framework.[11] Within three generations, artefacts – or telltale AI distortions associated with primitive AI image generators – dominate, in effect rendering the outputs useless.

Figure 1. The above diagram from ‘Self-Consuming Generative Models Go MAD’ shows the rapid decline in AI output
All courts will be playing a very fine balancing act in determining the future of copyright in the context of AI training. Interestingly, the main concern for courts has shifted towards the impact on creators, rather than questioning whether the process behind AI image generation is sufficiently ‘creative’ to avoid copyright breaches. This gives additional credence to the idea that courts may begin to explicitly favour creators. Whatever the legal position ultimately becomes, both approval and rejection risk undermining AI – we may very well be in the golden age of AI development in the creative industries – and an internationally consistent regulatory framework is urgently needed for both creators and AI developers. In the meantime, those creating AI datasets and the models that use them will need to consider the possibility that the regulatory headwinds will shift. Companies should focus on pre-emptively planning safeguards to avoid using copyrighted material. For creators unwilling to allow their work to be used to populate AI datasets, it appears there are few if any effective legal safeguards once work has been uploaded to the internet. For creators, it may be necessary to mitigate the risk by using restricted access forums and digital hubs, albeit adopting this measure will likely reduce audience reach.
[1] Beaulier v Roblox Corporation, No 3:26-cv-02642, Class Action Complaint (ND Cal, filed 26 March 2026)
[2] Scraper.io, ‘Web-Scraping Lawsuits’ (last updated 31 August 2026) [https://scraper.io/feeds/web-scraping-lawsuits] accessed 16 September 2026
[3] Beaulier v Roblox Corporation, para 60
[4] Ibid, para 87
[5] Bartz v Anthropic PBC, 787 F Supp 3d 1007 (ND Cal 2025)
[6] Kadrey v Meta Platforms Inc, 788 F Supp 3d 1026 (ND Cal 2025)
[7] Creative Commons, ‘State of the Commons 2022’ (11 April 2023) [https://creativecommons.org/state-of-the-commons-2022], accessed 16 September 2026
[8] Stella Jia and Abhishek Nagaraj, ‘Cloze Encounters: The Impact of Pirated Data Access on LLM Performance’ (NBER Working Paper No 33598, 2025)
[9] Sijie Lin, ‘Hiding From Generative AI’ (working paper, February 2026)
[10] Queen Mary University of London, ‘Creative Industry Workers Feel Job Worth and Security Under Threat from AI’ (23 January 2025) [https://www.qmul.ac.uk/news/latest-news/2025/queen-mary-news/pr/creative-industry-workers-feel-job-worth-and-security-under-threat-from-ai-.html], accessed 16 September 2026
[11] Sina Alemohammad and others, ‘Self-Consuming Generative Models Go MAD’ (International Conference on Learning Representations 2024)
‘MAD’ is a reference to mad cow disease.