Can You Use Museum Open-Access Images for AI Art?
Copyright-wise, CC0 museum images are about the cleanest input you can feed a model. But the output side, the platform terms, and the cultural-sensitivity question are all less obvious.
Quick answers
- Can I use CC0 museum images as AI training data or prompt inputs?
- As a matter of copyright, yes. A CC0 release waives the rights that would otherwise restrict copying, and a public-domain artwork has no rights holder. These images are among the cleanest inputs available for training, fine-tuning, or image-to-image prompting.
- Do I own the copyright in an image an AI generates from public-domain art?
- In the United States, purely machine-generated output is not eligible for copyright, because protection requires human authorship. Substantial human creative contribution can be protected, but the protection covers that contribution rather than the whole image.
- Does using public-domain inputs protect me from AI copyright disputes?
- It removes one category of risk, the rights in your inputs, which is the category most disputes have concerned. It does not address the rights in a model's own training data, platform terms of service, or how closely an output resembles a protected work.
- Are there non-copyright reasons not to use certain museum images?
- Yes. Museums flag culturally sensitive material such as ancestral remains, sacred objects, and community-restricted images. These can be legally unrestricted while still carrying real obligations, and many institutions publish guidance asking that they not be used casually.
Of every question in this subject, this one gets the most confident wrong answers in both directions. So, precisely:
On the input side, CC0 museum images are about the cleanest material in existence. On the output side, the answer is more interesting than most people expect. And there's a third consideration that isn't legal at all.
Inputs: genuinely clear
The copyright question about AI inputs is whether copying a work to train on or prompt with it infringes someone's rights. For an open-access museum image, there is no someone:
- The artwork is out of copyright — nobody holds rights in a 1640 painting.
- The reproduction is covered by the museum's CC0 waiver, which explicitly surrenders any claim in the photograph.
That double clearance is why these collections turn up constantly in research datasets and fine-tuning sets. Training a style model on a museum's holdings, running image-to-image from a public-domain portrait, building a conditioning set from a textile collection — all clear on inputs. Compare a scraped image board, where the rights position is unknown for every file, and the difference is the entire ballgame.
Outputs: the surprising part
Here's where people's assumptions break. In the United States, purely machine-generated output is not eligible for copyright — protection requires human authorship, and the Copyright Office has consistently held that a prompt alone doesn't supply it. Courts have agreed.
What that means in practice:
- Your AI-generated image may be something nobody can own, including you.
- Human creative contribution can be protected — your arrangement, your selection, your substantial edits — but the protection covers that contribution, not the machine-made parts underneath.
- If you're selling AI-assisted work, this is a commercial fact, not a trivia point: you may not be able to stop a competitor reusing your output.
- Other jurisdictions differ, and the position is moving. Check where you publish.
The irony is worth noting: you start from art that is free for everyone to use, and you may end up with a result that is also free for everyone to use.
The risks that public-domain inputs do not remove
Clean inputs solve one problem and leave three:
- The model's own training data. You did not train Midjourney or Stable Diffusion. Whatever went into them is a separate question from what you put in your prompt, and most of the disputes so far have been about that.
- Platform terms of service. Commercial rights, output ownership, and usage limits are set by your contract with the tool, not by copyright law. A free tier often reserves rights a paid tier grants.
- Substantial similarity. Prompt a model hard enough toward a living artist's style and the output can resemble a protected work regardless of how clean your reference image was. Starting from a public-domain input is not a defence against arriving at a protected destination.
The question that isn't legal
Open access made millions of images legally frictionless. It did not make all of them ethically neutral, and museums have been fairly direct about this.
Collections contain ancestral remains, sacred and ceremonial objects, images of identifiable people held under colonial conditions, and material that originating communities have asked be treated with restraint. The Smithsonian and other institutions publish cultural-sensitivity guidance for exactly this reason. Such items can be CC0 and still carry obligations — a released copyright doesn't relinquish a community's interest in its own heritage.
None of this bears on a Dutch still life. It bears a great deal on feeding an ethnographic collection wholesale into a style model. The rule of thumb is unglamorous and works: if the object record carries a sensitivity note or names a living community, read it before you train on it.
A practical workflow
- Source from open access only, so your inputs are documented rather than assumed.
- Record provenance per image — object ID, institution, rights status — before it goes into a set. Reconstructing this afterwards is miserable.
- Read the object record, not just the rights badge, on anything ethnographic or human-remains adjacent.
- Know your platform's output terms before commercial use.
- Keep the human contribution visible and documented if you want any claim to the result.
Building a clean set
Musist is useful here precisely because it makes the rights status structural rather than something you go looking for: every object across The Met, the Rijksmuseum, and the Smithsonian carries a rights badge, and a Download image action appears only for public-domain works with a full-resolution file. Assembling a set where every item is documented CC0 becomes browsing rather than auditing.
Full metadata — artist, date, medium, credit line, source — sits on each object page, which is what you'll want in your provenance record. Start from the collections, or let build you a set that is more interesting than a keyword search would have been.
Practical guidance, not legal advice. The law here is moving faster than almost anywhere else in copyright — verify the current position before you build a business on it.
- licensing
- open access
- ai
- cc0