How it works

Where Our Data Comes From

No scraping, no guessing: ART WIKI is built on museums' own open-access APIs, public-domain imagery, and editorial records that cite their sources. A tour of the workshop.

16 augustus 2026 · 7 min lezen · Redactie ART WIKI

Every fact on ART WIKI is traceable. Open any artwork page and scroll to Sources: you'll find the providing institution, when we fetched its record, and the license of the image you're looking at. This isn't a footnote aesthetic — it's the product.

The providers

The Art Institute of Chicago publishes a CC0 API that includes live display state per gallery. The Cleveland Museum of Art's Open Access program covers both data and imagery. The Met's Open Access does the same for public-domain works, down to the gallery number. Each museum becomes a connector — around two hundred lines of code — and everything downstream of it, from entity resolution to notifications, is shared machinery.

Mand met appelen
Cézanne's Basket of Apples — CC0 metadata from AIC; public-domain image via Wikimedia Commons. Naar het werk →

When sources disagree or disappear

Different museums spell the same artist differently; we resolve identities through authority records — Wikidata, Getty ULAN — and never merge on names alone. Suspected duplicates are flagged for human review, not auto-merged. And when a museum's image server went behind a bot wall this summer, we replaced imagery only with hand-verified public-domain files, checked one by one against the actual object, and said so in the record.

“The moat is not the interface. The moat is the data being right.”

— ART WIKI working notes

Editorial records — the stolen works, the private collections — carry a lower confidence grade by design, with citations to public documentation. When we don't know, the interface says so. That's the whole trick.

Op de omslag: Mand met appelen — Paul Cézanne. Public domain (PD-Art), via Wikimedia Commons

Where Our Data Comes From · Journal · ART WIKI