About the Tinbergen Project
How a collection of Jan Tinbergen's correspondence became one of the Library's first Linked Open Data datasets — and what we learned along the way.
Project goal
We have valuable collections that are digitally accessible and available via ContentDM. Some researchers already find the collections through our social media channels. We want more people to discover our collections via search engines, which can be achieved by publishing the metadata as Linked Open Data.
There is already some knowledge about this within our organization, but not yet sufficient to realize it. By collaborating with NDE (Network Digital Heritage) and other libraries participating in the PICA Foundation's program we can achieve our goal more easily. We aim to enrich our (meta)data within the workflows and publish it as Linked Open Data. We format the available data from the open digital collections into LOD format, publish it as Linked Open Data, and deliver it as a dataset. In a subsequent phase, we intend to explore the possibilities of digitising and making more collections available.
Target audience
Our special collections are of great importance to researchers, historians, economists, and students. By making these collections more accessible, we expect to reach a wider audience and strengthen our contribution to academic research.
Project results
- Enriched and published metadata: The metadata of our digital collections will be enriched and published as Linked Open Data, significantly improving findability and accessibility.
- Signed NDE Manifest: By signing the manifesto, we confirm our commitment to the principles of the NDE and our intention to implement them.
- New digital strategies: We will implement new workflows and cataloguing methods focused on persistent identifiers and the use of LOD.
- Basis for future collaboration: The project lays the start for further digitization and collaboration with other institutions, which could potentially lead to the establishment of a thematic portal.
Scope
Proposed digitised collections:
-
Tinbergen correspondence
Source data
The collection was previously disclosed (Dublin Core) and published in ContentDM. The source data was exported to Excel, from which we enriched the metadata.
Process
Phase 1. Validating and linking entities
Model and vocabularies
We have chosen Schema.org as our model and applied various vocabularies such as NTA, VIAF, ROR, Geonames, and Wikidata.
- Organisations, countries, cities, languages — 95% found (484). These entities were linked to Wikidata or specialised sources: Geonames for geographic names, and ROR for research institutions. As these types of entities are generally well-documented, the match rate was high.
- Persons — 63% have a URI (1165). This proved more challenging for several reasons:
- The letters mostly date from the 1950s and 1960s. Some addressees are internationally known (academic peers of Tinbergen), while others were active in Dutch politics or the business world and did not yet have an existing Wikidata record.
- For the latter group, new Wikidata records were created where possible, based on Delpher and other open archives.
- In cases of doubt, the field was deliberately left blank rather than forcing an uncertain match — particularly for Dutch names where only initials and a surname were known, as multiple family members could share the same name at the time.
Phase 2. Creating PIDs and linking to ContentDM
We first investigated whether it was feasible to create a PID per item. The outcome was that this can be achieved via a DataCite licence, which enables the generation of DOIs (Digital Object Identifiers).
Each DOI points to the landing page of the relevant item in ContentDM — i.e., the location where the digitised letter or manuscript can actually be viewed. A DOI thus functions as a permanent, unchanging address, even if the underlying URL were ever to change. The DOIs will be included in the code for conversion into RDF.
Phase 3. Processing corrections in ContentDM via the Catcher API
Using a script, all data in ContentDM were overwritten, including the URIs (the Wikidata/Geonames/ROR/VIAF/NTA links from Phase 1). This ensures that the source data system itself is already enriched prior to the RDF conversion in Phase 4.
By taking this step, not only is the final RDF dataset (Phase 4) enriched, but the source itself — ContentDM, the system in which the collection is managed and searched on a daily basis — is also brought up to standard. This prevents the enriched metadata from existing only in a separate, detached dataset, disconnected from the system that users and administrators work with every day.
Phase 4. RDF conversion and publishing the dataset
Conversion to RDF
This phase starts with converting the data to RDF (Resource Description Framework), the data model in which information is expressed as subject-predicate-object triples — the standard form for linked data. Only in this format do the previously established links (entities to Wikidata, items to DOIs) become machine-usable and linkable to other datasets worldwide.
Publication at DataLegend
The resulting dataset was published at Druid/DataLegend, the linked-data platform of Clariah and partners. This is a specialised infrastructure for hosting and making RDF datasets searchable, which aligns with the LOD-oriented setup of the Tinbergen Letters project.
Registration with the NDE Datasets Register
In addition, the dataset was submitted to the Datasets Register of the Network Digital Heritage (NDE). This register makes datasets findable and reusable within the broader Dutch heritage sector and ties in with the manifesto that will be signed in Phase 6.
Public sharing via GitHub
Finally, both the RDF code and the AI prompts used were made available via the project's GitHub account. The latter is particularly valuable: it not only makes the technical conversion transparent but also reveals how AI was used in the process (e.g., for metadata extraction or transcription), which ties into your broader interest in responsible AI integration in cataloguing work.
Phase 5. Implementation and data in Omeka S
Phase 5 focuses on the presentation layer of the project: while Phase 4 makes the data linked and machine-readable, Omeka S ensures that the collection also becomes accessible and meaningful to human visitors.
Upload to Omeka S
The team is also uploading the data to Omeka S, a content management system aimed at heritage collections. This happens alongside the publication at DataLegend — Omeka S thus fulfils a different role from the RDF publication: it is the place where visitors can actually browse and read the letters and manuscripts, with the enriched metadata as the underlying structure.
Why Omeka S
The presentation layer of the Tinbergen Letters project chose Omeka S for a number of reasons that are not available in ContentDM:
- Native support for linked data. Omeka S works natively with RDF and controlled vocabularies such as Dublin Core and allows fields to be directly linked to URIs.
- Open source and embraced by the heritage sector. As an open-source system with an active community within libraries, archives, and museums, it also offers a universe of extensions (plugins, modules) designed by developers working with GLAM collections.
Re-transcription with Google Studio AI
The letters were re-transcribed using Google Studio AI and subsequently included in Omeka S. This is a separate AI application within the project — not aimed at entity recognition or metadata extraction, but specifically at making the historical correspondence readable. This makes the letters full-text searchable.
Page with biography
The Omeka S includes a page with a biography of Professor Jan Tinbergen. This gives visitors unfamiliar with his work immediate context for the collection and links the individual letters to the broader story of his life and thinking.
Phase 6. Signing the NDE Manifest and communicating with stakeholders
The NDE Manifest
The Network Digital Heritage (NDE) has drawn up a manifesto through which heritage institutions commit to shared principles for digital heritage — such as guidelines for making collections available as Linked Open Data, using shared standards, and contributing to a national infrastructure (including the Datasets Register to which the dataset was already submitted in Phase 4). By signing this manifesto, the institution formally confirms that it endorses these principles and actively wishes to contribute to the network.
Communication to stakeholders
In addition to the signing itself, this involves communication with the relevant stakeholders. This can serve several purposes:
- Internal visibility: informing colleagues and management of this commitment, aligning with your broader strategic goals.
- External visibility: informing collaborative partners.
- Accountability to the funder: the PICA Foundation, which funded the project through the Connected Digital Heritage programme, has an interest in knowing that the results are also anchored at the national level.
Phase 7. Evaluation, management, and impact
Evaluation
In the evaluation, we want to review what went well and what could be improved. We received advice from the steering committee at various points, but we will also ask colleagues for feedback. To that end, the project and its results will be presented during a lunch meeting in the library.
In this phase, we will also consider how we will monitor data quality and safeguard sustainability through version control.
Impact
Ultimately, the goal is to increase visibility. We want to encourage the use of the dataset for research and education. The most visible way to do this is by exploring the collection through a data story based on SPARQL queries. At the local level, the digital presentation of Professor Tinbergen's letters offers a valuable opportunity, especially in the context of the renovation of the building on the Erasmus University campus that bears the same name.