Community is key

Interview Elena Müller

When you think of research data, you might first picture endless files stored somewhere in the cloud. But Michael Goedicke, Professor of Software Engineering at the University of Duisburg-Essen and spokesperson of NFDIxCS, sees more than just storage. In this interview, he shares how sustainable infrastructures and community-driven standards can turn data into a reliable foundation for future science.

Quote character

“My advice to myself, and to others, would be: safeguard your data, preserve it together with related research software, and make it accessible.”
Michael Goedicke

What services are you building for researchers?

One of our core elements are Research Data Management Containers, which package different types of data together with their context. There are two important elements in this context which are meta data and the related research software including an execution environment to allow the access to the data via the research software even years after the publication of the data. Our portal enables researchers to create, store, retrieve, and share these containers while managing access rights. We also provide a helpdesk, consulting, training, and workshops. The idea is to combine technical services with community support.

Why is community involvement so important?

Infrastructure alone is not enough. We need to agree on shared standards and quality criteria, otherwise research data cannot be reused. Our goal is to involve researchers directly in shaping these standards – only then will they be accepted. Just as important is working closely with the community to ensure that what we build is genuinely needed.

Building sustainable research data infrastructures

Step 1 — Data + software belong together

Research data should not be stored alone, but bundled with:

  • metadata

  • research software

  • execution environment

 in Research Data Management Containers that can still be used years later.

Step 2 — Technical services must be combined with human support

Alongside the portal for creating and sharing containers, they provide:

  • helpdesk
  • consulting
  • training & workshops

 

 

Step 3 — Community-driven standards are essential

Infrastructure alone is not enough.

Researchers must:

  • jointly define standards

  • agree on quality criteria

otherwise data cannot be reused or accepted.

Step 4 — Sustainability has three major challenges

  1. Technical — building stable long-term infrastructure

  2. Community — developing standards together

  3. Political/financial — securing long-term funding

Step 5 — Long-term preservation improves science itself

Preserving data together with software:

  • strengthens reproducibility

  • improves publication quality

What are the biggest challenges – and how do you ensure long-term sustainability?

We are creating an overarching architecture that ensures all services fit together and can be operated in the long term. But a blueprint alone is not enough – the system must be built, maintained, and kept reliable over many years. As a result, we are facing three main challenges: First, the technical one – building a stable infrastructure. Second, the community challenge – developing standards and quality criteria together with researchers. And third, the political and financial challenge – ensuring that core systems can be operated sustainably. For long-term sustainability, we need operational models, monitoring, and reliable baseline funding, much like university libraries. Basic services should remain free, while additional funding models will be required for specific needs.

Looking back, what advice would you give your younger self about research data?

In the 1980s, I generated data that took great effort to produce – and then discarded them. Today, I would do it differently. Preserving data greatly improves reproducibility and strengthens the quality of publications. My advice is clear: safeguard your data, preserve it together with related research software, and make it accessible.

 

Michael Goedicke

is a Professor of Software Engineering at the University of Duisburg-Essen and has been working there since 1990. After earning his PhD and habilitation in Dortmund, he spent time at Imperial College London, focusing on view based software engineering and modularity in software specification. This also led to the development of an automated system for checking assignments and exams across various scientific disciplines. He joined and helped to organize the forces to build the Research Data Management Infrastructure for Computer Science since 2018.

Since March 2023, he has been serving as the spokesperson of NFDIxCS, coordinating the scientific work and overseeing administration, finances, and communication with the German Research Foundation (DFG).

About the project

The main goal of the consortium NFDIxCS is to identify, define and finally deploy services to store complex domain specific data objects from the specific variety of sub-domains from Computer Science (CS) and to realize the FAIR principles across the board.

This includes to produce re-usable data objects specific to the various types of CS data which contain not only this data along with the related metadata, but also the corresponding software, context and execution information in a standardized way. These data objects can be of any size, structure and quality. 

Follow-up