The Broader Picture

Ulrike Lucke has a comprehensive view of research data and software, guiding the community toward transparency and sustainability. (© Gesellschaft für Informatik)
Ulrike Lucke
is the co-spokesperson of the NFDIxCS project and Professor for Complex Multimedia Application Architectures at the University of Potsdam. With a comprehensive overview of the computer science research landscape, she advocates for sustainable and transparent management of research data and software.
Ms. Lucke, the NFDIxCS project was officially launched just five years ago. What changes have you noticed since then?
The first projects have been running for quite some time now. But of course, the community has been active for much longer. I am thinking, for example, of the Council for Information Infrastructures, which initiated important changes toward a more systematic approach to research data at an early stage. Some funding bodies also introduced policies on the handling of research data years ago. Many have recognized that data does not always have to be collected again and again. This does not make sense scientifically or financially. In this respect, the framework has been in place for a long time. Now that the NFDI is up and running, its operationalization is slowly becoming more visible, not only in terms of the political framework conditions, but also in practice.
Do you have any exaples?
We are seeing changes in research and teaching. What makes me very happy is that aspects of research data management and research software development are increasingly appearing in curricula. For me personally, the discussions going on behind the scenes are even more exciting.
What are these discussions about?
Often, it's about transparency—in other words, the question of why we're actually doing this. But the difference between replicability and reproducibility is also much discussed: Can a result really be reproduced? If I apply a certain method a second time to a specific data set, will I end up with the same result as the researchers before me? This is a seal of quality that did not exist in research when ideas could only be traced on paper. And, of course, this also leads to different methodological approaches.
In what way?
Depending on whether I am considering replicability or reproducibility, I have to prepare data and software differently, i.e., make them usable in a different environment. This is where exciting differences between the disciplines emerge. Theoretical sciences such as mathematics, where reproducibility has always been inherent in their own logic, have so far simply written down the proof. This then appeared clearly comprehensible. Nowadays, machine theorem provers are used for this purpose. In contrast, in subjects such as psychology, where large data sets are used, even a small change in the code can lead to completely different results. In this case, the software must also be taken into account. If we look at computer science, for example, the entire execution environment is also important for performance measurements. The archiving requirements are therefore different. Ultimately, such discussions also sharpen the self-image of the respective discipline: Why are we doing this research? And what are society's expectations of us in terms of traceability?
That sounds like a lot to talk about! Are there any good arguments against transparency?
Absolutely. This concerns issues such as intellectual property or the fear of misuse of research data—for example, if it is commercialized too early without ethical guidelines. And then there is the financial factor. At universities, we have the luxury of being largely funded by taxpayers' money. So we don't have to generate income, but we do have to justify it. We have to build credibility in society for the work that this money goes towards.
“We generate credibility in society by disclosing how we arrived at our findings.”
Ulrike Lucke
Speaking of credibility: What role do the latest developments in artificial intelligence play in this context?
We generate credibility in society by disclosing how we arrived at our findings. With AI, often referred to as a black box, this is not always easy, even for computer science experts. In addition, there is a very interesting interplay between technical limitations, which can be extended to a certain extent, for example through so-called explainable AI. But on the other hand, these are simply power structures in which people—not machines or algorithms—hold positions of power and transparency would damage business. In my opinion, this is a challenge that can only be addressed through regulation. This requires our technical expertise as a computer science community and clear statements about why we consider this to be harmful to society.
Which brings us to the much-discussed dual role that computer science plays in the debate surrounding open science. Does having technical expertise make the effort easier?
On the one hand, it makes things easier, but on the other hand, it also makes them more difficult. It is often said that the cobbler's children have no shoes, and rightly so. I believe this also applies to computer science and the way we design our own IT systems and processes. One thing is certain: for us, software is not just a tool that we use or build, but also the subject of our research. Our research data also includes the software and its execution environment. I observe a pronounced tendency toward in-house development in our discipline: people prefer to build things themselves rather than build on something that already exists.
Why ist that a Problem?
During the design process, many people often only think about short-term use without asking themselves what would need to be changed in the design to ensure that software is still operable in 3, 5, 10, or 20 years—as would be in line with the principle of open science. We need a cultural shift in our discipline, one that allows us to consider long-term operation and runnability in the architecture from the outset. When other disciplines tinker around, it may be excusable. But in computer science, the sustainability of software should be a matter of course.
How can this cultural change be initiated?
It is a large and heterogeneous community that we need to reach. We do this primarily through the structures of the German Informatics Society. We contact the heads of all departments to understand the internal logic of these subcommunities. This helps us to address them in a targeted manner. We also always involve different hierarchical levels: depending on where they are in their careers, people have different approaches to the topic of openness, different amounts of leverage to make a difference, and, of course, different amounts of resources they can contribute. We address this with different formats: from short evening WebTalks to participatory workshops or hackathons to summer schools. We empower the community to take action themselves. And I also notice that we no longer need to be overly evangelical in our approach.
How do you determine that?
I see a great deal of openness to this topic all around me. The key is to provide people with methods and tools that enable them to integrate open science into existing routines at a reasonable cost.
And then, of course, you need people with the right expertise.
I don't think we can afford to have people leaving university in any subject without having learned how to handle research data. Research software is another matter, but in my view, research data is generated everywhere. So it's essential to teach prospective researchers the skills they need to handle research data.
How open science culture is changing in computer science
Step 1 — The groundwork for research data management was laid years ago
Before NFDI even launched:
- policy frameworks already existed
- funding bodies introduced data management rules
- communities realized re-collecting data makes no sense
NFDI made this operational and visible in practice.
Step 2 — Transparency became a core quality discussion
Behind the scenes, debates now focus on:
- why openness matters
- replicability vs. reproducibility
- how differently disciplines must prepare data & software
traceability becomes a scientific quality seal.
Step 3 — Software and execution environments became part of research data
Especially in computer science:
software is not just a tool, but research output
small code changes can completely alter results
environments matter for performance and reproducibility
archiving requirements differ by discipline.
Step 4 — A cultural shift toward long-term sustainability is needed
Current problem:
- researchers design software for short-term use
- long-term operability is rarely planned
What’s needed:
- sustainability built into architecture from the start
- open science thinking embedded in workflows
Step 5 — The community itself must drive change
Change happens through:
- targeted outreach to subcommunities
- workshops, hackathons, summer schools
- empowering different career levels
not preaching, but practical tools & routines.
Step 6 — Best-case future for open science
Preserving data together with software:
- strengthens reproducibility
- improves publication quality
Curricula are already very full, especially in computer science. How do you deal with that?
Yes, it's definitely not easy. On the one hand, diversification can help here, i.e., specific training that leans more toward data stewardship or similar areas. The other solution could be to integrate such content into existing educational programs: for example, where students are required to program anyway, such as in exercises or internships. Instead of starting from scratch, existing data and software could be used to directly convey the added value of open science.
Finally, let's take a look into the future, ten years after this conversation: What does your best-case scenario look like?
My idea would be for important conferences and journals to have an artifact evaluation track. This would mean that, in addition to articles, software and data would also be submitted, reviewed, and published as a matter of course. To achieve this, all project teams would need to work with data and software management plans from the outset as a matter of course—and include the necessary resources for this in their budgets. In addition, we as a scientific community should also make this issue a decisive factor, for example in selection procedures for professorships. This is not a nice-to-have, but a must-have.