“It often comes down to:
Just trust us”
What is the current situation in computer science regarding research data and software management? Is there still a lot that needs to be done?
Yes, definitely—there is still a lot to be done. Data and software are rarely truly referenceable. There is often a lack of persistent identifiers such as DOIs for data sets or software versions. Even as a reviewer, it is often time-consuming and frustrating to get software up and running in the first place. In addition, there are hardly any central, subject-specific repositories for computer science. Although the discipline is very broad, ranging from technical to political to economic computer science, there is currently no standard location where data, metadata, and software can be systematically uploaded, stored, and made accessible. All in all, there is a lack of infrastructure, standards, and processes for the sustainable management of research data and software in computer science.
Why is that? Are we running a little behind?
I wouldn't see it as pessimistic. Rather, the requirements for scientific work have changed significantly over the past 20 years. In the past, it was quite common to submit papers in which software was only shown as a mock-up in PowerPoint – sometimes without a working implementation. Today, that would be unthinkable. Expectations of research have risen; the software must run, data must be accessible and traceable, and results must be reproducible. At the same time, data and software management is becoming more complex—for example, when it comes to data protection, data quality, or technical dependencies. And although there is a growing need for standards, computer science has traditionally been strongly shaped by individual solutions. This does not really make it easier to establish uniform infrastructures that are accepted by the entire community. An important step in the right direction is that the German Informatics Society is on board with NFDIxCS. This creates a solution that is developed with the community—and not without it. This is crucial for creating acceptance and establishing functioning research data and software management in computer science in the long term.

Agnes Koschmider is committed to systematically versioning complex research data. (© Jürgen Haacks, Uni Kiel)
Agnes Koschmider
ist Professorin für Wirtschaftsinformatik an der Universität Bayreuth und verantwortlich bei NFDIxCS für Version Management. Sie entwickelt Richtlinien zur Versionierung komplexer Forschungsdaten, überträgt bestehende Versionsinformationen wie aus Git auf unterschiedliche Datentypen und berücksichtigt dabei auch Varianten für verschiedene Zwecke.
"Successful systems are often those that are simple – but it is not always the best solutions that prevail."
Agnes Koschmider
What are the biggest challenges in versioning research data—and how does this differ from software versioning with Git, for example?
Git is great for code. But when it comes to research data, we quickly reach its limits. In research, we deal with a wide variety of artifacts: raw data, processed data sets, software, user studies, protocols—and they all belong together. Git cannot map these connections. There is a lack of metadata, referability, and a structured approach to dealing with different degrees of data maturity. With software versions, it is clear what changes. With data, it is more complex: it evolves, often goes through several processing steps—we are talking about a real life cycle. And especially in larger projects with multiple participants, it is important to be able to track exactly who changed what and when. Another issue is acceptance in the community. Just because a solution exists does not mean that it will be used. Successful systems are often those that are simple—but it is not always the best solutions that prevail.
Why is it so important to collect different versions of data for different research purposes?
In research, different roles have very different requirements for data. This starts within a research group: one person develops software, later leaves the project, and the next person tries to build on it. Then a specific version or proof that something worked in a specific environment may be missing. Version control helps enormously here: you can see who changed what, whether these changes made sense, and, if necessary, you can also go back. This is not only important for your own overview, but especially in collaborative projects where many people are working on the same data or tools at the same time. It must be clear which version is the valid one that everyone is referring to. Versioning therefore ensures transparency, traceability, and scientific quality.
What does it mean for scientific collaboration if a standardized solution for versioning research data can be developed?
Above all, this brings credibility. From the perspective of experts, it is often difficult when, for example, software is released but the associated data is missing—especially when this data is not openly accessible for data protection reasons, as is the case with company data. Then it often comes down to : “Just trust us.” And that is, of course, problematic. At NFDIxCS, we also pay close attention to data protection and how data can be anonymized so that it can still be used and verified. This makes science as a whole more transparent and credible. When all research artifacts are published in a uniform and traceable manner, reproducibility increases enormously. Experiments can be reproduced, software can be made to work, and results can be verified. This is important not only for science, but also for transfer to industry, because it means that solutions are reliably and robustly tested.
How could this influence the future of research work and publication in concrete terms?
I see some really promising opportunities here. For example, conferences and meetings could introduce additional tracks—not just for traditional papers, but also for data papers. These are contributions that focus solely on creating, preparing, and providing data sets. This often involves a tremendous amount of work, sometimes even more than writing a paper itself, but it is currently hardly recognized. Establishing this as a separate category would give researchers more recognition for their data work—and that would naturally motivate them to prepare and share data carefully. It would also create more transparency and better traceability. NFDIxCS could lay an important foundation here for anchoring such new concepts in science. The aim is to enable reproducibility, traceability, and reusability in a truly systematic way, thereby making research more sustainable overall.
What is the focus of your discipline the business informatics?
The data is very diverse: we are talking about survey results, online questionnaires, eye-tracking data, process models, text files, and software. In business informatics in particular, we urgently need a structured infrastructure that enables access to this data. When more researchers can work together on the same data sets, not only does the quality of research increase, but errors are also detected more quickly. With NFDIxCS, we want to achieve exactly that: make the community's artifacts visible and usable, create new paths, and thus take a big step toward open science and FAIR data. This is very important to me personally in order to further advance business informatics.
What are your current priorities at NFDIxCS and how do you see the future of the project?
We deal specifically with semantic aspects—for example, knowledge graphs. We try to link research artifacts such as software, data sets, questionnaires, results, and research questions in such a way that they can be searched more easily and stored for the long term. This not only helps reviewers to check, for example, whether any relevant preliminary work has been overlooked, but also helps researchers to more easily understand what others have been working on. My doctoral supervisor always said: Just because two people are working on the same topic doesn't mean they will come up with the same solution.
Is there anything else you would like to share with our readers?
The success of NFDIxCS clearly depends on the community. The project will only be truly successful if many people participate and see the benefits. It's a classic chicken-and-egg problem: the community often says that it needs to recognize the added value before it will use the system—and that's exactly what we're working hard on. We are very glad that the German Informatics Society is on board. It ensures that the project is visible and disseminated throughout the entire community. This is an important step toward achieving something great together.