An alliance to rely on

Jan Bernoth, Firas Al Laban, Ulrike Lucke

Ramin Yahyapour has many titles. He is a professor at University of Göttingen and, simultaneously, his university’s Chief Information Officer. Moreover, he acts as managing director of GWDG, a joint data center and IT competence center by the University of Göttingen and the Max Planck Society. He shares insights on the work with complex international projects and finds a surprising analogy for how scientists approach research data management services.

Ramin, apart from all your other titles, you are also one of the founders of the E-Research Alliance in Göttingen. What was the idea behind it?

Here in Göttingen, we have been looking at the topic of research data management for many years. I think we were one of the pioneers in dealing with it structurally, in terms of how to support scientists on a large research campus. We found that we have different players on campus who can contribute to research data management and who already have services and skills. In Göttingen, this is, for example, the State university library, which is very active in various fields. Then, we have the GWDG as a large data center, offering many different services. Of course, we also have the administrative unit for research, as well as the institute for medical informatics. It was great to see that there are so many different groups on campus that are on the move. But for those who are looking for help and support, it can sometimes be unclear what is already there and who they should talk to. So, we founded the e-Research Alliance to create a structure that covers the topic of research data management as well as the topics surrounding it.

How does this alliance operate? 

It consists of people from the GWDG, from the library and from other groups. This serves as a “one-face to the customer” for questions on research infrastructures. As an individual, you can approach it with your questions and the team will coordinate the services you can use and the people who might be able to help you. We also established structurally that includes all large projects you might find on such a campus, like collaborative research centers, even before they are submitted.

Do you have an example for such a project? 

A project that I personally find very interesting is the collaborative research center 990. The project is about analyzing ecological as well as socio-economic effects in the tropical rainforest regions of Indonesia. And as you can imagine, a lot of data is being collected in this project. There are many international partners involved and a lot of activities take place on site. The scientists are out in the field, researching the impact of palm oil or rubber production – on the environment as well as on the local population and economy in these regions. Many different disciplines are involved. In Göttingen, researchers from the Agricultural, Biological and Geophysical Faculty are part of the project team. Additionally, there are colleagues from both the Faculty of Economics and the Faculty of Social Sciences.

Quote character

“I think one of the strengths of institutions that do a lot of research data management is to identify recurring basic services: things that don't always need to be rebuilt individually.”
Ramin Yahyapour

Ramin Yahyapour

is the managing director of the Göttingen Society for Scientific Data Processing (GWDG) and a professor at the University of Göttingen. His expertise lies in IT infrastructure and scientific data management. He focuses on promoting digital research opportunities and interdisciplinary collaboration.

What was the role of the GWDG in this project?

We actually covered the entire lifecycle of research data. At the beginning, we helped with questions such as: Which data will be collected in the context of the project? And what do you want to do with it? This led to a research data policy. That's what is desired and demanded these days, for a project to have a research data management plan, which you also authorize and try to implement within the framework of the project and the training that goes along with it. The participants of the e-Research Alliance are now taking care of training the team members. In the context of the project, data is collected on site, often geo-referenced data, data from very different sources such as photos or films, but also soil samples and other samples from nature, which need to be deposited and later assigned accordingly. In the course of the project, data is also collected from the same sites again and again to see changes.

Sounds like a very complex, but also interesting case for research data managers. 

Absolutely. What makes this project interesting in that respect is also that it takes place in the distance at a remote location. You can’t just go there yourself if there is a problem. So, we had representatives of the library or GWDG on site, who were there with a laptop, to see for themselves how the data collection takes place. It is a project in cooperation with partner institutions in Indonesia, where the great desire was that, within the framework of data sovereignty, such data is also available in Indonesia. That means that a so-called mirroring was set up: We collect the data first and make it available within our structures, and then we provide a copy of this data for Indonesia on site. We don't always have that in projects and I think that's a really nice component. What I also find fascinating are the very different types of data. The range spans from economic aspects over images to biodiversity data, which of course all have very different requirements. Some data is very large, some data is rather tabular, textual. On top of that, we also paid close attention to identity management, access control and long-term availability.

What does that entail?

You should never underestimate the importance of well-established identity management, particularly if you work with partners from outside your institution. A key element is to assign clear rights and duties to certain people and institutions within the project.

Time is another key factor in the field of research data. What challenges were you facing in that respect? 

The question of who needs access to the data at what point is a very interesting one. We often assume that we collect a certain sample and then work with it right away, but most of the time, the process is done in various stages. In our case, the group who collected the data worked on it first, then made it available as part of the project and six months later the data was published. This means that we had to work with technically established lag times and embargo times, which in the end lead to the data being secured, archived, but also publicly available. In terms of reusability, we are working with persistent identifiers.

What are persistent identifiers?

Usually, people integrate URL in their papers and think that will be enough. But reality has shown that many of those URL are no longer referenceable with increasing age. Servers may have changed, or they are just not reachable anymore. Persistent identifiers provide a solution for this problem. They are part of a global network system where you can go from, let's say, abstract persistent identifiers, which can solve the problem, to where the data really is, and also have metadata information.

What are your takeaways from working on projects like this?

I think one of the strengths of institutions that do a lot of research data management is to identify recurring basic services: things that don't always need to be rebuilt individually. I always compare that with the production of cars. Everyone wants to have an individual model, but of course it is only economical if things like the brakes or the engine are the same as in existing cars. My personal experience is that you can offer a lot of generic solutions, but in the end, there is a great desire and need of the experts to give the whole thing its own identity and some individualization. They all want to have their own portal with their own colors, their own logo, slight variations. The things that are in the portal may be 80–90% similar. But it’s important to understand that a portal framework or a web service needs to be able to cover this individuality, but at the same time you need to find economic ways to support these solutions in the long run. As such, you need economy of scale through common basic services that are maintained by large institutions like a library or computing center. I think that's an experience that you can take away from such projects.

What are the biggest hurdles you have encountered in this field?

Most of these projects run very well during the project period. But research data management becomes particularly relevant after the official end of a project. These are times that are much more difficult to cover, in my opinion, and we put a lot of thought into how we can ensure that results and data are available in a timely manner. A big challenge here is certainly personnel. Those who built or developed something may not be there in 10 or 15 years. How can we ensure that this software stack, this service stack, is kept alive and maintained? And it’s not just the teams that change. As a computer scientist, you tend to be surprised by everything older than seven years. It is relatively well understood how we go from one storage solution to the next, from one tape library to the other, from one file system provider to the next. We can do that relatively well. We can also get from the virtual machine to the container. But when it comes to how we go from software stack A to a software stack B at a later point in time, these migration projects are usually still the big challenges. At NFDIxCS, we are particularly concerned with the software lifecycle. In many projects, we need to ask ourselves: How can we make this software operable in 5, 10, 15, 20 years? Do we have to archive the lifecycle environment? Do we have to establish migration processes for software? Do we need people to look at the software? I think this is also a relatively big challenge – and a very interesting one!

About the project 

The eResearch Alliance is an initiative of the University of Göttingen that supports researchers in effectively integrating modern digital technologies and tools into their scientific work. The aim of the project is to provide a sustainable infrastructure for data management, computing resources and scientific software. The alliance offers best practice trainings and consultations on research data management and eResearch topics.