Questions of life,
the Universe and everything

Jan Bernoth, Firas Al Laban, Ulrike Lucke

You might think that you study very dense sets of data, but Tim Dietrich certainly works on more dense objects. He specializes in black holes and neutron stars. In his research he runs complex simulations which lead to potential answers to the big questions of our Universe – and to endless amounts of data. Since not everyone has a supercomputer to run it, Tim and his colleagues came up with a database that presents simulation in a compromised form, in megabytes, not terabytes.

Tim, your research focus is on black holes. How exactly can you conduct such studies?

We ask ourselves questions such as what happens when black holes or neutron stars collide and try to draw conclusions about our Universe and matter at high densities. In doing so, we try to decipher how the cosmos is structured. As you can imagine, this is mostly work done in the theoretical area. That means we are developing models and looking for ways to simulate the whole thing on high-performance computers. Recently we have also been evaluating data, but mainly our goal is to make predictions based on our theories.

What role does data play in your work?

We carry out the simulation ourselves and this generates very, very large amounts of data—several terabytes in a short amount of time. We save at least some of it, but we'd actually like to save several hundred terabytes a year. That's a problem, of course. Recently there have also been experiments that actually measure what we simulate, such as gravitational waves or electromagnetic radiation.

Two very different sets of data, I assume.

Absolutely, and it’s very important to differentiate between simulation data on the one hand, which we create ourselves, save and make available as far as possible, and the observation data on the other hand. The latter is data that is collected as part of collaborations with other institutes. We interpret it and also provide tools and manuals on how this interpretation can be carried out as quickly and accurately as possible – so that, when you observe something, you know what exactly you have seen. And we develop the models for that.

Quote character

“One argument against the amount of time it takes is the visibility you get from others working with your research and knowing you as the person or team to talk to.”
Tim Dietrich

You have already voted for this poll.

How much effort should researchers invest in preparing and curating data for others (vs. publishing results quickly)?

1: Minimal effort

5: Significant effort

What does such data look like?

The simulation data is mainly arrays of numbers. Such simulations can take several weeks or even months on high-performance computers. We do several hundreds of such simulations and they involve petabytes of data. That’s also why we cannot output everything. For that reason, our solution is to store only a reduced amount of data: What is the total mass of the system? What does it look like along the x-plane? Unfortunately, we can only very, very rarely output 3D data, because it is simply too expensive, in terms of storage space alone. We simply don’t have the space for such outputs. So instead, we reduce the output data during the simulation and use it to find the answers we are looking for: What happens when two neutrons collide? How does the matter behave in a certain situation? What waves are emitted? Together with various colleagues, mainly from the University of Jena, we have created a database in which we provide a lot of reduced information on gravitational waves. We are not providing the entire 3D data, but simply the gravitational-wave signal from this simulation, with different parameters such as resolutions or grid discretization. We are working on this in a collaboration called Computational Relativity Collaboration.

What is the goal of this collaboration?

We started the database in 2017 to make gravitational-wave signals from our simulations available to other researchers. Today, we have uploaded about 600 simulation data sets. They come from simulations that took about half a billion of CPU hours in total. To rerun one of those simulations alone, you would need a very powerful computer in the basement. So that, our database provides only a fraction of the data, which amounts to only a few megabytes per simulation. As you can imagine, a lot of the work goes into extracting what we consider the important data. The database also includes metadata for every simulation: What are the masses, what was simulated and what was the code setting that was used? It also provides information on publications it refers to, so that people can read what the point of it was, why it was simulated in the first place. Already, many researchers use the database to simply take this data and do something different with it, to develop new models or to look at particular patterns, such as how the signal depends on the parameters that were simulated.

You have already voted for this poll.

How valuable is data reduction (e.g., compressing terabytes into megabytes) compared to storing full raw data?

1: Not valuable 5: Extremely valuable

Do you and your colleagues already have the database in mind when working on new simulations?

So far, we have done it in such a way that we have published all the simulation data, or almost all the simulation data, in two large blocks. Now we want to switch to a system where the data should be fed in again for each individual publication that comes out separately. Working it out properly and making the data usable in each case will mean that scientific work progresses a little more slowly. However, just this morning I had a conversation with a colleague at another institute who basically pointed to this database and said, yes, he had looked at something here and compared his simulations with it. He was happy that he could compare his results and wanted to get in touch when his data is ready so that we can take a look at it to make sure he did not overlook anything. I think that releasing the data can be a big benefit for a larger community.

Is the code archived too?

Yes, but only internally, it is not public. There are various reasons for this. Mainly, the code has grown over the course of 20 years or so and it stretches over one and a half million lines of code. To document it properly now, retrospectively, is a very large amount of work and whether it would be worth it is an ongoing discussion. Even the code being just available internally isn't trivial. There are always master's students, doctoral students or even professors who make a code change and implement something, but then forget to share this with the entire internal collaboration. Then, sometimes information and work get lost. That happens, it has happened and it will happen again and again.

 

Tim Dietrich

is an astrophysicist focused on black holes and neutron stars. His work aims to unravel fundamental questions about the universe and the behavior of cosmic matter under extreme conditions. He uses theoretical models and high-performance computer simulations to explore what happens when dense cosmic objects collide.

You have already voted for this poll.

How important is it to make massive astrophysical simulation data accessible to the wider research community?

1: Not important 5: Absolutely essential

About the project

Together with numerous collaborators around the world Tim Dietrich launched the “Computational Relativity Collaboration”, a database to make key gravitational wave simulation data accessible. By compressing large datasets, they enable researchers without supercomputers to study cosmic collisions more easily.

Back to the database: What do you know about its users?

It is used a lot by people who can't do the simulation themselves because they don't have a supercomputer—or the resources to develop the code needed for such operations. As I said, it took us 10 to 15 years to work on the code for our simulations. I had already taken it over from my supervisor as a doctoral student and so on. If you just want to have the simulation data to do physics with, then it's of course very convenient to simply have the respective data block and then work with it. That way, you can skip the mess of things that have to be done beforehand. However, if you skip it, you might not even recognize problems in the data because you think it has to look like that.

Do you see any other hurdles that are currently preventing a smooth process for research data?

I think that it is less now, but the time that has to be invested is still a huge factor. It naturally leads to people saying that they will do it later or not at all to keep up with other groups who publish faster because they neglect this extra step. One argument against the amount of time it takes is the visibility you get from others working with your research and knowing you as the person or team to talk to.

Last question: What advice would you give your younger self to be better prepared for such data-intensive challenges?

Simple! I would tell myself to assign names better and more sensibly. Test 1, Test 2, Test 3 to Test 72 is not a good naming scheme!
 

Quiz
/

Frage 1

What is the most important principle for planetary data archives?

Longevity and simplicity of formats
 

Quiz
/

Auswertung

Sie haben von Fragen
richtig beantwortet!