Open Proteins

It's not just graphs and tables: Proteins can be data too. (© Corona Borealis Studio/Shutterstock)
Melissa, you manage the Literature Team at the European Bioinformatics Institute (EBI), which is part of the European Molecular Biology Laboratory (EMBL). Could you tell us about your work?
My team’s main focus is running Europe PMC, a comprehensive, open-access database that offers global access to scientific literature. We also text and data mine the content to extract relevant information, such as data accession numbers. EBI hosts over 40 specialized databases for molecular biology, which include deposition and curated databases for biological research data.
How are they related to the principles of research data management or research software management?
For a long time, researchers have focused primarily on publishing research articles to share their research, hence the need for databases such as Europe PMC, a literature database. However, these articles don’t represent the full output. The actual data associated with these articles must be accessible for others to examine or reuse, and specialized databases provide a home for this data. Linking the outputs deposited in these data repositories to the literature that cite them is crucial. When someone reads an article and sees that data has been produced, they need to find it, and vice versa when looking at data in repositories they want to see the literature that contextualizes it. We are in the early stages of identifying and linking available software in the same way, but oftentimes the software is required to run the data, so is just as important for reproducibility.
Can you give an example?
If a large dataset has been processed using specific code, access to that code is essential to replicate the results, the data and the software are crucial. Releasing software code on platforms like GitHub is a great start, but proper archiving is necessary to ensure long-term preservation and access. We are collaborating with Software Heritage to link archived GitHub repositories using persistent identifiers to the literature. A critical issue is the licensing of software and data, if it’s not openly licensed it cannot be reused.
“When someone reads an article and sees that data has been produced, they need to find it, and vice versa, when looking at data in repositories they want to see the literature that contextualizes it.”
Melissa Harrison

Profile photo: Melissa Harrison (© privat)
Melissa Harrison
leads the Literature Services team at European Bioinformatics Institute (EBI), which is part of the European Molecular Biology Laboratory (EMBL), and manages Europe PMC. With a background in biology and scientific publishing, she focuses on advancing open science by improving access to research articles, data, and software.
Which services from the EBI do you consider well-established in terms of supporting research data and research software management?
A prime example of a well-established service is the European Protein Data Bank (PDBe). Researchers submitting their protein structures to this database receive an accession number, making the data openly accessible with a persistent identifier. Citing these accession numbers in research articles creates a direct link between the literature and the data. This dual access allows users to either search the database directly or discover related literature and delve deeper from there. Persistent identifiers, which align with the FAIR principles, ensure this content remains accessible and interconnected, enhancing discoverability and usability. A significant collaboration that highlights the importance of open data is the partnership between Google DeepMind and EMBL-EBI, resulting in the AlphaFold database. Launched in July 2021, AlphaFold offers open access to over 20 million protein structure predictions and this would not have been possible without the underlying protein structures of PDBe.
How do you enhance data accessibility?
Reflecting on my previous career in publishing, I spent considerable time promoting data accessibility statements and proper citation of software and data in reference lists. When I transitioned to EMBL-EBI and joined Europe PMC, I realized that, despite the dedication of many publishers, proper citation practices were often neglected. At Europe PMC, we perform text and data mining of accession patterns for approximately 45 databases. This means we can recognize database accessions when cited in the text of research articles, create the link between data and research articles and make the data more accessible to scientists. On our platform, users can access the data behind an article through this linkage. It's not straightforward to determine if the data was generated from the research or reused by the research of the article, but we are striving to make this distinction possible through new projects.
How can users access this data?
They can access it by clicking on the data link associated with the article. All the information is also available via our API, enabling those conducting meta-analyses to track how accession number citations grow over time and across different databases. This also benefits the databases by providing them with usage statistics. The literature is the most recognized output of scientific research, and our goal is to ensure that data is given prominence, making it easier to find when associated with research articles. An important aspect is the use of persistent identifiers within life sciences databases, specifically accession numbers. While DOIs for research articles are widely recognized, accession numbers are equally critical in our ecosystem. EMBL-EBI provides a service called identifiers.org, which resolves these accession numbers. This service functions similarly to the DOI infrastructure but is specific to our field and is free of charge.
How do researchers benefit from these services?
The great thing about EMBL-EBI databases is that there's no fee to submit or access data, making it accessible to anyone worldwide. Even if research articles aren't open access, all the data stored in our databases is free and available globally. This means anyone can access or submit data from anywhere in the world. In research data management, we face many challenges, and it's crucial to work together and support each other. Improving research data management requires understanding the barriers and finding solutions. One significant issue is making researchers understand the importance of submitting data and how it can be reused to the benefit of everyone. Metadata is a big challenge. Without detailed metadata, reusing data becomes much more challenging.
What changes would you make if you had the opportunity to reshape your early research?
Reflecting on my early years, I would have engaged with the open science movement earlier—open access of journal articles, open data and open software. That's the main takeaway. The journey of open science so far has been long, with many twists and turns, and we're still progressing; we’re not there yet. In the past, there were concerns about being scooped, that releasing your data sets early might allow others to build on your work and gain credit. However, making things open and attributing them to yourself is crucial. When others reuse your data, you receive credit for your contributions, and that's key. Open access to research and literature has been around for a while now, but there's now significant momentum to make data and software openly available, not lost as hard copy in a locked drawer or on a disc somewhere. This accessibility truly benefits the research community.
Frage 1
What do you see as the main barrier to data sharing?
"releasing datasets early might allow others to build on your work and gain credit"
“if it’s not openly licensed it cannot be reused”
She emphasizes that many researchers still focus on publishing papers and do not fully understand the importance and benefits of submitting and reusing data.
She highlights problems with proper software archiving, long-term preservation, and especially insufficient metadata, which makes reuse difficult.
Auswertung
richtig beantwortet!
Frage 1
What is the most important step to improve open science?
She notes that open access to journal articles was an early and important part of the open science movement.
She explicitly argues the opposite:
Articles do not represent the full research output
Data and software must be accessible and linked to literature for reuse and reproducibility
(“the actual data associated with these articles must be accessible… software is just as important for reproducibility”)
She stresses that all data in EMBL-EBI databases is freely accessible worldwide and that open data is essential for the research community.
She explains that simply releasing code on GitHub is not enough and that proper archiving with persistent identifiers is needed for long-term reuse and reproducibility.
Auswertung
richtig beantwortet!
About the project
The Europe PMC project, managed by EMBLEBI, is a comprehensive open-access database dedicated to life sciences and biomedical research. It connects scientific articles with associated datasets, improving data accessibility and supporting research reproducibility.