Behind the scenes
of educational data

Jan Bernoth, Firas Al Laban, Ulrike Lucke

Daniel Schiffner is leading the educational computer science unit at DIPF, Leibniz Institute for Educational Research and Information. In the interview, he discusses the efforts in research data and software management, the importance of FAIR principles, and the challenges and strategies for effective data lifecycle management. He emphasizes the need for collaboration, continuous improvement, and the adoption of new technologies to support the research community.

How is your work related to research data and software data management?

At our institute, we strive to make data accessible and reusable, working at the intersection of computer science and data science. This effort extends to the educational technology community, highlighting the importance of returning data to the community and emphasizing its relevance. This comprehensive approach is a significant aspect of my work, supported by my colleague Natalie Kiesler.

What are you specifically focusing on in your work?

The data lifecycle can be crucial. Particularly in open science. Our focus is on data acquisition, setting up and populating a research data center, and simplifying this process for data providers. The subsequent steps include internal lifecycle processes like data processing, metadata curation and ensuring reusability. A significant challenge is representing information accurately, considering that we often cannot consult the original data providers due to their evolving careers. This emphasizes the importance of preserving knowledge and making data, software, and publications accessible. A holistic approach is necessary, possibly including unpublished notes, to ensure comprehensive data management and usability.

A research data center at your institute is part of a network where multiple research data centers collaborate. What is the advantage of such a network?

Each center specializes in different types of educational data like qualitative, educational trajectories and quantitative. But researchers can just reach out to the network, which then directs them to the appropriate center. This simplifies the process and addresses data heterogeneity. Another key service is the technical process for data registration and curation, ensuring high data quality. This involves transferring data to an archive with proper legal frameworks in place, beyond just uploading data to platforms like GitHub. Our research data center also handles personal data, ensuring anonymization where possible and applying usage restrictions as needed. Additionally, we are part of the KonsortSWD project, defining APIs for data exchange at the metadata level and determining suitable metadata schemas.

How do you ensure efficient data use and collaboration in large studies?

A relatively large study that many people know about is the PISA study. We also collected data for it at the institute. And a very big problem that I have spoken to many researchers about is that, well, we just have the data. Many researchers access the data, understand it through the codebook, but then need to write their own scripts for cleanup. To address this, once a script is written, it should be shared to prevent others from repeating the same work. This enhances the data lifecycle and promotes efficiency. When publishing data or findings, it’s essential to include all relevant steps taken, ensuring transparency and enabling deeper discussions beyond the typical length of research papers. Best practices in research and data management, including proper documentation and metadata, are key to effective collaboration and reuse of data.

Quote character

“When publishing data or findings, it’s essential to include all relevant steps taken, ensuring transparency and enabling deeper discussions beyond the typical length of research papers.”
Daniel Schiffner

Daniel Schiffner

is head of educational computer science at the Leibniz Institute for Educational Research and Information (DIPF). He holds a PhD in computer science and is a visiting professor at the PH Weingarten. His research and development focus is on information systems, visualizations, and data collection.

What tools and strategies have you found to be particularly helpful in this area?

For research data management itself, the Research Data Management Organizer (RDMO) is a relatively big topic. It is a significant tool that uses standardized questionnaires to guide researchers through important questions related to their projects and data sources. This helps ensure compliance with requirements and thorough documentation. High usability is crucial, overwhelming researchers with numerous metadata fields is not effective. Instead, simplifying processes, such as using ORCID for researcher identification, enhances efficiency. Best practices change over time, reflecting advances in data collection methods. Exploring new technologies like GenAI for data collection could improve efficiency, demonstrating the importance of open standards in fostering long-term opportunities. DIPF's services are designed to support researchers not just within the institute but broadly, emphasizing usability and quality in research data management.

How do you make sure your work benefits the wider research community?

The institute's main task is to manage research data, promote open science, and provide educational resources and do research in the domain. To have an impact beyond the institute, active engagement with the research community is essential. This involves interacting with users, staying updated, and participating in community activities to understand and address their needs. Developing new tools and solutions, such as interfaces for qualitative research data, can be beneficial if there's demand. Improving internal processes to make them more efficient and user-friendly can set standards that influence broader practices. It might also be helpful to create incentive systems for researchers, like badges or certificates for best practices. By fostering a culture that values data sharing and proper citation, and continuously engaging with the community, we can advance knowledge and improve research practices. There is a lot to do, and working together is crucial to improving research data management services.

What are your ideas for improving research data management services?

Improving research data management services is challenging because it's a continuous cycle requiring constant adjustments. Both service providers and users need to work together, sharing what information is needed and what can be provided.

What are the challenges for such collaboration in your field?

One major barrier is related to software development in educational technology. Researchers often develop software to collect and analyze data, but sharing this software can be challenging due to access restrictions and the imperfect nature of the code. The aim isn't to create perfect software but to demonstrate that an idea works. We're not selling a product but proving effectiveness and solving specific problems. However, publication formats often limit this, making it hard for others to replicate or build on our work. To address this, we need to make our software public, encouraging open exchange and discussion about the data and methods used. This can help create new synergies and advance knowledge collectively.

If you could start over your academic career, what would you do differently?

If I were to study again, I would focus on understanding the fundamental principles of computer science rather than just learning programming languages. It's essential to grasp how technology processes information. Programming languages like C++, Python or JavaScript are tools for machine communication, but the real goal is effective communication with others through good documentation and clean code. Collaboration, understanding core concepts, and developing effective communication skills are vital. It's about engaging in discussions, learning from diverse perspectives, and generating new insights independently, not relying on shortcuts like AI or copying others' work. These are the values I would prioritize in my studies to achieve a deeper and more meaningful education.

Multiple Choice
/

Frage 1

What is the biggest challenge in educational data management?

This is clearly false because he emphasizes quality assurance and curation as key services and ongoing work.

He repeatedly stresses metadata curation, documentation, and preserving knowledge (“metadata curation and ensuring reusability”; also “proper documentation and metadata” are “key”).

He explicitly says their data center handles personal data, anonymization, and usage restrictions, and needs proper legal frameworks.

He calls out the problem that researchers repeatedly write cleanup scripts and should share them; he also mentions barriers to sharing software (imperfect code, access restrictions) and that openness helps replication.

He proposes incentive systems like “badges or certificates” to encourage best practices and sharing—framing incentives as something that could improve the system.

Multiple Choice
/

Auswertung

Sie haben von Fragen
richtig beantwortet!

About the project 

The DIPF maintains two research data infrastructures, one being the research data center education and the other being the German Network of Educational Research Data (Verbund Forschungsdaten Bildung). In both, the collection, archiving and dissemination of research data is the main task. Other services by the DIPF are tools to support teaching and education, providing access to research papers or generating overviews and reviews in the context of the German educational system.

Multiple Choice
/

Frage 1

What would motivate you most to share your own research data?

He explicitly suggests incentive systems like “badges or certificates” to encourage best practices and sharing.

This is false because he says the opposite: overwhelming researchers with many metadata fields is not effective, and simplifying processes is important.

He repeatedly emphasizes returning data to the community, collaboration, reuse, and shared scripts to avoid duplicated work (“working together is crucial”; share scripts; benefit wider community).

He explicitly mentions compliance/requirements and tools that “ensure compliance with requirements” (RDMO guiding researchers through requirements).

Multiple Choice
/

Auswertung

Sie haben von Fragen
richtig beantwortet!