Lessons learned from educational tech
What role does research software play in your current field of research?
SVEN: Software plays a central role for us, especially in the development and testing of new scenarios. We often either write new programs or adapt existing software for this purpose. An important part of my work is supporting peer feedback processes. Both specially developed programs and standard software such as Excel are used, particularly in the analysis of homework, programming projects and tutor assessments.
MICHAEL: In my research, I differentiate between two main groups of data. The first concerns classic e-assessment research, which focuses on software artifacts. Here, systems or extensions are analyzed and compared, for example to determine which algorithm provides better feedback. The second group comprises data from studies, such as student submissions and the associated metadata. These are used to apply different methods to the same submissions and compare their results. I am also investigating other systems on a software level to analyze architectures and interactions between educational technology systems.
How exciting that you both pursue different approaches in your research. What is your current project about?
SVEN: Our project aims to create and maintain a detailed, versioned database of e-assessment systems. The goal is to offer a comprehensive and regularly updated overview to support researchers answering various research questions. Unlike individual paper reviews, our database provides systematic information on multiple systems, often composed of several papers. This allows for precise tracking of developments and supports reproducibility, enabling other researchers to build on and compare results from our data. A key advantage of our approach is the possibility of reproducibility: researchers can use our versioned data to conduct their own analyses and replicate results.
MICHAEL: Our goal was to create a dataset that exists independently of a specific research question. Conventional reviews often exclude systems that are not relevant to the research question, which limits the reusability of the data. To counteract this problem, we have developed a corpus that is structured according to formal criteria, independent of a specific research question. This corpus is regularly updated and expanded so that it remains a living, evolving artifact that continuously integrates new data sources.
Have other researchers already used your dataset for new research questions?
MICHAEL: Yes, our corpus has been used in various ways. For example, one paper identified a gap in our corpus, which led to a new study to investigate it. When our corpus is cited, its visibility increases, encouraging others to use it in their own research. This kind of scholarly word of mouth ensures our corpus continues to be utilized and disseminated. Publishing our corpus on the website has several advantages over static repositories. While static repositories are constrained by fixed data formats and lack interactivity, our website offers dynamic use with an interactive interface that allows easy filtering and searching, making specific information more accessible. For instance, I was able to quickly determine how many systems handle programming tasks in C or C++ without needing complex data analysis. This ease of use lowers the barrier to reusing the data. Additionally, the website can be continuously updated, promoting ongoing interaction and the development of new research questions, even without a DOI.
SVEN: Researchers have already contacted us directly to add their own systems to the corpus, which is a good sign of the relevance and usefulness of our work. Building the corpus was a big effort at the beginning, as we took a comprehensive approach by targeting relevant venues and conducting a detailed analysis. Our corpus currently comprises 178 systems and is constantly being expanded to include further relevant systems. The advantage of this dynamic approach is that the corpus can be flexibly adapted to new developments. A major problem was the publication and accessibility of the corpus. The need for extensive tables and analyses made it challenging to find a suitable platform for publication. We currently use a subdomain on our website to archive the corpus and provide references. Nevertheless, we are looking for better solutions for archiving and versioning, as the existing platforms such as Zenodo do not yet meet all our requirements.
“To improve the quality and usefulness of research data, it is crucial that we not only rely on formal committees, but also actively think from the user's perspective. It is important that researchers themselves clearly define their needs and bring this information into the community process.”
Michael Striewe

Profile Photo: Michael Striewe
Michael Striewe
a professor for Computer Science at Trier University of Applied Sciences. His teaching and research areas are software engineering and educational technology. He is particularly engaged in e-assessment, covering its organizational, technical, and didactic aspects. His expertise extends to the systems involved, the data they generate, and the scenarios they support.
“Researchers have already contacted us directly to add their own systems to the corpus, which is a good sign of the relevance and usefulness of our work.”
Sven Stickroth

Profile Photo: Sven Strickroth
Sven Strickroth
is a Professor for Technology-Enhanced Learning in Computer Science at LMU Munich. His research focuses on supporting teaching and learning through digital media and technologies in connection with (semi-)automatically generated feedback, cooperation, and social interaction. This includes developing new scenarios and technologies for instant feedback fostering student collaboration in large lectures, and providing analytics overviews for educators and learners.
How has creating and publishing the corpus changed the way you prepare data and literature?
MICHAEL: The effort involved in the corpus has shown us that we still need to work on our workflow, especially with regard to regular updates, the integration of contributions from third parties and the processing of correction reports. We still face challenges in how we deal with changes and errors in the data. At the same time, the successful publication of the corpus has given us positive impetus for future research projects and provided us with new ideas for the further development and publication of datasets.
SVEN: I agree with the assessment that open source and open data offer great advantages, but in practice there are often difficulties, especially when it comes to agreeing to the publication of data that originates from programming tasks and similar projects. This data is often not available or difficult to publish. Cooperation with students who agree to make their data available is also challenging and time-consuming. We have found a model that works, but it is not yet scalable when more people are involved.
What challenges do you face in research data management, and what opportunities for improvement do you see to overcome these problems?
MICHAEL: A major annoyance in research data management is the uncertainty about what is currently considered 'state of the art'. It is often not clear which standards and practices are accepted in the community, which leads to uncertainty and frustration. Even when efforts are made to publish data correctly, different expectations and assessments of formats or metadata can lead to problems. This disagreement within the community can make researchers reluctant to take action for fear of not meeting expectations. It is therefore important to develop clear guidelines and consensus within the community to reduce these uncertainties.
SVEN: A key problem in research data management is how to effectively communicate important information that is not necessarily the main contribution of a paper. Often the focus is on having a paper stand on its own, which makes it difficult to include additional data or code. This can make researchers reluctant to fully share their data, especially if the code provided is imperfect or in an unfinished state. The fear that incomplete or 'hacky' scripts could negatively impact one's expertise is another obstacle. So how can we ensure that important additions to the reproducibility and reusability of research findings are appropriately communicated and acknowledged without compromising the quality or reputation of the authors?
That’s an important question. Who should take the lead in improving research data management?
Sven: Ultimately, it is up to the community to work together to develop standards and consensus for research data management. Institutions and conferences can contribute by setting clear requirements and guidelines. Personally, I try to increase visibility by providing source code and data in my publications. By including links to repositories and datasets, I encourage the use and further development of research data, motivating others to do the same.
Michael: To improve the quality and usefulness of research data, it is crucial that we not only rely on formal committees, but also actively think from the user's perspective. Researchers themselves should clearly define their needs and bring this information into the community process. Visibility is key: creating more transparency about what data and metadata is actually needed helps the community recognize and acknowledge these requirements. Positive examples and recognition for contributions – such as working with and publishing the data of others—can encourage the willingness to actively share and use research data. This not only encourages sharing, but also demonstrates the value of publishing and using research data.
What tip would you give your younger self to be better prepared for the current challenges?
Michael: I wish that during my studies I had been made more aware of how important it is to publish research data and in what form. A simple guide on how to do it better would have been a big step forward.
Sven: I would have liked more emphasis on the importance of publishing research data. Although I was aware of the importance of transparency and traceability through my early activity in the open source community, targeted training in handling and publication of research data could have been of great benefit to others.
Frage 1
Who should take the lead in developing clear standards?
Auswertung
richtig beantwortet!
About the project
The “Corpus of Task-Based Grading and Feedback Systems” focuses on building a comprehensive corpus of task-based grading and feedback systems for programming education. By collecting and analysing systems with source code that have been published in recent years, the aim is to provide the research community with a solid, continuously updated basis for future studies.