Naujoks-Schober, Nick; Händel, Marion (2026)
Vortrag auf der EARLI SIG 6 & 7 Conference, August 2026.
What do students know about self-regulated learning with GenAI? A situational judgment test
Authors: Nick Naujoks-Schober & Marion Händel
Abstract
The rapid diffusion of generative artificial intelligence (GenAI) tools such as ChatGPT is reshaping students’ self‑regulated learning (SRL) practices, yet little is known about learners’ conditional knowledge for strategically employing these technologies. To address this gap, we developed and preliminarily validated a situational judgment test that measures students’ strategic knowledge of interacting with GenAI across two typical higher‑education learning scenarios. Following a multi‑step design, we (1) conducted a literature review to identify the most frequent cognitive, metacognitive, and resource‑oriented learning situations involving GenAI; (2) performed problem‑centered interviews with 18 students from eight disciplines to elicit real‑world strategies; and (3) refined these strategies into 36 items spanning the forethought, performance, and self‑reflection phases of SRL for two generic scenarios—planning a study project and creating a summary. Experts (N = 48 SRL scholars and AI specialists) rated the usefulness of each strategy on a six‑point Likert scale, yielding significant pairwise discriminations for the majority of items (28/36 for the summary scenario, 21/36 for the planning scenario). The situational judgment test thus captures conditional knowledge of SRL‑GenAI interaction, offering a robust instrument for assessing metacognitive competence beyond self‑report AI‑literacy measures. Findings will inform educators on students’ knowledge about SRL with GenAI and may guide interventions to foster effective, reflective AI‑enhanced learning.
1. Introduction and aims
According to models of self-regulated learning (SRL), students should actively manage their learning processes using situation-specific learning strategies (Panadero, 2017). To employ the most effective strategy in a given context, students require conditional knowledge regarding their strategy use. To assess this knowledge, situational judgment tests compare students’ evaluations of strategy usefulness in realistic, hypothetical learning scenarios against expert benchmarks (Dörrenbächer-Ulrich et al., 2024; Pfost & Hübner, 2025).
However, the rapid rise of generative artificial intelligence (GenAI) applications, such as ChatGPT, is transforming learning scenarios and strategies associated with student SRL (Chiu, 2024). Current research discusses potential benefits of integrating GenAI into higher education for individualizing learning, for instance, through quizzes generated with students’ lecture notes (Dillon, 2024; Gärtner et al., 2024). Conversely, first studies highlight challenges in SRL with GenAI as students may delegate tasks too extensively to the AI (cognitive offloading; Lee et al., 2025) or insufficiently monitor their own learning and interaction with the AI (metacognitive laziness; Chardonnens, 2025; Fan et al., 2025).
According to the current state of research it is unclear whether students possess the strategic knowledge necessary to successfully interact with GenAI during self-regulated learning. Previous measurement instruments to assess AI literacy have focused either on students' self-reports (Almatrafi et al., 2024) or on their knowledge about GenAI specifications (Markus et al., 2025). We therefore aimed to develop and validate a situational judgment test to assess students’ knowledge of how to strategically interact with GenAI during their self-regulated learning.
2. Method
To develop a situational judgment test, we followed a multi-step procedure: (1) We performed a literature search to identify scenarios of leaning with GenAI, (2) we conducted problem-centered interviews with students to select relevant scenarios and to identify strategies to interact with GenAI across different disciplines, and (3) we obtained expert ratings on the consequently developed situational judgment test.
Based on 60 publications, we extracted the three most frequent cognitive, metacognitive, and resource-oriented scenarios. Those scenarios were used for problem-centered interviews with students to analyze their strategy use during GenAI interactions. Using a maximum-variation sampling approach, interviews were held with 18 students across eight different study programs. Students were interviewed regarding their learning behavior when interacting with GenAI. Drawing on the Self-Regulated Learning Interview Schedule (Zimmerman & Pons, 1986), participants first rated the relevance of nine given scenarios (e.g., using GenAI to create a summary) regarding their own learning. Next, students described their strategic approach for those scenarios that they had rated as relevant. Both the content and usefulness of the reported strategies were coded using qualitative content analysis (Mayring & Fenzl, 2019), employing both inductive category development and deductive assignment to overarching learning strategy categories.
In the next step, we extracted strategies for the most relevant scenarios from the interviews. These strategies were iteratively refined according to low, middle, and high usefulness based on the degree of students’ (meta-)cognitive activity and the suitability of the GenAI. This final test was given to experts from the fields of self-regulated learning (N = 31) and/or artificial intelligence (N = 17) who rated the usefulness of the strategies on a 5-point Likert scale from 1 (very low) to 5 (very high).
3. Results
Identification of scenarios. From the literature review, both discipline-specific and cross-disciplinary scenarios and strategies were identified. Especially the generic scenarios provided a solid foundation for the development of the situational judgment test, which is intended to be applicable to students across all fields.
Development of the situational judgment test. Two of the most relevant and generic scenarios for learning with GenAI were selected for the test, namely planning a study project and creating a summary. For each scenario, strategy options were identified from the interviews that differed in terms of their quality and usefulness for learning purposes. In line with Zimmerman's model (1986), each scenario was structured along the three phases of self-regulated learning. That is, the forethought phase before planning (p)/summarizing (s), the performance phase during p/s, and the self-reflection phase after p/s. For each scenario and each phase, six strategies were developed, resulting in 36 strategies differing in usefulness.
Usefulness of strategies (expert ratings). Initial descriptive observations of the expert ratings and two-factor variance analyses for Friedman ranks in connected samples indicate theoretically valid and significant pair comparisons of the strategies. Across all phases of the summary scenario, the expert ratings confirmed 28 out of 36 pair comparisons. In the planning scenario, however, the expert judgments showed only 21 valid pair comparisons. At the time of the conference, data from a student sample from different disciplines will be available, which allows for an initial validation of the test instrument.
4. Theoretical and educational significance of the research
This research addresses the rapidly changing learning behavior shaped by GenAI, where SRL competencies are becoming increasingly vital for actively engaging in and maintaining an overview of one's own learning process. Additionally, by using typical scenarios in higher education and comparing student and expert ratings, the developed test goes beyond current self-reports of AI literacy and knowledge tests about GenAI. In doing so, the situational judgment test also captures an aspect of higher metacognitive skills through conditional knowledge of strategy use, which is often overlooked in other operationalizations of AI literacy (Almatrafi et al., 2024). As metacognitive monitoring and learners’ cognitive engagement seems central for self-regulated learning with GenAI interaction, the situational judgment test should provide a robust measure even against the rapid development of GenAI models.
Assessing conditional knowledge for strategic learning with GenAI also provides educators with an estimate of the learners’ existing competence in this regard. Based on students' answers, it will also be possible to identify how students consider GenAI useful across different learning scenarios. These insights reveal, for example, whether pure task outsourcing is considered more useful than targeted co-constructive processes, and whether students recognize their own (meta-)cognitive activity as a central element. Future research will focus on how students can be supported in their interactive process with GenAI.
5. References
Almatrafi, O., Johri, A., & Lee, H. (2024). A systematic review of AI literacy conceptualization, constructs, and implementation and assessment efforts (2019–2023). Computers and Education Open, 6, 100173. https://doi.org/10.1016/j.caeo.2024.100173
Chardonnens, S. (2025). Adapting educational practices for Generation Z: Integrating metacognitive strategies and artificial intelligence. Frontiers in Education, 10, Article 1504726. https://doi.org/10.3389/feduc.2025.1504726
Chiu, T. K. F. (2024). A classification tool to foster self-regulated learning with generative artificial intelligence by applying self-determination theory: A case of ChatGPT. Educational Technology Research and Development, 72(4), 2401–2416. https://doi.org/10.1007/s11423-024-10366-w
Dillon, T. (2024). Korean university students’ prompt literacy training with ChatGPT: Investigating language learning strategies. English Teaching, 79(3), 123–157. https://doi.org/10.15858/engtea.79.3.202409.123
Dörrenbächer-Ulrich, L., Sparfeldt, J. R., & Perels, F. (2024). Knowing how to learn: Development and validation of the strategy knowledge test for self-regulated learning (SKT-SRL) for college students. Metacognition and Learning, 19(2), 1–45. https://doi.org/10.1007/s11409-024-09379-w
Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
Gärtner, C., Moraß, A., Koss, S., Garbusa, S., Matern, S., Innermann, I., & Forschungs- Und Innovationslabor Digitale Lehre, F. (2024). Einsatz, Nutzen und Grenzen von ChatGPT und anderen Large Language Modellen an den bayerischen HAWs (No. 5; Die Studien- und Schriftenreihe des Forschungs- und Innovationslabors Digitale Lehre – FIDL). FIDL – Forschungs- und Innovationslabor Digitale Lehre. https://doi.org/10.34646/THN/OHMDOK-1466
Lee, H.-P. (Hank), Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Article 1121. https://doi.org/10.1145/3706598.3713778
Markus, A., Carolus, A., & Wienrich, C. (2025). Objective measurement of AI literacy: Development and validation of the AI competency objective scale (AICOS). Computers and Education: Artificial Intelligence, 9, 100485. https://doi.org/10.1016/j.caeai.2025.100485
Mayring, P., & Fenzl, T. (2019). Qualitative Inhaltsanalyse. In N. Baur & J. Blasius (Eds.), Handbuch Methoden der empirischen Sozialforschung (pp. 633–648). Springer Fachmedien Wiesbaden. https://doi.org/10.1007/978-3-658-21308-4_42
Panadero, E. (2017). A review of self-regulated learning: Six models and our directions for research. Frontiers in Psychology, 8, Article 422. https://doi.org/10.3389/fpsyg.2017.00422
Pfost, M., & Hübner, V. (2025). Assessment of strategic knowledge of learning from errors in higher education. Diagnostica, 71(2), 53–63. https://doi.org/10.1026/0012-1924/a000341
Zimmerman, B. J. (1986). Becoming a self-regulated learner: Which are the key subprocesses? Contemporary Educational Psychology, 11(4), 307–313. https://doi.org/10.1016/0361-476x(86)90027-5
Zimmerman, B. J., & Pons, M. M. (1986). Development of a structured interview for assessing student use of self-regulated learning strategies. American Educational Research Journal, 23(4), 614–628. https://doi.org/10.3102/00028312023004614
Händel, Marion; Naujoks-Schober, Nick; Kamath Barkur, Sudarshan (2026)
DGPs-Veranstaltung "Künstliche Intelligenz menschzentriert gestalten" in Berlin.
Biller, Simon; Händel, Marion (2026)
Poster auf der Abschlusstagung lernen:digital in Berlin.
Naujoks-Schober, Nick; Beatrix, Getze; Händel, Marion (2026)
theansweringmachine.
Händel, Marion (2026)
Diskutantin im Symposium von L. Dörrenbächer-Ulrich auf der 13. GEBF Konferenz in München.
Biller, Simon; Groß-Mlynek, Lena; Bastian, Jasmin; Händel, Marion (2026)
Education and Information Technologies.
DOI: 10.1007/s10639-026-13918-0
Digital communication has played an increasingly important role in schools around the world, especially since the COVID-19 pandemic. For professional communica tion and collaboration in particular, digital tools have provided teachers with the op portunity to collaborate location-independently and to easily exchange information and teaching materials. Hence, this study aimed to explore how teachers communi cate and collaborate digitally by examining differences in the use of instant messag ing and videoconferencing and the attitudes of teachers towards these technologies. Therefore, an online survey was conducted with primary and secondary school teachers from Germany (N = 250, 72.0% female). The analysis showed that mes sengers were used significantly more than videoconferences and that they differed in their usefulness for occasions of communication and collaboration. Structural equation modeling indicated that the self-assessed digital communication compe tence of teachers is both a significant predictor for the behavioral intention to use messengers as well as videoconferences, although the behavioral intention to use messengers in the future was significantly higher than for videoconferences. While a higher threshold to use videoconferences might be a reason for the differences that were identified in this study, further research into the communication and col laboration among teachers is still needed to understand the reasons for the differ ences in use.
Biller, Simon; Händel, Marion (2026)
In: Kallenbach, C., Karnebogen, M., Serpemen, A., Seufert, P. (eds) Fortbildungs- und Professionalisierungsangebote: Schulentwicklung. Kompetenzverbund lernen:digital, Potsdam, 52-53.
Händel, Marion (2026)
xplr-media: in Bavaria .
Zentner, Michelle; Stang, Philipp; Händel, Marion (2025)
In: Stang, P., Weiss, M., Köllner, M. (eds) Health psychology. Applications in clinical and sports contexts, 1. Auflage, Nomos, Baden-Baden, 197–205.
Biller, Simon; Händel, Marion (2025)
Videokonferenz Fachforum Schulentwicklung.
Naujoks-Schober, Nick; Händel, Marion (2025)
1. Science Slam der Hochschule Ansbach.
Naujoks-Schober, Nick; Händel, Marion (2025)
Ansbacher Science Slam.
Händel, Marion (2025)
Campus Wissen.
Biller, Simon; Händel, Marion (2025)
Herbsttagung der Sektion Medienpädagogik der Deutschen Gesellschaft für Erziehungswissenschaften DGfE in Nürnberg.
Der Beitrag stellt den Entwicklungsprozess des im Projekt LeadCom entstandenen Selbstlernkurses Videokonferenznutzung zur kollegialen Kommunikation und Kooperation dar und zeigt auf, warum diese Fortbildung zum Thema Videokonferenzen in außerunterrichtlichen Situationen für eine zeitgemäße Schulentwicklung relevant ist. Einerseits wurde der Kurs mit dem Ziel konstruiert, durch eine kompetente Nutzung von virtuellen Meetings im Schulkontext, eine moderne und effektive Zusammenarbeit im Lehrerkollegium zu fördern und so den sich verändernden Ansprüchen einer zunehmend digitalisierten Welt gerecht zu werden. Andererseits soll die Nutzung des Selbstlernkurses auch dazu beitragen, dass Lehrkräfte und Schulleitungen sich digital kompetent und selbstwirksam im Umgang mit digitalen Kommunikationsmedien erleben. Als Transferziel sollten die Einstellungen und Erwartungen von Lehrkräften gegenüber Informations- und Kommunikationstechnologien (ICT) und die Nutzung von ICT für schulinterne Kooperation die Nutzung von ICT im Unterricht sowie die digitale Kompetenzentwicklung von Schülerinnen und Schülern positiv beeinflussen (u.a. Drossel et al., 2016; Fraillon et al., 2020). Die Entwicklung des Selbstlernkurses fußte dabei unter anderem auf dem weiterentwickelten Technologieakzeptanzmodell von Venkatesh et al. (2003, 2012), der Unified Theory of Acceptance and Use of Technology (UTAUT). In einer eng mit dem Entwicklungsprozess der Fortbildungseinheit durchgeführten quantitativen Studie wurden anhand des UTAUT-Modells Einstellungen und Erwartungen von Lehrkräften in Bezug auf die Nutzung von Videokonferenzen in außerunterrichtlichen Kommunikations- und Kooperationssituationen untersucht. Der Beitrag wird aufzeigen, wie sowohl die theoretischen Überlegungen aus dem UTAUT-Modell als auch in Teilen die Ergebnisse der Studie in die Konstruktion des Selbstlernkurses eingeflossen sind.
Händel, Marion (2025)
Konferenz des Bayerischen Forschungsinstitutes für Digitale Transformation bidt 2025.
Händel, Marion; Naujoks-Schober, Nick (2025)
Forschungs- und InnovationsTag (FIT) 2025 der Hochschule Ansbach.
Generative KI-Tools wie ChatGPT eröffnen neue Möglichkeiten für das Lernen – von schnellen Erklärungen bis hin zu personalisierter Unterstützung. Für erfolgreiche Lernprozesse sollten Lernende aber nicht alle Denkprozesse auslagern (cognitive offloading), den eigenen Lernfortschritt im Blick behalten (metacognitive laziness) und die KI als aktiven Lernpartner nutzen – nicht nur als Suchmaschine (co-creation). Im Vortrag wird vorgestellt, wie diese KI-Interaktionen im Projekt SekoKI beforscht werden.
Bastian, Jasmin; Biller, Simon; Groß-Mlynek, Lena; Händel, Marion (2025)
Wissenschaftliches Poster auf der Herbsttagung der Sektion Medienpädagogik der Deutschen Gesellschaft für Erziehungswissenschaften DGfE in Nürnberg.
Händel, Marion; Berges, Marc-Pascal; Gläser-Zikuda, Michaela; Kammerl, Rudolf; Kudlich, Hans; Martschinke, Sabine; Pirner, Manfred (2025)
Händel, Marion; Berges, Marc-Pascal; Gläser-Zikuda, Michaela; Kammerl, Rudolf...
Education and Information Technologies 30, 25177–25196.
DOI: 10.1007/s10639-025-13714-2
Learning in the digital world requires not only technological skills for using digital tools but also ethical skills to critically reflect on (in)adequate digital media use and potential negative consequences. These skills are particularly crucial in professions dealing with public welfare and societal issues. Inter alia, those are teachers who educate the youth, legal professionals who judge (il)legal behavior regarding media, or computer scientists who bear responsibility when developing algorithms. Accordingly, higher education students studying teacher education, law studies, or computer science studies should develop ethical skills for the digital world. This study examined how higher education students perceive problematic media behaviors and which digital competences they regard relevant for ethical issues. In addition, the study investigated whether students of teacher education, law studies, and computer science studies differ in their perceptions. To this end, an online survey with N = 461 participating students was conducted. Study results indicated that higher education students perceived problematic media behaviors as such with posting inappropriate content identified as the most problematic. Furthermore, students considered several digital competences as relevant for ethical issues with protecting and acting safely as most relevant. In-depth analyses uncovered subject-specific differences with computer science students being most ethically savvy, albeit differences were only of small effect size. The study provides valuable insights into the intersection of digital competences, ethical considerations, and academic disciplines. In the future, longitudinal and training studies will help to understand how differences emerge and whether students of different study subjects benefit from digital ethics training.
Naujoks-Schober, Nick; Reinhold, Lhea; Händel, Marion (2025)
Wissenschaftliches Poster auf der 14th Conference of the Media Psychology Division (DGPs) in Duisburg.
Does the AI agree? Inter-rater agreement in learning diary evaluation
Theory
Learning diaries as formative assessments are promising to support the learning process and stimulate reflection. In an open learning diary, learners apply learning strategies to reflect on their learning and deepen their knowledge. To guide learners, learning diaries can be structured according to different learning strategies. However, grading and feedback on learning diaries is effortful for teachers. Artificial intelligence (AI) may assist teachers in the evaluation process. A prerequisite is that AI and teachers show high inter-rater agreement.
Research Questions
The current study aimed to analyze the agreement between teachers and ChatGPT-4o by examining four separately assessed learning strategy categories of a learning diary in adult education (organization, in-depth elaboration, transfer-supporting elaboration, and metacognition). The two research questions were:
• Q1: How accurate is the overall agreement between teachers and ChatGPT-4o, and are there differences across different learning strategy categories?
• Q2: Does the inter-rater agreement differ between teachers across the four learning strategy categories?
Method
Seven different adult education teachers and ChatGPT-4o evaluated a total of 540 learning diary entries. Each teacher assessed approximately 65 entries. Teachers were trained in criteria-based evaluation per learning strategy category. An engineered prompt supported the ChatGPT-4o model.
Teacher ratings served as the reference for the inter-rater agreement. Absolute accuracy and under-/overestimation (bias) were calculated for each learning strategy category as accuracy measures. Furthermore, overall accuracy values were calculated across the four categories for absolute accuracy and bias.
A doubly multivariate repeated measures ANOVA was conducted with the four learning strategy categories as repeated measures and absolute accuracy and bias as measures to test for accuracy differences between the learning strategy categories (Q1). Additionally, teacher was used as a between-subjects factor. Thus, the interaction of the learning strategy category and rating teacher regarding accuracy could be tested statistically (Q2).
Reinhold, Lhea; Händel, Marion (2025)
MedienPädagogik: Zeitschrift für Theorie und Praxis der Medienbildung MEDIDA24 (65), 227-250.
DOI: 10.21240/mpaed/65/2025.08.03.X
Künstliche Intelligenz (KI) kann im Prozess der Leistungsbewertung assistieren und diesen transformieren. Besonders lohnend scheint eine KI-Assistenz bei der Bewertung von komplexem, geschriebenem Text. Jedoch ist der Einsatz von KI im Bewertungsprozess «hochriskant» (EU 2024) und bedarf umfangreicher Analysen. Die vorliegende Studie untersucht, inwiefern ChatGPT-4o die Auswertung und Interpretation von Lerntagebucheinträgen objektiv vornehmen kann. Dafür werden 757 Lerntagebucheinträge aus der geförderten Weiterbildung in Deutschland von Mensch und Maschine bewertet. Sowohl Mensch als auch Maschine erhalten hierzu Kriterien, nach denen die Bewertung vorzunehmen ist; ChatGPT-4o wird diesbezüglich mit einem Prompt unterstützt. Die Übereinstimmung der Bewertungen wird anhand der Masse Sensitivität und Spezifität gemessen. Die Ergebnisse zeigen, dass die Bewertungsvorschläge von ChatGPT-4o eine moderate bis hohe Übereinstimmung mit den menschlichen Bewertungen aufweisen; gleichzeitig neigt ChatGPT-4o jedoch zu einer optimistischen Bewertung der Lerntagebucheinträge. Die Ergebnisse weisen darauf hin, dass eine hybride Intelligenz, also eine Kombination der Stärken von Mensch und Maschine, gewinnbringend für Bewertungsprozesse sein kann. Künftig denkbar sind halbautomatisierte Bewertungsprozesse von Lerntagebucheinträgen, in denen die KI die Bewertung der Lerntagebucheinträge übernimmt und Lehrkräfte bei kritischen Fällen regulierend eingreifen. So könnte die Korrektureffizienz ohne bedeutende Qualitätsverluste gesteigert werden.
Fakultät Medien
Residenzstr. 8
91522 Ansbach
marion.haendel[at]hs-ansbach.de
ORCID iD: 0000-0002-3069-5582