Paper deep dive
Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study
Hansika Ekanayake Mudiyanselage, Rohan Jai Dharmaraj, Malik Abdul Sami, Zheying Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/3/2026, 9:48:53 AM
Summary
This study investigates the integration of a multi-agent AI tool into a master-level Requirements Engineering (RE) course at Tampere University, guided by the TPACK framework. Using a mixed-methods design with 100 students, the research analyzes how structured assignment design influences students' use of AI for requirements quality analysis. Results indicate that students used AI selectively for analysis and evaluation rather than automation, showing improved understanding of concrete quality dimensions like value articulation and testability, while negotiability showed mixed effects. The findings suggest that TPACK-guided scaffolding aligns AI affordances with pedagogical goals, fostering critical evaluation and conditional trust.
Entities (10)
Relation Signals (8)
Tampere University â affiliation â Rohan Jai Dharmaraj
confidence 98% ¡ Rohan Jai Dharmaraj... Tampere University
Tampere University â affiliation â Malik Abdul Sami
confidence 98% ¡ Malik Abdul Sami... Tampere University
Tampere University â affiliation â Zheying Zhang
confidence 98% ¡ Zheying Zhang... Tampere University
Tampere University â affiliation â Hansika Ekanayake Mudiyanselage
confidence 98% ¡ Hansika Ekanayake Mudiyanselage... Software Engineering Research Center (TASE), Tampere University
Multi-Agent AI Tool â usedin â Requirements Engineering
confidence 97% ¡ examines a TPACK-guided integration of a multi-agent AI tool into a master-level RE assignment
TPACK Framework â guidesintegrationof â Multi-Agent AI Tool
confidence 96% ¡ This study addresses this gap by examining how a multi-agent AI tool can be integrated into an RE course assignment using the Technological Pedagogical Content Knowledge (TPACK) framework
INVEST Framework â definesqualitycriteriafor â Requirements Engineering
confidence 95% ¡ requirements quality frameworks such as INVEST [10]
Multi-Agent AI Tool â developedby â Software Engineering Research Center
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid adoption of generative Artificial Intelligence (AI) in software engineering (SE) practice creates a need for pedagogically grounded approaches to AI integration in SE education, especially in conceptually intensive subjects such as requirements engineering (RE). This study examines a TPACK-guided integration of a multi-agent AI tool into a master-level RE assignment on requirements quality analysis. Using a mixed-methods design (N=100; 72 submissions analysed), we examine how structured assignment design shaped students' AI use, affected their understanding of user story quality criteria, and influenced their perceptions of AI's benefits and limitations. Results show that students used the AI tool selectively, mainly as support for analysis and evaluation rather than automation. Alignment improvements were most evident for structurally concrete requirements quality dimensions, such as value articulation and testability, while negotiability showed mixed effects. Students reported conditional trust, active refinement, and increased awareness of quality criteria, alongside moderate usability challenges. The findings show that TPACK-guided scaffolding can align AI affordances with pedagogical goals and RE content, offering design guidance for responsible AI integration in RE education.
Tags
Links
- Source: https://arxiv.org/abs/2607.28176v1
- Canonical: https://arxiv.org/abs/2607.28176v1
Trouble viewing inline? Open PDF directly â
Full Text
55,193 characters extracted from source content.
Expand or collapse full text
Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study Hansika Ekanayake Mudiyanselage, Rohan Jai Dharmaraj, Malik Abdul Sami, and Zheying Zhang â Software Engineering Research Center (TASE), Tampere University, Finland hansika.ekanayakemudiyanselage, rohanjai.dharmaraj, malik.sami, zheying.zhang@tuni.fi *corresponding author AbstractâThe rapid adoption of generative Artificial Intel- ligence (AI) in software engineering (SE) practice creates an urgent need for pedagogically grounded approaches to AI inte- gration in SE education, particularly in conceptually intensive subjects such as requirements engineering (RE). While prior studies report studentsâ perceptions and concerns regarding AI use, systematic investigations grounded in established educa- tional theory remain limited. This study addresses this gap by examining how a multi-agent AI tool can be integrated into an RE course assignment using the Technological Pedagogical Content Knowledge (TPACK) framework as both a design principle and an analytical lens. Using a mixed-methods design (N=100; 72 submissions anal- ysed), we analyze (i) how structured assignment design shapes studentsâ AI use, (i) how AI-supported assignment affects studentsâ understanding and application of user story quality criteria, and (i) how students perceive the benefits and limitations of AI use in the course assignment. The results show that students engaged selectively with the AI tool and predominantly used it as a support for analysis and evaluation rather than as an automation mechanism. Learn- ing gains were most evident in structurally concrete quality dimensions such as value articulation and testability, while attributes such as negotiability showed mixed effects. Analysis through a TPACK lens revealed that structured pedagogical scaffolding mediated technological affordances, guiding stu- dents to engage with core content knowledge rather than relying on automation. Students reported conditional trust, active refinement, and enhanced awareness of quality criteria, alongside moderate usability challenges. This study provides empirically grounded design guidance for responsible AI integration in RE education and demonstrates how TPACK-guided assignment design can align technological affordances with pedagogical intent and established require- ments quality criteria. Keywordsârequirements engineering education; AI tool; user story quality; TPACK; INVEST framework 1. INTRODUCTION Generative artificial intelligence (AI), particularly large lan- guage models (LLMs), is rapidly transforming software en- gineering (SE) practice. Requirements engineering (RE) is especially affected, as its core activities, such as elicitation, analysis, specification, and validation, rely heavily on natu- ral language articulation and interpretation. Recent research demonstrates the expanding use of LLMs to generate, re- fine, and even prioritize user stories and other requirements artifacts [1][2][3]. While such advances offer opportunities for productivity and support, they also highlight unresolved challenges related to trust, humanâAI collaboration, and con- textual alignment [4][5]. These developments signal not only a technological shift in RE practice but also a pedagogical challenge for RE education (REE). Early empirical studies [6] report that students perceive AI tools as helpful for comprehension and productivity, yet con- cerns remain regarding over-reliance, superficial adoption, and limited critical evaluation of AI-generated outputs. Broader reviews of AI integration in SE and engineering education indicate that many implementations remain exploratory and lack systematic pedagogical grounding and rigorous empirical validation [7] [8]. There is still limited evidence on how AI can be deliberately integrated into conceptually intensive subjects such as RE in ways that preserve and strengthen core analytical competencies. The challenge, therefore, is not whether AI can support RE tasks, but how it should be pedagogically orchestrated. Effective integration requires alignment between technological affordances, disciplinary knowledge, and instructional design. To address this challenge, we adopt the Technological Peda- gogical Content Knowledge (TPACK) framework [9] as both a design principle and an analytical lens. TPACK concep- tualizes effective technology integration as the alignment of content knowledge (CK), pedagogical knowledge (PK), and technological knowledge (TK). In the context of REE, CK corresponds to principles, techniques, and practices for RE, including requirements quality frameworks such as INVEST [10], ISO/IEC/IEEE 29148 [11] etc., PK involves structured learning mechanisms such as contrastive analysis and peer review, and TK refers to AI tools supporting RE activities. Only through deliberate alignment of these dimensions can AI function as a scaffold for analytical reasoning rather than as an automation substitute. This study reports on a TPACK guided integration of a multi-agent AI tool into an assignment in a master-level RE course. The assignment was deliberately designed to sequence manual reasoning before AI use, require comparison between students- and AI-generated requirements, and include itera- tive refinement, peer review, and structured reflection. Rather arXiv:2607.28176v1 [cs.SE] 30 Jul 2026 than evaluating the AI tool in isolation, we investigate how pedagogical sequencing shapes studentsâ enactment of AI, how such integration influences their application of require- ments quality criteria, and how students critically evaluate AI-generated outputs. Specifically, we address the following research questions: RQ1 How can an AI tool be effectively integrated into course assignments to support structured analysis, reflection, and iterative refinement in studentsâ learning? RQ2 How does the AI tool influence studentsâ understanding and application of requirements quality criteria? RQ3 How do students perceive the usefulness, trustworthi- ness, and limitations of the AI tool, and to what extent do they critically evaluate its outputs? The study uses a mixed-methods design in the RE course (N = 100; 72 analyzed submissions), combining quantitative analysis of requirements quality assessments with qualitative analysis of interaction patterns and reflective responses. By triangulating observable behavior, performance measures, and self-reported experiences, we provide empirical evidence of how structured AI integration shapes learning processes in RE. This study contributes to SE education research in three ways. First, it applies TPACK in the context of AI-supported REE through a replicable assignment design. Second, it provides empirical evidence that structured instructional sequencing can shape studentsâ AI enactment toward evaluative and reflective engagement. Third, it identifies differential effects of AI sup- port across requirements quality dimensions, offering design implications for integrating generative AI into conceptually intensive SE subjects. By shifting attention from AI capability to pedagogical or- chestration, this work advances evidence-based guidance for responsible and analytically grounded AI integration in soft- ware engineering education. 2. BACKGROUND AND RELATED WORK Requirements engineering (RE) is a core subject in software engineering curricula, focusing on eliciting, analyzing, speci- fying, validating, and managing stakeholder needs. REE places strong demands on studentsâ abilities to reason about ambigu- ity, communicate with stakeholders, and work effectively with textual description. These make RE conceptually demanding and particularly sensitive to changes introduced by generative AI technologies. While AI tools can generate syntactically plausible requirements for a given software project, the devel- opment of analytical judgment and quality reasoning remains a central educational objective. Recent advances in LLMs have influenced RE practice. A systematic literature review by Cheng et al. [1] synthesizes the application of AI across RE activities and highlights persistent challenges related to trust and humanâAI collaboration. These developments have motivated initial research on exploring AI use in REE. Guardado et al. [6] reported that guided LLM use can improve studentsâ comprehension of RE practices, yet also raise concerns regarding academic integrity, over-reliance, and insufficient critical evaluation of AI outputs. Similarly, Tiwari and Rathore [12] propose structured approaches for integrating LLMs into REE, which emphasize the need for deliberate in- structional design. Furthermore, practice-oriented RE research further demonstrates the increasing capability of LLMs to support RE workflows [13]. Such studies reinforce the need to prepare students for AI-supported RE practice. Neverthe- less, systematic investigations of pedagogically grounded AI integration in REE remain limited [14]. 2.1 AI in Software Engineering Education Beyond RE, a growing body of research examines AI integra- tion in SE and engineering education. Sah et al. [7] provides a comprehensive review of AI adoption in SE education. Their findings indicate that while experimentation is widespread, many implementations lack robust pedagogical grounding. Identified challenges include instructor readiness, ethical con- cerns, assessment validity, and alignment between AI use and learning objectives. Similarly, Filippi and Motyl [8] present a growing interest in LLM applications across engineering education, while noting a lack of structured guidelines and pedagogical frameworks for effective AI integration. Studies in broader engineering education contexts, such as [15], report generally positive student perceptions of AI-supported learning but stress the importance of careful instructional design. Researchers argue that AI fundamentally challenges educa- tional practices. Kirova et al. [16] contend that SE education must adapt assignment and assessment strategies to LLM-rich environments, while early exploratory studies [17] highlight potential benefits of AI-supported learning while emphasizing the need for more systematic and theory-driven research. Over- all, these findings suggest that AI adoption in SE education is advancing faster than the development of pedagogically grounded integration strategies. While AI tools demonstrate technical capability, less attention has been devoted to how instructional design mediates studentsâ interaction with AI- generated artifacts, particularly in conceptually intensive sub- jects such as RE. 2.2 Pedagogical Framework for AI Integration and Research Gap The Technological Pedagogical Content Knowledge (TPACK) framework [9] provides a well-established theoretical founda- tion for analyzing technology-enhanced learning. TPACK con- ceptualizes that effective educational use of technology arises from the integration of content knowledge (CK), pedagogical knowledge (PK), and technological knowledge (TK). Rather than treating technology as an isolated enhancement, TPACK emphasizes that meaningful learning emerges when techno- logical affordances are deliberately aligned with disciplinary content and instructional strategies. Recent extensions of TPACK to AI-enabled contexts empha- size the importance of AI literacy, trust, and ethical awareness in shaping educatorsâ acceptance and effective use of AI tools [18]. L. Eyal [19] emphasizes that effective AI integration requires educators to develop AI-specific technological, ped- agogical, and content knowledge and to purposefully align AI use with instructional objectives. However, existing AI- TPACK research primarily focuses on general or K-12 educa- tional contexts. Its application in higher education, particularly in technically intensive domains, such as SE, remains under- explored. Despite increasing experimentation with AI tools in SE edu- cation, two limitations remain evident. First, AI integration in SE education remains largely exploratory and insufficiently grounded in pedagogical design [7], [8], [14]. Second, al- though TPACK offers a useful lens for technology integration, its application to AI-supported learning in higher education, particularly in REE involving requirements quality framework such as INVEST [10] and ISO/IEC/IEEE 29148 [11] is underexplored. From a TPACK perspective, this gap reflects limited alignment among technological affordances (AI capability), pedagogical scaffolding (structured learning and assessment strategies), and disciplinary content knowledge (requirements quality criteria). In REE specifically, there is little empirical evidence on how AI tools can be embedded into assignments in ways that support structured analysis, reflection, and iterative refinement without undermining studentsâ analytical competencies. To address the gaps, this study explicitly designs and eval- uates a TPACK-guided integration of a multi-agent AI tool into an RE course assignment, examining how technological, pedagogical, and content elements interact to shape studentsâ AI use and learning outcomes. 3. RESEARCH CONTEXT This section introduces the context and the AI tool used in the study. It provides the necessary background to motivate the subsequent research design, including the conceptual integra- tion of the AI tool into course assignments and the associated data collection and analysis procedures. 3.1 Course Description The study was conducted in the Requirements Engineering course, a 5 ECTS course in the Masterâs Degree Programme in Computing Sciences and Electrical Engineering at Tampere University, with approximately 100 students enrolled annually. The course covers core RE activities in software system devel- opment. It integrates both principles and practical approaches to requirements elicitation, analysis, specification, validation, and requirements management. Students are expected to pos- sess prior knowledge of fundamental software engineering concepts. The course comprises 20 hours of lectures, ten weekly indi- vidual assignments, two mastery exercises, and a collaborative group work. All learning materials, including lecture notes, supplementary readings, assignment descriptions, and project guidelines, are made available through the universityâs Moodle platform. Students complete individual assignments based on their understanding of lectures and textbooks. After six weeks of lectures, students start collaboratively working on a project in groups of three or four, selecting a predefined or self-proposed topic. The project allows students to apply RE knowledge in a research or practice-oriented context. In addition, a mid-term and a final mastery exercise, both given in forms of online quizzes, are used to assess studentsâ understanding of core RE concepts. Although the assignments, group project, and mastery ex- ercises are non-compulsory, active participation is essential for achieving a good final grade. This structured yet flexible course design aims to foster analytical reasoning, critical think- ing, and collaboration, and prepares students to address the inherent complexity and ambiguity of real-world RE practice. 3.2 A Multi-Agent AI Tool To support studentsâ learning in the REQ, the course inte- grated a multi-agent AI-based requirements assistance tool designed to facilitate requirements generation and analysis through iterative refinement. The tool [3][5] was developed as a research prototype in the Software Engineering Research Center (TASE) 1 at Tampere University. Its development was motivated by ongoing research on how LLMs can be orches- trated through multiple specialized agents to assist RE tasks. As illustrated in Figure 1, the tool implements four features: F1 Agent configuration: Users can select agents represent- ing distinct stakeholders and specify their roles for task completion. F2 Requirements generation: Based on project-specific in- put such as product vision, minimal viable product de- scription, and target users, selected agents collaboratively generate an initial set of user stories. F3 Requirements analysis and refinement: Users can iter- atively refine the generated user stories through feedback cycles and manual edits. F4 Requirements prioritization: Once requirements are finalized and approved, selected agents collaboratively prioritize requirements. It is worth noting that this feature is not used in the assignment investigated in this study The integration of this AI tool into the course was motivated by its explicit support for requirements generation and re- finement. Writing high-quality requirements is one learning objective of the course and is also an area where students commonly experience difficulties, such as formulating clear, consistent, and well-structured requirements. The toolâs multi- agent design, which exposes students to multiple stakeholder perspectives, aligns closely with the pedagogical goal of supporting students in reasoning about requirements quality, completeness, and trade-offs, rather than merely producing final student assignment submissions. 4. RESEARCH DESIGN This study adopts a mixed-methods design to investigate (i) how a multi-agent AI tool can be pedagogically integrated into an RE course assignment (RQ1), (i) its impact on stu- dentsâ understanding of requirements quality (RQ2), and (i) studentsâ perceptions and evaluative engagement (RQ3). The 1 https://research.tuni.fi/tase/ NoCategoryRequirementDescription 1Seamless Integration I want to connect my WhatsApp account with my email account, in order to manage all my communications. The user can authenticate both platforms and connection is established. 2User-Friendly Automation I want to automate the forwarding of WhatsApp messages to my email The feature should allow user specified criteria (e.g., sender, ...) for setting up these rules F 1 F 3 F 2 F 4 Figure 1. Screenshots of features implemented in the multi-agent AI tool design combines quantitative analysis of assignment submis- sions with qualitative analysis of student reflections, enabling triangulation between observed performance and self-reported learning experiences. 4.1 Pedagogical Design of AI Integrating The AI-integrated assignment was designed using the TPACK framework to explicitly align content knowledge (requirements quality criteria), pedagogical knowledge (contrastive and re- flective learning activities), and technological knowledge (AI- based requirements generation and refinement). The design aimed to ensure that AI use supported structured analysis and iterative improvement rather than automation. The AI tool was integrated into an individual assignment focusing on reviewing and improving requirements. Students reviewed requirements elicited in earlier assignments for the same project context and aligned them with quality criteria such as INVEST framework [10] and related guidelines. This assignment was selected for AI integration because prior course offerings revealed recurring learning challenges. De- spite explaining the concepts with examples through lectures, students tended to specify requirements that were vague, solution-oriented, insufficiently testable, or lacking explicit justification of value, and struggled to systematically assess and improve the requirements. The AI tool was introduced only after manual revision, al- lowing students to generate alternative user stories for the same functional goals and compare them with their own revisions. This contrastive design encouraged analytical eval- uation and iterative refinement. The expected workload was approximately 2â3 hours. 4.1.1 Task Workflow The assignment followed a staged workflow designed to support contrastive learning and to discourage automation- oriented use of the AI tool. As illustrated in Figure 2, each stage required students to actively engage in evaluation, com- parison, and refinement of requirements rather than relying on AI-generated outputs. Step 1 Manual improvement: Students selected a subset of requirements from earlier assignments and manually revised them as user stories, following the INVEST quality framework that includes quality attributes of Independent, Negotiable, Valuable, Estimable, Small, and Testable [10]. Step 2 AI-assisted user story generation: Students used the AI tool to generate user stories by configuring persona agents representing different stakeholders and providing project-specific inputs, including the product vision, user groups, and other relevant descriptions Step 1: Manual improvement (Student-only, INVEST- based) Step 2: AI-assisted user story generation (Multi-agent perspectives) Step 3: Iterative refinement (Update with AI + manual edits + approve) Step 4: Comparison (Manual vs AI-assisted) Step 5: Peer review (INVEST scoring + justification) Step 6: Reflection (Evaluation & learning) Figure 2. Assignment workflow for AI tool integration in Assignment 4 documented in prior assignments. Tool instructions 2 were provided on the course Moodle page. The instruc- tions were intended to familiarize students with the tool while leaving configuration choices to the studentsâ discretion. Step 3 Iterative refinement: Students reviewed and refined AI-generated user stories through feedback cycles, manual edits, addition or removal of stories, and ex- plicit approval actions. Step 4 Comparison: Students compared manually revised and AI-generated user stories, focusing on differences in clarity, specificity, and alignment with INVEST at- tributes. The user stories were documented in a shared spreadsheet 3 , which is used for the subsequent peer- review step. Step 5 Peer review: Students evaluated their peer studentâs requirements documented on the shared spreadsheet, using structured INVEST-based scoring with justifi- cations. The authors of the reviewed stories were expected to respond with agreement or disagreement. Step 6 Reflection: After completing the peer-review step, students completed a reflection questionnaire 4 to docu- ment their interactions with the tool and their perceived strengths and limitations of the AI tool. In addition to the main assignment workflow, a pre-assignment and a follow-up mastery exercise were deliberately integrated. In both tasks, students analyzed four user stories (US1âUS4) to identify quality violations against the six dimensions of the INVEST framework. To provide a benchmark, the instructor and teaching assistants together established a reference anal- ysis, i.e. evaluation of each user story across the INVEST dimensions 5 . The repeated tasks, compared against the prede- fined reference values, allowed us to compare studentsâ ana- lytical reasoning before and after the AI-integrated assignment workflow. 4.1.2 Roles of the AI tool and Students Within this design, the AI tool assumed the role of a supporting assistant or alternative analyst, comparable to a stakeholder or domain expert capable of proposing candidate requirements 2 https://doi.org/10.6084/m9.figshare.31440988 3 https://doi.org/10.6084/m9.figshare.31430224 4 https://doi.org/10.6084/m9.figshare.31430221 5 https://doi.org/10.6084/m9.figshare.31442422 based on provided context and feedback. The tool generated alternatives and responded to refinement prompts but did not make final decisions. Students, in contrast, assumed multiple active roles aligned with the learning objectives of the assignment. They acted as product owners when manually articulating requirements, as a project manager when engaging multiple stakeholder perspectives through agent selection, as requirements analysts when refining and approving AI-generated outputs, and as reviewers during comparison and peer-review phases. This role-based design positioned students as responsible evaluators and decision-makers, ensuring that learning outcomes related to structured analysis, reflection, and iterative refinement re- mained firmly under student control. 4.2 Data Collection The RE course was delivered from September 2 to December 1, 2025, with a total enrollment of 100 students. The designed assignment was released on September 24 as Assignment 4 within a sequence of ten individual assignments and remained open for a two-week completion period. Data were primarily collected from student assignment submissions in Steps 1- 5 and the reflection questionnaire completed in Step 6. In addition, the data on studentsâ quality assessment of US1-US4 were collected in the pre-assignment and the mastery exercise. Student assignment submissions were collected using a shared spreadsheet. The collected data included manually revised requirements, AI tool generated and subsequently refined user stories, peer-review quality scores, and studentsâ responses indicating agreement or disagreement with peer feedback. These data provide observable traces of student interaction with the AI tool, including selection, refinement, evaluation, and justification behaviors. The reflection questionnaire was designed to collect studentsâ perspectives on their interaction with the AI tool. It served as complementary data that are not directly observable from the assignment submissions alone, such as reasoning behind refinement decisions, criteria used for acceptance or rejection of AI outputs, and perceived learning benefits or limitations. The questionnaire addressed: (i) studentsâ interaction patterns with the AI tool, e.g., time spent, number of generated and approved user stories, refinement actions, etc.; (i) studentsâ choices and rationale for configuring AI agents in different stakeholder roles; (i) perceived impact of the tool on stu- dentsâ understanding and application of requirements quality attributes; (iv) studentsâ evaluation of AI-generated outputs, including editing, acceptance, or rejection decisions; and (v) perceptions of usability, usefulness, confidence, and limitations of the AI tool. In total, the questionnaire comprised 20 mandatory close-ended items and five optional open-ended questions. 4.3 Data Analysis Descriptive statistics and qualitative thematic coding were applied to examine learning impact and student perceptions. Student interaction with the AI tool was analyzed using met- rics such as generation counts, approval rates, and refinement frequencies to characterize engagement patterns. To evaluate changes in requirements quality reasoning, stu- dentsâ INVEST-based assessments of US1âUS4 were com- pared before and after the AI-integrated assignment. Agree- ment proportions relative to predefined reference analyses were calculated, and changes in alignment (â = p after âp pre ) were examined for each INVEST quality dimension. Questionnaire data were analyzed using distributional sum- maries of Likert-scale responses. Open-ended responses were coded thematically following established approaches to the- matic analysis [20] to identify patterns related to perceived usefulness, evaluative behaviors, trust, and reported limita- tions. Cross-tabulation analyses explored relationships be- tween reported refinement behaviors and perceived improve- ments across quality dimensions. This triangulated approach [21] enabled an integrated interpretation of pedagogical de- sign, observed interaction patterns, learning outcomes, and student perceptions. 5. RESULTS Although the assignment was non-compulsory, 74 out of 100 students completed it, of whom 72 provided informed consent and were included in the analysis. The average completion time for the reflection questionnaire was 27 minutes and 45 seconds. This participation level indicates substantial voluntary engagement with both the assignment and reflection compo- nents, providing a sufficiently rich dataset to examine the pedagogical integration of the AI tool and studentsâ enactment of the intended assignment design in practice. 5.1 RQ1 - Studentsâ Engagement and Interaction Patterns The staged assignment workflow reflects an alignment of technological, pedagogical, and content elements as concep- tualized in the TPACK framework. The assignment design includes several explicit pedagogical mechanisms that guide studentsâ engagement with the tool. These included (i) a se- quenced task structure requiring manual requirement improve- ment prior to AI use, (i) explicit grounding in the INVEST quality framework, (i) an iterative refinement loop with explicit approval actions, (iv) multi-agent personas represent- ing diverse stakeholder perspectives, and (v) mandatory peer review requiring justification of agreement or disagreement. These mechanisms were aligned with three intended learning processes. Structured analysis was supported through require- ments quality evaluation criteria and peer-review. Reflection was embedded via comparison tasks, disagreement justifi- cation, and post-task reflection prompts. Iterative refinement was encouraged through repeated AI-assisted updates and selective approval requirements. Together, these defined the intended pedagogical role of the AI tool as a support for analytical reasoning and evaluative judgment rather than as an automation mechanism. Table I summarizes key interaction metrics describing stu- dentsâ engagement with the AI tool during the assignment. The tool generated a median of 16.5 user stories, of which a median of 9.5 stories were approved by students, corre- sponding to a median approval rate of 56%. This pattern indicates selective filtering rather than general acceptance of AI-generated outputs. Students completed a median of 1.5 refinement rounds, indicating iterative engagement with the generated requirements. Additionally, 24% of students (n = 17) explicitly reported editing AI outputs, providing further evidence of active refinement behavior. TABLE I SUMMARY OF STUDENT INTERACTION WITH THE AI TOOL (N=72) MetricValue Median AI-generated stories per student16.5 Median AI-approved stories per student9.5 Median approval rate56% Median refinement rounds1.5 Students who edited AI outputs24% Figure 3 illustrates the distribution of time spent using the tool. 52 students (72%) reported 15â60 minutes of use, aligning with the expected workload. A smaller group of 17 students (24%) spent less than 15 minutes, suggesting more surface- level interaction, wheras 3 students spent more than 60 min- utes, often associated with experimentation or troubleshooting described in the reflection responses. < 5 min 5â15 min 15â30 min30â60 min > 60 min 0 10 20 30 40 1 16 31 21 3 Time spent using the AI tool Number of students Figure 3. Distribution of time spent using the AI tool during Assignment 4 (N=72). Figure 4 presents the refinement frequency. Overall, 75% of students modified at least one AI-generated user story, with 29 students performing 1â2 refinement rounds and 25 students performing 3 or more. Only 18 students made no refinements. Together with the 56% approval rate of the AI tool generated user stories, these results indicatethat most students engaged in evaluative selection and modification rather than passive adoption. 5.2 RQ2 - Learning Impact on Requirements Quality Assess- ment Learning impact was assessed by comparing studentsâ eval- uations of four user stories (US1âUS4) before and after the 01â23â5 > 5 0 10 20 30 18 29 19 6 Number of refinements made Number of students Figure 4. Distribution of the number of refinements made by students during Assignment 4 (N=72). AI-integrated assignment. For each INVEST dimension, agree- ment with instructor-defined reference values was calculated, and alignment changes (â = p after â p pre ) were computed. Table I shows that alignment changes are quality dimension- and story-specific rather than uniformly positive. The most consistent improvements appear for Valuable, which is positive in US2 to US4, and Testable, which is positive in US1 to US3, including a notable increase of 0.108 in US1. On the other hand, Small, Independent, and Estimable exhibit mixed patterns across stories. Negotiable shows non-positive change across all four stories, with the largest decrease in US2 at -0.292. TABLE I AGREEMENT CHANGE IN INVEST EVALUATION (â =p after âp pre ) INVEST dimensionâ US1 â US2 â US3 â US4 Independent0.000-0.1530.194-0.014 Negotiable0.000-0.292-0.042-0.127 Valuable-0.0410.0140.0140.042 Estimable-0.0810.0690.028-0.113 Small-0.0140.1250.028-0.155 Testable0.1080.0140.028-0.014 The contrast between US3 and US4 provides additional in- sight. US3, which violated multiple INVEST dimensions in the reference analysis, shows positive alignment changes across most dimensions, indicating increased convergence toward the reference evaluation. In contrast, US4, which did not violate any dimension in the reference evaluation, exhibits limited improvement and slight divergence in some dimensions. Given the absence of clear flaws in US4, these shifts reflect greater variability in studentsâ evaluations. Figure 5 complements these findings by presenting studentsâ perceived AI tool support of understanding across the IN- VEST dimensions. The Likert distributions are predominantly positive, with roughly two-thirds to three-quarters of students agreeing that the AI tool supported improvements in most dimensions, including Negotiable. A divergence arises between the measured alignment and the perceived improvement. While Valuable and Testable display both positive alignment shifts and strong perceived support, Negotiable shows high perceived support despite declining alignment with the reference evaluation. This pattern indicates that perceived improvement does not necessarily correspond to convergence with the instructor-defined interpretation of specific quality attributes. Testable Small Estimable Valuable Negotiable Independent 0 20 40 60 Number of students (N=72) 1 Strongly disagree2 Disagree3 Neutral4 Agree5 Strongly agree Figure 5. Distribution of student Likert-scale responses for perceived support across INVEST dimensions. 5.3 RQ3 - Student Perceptions: Usefulness, Trust, Critical Thinking Building on the learning effects reported in RQ2, we next examine how students perceived the AI tool and the extent to which they critically evaluated its outputs, using Likert-scale responses and thematic coding of open-ended reflections. Figure 6 presents distributions for studentsâ perceived usabil- ity, critical-thinking support, and confidence in AI outputs. Perceived usability was moderately positive, with 36 students selecting agreement categories and 16 expressing negative per- ceptions, indicating variability in user experience. Perceived support for critical thinking received the highest ratings, with 52 students selecting agreement categories. Confidence in AI- refined outputs was also generally positive, with 40 agreement responses. These quantitative patterns suggest that students largely viewed the tool as cognitively supportive, though not uniformly smooth in interaction. Table I summarizes themes identified from open-ended re- sponses, with each theme counted once per student. Usefulness was most frequently associated with idea gen- eration and increased specificity. 47 students described the tool as helpful for generating or structuring ideas, 32 re- ported increased detail in their requirements, and 30 indi- cated heightened awareness of missing INVEST attributes. Thirteen explicitly mentioned support for defining acceptance criteria. One student summarized this perceived benefit: âThe AI-generated stories provided structured language, clearer acceptance criteria, and generally better adherence to the INVEST frameworkâ. These perceptions align with the high Ease of use Critical thinking Confidence in AI outputs 0 20 40 60 Number of students (N=72) 1 Strongly disagree2 Disagree3 Neutral4 Agree5 Strongly agree Figure 6. Distribution of student Likert-scale responses for perceived usability, critical-thinking support, and confidence in AI outputs). TABLE I REFLECTION THEMES (N=72) Construct and Representative AspectsStudents (n) Usefulness Idea generation / helpful suggestions47 Added detail / increased specificity32 Increased INVEST awareness30 Acceptance-criteria support13 Confidence in AI outputs Explicit distrust toward AI outputs4 Active verification against INVEST14 Explicitly challenged or rejected outputs10 Critical evaluation and refinement Edited for improved testability21 Edited for improved value alignment17 Edited for smaller, manageable stories13 Edited for independence or negotiability16 Perceived limitations Usability / interface issues22 Context mismatch or vague outputs13 Overly large or duplicated requirements8 Likert ratings for improvements in Valuable, Independent, and Small dimensions. Confidence in AI outputs was generally positive but condi- tional. Explicit distrust was rare (n = 4). However, trust was not unconditional: 14 students reported actively verifying outputs against the INVEST framework, and 10 explicitly rejected or questioned AI-generated requirements. As one student noted: âOccasionally, the AI-generated stories included generic lan- guage that was not specific to drone delivery ... I had to manually refine these outputs to make them fully relevant to the project scenarioâ. Such responses indicate cautious reliance, where AI suggestions were treated as drafts requiring further validation. Critical thinking and evaluation was demonstrated by ev- idence of active refinement of requirements generated by the tool. 21 students reported modifying outputs to improve testability, 17 strengthened value alignment, and 13 reduced overly broad requirements. These behaviors correspond with the strong Likert ratings for critical-thinking support (52 agree- ment responses). Cross-tabulation further indicates alignment between editing and perceived improvement; for example, 15 of 21 students who edited for testability (71%) agreed that the tool improved testability, compared to 47 of 72 students (65%) in the overall sample. One reflection illustrates this process: My edits were mainly focused on making the stories more Testable and Valuable by adding measurable success criteria ... I also improved Independence by narrowing broad stories ... and removed vague terms like âusefulâ or âengagingâ. Usability and contextual mismatch were also reported. 22 stu- dents reported interface or workflow issues, and 13 described vague or contextually inappropriate outputs. 8 reported overly broad or duplicated requirements. Despite these limitations, distrust remained limited, as one student observed: âThe quality of answers was generally okay, but the tool tended for some reason to favour the Product owner ... I had to run the tool several times (after the first attempt) before developer requirements started appearingâ. Such statements indicate frustration with tool behavior rather than rejection of its conceptual value. Overall, students perceived the AI tool primarily as a scaffold for structured reflection and quality-oriented refinement rather than as an automated solution generator. Their responses indicate conditional trust, active verification, and dimension- specific engagement with requirements quality criteria, even when usability challenges were present. 6. DISCUSSION This study examined how a multi-agent AI tool can be ped- agogically integrated into a REQ course, through a TPACK- guided pedagogical design, and how such integration shapes student interaction, learning outcomes, and evaluative engage- ment. 6.1 Pedagogical Design Shapes AI Use The study shows that studentsâ AI use was shaped primarily by instructional design rather than by technological capability alone. Although the multi-agent AI tool enabled requirements generation and refinement, students demonstrated selective approval (median approval rate: 56%), iterative revisions, and explicit verification against the INVEST framework. These indicate evaluative filtering rather than automation-oriented acceptance. From a TPACK perspective [9], technological affordances were enacted within a pedagogically structured workflow, i.e. manual-first revision, approval, peer review, and reflection, and anchored in the requirements quality framework. This alignment of technological, pedagogical, and content knowl- edge preserved student responsibility for quality judgment and positioned AI-generated requirements as artifacts for analysis rather than final solutions. Where prior research has observed that AI integration in SE education often remains exploratory or insufficiently structured [6][7], the present findings provide empirical evidence that deliberate pedagogical orchestration can meaningfully shape AI enactment. Rather than evaluating the AI tool in isolation, this study demonstrates how workflow design mediates the relationship between AI capability and student behavior. In the context of REE where analytical reasoning about requirements quality is essential, these results imply that effective AI integration depends less on tool sophistication than on instruc- tional mechanisms that structure comparison, justification, and iterative refinement. 6.2 Learning Impact The alignment changes in requirements quality assessment re- veal differentiated learning effects across INVEST dimensions rather than uniform improvement. Notably, the results reveal a distinction between structurally explicit quality attributes such as Testable, Valuable and interpretive attributes such as Negotiable. Structurally explicit attributes lend themselves to clearer textual articulation and measurable criteria, whereas interpretive attributes require contextual reasoning about stake- holder flexibility and solution openness. In particular, the divergence observed for Negotiable suggests that students may have developed increased awareness of the attribute without fully converging toward the instructor-defined reference anal- ysis. This divergence does not necessarily indicate regression. Instead, it may reflect increased attention to interpretive ambi- guity, where students move from implicit assumptions toward more varied evaluative interpretations. In AI-supported con- texts, plausible textual formulations may amplify awareness of quality dimensions without ensuring convergence toward a single normative interpretation. These findings indicate that AI tools may more readily support structurally explicit quality attributes than interpretive ones requiring contextual judgment. From a TPACK perspective, this emphasizes that technological affordances can stimulate reflection, but sustained alignment with the requirements quality criteria depends on targeted pedagogical mediation. These findings are consistent with prior research showing that LLMs are effective at enhancing textual structure and clarity, yet less consistent in supporting nuanced conceptual reasoning and context-sensitive judgment [8]. In educational contexts, LLM-generated outputs often improve articulation and orga- nization but do not automatically ensure convergence with deeper evaluative criteria without instructional scaffolding. For practice, this suggests that interpretive quality dimensions, such as negotiability or independence, require explicit exem- plars, structured comparison tasks, and guided evaluative cri- teria. Future research should examine how tool configuration and scaffolding strategies can better support alignment across complex quality attributes in REE. 6.3 Student Perceptions and Conditional Trust Studentsâ perceptions of the AI tool reveal the usefulness of the tool and conditional trust. While students generally reported positive perceptions, particularly for idea generation, structuring requirements, and increasing awareness of qual- ity criteria, qualitative evidence revealed conditional trust: students verified AI outputs against INVEST criteria, edited unclear suggestions, and rejected contextually inappropriate requirements. AI-generated content was thus treated as provi- sional drafts rather than authoritative solutions. The discrepancy between perceived improvement and mea- sured alignment further sharpens this interpretation. Although students frequently reported that the tool enhanced their un- derstanding of quality attributes, convergence with instructor- defined evaluative standards was uneven. This suggests that AI integration may initially enhance awareness and reflective attention before producing consistent normative convergence. Rather than demonstrating blind reliance, students appeared to treat AI suggestions as drafts requiring validation. This indicates that structured integration can promote critical en- gagement rather than over-reliance. From a pedagogical standpoint, these findings highlight that AI can function as a scaffold for reflective practice when integrated into workflows that require explicit verification, justification, and iterative refinement. Rather than displacing analytical reasoning, structured AI integration can reinforce it. Future research should examine whether such guided engagement produces sustained development of analytical competence and judgment in RE contexts beyond short-term assignment performance. 6.4 Implications for RE and SE Education The findings collectively suggest that effective AI integration in RE and SE education requires deliberate orchestration rather than permissive tool adoption. This study suggests several design principles for AI integration in RE and SE education: ⢠Sequence human reasoning before AI use Manual- first workflows preserve ownership of analysis and reduce over-reliance and automation-oriented behavior. ⢠Structure AI use around contrastive analysis AI- generated alternatives are most effective when used as stimuli for comparison rather than as final outputs. ⢠Anchor AI use in explicit quality frameworks Stan- dards such as INVEST provide evaluative reference points for assessing AI outputs. ⢠Incorporate peer review and reflection Peer review and reflection promote active evaluation and mitigate over- reliance. Importantly, these principles indicate that AI integration should be differentiated according to the epistemic nature of the target learning objectives. For structurally explicit competencies, AI can serve as a productive accelerator. For interpretive and context-sensitive competencies, stronger scaf- folding and explicit exemplars remain necessary. 6.5 Threats to Validity We structure potential threats following the validity framework [22], considering internal, construct, conclusion, and external validity. Internal validity concerns whether observed effects can be attributed to the AI-integrated assignment rather than alter- native factors. The absence of a control group limits causal attribution, as improvements between the pre-assignment and mastery exercise may partly reflect course progression or repeated exposure to the INVEST framework in lectures and other assignments. Furthermore, students completed assign- ments independently, and full adherence to the prescribed as- signment workflow could not be verified. Triangulation across assignment submissions, interaction metrics, and reflective responses partially mitigates these concerns. Construct validity may be affected by operationalizing learning impact as agreement with instructor-defined reference analy- ses. While this provides a consistent benchmark, it may not capture alternative defensible interpretations or deeper con- ceptual understanding, particularly for interpretive dimensions such as Negotiable. Additionally, students may emphasize different quality aspects in their evaluations, and the use of only four user stories limits the breadth of assessment. Likert- scale responses capture perceived support rather than objec- tive learning gains; combining the questionnaire data with assignment-based evaluation and thematic coding of open- ended reflections helps mitigate this limitation. Conclusion validity is constrained by reliance on descrip- tive statistics and agreement proportions, which support ex- ploratory interpretation but limit strong causal inference. Moreover, the AI tool was developed within the same research environment as the course, potentially introducing contextual bias despite predefined evaluation criteria and structured cod- ing procedures. External validity is limited by the studyâs conduct within a single RE course using a research prototype. Findings may not generalize to other institutional contexts, undergradu- ate cohorts, commercial AI systems, or team-based settings. Nonetheless, the course structure reflects common REE prac- tices, and the high voluntary participation rate and multi- source data provide a meaningful basis for interpretation within similar contexts. 7. CONCLUSION Integrating multi-agent AI tools into REE presents both op- portunities and pedagogical challenges. This study examined a TPACK guided integration of a multi-agent AI tool into an RE course, focusing on how assignment design shaped studentsâ AI enactment and requirements quality reasoning. The findings demonstrate that AI can support studentsâ an- alytical engagement when it is embedded within a carefully designed pedagogical structure. Students did not treat the tool as an automated generator but instead selectively approved, refined, and evaluated AI-generated requirements. These results suggest that the effective AI integration depends less on a toolâs technical sophistication and more on how it is incorporated into assignment design. Sequencing manual work before AI interaction, embedding structured comparison and peer review, and anchoring evaluation in formal quality frameworks appear to be key factors in promoting critical thinking rather than passive reliance on automation. However, usability limitations and interpretive ambiguity in certain qual- ity dimensions highlight areas that require further instructional refinement. Future research should investigate the long-term development of analytical competence in AI-integrated RE environments and explore how orchestration strategies can be adapted across diverse institutional contexts and AI configurations. By refin- ing pedagogically grounded integration strategies, educators can leverage AI as a scaffold for reflective practice while preserving the analytical rigor central to SE education. REFERENCES [1] H. Cheng, J. H. Husen, Y. Lu, T. Racharak, N. Yoshioka, N. Ubayashi, and H. Washizaki, âGenerative ai for re- quirements engineering: A systematic literature review,â Software: Practice and Experience, 2025. [2] Z. Zhang, M. Rayhan, T. Herda, M. Goisauf, and P. Abrahamsson, âLlm-based agents for automating the enhancement of user story quality: An early report,â in International conference on agile software development. Springer, 2024, p. 117â126. [3] M. A. Sami, Z. Zhang, M. Waseem, K.-K. Kemell, Z. Rasheed, T. Herda, M. T. Hasan, J. Rasku, and P. Abrahamsson, âA multi-agent llm system for au- tomated requirements analysis: a study on user story generation and prioritization,â in Euromicro Conference on Software Engineering and Advanced Applications. Springer, 2025, p. 178â187. [4] E. Parra and S. Willingham, âTowards implementing and evaluating ai-assisted pull requests in software en- gineering education,â in 2025 IEEE/ACM 37th Interna- tional Conference on Software Engineering Education and Training (CSEE&T). IEEE, 2025, p. 13â18. [5] M. A. Sami, Z. Zhang, M. Waseem, K.-K. Kemell, Z. Rasheed, T. Herda, and P. Abrahamsson, âBridging humans and llms: Investigating human-ai collaboration in multi-agent requirements analysis for organizational ai adoption,â e-Informatica Software Engineering Journal, vol. 20, no. 1, p. 260103, 2026. [6] S. Guardado, R. Parveen, Z. Zhang, M. Rayhan, and N. Tripathi, âStudentsâ perceptions of the use of LLMs in requirements engineering education: A cross-university empirical study,â in Proceedings of the 33rd IEEE In- ternational Requirements Engineering Conference (RE). IEEE, 2025, p. 130â141. [7] C. K. Sah, L. Xiaoli, M. M. Islam, and M. K. Islam, âNavigating the AI frontier: A critical literature review on integrating artificial intelligence into software engi- neering education,â in Proceedings of the 36th Interna- tional Conference on Software Engineering Education and Training (CSEE&T). IEEE, 2024, p. 1â5. [8] S. Filippi and B. Motyl, âLarge language models (llms) in engineering education: A systematic review and sugges- tions for practical adoption,â Information, vol. 15, no. 6, p. 345, 2024. [9] M. Koehler and P. Mishra, âWhat is technological ped- agogical content knowledge (tpack)?â Contemporary is- sues in technology and teacher education, vol. 9, no. 1, p. 60â70, 2009. [10] âWake,b:Investingoodstories,and smarttasks,âhttps://xp123.com/articles/ invest-in-good-stories-and-smart-tasks/. [11] âIso/iec/ieee international standard - systems and soft- ware engineering â life cycle processes ârequirements engineering,â ISO/IEC/IEEE 29148:2011(E), p. 1â94, 2011. [12] S. Tiwari and S. S. Rathore, âLeveraging llms for re- quirements engineering education: How to approach?â in 2025 IEEE 33rd International Requirements Engineering Conference (RE). IEEE, 2025, p. 458â466. [13] B. Wei, âRequirements are all you need: From require- ments to code with llms,â in 2024 IEEE 32nd In- ternational Requirements Engineering Conference (RE). IEEE, 2024, p. 416â422. [14] M. Vierhauser, I. Groher, T. Antensteiner, and C. Sauer- wein, âTowards integrating emerging ai applications in se education,â in 2024 36th International Confer- ence on Software Engineering Education and Training (CSEE&T). IEEE, 2024, p. 1â5. [15] S. M. Vidalis, R. Subramanian, and F. T. Najafi, âRevo- lutionizing engineering education: The impact of ai tools on student learning,â in 2024 ASEE Annual Conference & Exposition, 2024. [16] V. D. Kirova, C. S. Ku, J. R. Laracy, and T. J. Marlowe, âSoftware engineering education must adapt and evolve for an llm environment,â in Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1, 2024, p. 666â672. [17] O. Levy, I. Dikman, N. Levy, and M. Winokur, âWork in progress: Ai-powered engineering-bridging theory and practice,â in 2025 IEEE Engineering Education World Conference (EDUNINE). IEEE, 2025, p. 1â4. [18] A. M. Al-Abdullatif, âModeling teachersâ acceptance of generative artificial intelligence use in higher education: The role of ai literacy, intelligent tpack, and perceived trust,â Education Sciences, vol. 14, no. 11, p. 1209, 2024. [19] L. Eyal, âDeveloping and validating an ai-tpack assess- ment framework: Enhancing teacher educatorsâ profes- sional practice through authentic artifacts,â Education Sciences, vol. 15, no. 11, p. 1452, 2025. [20] V. Braun and V. Clarke, âUsing thematic analysis in psychology,â Qualitative Research in Psychology, vol. 3, no. 2, p. 77â101, 2006. [21] J. W. Creswell and V. L. Plano Clark, Designing and Conducting Mixed Methods Research, 3rd ed.SAGE, 2018. [22] C. Wohlin, P. Runeson, M. H Ě ost, M. C. Ohlsson, B. Reg- nell, A. Wessl Ě en et al., Experimentation in software engineering. Springer, 2012, vol. 236.