Paper deep dive
A Simulation-Based Method for Testing Collaborative Learning Scaffolds Using LLM-Based Multi-Agent Systems
Han Wua, Lishan Zhang, Chunming Lu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 1:38:43 AM
Summary
This study proposes an LLM-based multi-agent simulation framework using MetaGPT and GPT-4o to model collaborative learning processes. By comparing 'Deep Think before Speak' and 'Direct Speak' scaffolding strategies across classical Chinese poetry tasks, the researchers demonstrate that the 'Deep Think' scaffold improves discourse diversity and interaction depth, aligning with the ICAP framework. The study validates the feasibility of using multi-agent systems to simulate authentic collaborative learning dynamics.
Entities (5)
Relation Signals (3)
MetaGPT â implements â Multi-Agent System
confidence 95% · The simulation system was implemented using the MetaGPT framework
Multi-Agent System â alignswith â ICAP framework
confidence 90% · These findings align with the ICAP framework
Deep Think before Speak â improves â Discourse Diversity
confidence 90% · The introduction of the 'Deep Think before Speak' scaffold significantly improved the agents' discourse diversity
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Traditional research on collaborative learning scaffolding is often time-consuming and resource-heavy, which hinders the rapid iteration and optimization of instructional strategies. LLM-based multi-agent systems have recently emerged as a powerful tool to simulate complex social interactions and provide a novel paradigm for educational research. Objectives: This study proposes an LLM-based multi-agent simulation approach to investigate collaborative learning processes and the effectiveness of instructional scaffolds prior to actual classroom deployment. The research specifically examines the feasibility of simulating group discussions and the alignment of these simulations with established learning science theories. Methods: The simulation system was implemented using the MetaGPT framework and GPT-4o, comprising one teacher agent and five distinct student roles (Leader, Supporter, Expounder, Rebutter, and Summarizer). Two scaffolding strategies, "Deep Think before Speak" and "Direct Speak", were compared across ten classical Chinese poetry appreciation tasks. Evaluation was conducted through discourse analysis of quality and behavior. Results and Conclusions: The introduction of the "Deep Think before Speak" scaffold significantly improved the agents' discourse diversity and interaction depth while notably reducing content repetitiveness. Behavioral analysis showed that the scaffold encouraged more complex interaction patterns, such as reflecting, rebutting, and explaining. These findings align with the ICAP framework, as the scaffold prompted agents to move from simple "Active" participation to "Constructive" and "Interactive" knowledge co-construction. This study demonstrates the feasibility and ecological validity of using LLM-based multi-agent systems to simulate authentic collaborative learning dynamics.
Tags
Links
- Source: https://arxiv.org/abs/2604.11161v1
- Canonical: https://arxiv.org/abs/2604.11161v1
Trouble viewing inline? Open PDF directly â
Full Text
73,912 characters extracted from source content.
Expand or collapse full text
ASimulation-BasedMethodforTestingCollaborative LearningScaffoldsUsingLLM-BasedMulti-AgentSystems HanWu a ,LishanZhang a *,ChunmingLu b a SchoolofEducation,BeijingInstituteofTechnology,Beijing,PeopleâsRepublicof China; b StateKeyLaboratoryofCognitiveNeuroscienceandLearning& IDG/McGovernInstituteforBrainResearch,BeijingNormalUniversity,Beijing, PeopleâsRepublicofChina *Correspondingauthor:LishanZhang,Address:No.5SouthZhongguancunStreet, HaidianDistrict,Beijing,100081,P.R.China. Abstract Background:Traditionalresearchoncollaborativelearningscaffoldingisoften time-consumingandresource-heavy,whichhinderstherapiditerationand optimizationofinstructionalstrategies.LLM-basedmulti-agentsystemshaverecently emergedasapowerfultooltosimulatecomplexsocialinteractionsandprovidea novelparadigmforeducationalresearch. Objectives:ThisstudyproposesanLLM-basedmulti-agentsimulationapproachto investigatecollaborativelearningprocessesandtheeffectivenessofinstructional scaffoldspriortoactualclassroomdeployment.Theresearchspecificallyexamines thefeasibilityofsimulatinggroupdiscussionsandthealignmentofthesesimulations withestablishedlearningsciencetheories. Methods:ThesimulationsystemwasimplementedusingtheMetaGPTframework andGPT-4o,comprisingoneteacheragentandfivedistinctstudentroles(Leader, Supporter,Expounder,Rebutter,andSummarizer).Twoscaffoldingstrategiesâ â DeepThinkbeforeSpeak â and â DirectSpeak ââ werecomparedacrossten classicalChinesepoetryappreciationtasks.Evaluationwasconductedthrough discourseanalysisofqualityandbehavior. ResultsandConclusions:TheintroductionoftheâDeepThinkbeforeSpeakâ scaffoldsignificantlyimprovedtheagents'discoursediversityandinteractiondepth whilenotablyreducingcontentrepetitiveness.Behavioralanalysisshowedthatthe scaffoldencouragedmorecomplexinteractionpatterns,suchasreflecting,rebutting, andexplaining.ThesefindingsalignwiththeICAPframework,asthescaffold promptedagentstomovefromsimple"Active"participationto"Constructive"and "Interactive"knowledgeco-construction.Thisstudydemonstratesthefeasibilityand ecologicalvalidityofusingLLM-basedmulti-agentsystemstosimulateauthentic collaborativelearningdynamics. Keywords Largelanguagemodel,Multi-agentsystems,Collaborativelearning,GenerativeAI, InstructionalScaffolds 1.Introduction Collaborativelearningiswidelyknownasaneffectivemethodtoimprovelearning efficiency.Insuchlearningcontext,studentsasgroupsareexpectedtoco-construct uponeachother,sothateverygroupmembercangetdeepunderstandingstowardsthe learningobjects.However,itisnotstraightforwardtocomeupefficientscaffoldingto leadgroupco-constructions.Itoftenrequiresextensivestudiesfordifferentlearning context.Traditionally,researchersraisedsomebrilliantideas(e.g.awaytoscaffolda discussionforagroupofstudents)andtestedtheseideasbyrunningexperimentswith humanparticipants,andhopefullygetsignificantimprovementsinteachingor learningefficiencies.Theresearchersmayimprovetheexistingdesigns,adjust differentexperimentvariables(e.g.differentversionsofonetypeofscaffolding),and runmoreroundsofexperimentstogetdeeperunderstandingregardingthelearning circumstancetobestudied.Thiskindofstudiesmaylastseveralmonthsandinvolve hundredsofhumanparticipantsifeverythinggetssmooth.Itisverycommonthat researchersfindascaffoldingunsuccessfulafterrunningtheactualexperiments.We couldnoteliminatesuchpossibilitiesforsure,butcouldwesomehowpredictthe resultsgivenadesignedscaffolding,soastoimprovethesuccessfulrate?Large languagemodels(LLMs)drivenagenttechnologycanprovideanoptiontothis question. Recently,researchershavebeguntoleveragethepowerfulgenerativeand reasoningcapabilitiesofLLMstoconstructmulti-agentsystemsthatsimulate complexandrealisticsocialinteractionprocesses(J.Zhangetal.,2025). Sophisticated,model-drivennarrativescaneffectivelycatalyzecooperativebehaviors withindigitalsocieties,particularlywhencombinedwithstrategicnetwork positioning(DeCurtĂČ&DeZarzĂ ,2025).Byleveragingpromptengineeringand fine-tuningtechnologies,systemssuchasS3enableagentstoperceiveenvironmental informationandauthenticallymimichumanbehavioralpatterns(Gaoetal.,2023). Evaluationsoftheseagent-basedmodelsshowthatthesimulatedgroup-level phenomenaarehighlyconsistentwithempiricalreal-worlddata(Gurcan,2024). However,LLMsandmultiagenttechnologiesaremerelyusedinsimulating interactionsinthefieldofeducation.Moreover,asfarasweknow,noonehas establishedsimulatedcollaborativelearningenvironmentstoexploretheeffectiveness ofdifferenttypesofinstructionalscaffolding,whichiswidelyrecognizedasacritical mechanismforpromotingdeepcollaborativelearning(VerdĂș&Sanuy,2014).By providingastructuralframework,scaffoldinghelpslearnersprogressfromlower levelsofengagementtohighercognitivelevels(Kollar,Wecker,&Fischer,2018).To promotelearners â deeplearning,variousscaffoldingstrategieshavebeendeployedto bridgethegapbetweenconcreteandabstractthinking(Pierce&Gilles,2021).Aswe notedbefore,makingeffectivescaffoldingtypicallyrequirestime-intensive Design-BasedResearch(DBR)orexpertpedagogicaldesign(Hsu,Lai,&Hsu,2015). Therefore,weproposedasimulatedcollaborativelearningsystembasedonLLMsand multi-agenttechnologies,andevaluatedtheeffectivenessofdifferenttypesof scaffoldingoncollaborativelearninginthesimulatedsystem. Oursimulatedsystemincludesoneteacheragentandseveralstudentagents.The agentssolvetasksbydiscussionjustlikehumanstudentsintypicalcollaborative learningtasks.Inspecific,weimplementedthesimulatedsystembasedonMetaGPT framework,employedthepromptengineeringtoconstructtwotypesofmetacognitive scaffolding,andassumedthatthestudentagentsstrictlyfollowedthetaught scaffoldingtoconducttheirdiscussion.Thefirsttypeofscaffoldingis â DeepThink beforeSpeak â andthesecondtypeofscaffoldingis â DirectSpeak â .Thefirsttype ofscaffoldingshouldleadmorehigh-qualityinteractionsthanthesecondone,because existingstudiesshowedthatindividual â sinternalself-explanationorknowledge construction(i.e.deepthink)canincreaseinteractionqualitiesduringcollaborative learning(Chi,DeLeeuw,etal.,1994;Chi&Wylie,2014;Ploetzneretal.,1999).By constructingthesimulatedcollaborativelearningsystemsandrunningexperiments withthesystem,weaimtoanswerthefollowingtworesearchquestions: (1)Howcanthegroupdiscussionincollaborativelearningbesimulatedbasedon LLMsandmulti-agenttechnologiessuchasMetaGPT? (2)Arethefindingsinthesimulatedcollaborativelearningsystemconsistentwiththe existingqualitativestudieswithhumanparticipants? Therestofthepaperfirstbrieflyreviewstheexistingliteratureandtheoretical foundationsinrelatedfieldssuchasComputerassistedCollaborativeLearning (CSCL),Multi-AgentSystems(MAS),andAIforSocialScience.Then,wedetailthe designmethodsandevaluationframeworkforsimulatedcollaborativelearning experiments.Next,wereportandanalyzetheresultsfromtheperspectivesof discoursequalityanddiscoursebehavior.Finally,weconcludewithremarks. 2.Relatedwork 2.1Computerassistedcollaborativelearning Collaborativelearningisdefinedasasituationinwhichtwoormorepeople attempttolearntogetherthroughcoordinated,synchronousactivitiesaimedat constructingandmaintainingasharedconceptionofaproblem(Dillenbourg,1999; Roschelle&Teasley,1995).Duringthisprocess,groupmembersworktogether throughcollaboration,in-depthdiscussion,andinformationexchangetocomplete learningtasks,sharelearningoutcomes,andultimatelyachievecommonlearning objectives(Miyake,2007).Theessenceofcollaborativelearningliesincooperation andcommunication,allowingstudentstodeeplyanalyzeandco-constructknowledge throughdialogue,discussion,anddebate.However,traditionalcollaborativelearning modelsoftenfacenumerouschallenges,suchaslackofcoordinationintaskdivision amonggroupmembers,difficultyinsynchronizingcollaborationtime,andsome studentsavoidingordisengagingfromdiscussions(Shea,1995).Toovercomethese obstacles,Computer-SupportedCollaborativeLearning(CSCL)emerged,rapidly becomingafocalpointforresearchandapplicationamongscholars.Forinstance, JĂ€rvelĂ€,S.andHadwin(2013)proposedtheuseofCSCLenvironmentsfromthe perspectiveofsharedregulation,emphasizingtheroleoftechnologyinsupporting learningregulationwithincomplexcollaborativetasks.TheimpactofCSCLon collaborativelearningandthebroaderfieldofeducationhasbeenempiricallytested. Jeongetal.(2019)conductedameta-analysisbasedonstudiesfrom2005to2014and foundanoveralleffectsizeof0.51forCSCLinSTEMfields.Thiseffectsizeis consideredmoderatebuthighlysignificantineducationalresearch,strongly supportingtheoveralladvantagesofCSCLinSTEMeducation. Withintherecentyears,AIagentsbasedonlargelanguagemodelshaveemerged asanewmediumtoofferpersonalizedanddynamicregulatorysupportinlearning environments.Thisdevelopmenthasfurtherexpandedtheresearchboundariesof traditionalCSCL.SeveralstudieshaveconfirmedthesignificantadvantagesofAI agentsinenhancingstudents â learningexperience.Forinstance,Wangetal.(2025) developedanAI-agenttoenhancestudents'programminglearning.Comparedto traditionalCSCLenvironments,thisagent-supportedcollaborativelearning significantlyimprovedstudents'self-efficacyandacademicperformancewhile effectivelyreducingcognitiveloadatthesametime.AIagentalsoshowedits effectivenessinsecondlanguagelearning,wherestudentsgainedhighersocialand cognitivepresenceduringinteractions,andachievedbetterlearningoutcomes(X. Wang,Pang,Wallace,Wang,&Chen,2024).Besidesactasfacilitators,AIagents canalsoactasteachableagent.Studentscanthenimprovetheirownknowledgeand abilitiesbyteachingtheAIagenthowtoprogram(Chen,Wei,Le,&Zhang,2024). Insummary,theresearchdevelopmentofcollaborativelearninghasdemonstrated thattheintelligenttechnologyisbecomingmoreandmoreimportantinthisfield. However,tothebestofourknowledge,existingstudiesonlyconsideredusing intelligenttechnologyorAIagenttofacilitatecollaborativelearninginsteadof simulatingtheprocessofcollaborativelearning.Toaddressthisresearchgap,this studyusedmulti-agenttechnologytosimulatecollaborativelearningprocess. 2.2LargeLanguageModelsandAIagent LLMshasprovidedkeytechnologicalfoundationsforthedevelopmentofAIagent technologiesincludingbothsingle-agentandmulti-agentones.Withveryfew prompts,LLMscanenableAIagentstoanalyzecomplexdataandprovide contextuallyrelevantresponses(Ărpek,Tural,&Destan,2024).Agentsdrivenby Multi-modalLargeModelslikeGPT-4ocanunderstandandgeneratedifferenttypes ofinformationliketext,voice,image,videoandetc(Islam&Moushi,2025).An increasingnumberofstudieshavebeguntoapplytheseLLMs-drivenagentsinthe fieldofeducation.Theseagentscaninteractwithstudents,withotheragents,orwith both.Studentscangetlearningbenefitsbyinteractingwiththeagentsandobserving theagents â interactions. Basedonthenumberofagentsinvolved,applicationsintheeducationaldomain canbedividedintosingle-agentandmulti-agentsystems.Single-agentsystemscanbe easilyusedtoconstructconversationalbasedintelligenttutoringsystemslike AutoTuor(Graesseretal.,2004),whichcanguidestudentstoengageindeepthinking throughnaturaldialogues.Suchsystemsusedtorequirealotofeffortsinbuildingthe appropriatedialogscriptsbeforetheadvanceofLLMs.Recently,researchersare morefocusonexploringtheeducationalapplicationsofmulti-agentsystems,which arecapableofdecomposingcomplextasksintosmallerandmoremanageableones thataresolvedbycorrespondingdedicatedagents(Jiangetal.,2024).Forexample, Mohamedhenetal.(2024)builtamulti-agentsystembasedonconvolutionalneural networksandmultilayerperceptrons,applyingtheFelder-Silvermanmodeltofirst accuratelyidentifystudents'learningstyles,thenprovidepersonalizedcontent recommendations.Viswanathanetal.(2022)designedseveralteachingagentswith differentfunctionsforanonlineeducationsystem,forminganadaptivesystem capableofprovidingmeaningfulandtargetedcoursecontenttolearners.Zhangetal. (2025)developedanovelintelligentclassroomframework,SimClass,andallowed reallearnerstointeractwiththissimulateddynamicteachingenvironmentacross multiplecourses,andshowedthatthemulti-agentarchitecturesignificantlyenhanced studentengagementandlearningoutcomes. IthasnodoubtthatLLMsandAIagenttechnologieshavebeenquicklyapplied ineducationafterthenotableachievementofGPT,andnowresearchershavestarted tobuildeducationalapplicationswithmulti-agenttechnology.Inthisstudy,we furtherexploredhowwecouldusethemulti-agenttechnologytosimulate collaborativelearning. 2.3AIforsocialscience Thegrowinghuman-likereasoningandcognitiveabilitiesofLLMsandAIagents havegainedsignificantattentionfromsocialscienceresearchers,promptingthemto explorehowthistechnologycanbeleveragedtoadjustorevenreshaperesearch practicesinthesocialsciences(Grossmannetal.,2023).Xuetal.(2024)assessedthe relationshipbetweenAIandsocialscience,highlightingthedualroleofLLMsin research:aspowerfulresearchassistantsandasreliableexperimentalagents simulatinghumanbehavior. TheapplicationofLLMsasresearchassistantshasbeenparticularlynotable, greatlyenhancingresearchefficiencyandtheprecisionofdataprocessing.Withtheir robustnaturallanguageunderstandingcapabilities,LLMscanautomatethethematic analysisandcodingofcomplexqualitativedata(Zhang,Wu,Duan,&Du,2025), significantlyimprovingtheefficiencyofdataanalysisandinter-coderreliability. BrydaandSadowski(2024)introducedasemi-supervisedcodingapproachthat appliesAIalgorithmstothecodingandthematicanalysisoffree-textinterviews, effectivelyenhancingtheprecisionofqualitativedataanalysisandreasoning efficiency.Farjametal.(2024)systematicallyguidedresearchersinusingLLMsfor automatedcodingincontentanalysis,emphasizingthenumerousadvantagesofthis approachinscalability,multilingualcoverage,andcost-effectiveness.LLMscanalso playacrucialroleingeneratingresearchhypothesesandconductingsystematic literaturereviews(Bazgir&Zhang,2025). Intermsofhumanbehaviorsimulation,LLMsandAIagenttechnologyareused tobuildplatformthatcanrunexperimentsforexploringandpredictinghuman behaviorsintherealworld.Aheretal.(2023)arguethatLLMsdrivenagentsarethe mostreliableones,capableoftestingtheoreticalhypothesesunderlarge-scaleand reproducibleconditions.Forexample,Parketal.(2022)designedasocialsimulation techniquebasedonLLMstomodelandgenerateresponsesfromcommunityusers,as wellassimulatesocialinteractions.Thisapproachhelpssocialcomputingdesigners conductbetteruserbehavioranalysisoridentifypotentialmarginalsituationsthat couldleadtosocietalcollapse.Horton(2023)analyzedthreeclassicbehavioral economicsstudiesthatusedGPT-3.5astheexperimentalsubject.Theirfindings showedthatthesimulationexperimentsqualitativelyreproducedconclusionssimilar tothosefromrealhumanexperiments,providingstrongsupportforthefeasibilityof usingagentsforsocialsciencesimulations.Thissimulationapproachislow-cost,can accommodatearbitrarysamplesizes,andallowsforthetestingofvariouspromptor parameterconfigurations,whilealsoavoidingethicalconcernsrelatedtohuman participants.Guo(2023)enabledGPTtounderstandtherulesofstrategicgame experimentsandperformdecision-makingreasoningwithhuman-likeresponses, uncoveringpotentialpatternsandlogicobservedinthegames.Furthermore,Parketal. (2023)developedaninteractivesandboxenvironmentconsistingof25agents, allowinguserstointeractwiththemandobserveandevaluatetheindividualand emergentbehaviorsproducedbytheagents. AsBail(2024)noted,thisLLMsandAIagent-basedsimulationinfrastructureis notonlyessentialforensuringbroadaccesstohigh-qualityresearchtools,butalso vitalforgainingadeeperunderstandingofthesocialforcesthatguidehuman behavior,particularlyinthecontextofAIadvancements.Inthisstudy,webuiltsuch simulation,andexploreditsusabilityinthecontextofcollaborativelearning. 3.Collaborativelearningsimulationdesign 3.1Thesimulatedcollaborativelearningmechanism ThisstudyusesMetaGPTframeworktoconstructthesimulatedcollaborativelearning system,integratingtheDeepSeek-v3modelforpromptoptimization,withGPT-4o servingasthecoredrivingforcetosimulatecollaborativelearningscenarios. Theoverallsystemarchitecturecomprisesthreecorecomponents:the CommunicationModule,theMemoryModule,andtheAgentModule.Thesemodules collaboratecloselytoestablishahighlyinteractiveandintelligentcollaborative learningenvironment,withthesystemarchitecturediagramshowninFig.1.During operation,theCommunicationModulemanagestheflowofinformationby employingamessageroutingmechanism,thusensuringaccurateinformation exchangebetweentheteacherandstudentagents.TheMemoryModulestoresand provideshistoricaldialogues,whichensuresthatagentsmaketheirdecisionsbasedon thecompletecontextualinformation.Furthermore,theAgentModuleenhances cognitiveprocessingbysettingChainofThought(CoT)andimplementspersonalized behaviorsthroughtheroleparametersettings. Fig.1Architectureofthemulti-agentcollaborativelearningsystem Thesystemoperatesinatask-drivencycle,guidingthediscussionbyfocusingon predefinedscoringpoints.Intheinitialphase,theteacheragentclarifiesthe collaborativetaskandkeyknowledgepoints,givesinitialinstructions,andsetsthe firstroundâsspeakingorderofthestudentagents.Oncereceivingtheinstructions, eachstudentagentreflectsontheinformationbasedontheirpersonaltraitsand previousinteractions,andgeneratesaresponse.Aftereachdiscussionround,the teacheragentevaluatesthecoverageofknowledgepoints,updatestheremaining knowledgepointstobeaddressed,andprovidesfeedbackalongwiththenext speakingorder.Thiscyclerepeatsuntilallscoringknowledgepointsarecovered. Additionally,studentagentisdesignedtointeractwitheachotherbasedon collaborativelearningpractices.Theycanobservethecontributionsofotheragents andadjusttheirbehavioraccordingly.Forinstance,whenoneagentoffersa high-qualityinsight,otheragentsmaychoosetosupport,challenge,orfurther elaborateonthepoint,forminganaturaldiscussionchain.Thenextsectionelaborates thedetaileddesignsoftheagents. 3.2Thedesignsoftheagents Thesystemisdesignedwithtwotypesofagentroles:theteacheragentandthe studentagent,asshowninFig.2Theteacheragentisresponsibleforcommentingon students'discussions,controllingthepaceoftheconversation,impartingkey knowledgepoints,andansweringquestions.Meanwhile,thestudentagentsemulate learnerbehaviors,includingposingquestions,participatingininteractivediscussions, andabsorbingknowledge.Thisdifferentiatedroleassignmentendowstheagentswith human-liketraits,significantlyenhancingtheinteractionexperienceandrealismin onlinecollaborativelearningcontexts. Fig.2Designofteacherandstudentagent Theteacheragentisresponsibleforguidingandmanagingthecollaborative learningprocess,ensuringthecompletionoftasks.Itsmainfunctionsinclude: initialization,thinking,behavioralchoosing,andspeaking.Theinitializationfunction ensuresthattheteacheragentisclearaboutitsidentity,taskobjectives,andgrading criteria.ThedatastructureofitsidentityinformationisshowninTable1.The thinkingfunctionallowstheteacheragenttoassessstudentparticipationandscoring pointsduringdiscussions,soastoanalyzethestudents'performance.Thebehavioral choosingfunctionprovidestheteacheragentwiththreetypesofoperationalchoices: 1.Theabilitytocommentandaskquestionsbasedonthediscussionintheprevious round,andadjustspeakingordersforthenextaround;2.Theabilitytoencourageor guidestudentstoexplorekeypointsofthecollaborativetasks;3.Theabilityto providefeedbackoneachstudent'sstrengthsandareasforimprovementaftermultiple roundsofdiscussion.Thespeakingfunctionintegratestheabovefunctionalmodules andgeneratesteacherspeechcontentwithawordlimitof150characters,takingthe resultsofthinkingastheguide,andexecutingthecorrespondingbehaviorselection. Table1.Theidentityinformationdatastructureoftheteacheragent FieldDatatypeExplanation idintUniqueidentifieroftheteacheragent. namestrThenameoftheteacheragent,usedinconversations. base_definitionstr Includingbasicinformationsuchasprofessionalfield,teaching experience,etc. learning_goalstrCollaborativetasksandobjectivesscoringcriteria. Thestudentagentsimulatesthebehaviorsofrealstudentsandparticipatesinthe collaborativelearningprocess.Itsfunctionsarealsodividedintofoursteps: initialization,thinking,behaviorchoosing,andspeaking.Theinitializationfunction allowsthestudentagentstoclarifyitsidentityinformation,priorknowledge,and specificroleofthepredefinedbehavior.Thedatastructureofstudents â identity informationisshowninTable2.Thethinkingfunctionenablesthestudentagentto generatethoughtsbasedonthetasksassignedbytheteacherandthediscussion history,andtomakebehaviorchoicesbycombiningthegroupmembers'opinionsand itsowninitializationparameters.Thebehavioralchoosingfunctionoffersfour operationalchoices:presentingviewpoints,questioning,raisingissues,and summarizing:1.Oncethestudentagentshaveformedtheirowninsights,theycan proactivelyexpressviewsandsharethoughtsandunderstanding.2.Whentheyhear viewpointsthatdifferfromtheirpriorknowledge,theycanraisequestionstoverify theaccuracyandcompletenessoftheinformation.3.Whentheyencountercontentor topicwhichtheywanttodiscussfurther,theycanaskquestionsandactivelyseek knowledge.4.Atacertainstageofthediscussionorattheendofthediscussion,they alsocansummarize,reviewthekeypointsandresultsofthediscussion,and consolidatetheirlearningoutcomes.Thespeakingfunctionissimilartothatofthe teacheragent.Basedonthethinkingandselectedbehavioralstrategies,thestudent agentsoutputspeechcontentthatconformstotheirrolesettings,withawordcountof about80words. Table2.Theidentityinformationdatastructureofthestudentagent FieldDatatypeExplanation idintUniqueidentifierofthestudentsâagent. namestrThenameofthestudentsâagent,usedinconversations. base_definitionstrIncludingbasicinformationsuchasgradeandmajor,etc. assigned_rolestrThespecificroleforstudentsâagentwithpredefinedbehaviors. 4.Evaluationmethods 4.1Collaborativelearningtasks ThecollaborativelearningtasksusedinthisstudywerederivedfromtheChinese poetryappreciationliteracycourseatauniversityincentralChina.Thesetaskswere basedonactualcollaborativediscussiontasksusedinthecourse,whichwerethen compiledintothecollaborativelearningtasksetfortheappreciationofclassical Chinesepoetryinthisstudy.Toensurethescientificvalidityandreliabilityofthe assessmentofthecollaborativelearningsimulation,LLMswereusedtogeneratefive referencescoringcriteriaforeachtask.Thesecriteriawerereviewedandrevisedby thecourseinstructors.Intheend,ahigh-qualitytasksetconsistingof10collaborative learningtasksforclassicalChinesepoetryappreciationwasformed(afullversionis attachedintheAppendix).Eachtaskincludestheoriginaltextofthepoem,one collaborativelearningtask,andfivescoringcriteria.Theexampleisillustratedin Table3. Table3.Exampleofcollaborativelearningtaskset PoetryCollaborative learningtask Scoringcriteria Anovernighteast windblows,spring returnstothe half-deadroots. Jadeflowers respondtofrostand snow,andthefruit Conceivesthe Universe. Howcouldtherebe suchbeautyinthe mountainsand rivers,yetso majesticastoattract phoenixesand cranes? Ionlycalluponthe hermittocomehere andclimbthisday. Pleaseanswerthe meaningofplum blossomsinthis poem. Tenaciouslife."SpringReturnstoHalf-DeadRoots"depictsthe resurgenceofadyingplumblossominthespringbreeze, symbolizingthetenacityoftheremnantsoftheMingandQing dynasties(QuDajunandothers)inclingingtolifeamidstthe tragediesoftheMingandQingdynasties,andimplicitly conveyingthehopeofnationalrestoration. Noblecharacter."Jadeflowersrespondtofrostandsnow" invokestheimageofdivinejadefromtheChuCi,imbuingthe plumblossomwiththecharacteristicofactivelyrespondingto sufferingwithitsnobleessence,breakingawayfromthe traditional"passive,resistanttothecold"paradigmofplum blossompoetry. TheintegrityoftheremnantsoftheMingandQingdynasties. "TheFruitConceivestheUniverse"borrowstheimageofasingle fruitremaininginthe"Peeling"hexagramfromtheBookof Changes,elevatingtheplumblossom'sfruitingintoa philosophicalstatementbytheremnantsoftheMingandQing dynasties(QuDajun)tosafeguardtheflameofcivilizationin troubledtimes. Thespiritofseclusion."Howcouldtherebesuchbeautyinthe mountainsandrivers"negatesthebeautyoffamousmountains andscenicspots,highlightingthespiritualheightsofthesmall templewheretheancientplumblossomresides."phoenixesand cranes"constructsamoralspacefortheremnantsoftheMingand Qingdynastiesthattranscendsthemundane. Culturalinheritance."Comehereandclimbthisday"alludesto themetaphorofpickingherbsin"ChuCi",transformingtheactof appreciatingplumblossomsintoacollectiveritualforthe remaininggroups(suchastheLingnanSchoolofPoetry)topass onChinesecivilizationthroughliterature. Tocomprehensivelyevaluatetheimpactofdifferenttypesofscaffoldingon collaborativelearning,thisstudyutilizedtwoapproachesfocusingonboththe structuralcharacteristicsandthecognitivedepthofstudentinteractions.Specifically, assessmentwasconductedacrosstwodistinctdimensions:discoursebehaviorand discoursequality.Discoursebehaviorassessmentwasemployedtoquantifyand categorizethefrequencyandpatternsofstudentactions.Discoursequalityassessment, incontrast,measuredtheepistemologicalandcognitivedepthofthecontributionsto determinetheintellectualvalueoftheresultingdialogue. 4.2Assessmenttools 4.2.1Discoursequalitycodingframework Thecodingframeworkwasdevelopedbysynthesizingandadaptingquantitative metricsfromrecentstudiesonLLMevaluationtothespecificcontextofsimulated collaborativelearning.DrawingontheworkofRaietal.(2024),whoestablished standardsforlinguisticnaturalnessandsyntacticcorrectness,wedefinedthefluency dimensiontoensurethereadabilityandgrammaticalprecisionofagentdiscourse.To makesurethelogicalconsistency,weintegratedLiuetal. â s(2024)metricsfor internallogicwiththecontradictiondetectionmethodsproposedbyMĂŒndleretal. (2023).Furthermore,informedbyRossietal. â s(2024)analysisofgenerativedata challengesinsocialsciences,weincorporatedthediversitydimensiontoexplicitly measuretheinnovationandvariabilityofviewpoints.Theproposedframeworkfor assessingthediscoursequalityisshowninTable4.Thisframeworkincludesfivecore dimensions:fluency,repetitiveness,contradiction,relevance,anddiversity. Specifically,fluencyfocusesontheaccuracyofgrammarandthenaturalnessof expression;repetitivenessassessesthesimilaritybetweencurrentutterancesand previousones;contradictionexaminestheconsistencyofanagent'sviewpoints, identifyingpotentiallogicalconflicts;relevanceevaluatesthealignmentofcontent withtaskobjectives;anddiversityemphasizesthenoveltyofviewpointsandthe breadthofknowledge.Sincetheteacheragentdoesnotgenerateviewpoints,diversity isnotusedforcodingitsdiscourses.Byanalyzingthesedimensions,thestudynot onlyevaluatesthefluencyofdialoguesbutalsoexploresthelogicalconsistencyof agentbehaviorandtheinnovationinviewpointexpression. Table4.Theframeworkofdiscoursequalitycoding QualitydimensionCoding rules JudgmentbasisExample FluencyYes-1 No-0 Isthesentence grammaticallycorrect?Is itinlinewiththedaily expressionhabitsof teachersandstudents? "'Theplumblossombloomsinresponsetothe frostandsnow,itsabundantfruitsbearfruitin theuniverse.'Canthisbeunderstoodasthe plumblossomsymbolizingresilienceand hope?"(HighFluency,1) "Theplumblossomsymbolizesresilience.It bloomsinwinter,soit'sstrong."(Disorganized Sentence,0) RepetitivenessYes-1 No-0 Aretheviewpoints, arguments,expression logicandteacher interventionshighly consistentwithhistorical statements? Whendiscussingthesymbolicmeaningof plumblossoms,afterthefirstmentionof "plumblossomssymbolizeperseveranceand hope,"eachsubsequentmentionofthesame argumentandevidencewillbecountedasone repetition(repeatinghistoricalspeechwillbe countedasonerepetition). ContradictionYes-1 No-0 Doesthespeechconflict withpreviousspeeches, divisionoftasksor objectivefacts? Thestudentfirstproposedthat"plumblossoms symbolizethehermitspirit,"butthendenied that"thehermitspiritistooidealistic" (self-contradiction,countedas1contradiction) RelevanceYes-1 No-0 Doesthespeechrevolve aroundthepoemandthe taskobjectives,anddoes itmakeasubstantial contributiontosolvingthe problem? TaskObjective:"Pleaseanalyzethemeaning ofthelotusimageinpoetry." SpeechContent:"Thelotus's'icyandjade-like' appearancesymbolizespurity."(Related) DiversityYes-1 No-0 Arestudentsâopinions innovative, multi-perspective,ordo theycite cross-disciplinary knowledge? "Theplumblossomsymbolizesanescapefrom realityandquotesalinefromTaoYuanming's poetry."(Highlyinnovative,countedas1 highlight) Eachdimensioninthetableisencodedindependentlyusingabinarysystem,with logicalparallelismratherthanexclusivitybetweenthedimensions.Thisallowsa singledatapointtobesimultaneouslymarkedas'1'acrossmultipledimensions.The corpusproducedbythestudentandteacheragentsforalltasksineachexperiment wascollected.Eachspeechwasthenevaluatedforqualityalongfiveorfour dimensions,andthenumberofsamplesforeachdimensionwasaccumulated. 4.2.2Discoursebehaviorcodingframework Incollaborativelearningresearch,discoursebehaviorsanalysisservesasaneffective evaluationtool,playingacrucialroleinunderstandingtheinteractionpatternsamong learners.Thecodingframeworkforthediscoursebehaviorsofstudentagentswas developedbysynthesizingestablishedcategoriesfromhumancollaborativelearning researchandtailoringthemtothecommunicativecharacteristicsofLLMs.We derivedtheirrelevantlearningbehaviorsdimensionfromAdamsetal.(2002), specificallyincorporatingtheirobservationson â watchingpassively â and â disengaged â behaviors.Tocapturethesocialdynamicoftheinteraction,we referredtothetaxonomyofTanetal.(2022),whichemphasizes â learningoutcomes â and â socialinteractionsandprocesses â asessentialapplicationindicatorsofAIin collaborativelearning.Furthermore,weadaptedfromtheverb-dominatedcoding schemeproposedbyWangetal.(2020),whichfocusesonstudyingbehavioral patternsinonlinecollaborativelearningenvironmentswithdifferentlearningmaterial formats.Thecodingframeworkforstudentagent â sdiscoursebehaviorsisshownin Table5.Itincludesfourkeydimensions:irrelevantlearningbehaviors,social relationalbehaviors,collaborativeanalyticalbehaviors,andviewpointconstruction behaviors. Table5.Theframeworkofdiscoursebehaviorcoding Behavioral Dimension Behavior Type CodingRulesExampleCoding irrelevant learning behaviors IneffectiveDonotspeakorareirrelevanttothe discussion. Donotspeakorare irrelevanttothetask objective. A1 social relational behaviors PlanThediscussionbeginsbybringing outtopicsrelatedtothetopicand taskunderdiscussionandproposing directions. Let'sstartthediscussion with...Whatare everyone'sthoughtson thisaspect? B1 MonitorDuringthediscussion,monitorthe completionoftaskobjectivesbased onhistoricalspeechesand supplementthediscussionwithnew directions. Wehadavery thorough/inspiring discussion.We covered...,andwecan moveonto... B2 collaborative analytical behaviors ReflectAstatementsummarizingtheresults ofthediscussionssofar. Wecovered...from... to...,coveringawide rangeoftopics. C1 viewpoint construction behaviors ElaborateSpeechthatconstructsone'sown opinionsthroughcitations, examples,etc.,orconstructsone's owncognitivesystembasedonthe speechesofothers. Ithink...,because...D1 SupportStartbyclearlystatingyour agreementwiththeotherperson's pointofviewanduseexamplesto illustrateit. Iagreewith..., because... D2 QuestionDuringthediscussion,raise questionsaboutpointsthatyoudo notunderstand. Ihavesomequestions aboutwhatyoumeant by... D3 RebutDuringthediscussion,question opinionsthatyoudisagreewith. Ithinkyourstatementis incorrect/one-sided, because... D4 ExplainAnswerandexplainotherpeople's questions. Regardingtheissue of...,Ithink/feel... D5 Becausetheteacheragentjustguidesthediscussioninsteadofreallyinvolvingit, onlythreetypesofcodeswereusedforcodingteacheragentâsdiscourses: Encouragement(A1),Guidance(B1),andSummarization(C1).Eachtypeofcode correspondstoadistinctteachinggoal.Encouragement(A1)fostersstudents' motivationandparticipationthroughpositivefeedback,especiallywhenstudentsoffer valuableinsightsorengageactivelyindiscussions.Guidance(B1)involvesdirecting thefocusofstudentsthroughquestioning,prompts,ortopicredirectiontohelpthem concentrateonlearningobjectivesandexplorekeyconcepts.Summarization(C1) typicallyoccursattheendofadiscussion,aimingtoconsolidateandsynthesize students'viewpoints,assistingintheformationofasystematicknowledgeframework andreinforcinglearningoutcomes. 4.2.3LLMassistedassessmenttool Conductoftraditionaldiscourseanalysisanddeductivecodingneedsconsiderable timeandlaborinvolved.Thecodingprocesscanalsobringresearchers'subjective biases.Therefore,wealsoexploredthefeasibilityofapplyingLLMstoautomatic deductivecodingtasks(L.Zhangetal.,2025).DrawingontheworkofTaietal. (2024),whoappliedLLMsintextdataanalysis,thisresearchfurtherinvestigatesthe specificpotentialofLLMsindiscourseanalysis.Intheinitialphaseofthestudy,two researcherspre-coded20%ofthedatasamplestoensurethereliabilityandvalidityof theevaluationframeworkused.Subsequently,GPT-4owasemployedasanauxiliary codingtooltoprovideautomatedcodingsupportforevaluatingdiscoursequalityand behavioralperformanceduringthecollaborativeprocess.Aftercompletingthe automaticdeductivecoding,amanualsamplingverificationisconducted.20%ofthe alreadycodeddataisrandomlyselectedagaintocalculatetheconsistencybetween themodel'scodingandthemanualcoding.Thisapproachnotonlyenhancesthe efficiencyofdataanalysisbutalsoprovidesreliablemethodologicalsupportfor subsequentcollaborativelearningresearch. ThepromptdesignforGPT-4o'sautomaticdeductivecodingconsistsoftwocore components:(1)adetailedcodingspecificationexplanation,whichdefinesthecriteria, judgmentstandards,andboundaryconditionsforeachcodingcategory,andincludes representativeexamplesaswellasedgecases;and(2)amandatoryself-explanation requirement,inwhichthemodelisinstructedtoprovideexplicitreasoningforits codingdecisions.Thepurposeofthisdesignistoenhancetheinterpretabilityofthe codingresultswhilealsofacilitatingsubsequentverificationandnecessary interventionsbyresearchers.ThecompletepromptcontentcanbefoundinAppendix. 5.Experimentdesign Remindthatstudentagentsinthesimulatedcollaborativelearningsystemwere assumedtostrictlyfollowthetaughtscaffolding(i.e. â deepthinkbeforespeak â or â directspeak â ).Weusedpromptengineeringtodefinehowthestudentagents processinformationandactdifferentlyinthetwoconditions.Specifically,astudent agentinthe â directspeak â conditiongeneratedutterancesindirectresponsetoother student/teacheragents â inputs,withoutengaginginadditionalprocessingmediated byexplicitpromptsorworkflows.Ontheotherhand,astudentagentinthe â deep thinkbeforespeak â conditiongeneratedutterancesguidedbyadeepcognitive processdescribedwithpromptengineering.Thecognitiveprocessmainly encompassesthefourstepsbelow: · Contentanalysis:Accuratelyunderstandingthecontentandthemeofthepoem; · Instructioninterpretation:Comprehendingtheteacher â sinstructionsandthe expecteddirectionofthediscussion; · Contexttracking:Payingcloseattentiontothekeypointsinthespeechofother studentagents; · Differentiatedcontribution:Formulatinguniqueinsightsorproblemstoaddress. Thedesignofthescaffold â spromptsisshowninTable6.,andthecomplete promptdesigncanbefoundintheappendix.Thespecificexperimentalprocessis showninFig.3. Table6.Designofthescaffoldâsprompts StagePrompt Think âBackgroundâ:âYouareinapoetryappreciationclassandneedtobrainstormbased onanassignmentgivenbytheteacher.Thereisoneteacherandfivestudents.Youneed toeachplayaroleinthediscussionandworktogethertocompletethegrouptask.â âPoetryâ:"poem", âDialogueHistoryâ:"context", âTeacher'sInstructionsâ:"latest_instruction", "ReflectionGuidelines":"Pleasereflectonthecurrentdiscussionandtheteacher's instructionsbasedonyourrole.Youneedto: 1.Understandthecontentandthemeofthepoem 2.Understandtheteacher'sinstructionsandexpectations 3.Payattentiontothekeypointsofotherstudents'statements 4.Considerwhatuniqueinsightsorquestionsyoucanoffer." âOutputTemplateâ: "UnderstandingofthePoem":"Basedonmyrole,myunderstandingofthis poemis...", "ReactiontoOthers'Comments":"Mythoughtsonotherstudents'comments are...", "PossibleContributions":"Consideringmyrole,theuniqueperspective orinsightIcanofferinthediscussionis...", "InnerThoughts":"Asname,mytruethoughtsatthismomentare..." Fig.3.Experimentalworkflowofcollaborativelearningsimulation Eachstudentagentwasassignedtoaspecificrolebasedontherolecollaboration theory(Benne&Sheats,1948).Thistheoryclassifiestherolesthatemergeduring groupinteractionsintothreebroadcategories:GroupTaskRoles(e.g.,Expounder, OpinionGiver),GroupBuildingandMaintenanceRoles(e.g.,Encourager, Coordinator),andIndividualRoles(e.g.,Dominator,HelpSeeker).Drawinguponthis theoreticalframeworkandconsideringthespecificneedsofcollaborativelearning,we identifyanddefinefivedistinctstudentagentroles:Leader,Supporter,Expounder, Refuter,andSummarizer.Thesefiverolesencompasskeyfunctionsrangingfrom taskprogressiontorelationshipmaintenance,aimingtosimulateadiverserangeof collaborativeinteractions.Thedetaileddescriptionsofallthefiverolesareshownin theTable7.Allthedialogdatageneratedbytheagentsarerecordedforthediscourse analysisafterward,sothatwecancomparehowtheagentsinthetwoconditions performdifferently. Table7.Fivetypesofstudentagentroles RoledesignRolefunctionExpectedperformanceoftherole LeaderStartatopicand maintaina discussion Raisethecoreissuesfordiscussion,guidethedirectionofthe dialogue,ensurethatthediscussiondoesnotdeviatefromthe topic,proposenewperspectiveswhenthediscussion stagnates,andpromotetheparticipationofallmembers SupporterCitation, explanation,and answeringquestions Providetheoreticalbasisortextualevidencetosupport viewpoints,explaincomplexconceptsindepth,clarifyother members'confusion,andexpandthedepthandbreadthof discussions ExplainerChooseapointof viewandsupportit withexamples Identifyandreinforcevaluableinsights,enhancethe persuasivenessofargumentsthroughspecificexamplesor analogies,createapositivediscussionatmosphere,and encouragememberstocontinuethinkingdeeply RebutterAskquestions activelyand questioncritically Challengeweaknessesorholesinexistingperspectives, presentcounterexamplesoralternativeexplanations,prompt theteamtore-examineassumptions,preventgroupthink,and promotemorecomprehensiveanalysis SummarizerSummarizekey pointsandreach consensus Regularlysummarizethekeypointsdiscussed,integrate differentviewpoints,clarifytheinterimresultsofthe discussion,helptheteamformacommonunderstanding,and pointoutunresolvedissues TheexperimentincludedtenChinesepoetryappreciationtasks.Eachtaskis dividedintofourmainphases.Thefirstphaseisthetaskinitiation,duringwhichthe teacheragentclearlydefinesthecollaborativetaskcontentandtheknowledgepoints involved,establishingbasicspeakingnormstolaythefoundationforthesubsequent in-depthdiscussions.Thesecondphaseisfreediscussion.Thestudentagentsgenerate utterancesbasedontheirrolesandcognitiveprocesslogic.Ineachturn,astudent agentmaychooseactionssuchasofferingopinions,raisingobjections,orremaining silent.Thestudentagentsusuallyengageinmultipleroundsofinteractionstoexplore variousissuesofthegiventask.Thenumberofroundsforsolvingataskdependson whenalltheknowledgepointsofthetaskarecovered.Inthedynamicintervention phase,theteacheragentfocusesonmonitoringtwocoreindicators:theactivation progressofknowledgepointsandthebalanceofstudentagents'participation.This phaseisconductedconcurrentlywiththesecondphase,withstudentagentsengaging indiscussionswhileteacheragentprovidesdynamicmonitoring.Finally,inthe conclusionphase,whenthepresetterminationconditionsaremet,theteacheragent willcomprehensivelyevaluatetheperformanceofeachstudentagentthroughoutthe discussionprocess,assessingtheir'cognitiveprogresslevel.'Theteacherwillalso summarize,organize,andanalyzeallthediscussedissues,providingappropriate 'commentsandencouragement'tohelptheagentsbetterunderstandandmasterthe learnedknowledge. 6.Results 6.1Descriptiveanalysisofagentdiscourse Wefirstcalculatedthewordcountsandtheratioofutterancesbetweenteacher agentandstudentagent.TheresultsaredepictedinFig.4.Thesizeofthebubbles representsthefrequencyofutterances.Themorecounts,thebiggerthebubble,andit isevidentthatthediscoursefrequencyofstudentagentssignificantlyexceedsthatof teacheragents,whichalignswithourexperimentsetting.Thedataindicatesthatin termsoftheproportionofdiscoursefrequenciesbetweenteachersandstudents,the teacher-studentratiointhe"thinkbeforeyouspeak"modeis1:3.9,whichisbetter thanthe1:3.6ratiointhe"speakdirectly"mode.Thissuggeststhattheformer providesstudentswithgreateropportunitiesforautonomousexpression.Atthesame time,thevisualizedresultsrevealthatinthe â deepthinkbeforespeak â mode,the linebetweenteachersandstudentsarelonger,withagreaterdisparityinthe"depthof speaking"betweenthem.Thissuggeststhatthemodeprovidesstudentswithmore spaceforautonomousexploration,whilereducingthefrequencyofimmediate, high-densityteacherinterventions. Fig.4.Comparisonchartofdiscoursetextdataofteacherandstudentagents Wealsousedanindependentsamplest-testtoexaminetheimpactoftwomodes onthelengthofdiscoursegeneratedbytheagents,asshowninTable8.Theresults indicatedthattherewasnostatisticallysignificantdifferencebetweentheaverage discourselengthinthe"deepthinkbeforespeak"mode(M=124.5,SD=22.3)and the"directspeak"mode(M=98.2,SD=18.7)(t=-1.239,p=.216).Additionally, theeffectsize(Cohen'sd=0.113)waswellbelowthe0.2threshold,further confirmingthatthedifferenceinoutputvolumebetweenthetwomodeswas negligible.Thisfindingsuggeststhattheadditionofcognitiveintervention scaffoldingdidnotsignificantlyaltertheoutputscaleoftheagents.Thereasonfor thismaylieintheunderlyingconstraintsoftheLLM,wherethemaximumlength limitofthepromptsandthepredefinedoutputstyleleadtorelativelystableoutput lengthsacrossdifferentscaffoldingstrategies.Fromaresearchvalidityperspective, thisconsistencyreflectsthatthepacingandinformationdensityoftheinteractions remainedhighlysynchronizedacrossbothmodes,thusvalidatingtherobustnessof themulti-agentsysteminsimulatingcollaborativelearningunderdifferent interventionstrategies. Table8.Comparisonofutterancelengthbetweentwomodes ModeNmeansdtPCohenâsd Deepthinkbeforespeak25391.7217.57 -1.23910.2160.1132 Directspeak22889.8615.00 Overall,thesecharacteristicsofthediscoursetextdatarevealthattheâdeep thinkbeforespeakâmodecontributestobuildingamoreefficientandautonomous collaborativediscussionenvironment,reducingineffectiveinteractions,andproviding asoliddatafoundationforsubsequentdiscoursequalityandbehavioralanalysis. 6.2Qualitativeanalysisofagentdiscoursequality Thisexperimentprimarilyfocusesonacomparativeanalysisoftheâdeepthink beforespeakâandâdirectspeakâcognitivemodes,usingdiscoursequalityasthe evaluationmetric.Initially,werandomlyselected20%ofthedataasthepre-coding sample,whichincluded100studentutterancesand25teacherutterances.After manualcodingbytworesearchers,theKappavaluereached0.75,demonstratinggood internalconsistencyofthecodingprocess.Subsequently,basedonaconsensus betweenthetworesearchersonthefinalcodingresults,weusedalargelanguage modeltore-codethesamedatasettoadjustandverifytheaccuracyofthemodel's coding.TheresultsshowedaKappacoefficientof0.80betweenthemanualand model-basedcoding,furtherconfirmingthehighconsistencyoftheevaluationsystem andvalidatingthemodel'sabilitytoperformautomateddeductivecoding.Finally,the largelanguagemodelwasusedtoautomatethecodingoftheremainingdata,witha finalmanualreviewandverificationconductedafterthecodingprocesswas completed. 6.2.1Between-subjectsanalysisofstudentagents Thestudyconductedanindependentsamplest-testonthediscoursequalitycoding datafromtentasksineachofthetwogroups,withtheresultsdetailedshowninTable 9.TocontrolfortheFalseDiscoveryRate(FDR)acrossmultipledimensions,the Benjamini-Hochberg(BH)procedure(Thissen,Steinberg,&Kuang,2002)wasused toadjustthep-values.Theanalysisrevealedthatthe â deepthinkbeforespeak â modesignificantlyoutperformedthe â directspeak â modeintermsofthe Repetitiveness(adjustedp=.012)andDiversity(adjustedp=.016).Therewereno statisticallysignificantdifferencesinFluency(adjustedp=.340)orRelevance (adjustedp=.340),andthezerovaluesfortheContradictionindicatorfurther confirmedthetechnicalstabilityoftheLLMinmaintaininggrammaticalnormsand contentcoherence. Table9.Thediscoursequalityofstudentagentsundertwothinkingmodes Discoursequality dimension âdeepthinkbefore speakâmode â directspeakâ mode tPAdj.p(BH) Fluency22.700 ± 5.25025.300 ± 6.549-0.9790.3400.340 Repetitiveness9.900±3.07116.300±5.165-3.3680.0030.012 Contradiction0.0000.000N/AN/AN/A Relevance22.700 ± 5.25025.300 ± 6.549-0.9790.3400.340 Diversity13.500±3.7499.100±2.8462.9560.0080.016 Note:Adjustedp-valueswerecalculatedusingtheBenjamini-HochbergproceduretoaccountfortheFDR. Focusingontheanalysisofthedifferentiationindicator,theâdeepthinkbefore speakâmodedemonstratedaclearadvantage.TheaveragevalueforRepetitiveness decreasedby6.40,whileDiversityincreasedby4.40.Thissuggeststhatthemulti-step iterativethinkingmechanismcaneffectivelydeepenthecognitiveprocessingofthe agent.Itspositiveeffectsaremainlyreflectedinthreeaspects:first,throughthe guidanceofthethoughtchain,itsuccessfullyreducedtheredundancyofthecontent, thusimprovingtherepetitivenessindicator;itactivatedabroaderrangeofknowledge associations,therebyenhancingthediversityofthediscourse;whileimprovingthe quality,itmaintainedtheinternalconsistencyofthesemanticsystem,preventingthe riskofgeneratingfalseinformation(hallucinations).Fromacognitiveperspective, thisframeworkismorealignedwithhumanâsdeeplythinkingprocesses,fosteringthe generationofmoreinnovative,logicallystructured,andinformation-richlanguage expressions. 6.2.2Within-subjectsanalysisofstudentagents Tofurtheranalyzethediscoursequalitydifferencesamongvariouscollaborativeroles inthe â deepthinkbeforespeak â mode,thestudyexaminedthedistinctionsin repetitivenessanddiversityindicatorsforeachroleandillustratedtheirpercentage distributionsinFig.5.It'simportanttonotethatbecausetheencodingofeach dimensionislogicallyparallelandnotmutuallyexclusive,apieceofdatacanbe encodedaseitherrepetitivenessordiversity.Forexample,inthecollecteddata,a studentagent'sstatementis,"IstronglyagreewithLiSiandWangMei'sviews! Furthermore,Ithink'Themountainairisbeautifulatsunset,andthebirdsflybackto theirnests'alsoperfectlyembodiesthemeaningofseclusion.Thetranquilityofthe mountainsandthesceneofbirdsreturningtotheirnestsechothereclusivelife symbolizedbychrysanthemums,showcasingTaoYuanming'stranscendenceand leisure."Thefirstpartofthisstatementsupportsandrepeatsthepreviousstudent's statement,sorepetitivenessisencodedas1.Thesecondpart,wherethestudent proposesacombinationofnaturalphenomenaandthepoet'spersonality,isan innovativeviewpoint,sodiversitycanalsobeencodedas1.Thedataanalysis indicatesthatthediscoursecharacteristicsofeachrolearehighlyalignedwithits predefinedfunctionalpositioning. Fig.5.Thediscoursequalityofstudentagentsintheâdeepthinkbeforespeakâmode Astheinitiatorandfacilitatorofthediscussion,theleaderexhibitedthehighest levelofinnovationandthelowestredundancy,withadiversityscoreof80.85%and theminimumrepetitiveness.Thisalignspreciselywiththeleader'scoreresponsibility ofinitiatingtopicsandmaintainingthedepthofthediscussion.Thediversitylevelof theexplainerwassimilartothatoftheleader,suggestingthattheviewpoints expressedbytheexplainerweresimilarlynovel,fittingtherole'sfunctionofoffering personalinsightsorbuildinguponothers'perspectives.Thecontrastbetweenthehigh repetitivenessrate(88.89%)andlowdiversityofsupportersreflectsthedual characteristicsoftheircontentfocusandconvergentargumentationwhenreinforcing specificviewpoints.Therebutter,withitscriticalthinking,demonstratedthehighest diversitywhilemaintaininglowrepetitiveness,highlightingtheroleâsdynamic advantageinchallengingexistingargumentsandpromotingdivergentthinking. Finally,thesummarizerpresentedanextremedistributionofdiscoursequality:ahigh repetitivenessof97.78%andalowdiversityofonly8.89%,visuallyreflectingthe mechanismofbuildingconsensusthroughfrequentrestatement. Thisrole-basedvariation,drivenbypresetbehaviors,effectivelyvalidatesthat the â deepthinkbeforespeak â modeactivatesthefunctionalattributesofeachrole. Bybalancingrepetitivenessandinnovation,thismodeenablestheconstructionofa diversifiedcollaborativediscussionecosystemwhileensuringcontentrelevance. Moreimportantly,itfacilitatesstrongerfunctionaldifferentiationinthequality indicatorsofdifferentroles,providingrobusttheoreticalandempiricalsupportforthe designandimplementationofefficientcollaborativelearningmodels. 6.2.3Between-subjectsanalysisofteacheragents Thestudyanalyzedtheperformanceoftheteacheragentunderbothmodes,asshown inTable10,andfoundnostatisticallysignificantdifferencesacrossanydimensions afterapplyingtheBenjamini-Hochberg(BH)procedure.WhiletheRelevance indicatorshowedanominaldifferenceintheunadjustedtest(p=.042),itdidnot remainsignificantaftertheBHcorrection.ThisobservedtrendinRelevanceislikely duetothefactthatinthe â deepthinkbeforespeak â mode,theteacheragentmakes decisionsonwhethertointerveneandguidebasedonthehistoricaldialoguerecords. Ifthesystemdeterminesthatnointerventionisneeded,theteacheragentmayprovide encouragementoraffirmation,inwhichcasethecontentofitsspeechisnotdirectly relatedtothecollaborativetask.Incontrast,theteacheragentinthe â directspeak â modedoesnotengageinsuchareflectiveorjudgmentalprocess,andtherefore,the contentofitsspeechismorecloselytiedtothetaskitself. Table10.Thediscoursequalityofteacheragentundertwothinkingmodes Discoursequality dimension âdeepthink beforespeakâ mode âdirectspeakâ mode tPAdj.p(BH) Fluency5.800±1.1357.000±1.885-1.7240.1020.102 Repetitiveness1.400 ± 1.0742.200 ± 0.919-1.7890.0900.102 Contradiction0.0000.000N/AN/AN/A Relevance5.200±1.3166.900±2.079-2.1850.0420.102 Note:Adjustedp-valueswerecalculatedusingtheBenjamini-HochbergproceduretoaccountfortheFDR. 6.3Qualitativeanalysisofagentdiscoursebehavior Intermsofthediscoursebehaviorcodingresults,thestudyalsoselected20%ofthe totaldataforevaluation.Theinter-raterreliabilitytestofthemanualcoding(Kappa= 0.73,>0.60)demonstratedgoodconsistency.Furthermore,theKappacoefficient betweenthemanuallyagreedcodingresultsandthelargemodel'scodingresults increasedto0.83,whichstronglysupportsthehighconsistencyoftheevaluation system. 6.3.1Between-subjectsanalysisofstudentagents Abetween-groupcomparativeanalysisofthediscoursebehaviordistributionin studentagentsrevealeddistinctpatterns,asdetailedinTable11.Tocontrolforthe inflationofTypeIerroracrosstheseindicators,aBonferroni-correctedalphalevelof 0.0063(0.05/8)wasalsoapplied.Underthisadjustment,theâdeepthinkbefore speakâmodedemonstratedstatisticallysignificantdifferencesinseveralkey categories.Notably,asignificantdecreasewasmaintainedinElaborate(D1)(adjusted p=.008),alongsidesignificantincreasesinPlan(B1)(adjustedp=.035),Explain (D5)(adjustedp=.035),Rebut(D4)(adjustedp=.044),andReflect(C1)(adjustedp =.045).TheBH-adjustedresultsconfirmthattheâdeepthinkbeforespeakâmode effectivelyencouragesagentstoadoptmorediverseinteractionstrategies,thereby broadeninganddeepeningthediscussion. Table11.Thediscoursebehaviorofstudentagentundertwothinkingmodes Discourse behavior dimension âdeepthinkbefore speakâmode â directspeakâ mode tPAdj.p(BH) Ineffective(A1)0.0000.000N/AN/AN/A Plan(B1)2.200 ± 1.3980.800 ± 0.7882.7570.0130.035 Monitor(B2)2.600±1.5783.200±1.932-0.7610.4570.457 Reflect(C1)3.100±1.4491.700±1.1592.3850.0280.045 Elaborate(D1)3.900 ± 2.23411.300 ± 4.218-4.9030.0010.008 Support(D2)4.700±2.0575.500±1.649-0.9590.3500.400 Question(D3)3.400±2.1712.200±1.7511.3610.1900.253 Rebut(D4)1.700 ± 1.4180.500 ± 0.5272.5080.0220.044 Explain(D5)1.200±1.2290.100±0.3162.7410.0130.035 Note:Adjustedp-valueswerecalculatedusingtheBenjamini-HochbergproceduretoaccountfortheFDR. Togainadeeperunderstandingofhowstudentagentbehaviorsaretransferred betweentheâdeepthinkbeforespeakâandâdirectspeakâmodes,thisstudy generatedabehaviortransitionmatrixthroughcomputationalanalysis,asshownin Fig.6.Thesetwoheatmapsillustratethetransitionprobabilitiesbetweenstudent agentbehaviorsunderdifferentthinkingmodes,wheredarkercolorsindicatehigher transitionprobabilities.Theverticalaxisrepresentsprecedingbehaviors,whilethe horizontalaxisrepresentssubsequentbehaviors. (a)âdeepthinkbeforespeakâmode(b)âdirectspeakâmode Fig.6.Heatmapofstudentagentâsbehaviortransferundertwothinkingmodes FromtheleftheatmapinFig6.3,itisevidentthattheâdeepthinkbeforespeakâ modeexhibitsadiverseandhighlyinterconnectedbehaviortransitionpattern.Notable transitionsinclude:theshiftfromD1toD2,whichshowsthatarticulatingaviewpoint oftentriggerssupportivebehaviors,thusformingapositivefeedbackloopof collaborativeknowledgeconstruction;thehigh-probabilitytransitionsfromD2toD3 andD2toD4revealthatsupportivebehaviorscaneithertriggerfurtherquestionsor leadtorebuttals,highlightingthemultidirectionalityofcognitivedevelopment;the hightransitionprobabilitiesfromC1toB1andC1toB2demonstrateanaturalflow fromreflectionandsummarytotheplanningandmonitoringofthenextround, completingafulldiscussionloop;finally,thetransitionsfromD3toC1andD4toC1 suggestthatquestioningandrebuttalbehaviorsoftenleadtoreflection,aidingdeeper cognitiveintegration.Thesecomplextransitionsequencescollectivelyformahighly interactivebehavioralnetwork,coveringtheentirecollaborativeprocessfromPlan (B1),throughmulti-facetedknowledgeconstruction(CategoryD),toReflect(C1). Incontrast,therightsideofFigure6.3revealsthatthe â directspeak â mode exhibitsasignificantlysimplifiedbehaviortransitionpattern.Themostnotable featureistheextremelyhightransitionprobabilityfromD2toD3,whichindicates thatsupportivebehaviorsaremostlyfollowedbyquestions.However,thesequestions oftenlacksubsequentin-depthfollow-up,makingitdifficulttoformacomplete cognitiveconflictresolutionchain.Afewothersignificanttransitionsinclude:D1to D2,whichreflectsalinearpatternofsupportfollowingthearticulationofapoint;B2 toD1andB2toD5,whichshowthatbehaviormonitoringcanstillleadtoknowledge expression;andD1toB1andD1toB2,suggestingthatafterthearticulationbehavior, theremaybeareturntoplanningandmonitoring,thoughthestrengthofthese transitionsisnotablyweakerthaninthe â deepthinkbeforespeak â mode.Inthe â directspeak â mode,transitionsrelatedtoRebut(D4)andExplain(D5)arealmost entirelyabsent,andthelowerhalfofthebehaviormatrixshowslargeareasofempty space,indicatingthatdeepcriticalinteractionsareseverelylimited. Bycomparingthebehaviortransitionmatricesofthetwomodes,wefindthatthe â deepthinkbeforespeak â mode â sbehaviortransitionnetworkismorecomplexand diverse,withmoreeffectivetransitionpointsandamorebalanceddistribution.In contrast,the â directspeak â mode â stransitionnetworkissimplerandmorelinear, concentratingprimarilyonafewlimitedpaths.Thiscomparisonstrongly demonstratesthatthe â deepthinkbeforespeak â modenotonlyhelpsbalancethe distributionofindividualbehaviortypesbutalsofacilitatestheconstructionofamore naturalandcompletebehavioralflownetwork,allowingtheentirediscussionprocess toexhibitthedynamicevolutioncharacteristicofrealcollaborativeenvironments. 6.3.2Within-subjectsanalysisofstudentagents Tomoreclearlyillustratethedistributioncharacteristicsofdiscoursebehaviorsacross differentrolesinthe â deepthinkbeforespeak â mode,thisstudyconstructeda Role-BehaviorProportionAnalysisChart(asshowninFig.7.)andconductedan in-depthanalysis.Thechartexplicitlyrevealsthedifferentiateddistributionpatterns offivecollaborativerolesacrossninecategoriesofdiscoursebehaviors,reflectinga distinctfunctionaldivisionoflabor. Fig.7.Thediscoursebehaviorofstudentagentsintheâdeepthinkbeforespeakâmode TheLeaderroleischaracterizedbyahighproportionofplanningbehaviors(B1) at43.48%,whichhelpsestablishtheframeworkforthediscussion,alongwith19.57% ofmonitoringbehaviors(B2)toadjustthediscussionflowinrealtime.Additionally, theLeaderroledisplays44.74%ofexplanationbehaviors(D1),effectivelyguiding thediscussiontowardtheintendedtopic,fulfillingitscorefunctionofsteeringthe conversation.TheExplainerroleleadsinbothexplanationandelaborationbehaviors, withtheproportionofexplanationbehaviorsbeingthreetofivetimeshigherthanthat ofotherroles,aligningperfectlywithitsfunctionofintroducingnovelideasand offeringin-depthexplanations.TheSupporterroleshowsasignificantlyhigher proportionofsupportingbehaviorscomparedtoallotherroles,stronglyemphasizing itsfunctioninreinforcingvaluableinsights.TheCriticrole,ontheotherhand, exhibitsthehighestproportionsofquestioningandrebuttalbehaviors,highlightingits criticalandinterrogativenature.TheSummarizerrolehasthehighestproportionof reflectionbehaviors,consistentwithitsfunctionofsynthesizingkeypointsand promotingconsensus.Additionally,theSummarizeralsodemonstratesarelatively highproportionofmonitoringbehaviors(B2),indicatingthatthisroleactively introducesnewdiscussionanglesduringreflectivephases. 6.3.3Between-subjectsanalysisofteacheragents Thebehavioraldifferencesintheteacherrolebetweenthetwothinkingmodesare alsonoteworthy.AsshowninTable12,afterapplyingtheBHprocedure,statistically significantdifferenceswereobservedinbothGuidance(B1)(adjustedp=.015)and Encouragement(A1)(adjustedp=.030).Specifically,theteacheragentinthe â deep thinkbeforespeak â modedemonstratedahigherfrequencyofencouragementanda lowerfrequencyofdirectguidance.Theseresultssuggestthatwhenstudentagents engageindiscussionsusingthisstructuredthinkingmode,theteacheragentcan reducethefrequencyofdirectinterventionsandinsteadfocusonproviding encouragementandaffirmation,therebyeffectivelypromotingstudents'autonomous explorationanddeepercognitiveengagement. Table12.Thediscoursebehaviorofteacheragentundertwothinkingmodes Discoursebehaviordimensionâdeepthink beforespeakâ mode â directspeak â mode tPAdj. p(BH) Encouragement(A1)2.100±0.5671.100±1.1002.5540.0200.030 Guidance(B1)3.400 ± 1.1735.500 ± 1.715-3.1940.0050.015 Summarization(C1)0.300±0.6740.400±0.516-0.3720.7140.714 Note:Adjustedp-valueswerecalculatedusingtheBenjamini-HochbergproceduretoaccountfortheFDR. Basedontheanalysisofbothdiscoursequalityanddiscoursebehavior,the âdeepthinkbeforespeakâmodeoutperformstheâdirectspeakâmodeinseveral coreindicators:(1)Contentnovelty:Thereisasignificantincreaseindiscourse diversityandanotabledecreaseinredundancy;(2)Rolealignment:Thebehaviorsof eachagentaremoreconsistentwiththeirpredefinedfunctionalroles;(3)Interaction diversity:Theproportionofbehaviorsreflectingdeepinteractions,suchassupport, rebuttal,andexplanation,hasincreased;(4)Teacherinterventionoptimization:The teacherâsencouragementbehaviorshaveincreased,whileguidancebehaviorshave decreased,furtherenhancingstudents'autonomy.Thesefindingsprovideimportant theoreticalguidanceforthefuturedesignofmulti-agentcollaborativelearning systems:byembeddingstructuredthinkingprocesseswithintheagents,thequality andefficiencyoftheirinteractionscanbesignificantlyimproved,makingthe simulatedcollaborativelearningprocessmorerealisticallyreplicaterealclassroom situations. 7.Discussion Thisstudyconstructsamulti-agentcollaborativelearningsystembasedonthe MetaGPTframeworkanddesignscomparativeexperimentstoinvestigatethesystem's performanceinsimulatinggroupdiscussionincollaborativelearning.Inthissection, wewilladdresstheresearchquestionsbasedonthefindings. RQ1:FeasibilityandCharacteristicsofLLM-basedMulti-AgentSimulation. Theresearchfindingssubstantiatethefeasibilityandeffectivenessofthe proposedmulti-agentsystemasanovelplatformforeducationalexperiments.Data analysisrevealsthatthesystemnotonlysuccessfullyreplicatesthecomplexand dynamicinteractiveenvironmentofreal-worldcollaborativelearning,butmore importantly,thestudentagentsdemonstrateremarkableconsistencywiththeir characters.Whetherastheleaderresponsibleforframingthediscussionorthe challengerstimulatingdepththroughcriticalthinking,theagents'discoursebehaviors aligncloselywiththeirfunctionalroles. Theintroductionofthe â deepthinkbeforespeak â modeprovedtobea determiningfactorinthissuccess.Bysimulatingthemetacognitiveprocessofhuman learners,thismodeenabledagentstoexhibitamorestructuredandsystematicpattern ofbehavioraltransitions.Unlikethe â directspeak â modeofchatinteractions,the agentsformedaself-sustainingdiscussionloopcapableofeffectivelygeneratingand resolvingcognitiveconflicts. Fromamethodologicalperspective,thisfindingaddressesthecallwithin computationalsocialsciencetoutilizeagentsforsimulatinghumansocialinteractions, suggestingthatLLM-basedagentspossesssufficientcognitiveabilitiestointernalize complexsocialroles.Comparedtotraditionalempiricalresearch,whichis time-consuminganddifficulttocontrolforvariables,thishigh-fidelityand controllablesimulationenvironmentoffersanefficientalternativeforexploring collaborativelearningmechanisms,enablingresearcherstorepeatedlytesteducational hypothesesataverylowcost. RQ2:AlignmentbetweenSimulatedBehaviorsandHumanCollaborative LearningPatterns. Comparativeanalysisrevealsasignificantpositiveeffectofthecognitive interventionscaffolding.Inthedimensionofdiscoursequality,theexperimental groupequippedwiththecognitivescaffolding(the â deepthinkbeforespeak â mode) exhibitedhighercontentdiversityandsignificantlyreducedredundancyintheir dialogues.Thissuggeststhatthemandatory"thinking"scaffoldeffectivelysuppresses variousnegativebehaviorscommonlyobservedinstudentcollaboration,encouraging themtoactivateabroaderknowledgenetworkbeforespeaking,thusgeneratingmore information-denseandinnovativeideas. Thismechanism-drivenenhancementofcognitivedepthalignswiththeICAP framework(M.T.H.Chi&R.Wylie,2014),whichclassifiesstudents'learning activitiesbasedontheirlevelofcognitiveinvolvement,rangingfromhightolow,into fourcategories:Interactive,Constructive,Active,andPassive.BasedontheICAP classificationcriteria,the â deepthinkbeforespeak â cognitiveinterventionscaffold designedinthisstudycanpromotestudents'agentstoshiftfromsimple â Active â to â Constructive â and â Interactive â knowledgeco-construction.Thisthinkingphase simulatestheknowledgeconstructionprocessofrealstudents,correspondingtothe "Constructive"dimensionintheICAPframework.Meanwhile,thediscussions, support,andexplanationsgeneratedbasedonthisprocessrepresentahigherlevelof "Interactive"participation. Incontrast,agentsinthe â directspeak â modepredominantlyremainedatthe â Active â level.Thefactthatthesimulatedagentsexhibiteddifferentiatedcognitive behaviorsconsistentwiththeICAPhierarchyvalidatesthesystem'sfidelity.It demonstratesthatLLM-basedagents,likehumanlearners,requireexplicit metacognitivescaffoldingtotranscendsuperficialparticipationandachievedeep knowledgeco-construction. Furtheranalysisofcollaborativeperformancerevealsthatthe â deepthinkbefore speak â modenotonlyoptimizedindividualoutputsbutalsoreshapedthegroup â s interactionecology.Thestudyobservedasignificantincreaseintheproportionof deeperinteractivebehaviors,suchasexplanationandrebuttal,formingamore complexandinterconnectednetworkofbehavioraltransitions.Thisindicatesthat studentagentsnolongermerelystatetheirviewsinisolation,butinsteadengagein morefrequentdebatesandnegotiationstoco-constructknowledge,resultinginashift from"shallowinteraction"to"deepcollaboration."Itisnoteworthythatthisincreased studentagencyalsopromptedadaptivechangesintheteacheragent â sbehavior â shiftingfromfrequentdirectinstructionstomoreemotionalsupportand encouragement.Thisphenomenon,wherestudentautonomyincreasesandteacher interventiondecreases,validatesthepositiveimpactofthecognitiveintervention scaffoldingonthequalityofdiscourseanddepthofdiscussionincollaborative learning.Itfurthersupportstheecologicalvalidityofthesimulationsystemin replicatingauthenticeducationaldynamics. 8.Conclusion Thisstudyisgroundedinthetheoreticalfoundationsofcollaborativelearning,aiming toexploretheapplicationvalueandpotentialoftheLLM-basedmulti-agentsystemin simulatingcollaborativelearning.Inspecific,wefirstpresentedhowtodesignand implementthesimulationsystembasedonthemulti-agentframeworknamed MetaGPT,andevaluateddifferentscaffolding(i.e.deepthinkbeforespeakv.s.direct speak)effectivenessonthesimulationsystem.Thecomprehensiveanalysisshowed thattheresultsofthesimulationsystemgenerallyalignedwiththelearningscience theoriessuchasself-explainandICAP.Thedetailedresultsalsoreportedthepossible fine-graineddifferencesofcollaborativelearningdiscourses,whichprovidedvaluable referencesinpractice. Whileourstudyvalidatesthetheoreticalefficacyofmetacognitivescaffoldingin thesimulatedcollaborativelearningsystem,futureworkisneededtoexplorehowthe acquiredknowledgefromthesimulationsystemcanguidetheteachingandlearning practicewithhumansubjects. Ethicsapprovalandconsenttoparticipate Notapplicable.Thisstudydoesnotinvolvehumanparticipantsoranimals,and thereforeethicalapprovalandinformedconsentarenotrequired. Funding Thisworkwassupportedby[BLINDED].Thefundingbodyhadnoroleinthedesign ofthestudy,collection,analysis,andinterpretationofdata,orinwritingthe manuscript. References Adams,J.P.,Brissenden,G.,Lindell,R.S.,Slater,T.F.,&Wallace,J.(2002).Observationsofstudentbehaviorincollaborative learninggroups.AstronomyEducationReview,1(1),25-32.https://doi.org/10.3847/AER2001002 Aher,G.V.,Arriaga,R.I.,&Kalai,A.T.(2023,July).Usinglargelanguagemodelstosimulatemultiplehumansandreplicate humansubjectstudies.InInternationalconferenceonmachinelearning(p.337-371).PMLR. Bail,C.A.(2024).CanGenerativeAIimprovesocialscience?ProceedingsoftheNationalAcademyofSciences,121(21), e2314021121.https://doi.org/10.1073/pnas.2314021121 Bazgir,A.,&Zhang,Y.(2025).Agentichypothesis:Asurveyonhypothesisgenerationusingllmsystems.TowardsAgenticAI forScience:HypothesisGeneration,Comprehension,Quantification,andValidation. Benne,K.,&Sheats,P.(1948).Functionalrolesofgroupmembers.SharedExperiencesinHumanCommunication,155. Bryda,G.,&Sadowski,D.(2024,January).Fromwordstothemes:AI-poweredqualitativedatacodingandanalysis.InWorld conferenceonqualitativeresearch(p.309-345).Cham:SpringerNatureSwitzerland. https://doi.org/10.1007/978-3-031-65735-1_19 Chen,A.,Wei,Y.,Le,H.,&Zhang,Y.(2024).LearningbyteachingwithChatGPT:TheeffectofteachableChatGPTagenton programmingeducation.BritishJournalofEducationalTechnology.https://doi.org/10.1111/bjet.70001 Chi,M.T.,DeLeeuw,N.,Chiu,M.-H.,&LaVancher,C.(1994).Elicitingself-explanationsimprovesunderstanding.Cognitive science,18(3),439-477.https://doi.org/10.1016/0364-0213(94)90016-7 Chi,M.T.,&Wylie,R.J.E.p.(2014).TheICAPframework:Linkingcognitiveengagementtoactivelearningoutcomes. EducationalPsychologist,49(4),219-243.https://doi.org/10.1080/00461520.2014.965823 DeCurtĂČ,J.,&DeZarzĂ ,I.(2025).LLM-DrivenSocialInfluenceforCooperativeBehaviorinMulti-AgentSystems.IEEE Access.https://doi.org/10.1109/ACCESS.2025.3548451 Dillenbourg,P.(1999).Whatdoyoumeanbycollaborativelearning?Collaborative-learning:Cognitiveandcomputational approaches.,1-19. Farjam,M.,Meyer,H.,&Lohkamp,M.(2024).APracticalGuideandCaseStudyonHowtoInstructLLMsforAutomated CodingDuringContentAnalysis.SocialScienceComputerReview,08944393251349541. https://doi.org/10.1177/08944393251349541 Gao,C.,Lan,X.,Lu,Z.,Mao,J.,Piao,J.,Wang,H.,...&Li,Y.(2023).S3:Social-networksimulationsystemwithlarge languagemodel-empoweredagents.arXivpreprintarXiv:2307.14984.https://doi.org/10.48550/arXiv.2307.14984 Graesser,A.C.,Lu,S.,Jackson,G.T.,Mitchell,H.H.,Ventura,M.,Olney,A.,&Louwerse,M.M.(2004).AutoTutor:Atutor withdialogueinnaturallanguage.BehaviorResearchMethods,Instruments,Computers&Education,36(2),180-192. https://doi.org/10.3758/BF03195563 Grossmann,I.,Feinberg,M.,Parker,D.C.,Christakis,N.A.,Tetlock,P.E.,&Cunningham,W.A.(2023).AIandthe transformationofsocialscienceresearch.Science,380(6650),1108-1109.https://doi.org/10.1126/science.adi1778 Guo,F.(2023).Gptingametheoryexperiments.arXivpreprintarXiv:2305.05516.https://doi.org/10.48550/arXiv.2305.05516 Gurcan,O.(2024).Llm-augmentedagent-basedmodellingforsocialsimulations:Challengesandopportunities.arXivpreprint arXiv:2405.06700.https://doi.org/10.48550/arXiv.2405.06700 Horton,J.J.(2023).Largelanguagemodelsassimulatedeconomicagents:Whatcanwelearnfromhomosilicus?(No.w31122). NationalBureauofEconomicResearch.https://doi.org/10.3386/w31122 Hsu,Y.-S.,Lai,T.-L.,&Hsu,W.-H.(2015).Adesignmodelofdistributedscaffoldingforinquiry-basedlearning.Researchin ScienceEducation,45(2),241-273.https://doi.org/10.1007/s11165-014-9421-2 Islam,R.,&Moushi,O.M.(2025,June).Gpt-4o:Thecutting-edgeadvancementinmultimodalllm.InIntelligent Computing-ProceedingsoftheComputingConference(p.47-60).Cham:SpringerNatureSwitzerland. https://doi.org/10.1007/978-3-031-92611-2_4 JĂ€rvelĂ€,S.,&Hadwin,A.F.(2013).Newfrontiers:RegulatinglearninginCSCL.EducationalPsychologist,48(1),25-39. https://doi.org/10.1080/00461520.2012.748006 Jeong,H.,Hmelo-Silver,C.E.,&Jo,K.(2019).Tenyearsofcomputer-supportedcollaborativelearning:Ameta-analysisof CSCLinSTEMeducationduring2005â2014.EducationalResearchReview,28,100284. https://doi.org/10.1016/j.edurev.2019.100284 Jiang,Y.-H.,Liu,T.-Y.,Zhuang,X.,Hu,H.,Li,R.,&Jia,R.(2024).Enhancingeducationalpracticeswithmulti-agentsystems: Areview.EnhancingEducationalPractices:StrategiesforAssessingandImprovingLearningOutcomes,47-65. https://doi.org/10.52305/RUIG5131 Kollar,I.,Wecker,C.,&Fischer,F.(2018).Scaffoldingandscripting(computer-supported)collaborativelearning.In Internationalhandbookofthelearningsciences(p.340-350):Routledge. Liu,Y.,Guo,Z.,Liang,T.,Shareghi,E.,VuliÄ,I.,&Collier,N.(2024).Aligningwithlogic:Measuring,evaluatingand improvinglogicalconsistencyinlargelanguagemodels.arXive-prints,arXiv:2410.02205. https://doi.org/10.48550/arXiv.2410.02205 Miyake,N.(2007).Computersupportedcollaborativelearning.TheSAGEhandbookofe-learningresearch,248-265. Mohamedhen,A.S.,Alfazi,A.,Arfaoui,N.,Ejbali,R.,&Nanne,M.F.(2024).Towardsmulti-agentsystemforlearningobject recommendation.Heliyon,10(20). MĂŒndler,N.,He,J.,Jenko,S.,&Vechev,M.(2023).Self-contradictoryhallucinationsoflargelanguagemodels:Evaluation, detectionandmitigation.arXivpreprintarXiv:2305.15852.https://doi.org/10.48550/arXiv.2305.15852 Ărpek,Z.,Tural,B.,&Destan,Z.(2024,September).Thelanguagemodelrevolution:Llmandslmanalysis.In20248th InternationalArtificialIntelligenceandDataProcessingSymposium(IDAP)(p.1-4).IEEE. https://doi.org/10.1109/IDAP64064.2024.10710677 Park,J.S.,O'Brien,J.,Cai,C.J.,Morris,M.R.,Liang,P.,&Bernstein,M.S.(2023,October).Generativeagents:Interactive simulacraofhumanbehavior.InProceedingsofthe36thannualacmsymposiumonuserinterfacesoftwareand technology(p.1-22).https://doi.org/10.1145/3586183.3606763 Park,J.S.,Popowski,L.,Cai,C.,Morris,M.R.,Liang,P.,&Bernstein,M.S.(2022,October).Socialsimulacra:Creating populatedprototypesforsocialcomputingsystems.InProceedingsofthe35thAnnualACMSymposiumonUserInterface SoftwareandTechnology(p.1-18).https://doi.org/10.1145/3526113.3545616 Pierce,K.M.,&Gilles,C.(2021).Talkingaboutbooks:Scaffoldingdeepdiscussions.TheReadingTeacher,74(4),385-393. https://doi.org/10.1002/trtr.1957 Ploetzner,R.,Dillenbourg,P.,Preier,M.,&Traum,D.(1999).Learningbyexplainingtooneselfandtoothers.Collaborative learning:Cognitiveandcomputationalapproaches,1,103-121. Rai,S.,Shapsough,S.,&Zualkernan,I.(2024,July).MeasuringFluency,CoherencyandLogicalityofGPT-4GeneratedEGRA ComprehensionStories.In2024IEEEInternationalConferenceonAdvancedLearningTechnologies(ICALT)(p. 201-203).IEEE.https://doi.org/10.1109/ICALT61570.2024.00064 Roschelle,J.,&Teasley,S.D.(1995,August).Theconstructionofsharedknowledgeincollaborativeproblemsolving. InComputersupportedcollaborativelearning(p.69-97).Berlin,Heidelberg:SpringerBerlinHeidelberg. https://doi.org/10.1007/978-3-642-85098-1_5 Rossi,L.,Harrison,K.,&Shklovski,I.(2024).TheproblemsofLLM-generateddatainsocialscienceresearch.Sociologica: InternationalJournalforSociologicalDebate,18(2),145-168.https://doi.org/10.6092/issn.1971-8853/19576 Shea,J.H.(1995).Problemswithcollaborativelearning.JournalofGeologicalEducation,43(4),306-308. https://doi.org/10.5408/0022-1368-43.4.306 Tai,R.H.,Bentley,L.R.,Xia,X.,Sitt,J.M.,Fankhauser,S.C.,Chicas-Mosier,A.M.,&Monteith,B.G.(2024).An examinationoftheuseoflargelanguagemodelstoaidanalysisoftextualdata.Internationaljournalofqualitativemethods, 23,16094069241231168.https://doi.org/10.1177/16094069241231168 Tan,S.C.,Lee,A.V.Y.,&Lee,M.(2022).Asystematicreviewofartificialintelligencetechniquesforcollaborativelearning overthepasttwodecades.ComputersandEducation:ArtificialIntelligence,3,100097. https://doi.org/10.1016/j.caeai.2022.100097 Thissen,D.,Steinberg,L.,&Kuang,D.(2002).QuickandeasyimplementationoftheBenjamini-Hochbergprocedurefor controllingthefalsepositiverateinmultiplecomparisons.Journalofeducationalandbehavioralstatistics,27(1),77-83. https://doi.org/10.3102/10769986027001077 VerdĂș,N.,&Sanuy,J.(2014).TheroleofscaffoldinginCSCLingeneralandinspecificenvironments.JournalofComputer AssistedLearning,30(4),337-348.https://doi.org/10.1111/jcal.12047 Viswanathan,N.,Meacham,S.,&Adedoyin,F.F.(2022).Enhancementofonlineeducationsystembyusingamulti-agent approach.ComputersandEducation:ArtificialIntelligence,3,100057.https://doi.org/10.1016/j.caeai.2022.100057 Wang,C.,Fang,T.,&Gu,Y.(2020).Learningperformanceandbehavioralpatternsofonlinecollaborativelearning:Impactof cognitiveloadandaffordancesofdifferentmultimedia.Computers&Education,143,103683. https://doi.org/10.1016/j.compedu.2019.103683 Wang,H.,Wang,C.,Chen,Z.,Liu,F.,Bao,C.,&Xu,X.(2025).ImpactofAI-agent-supportedcollaborativelearningonthe learningoutcomesofUniversityprogrammingcourses.EducationandInformationTechnologies,1-33. https://doi.org/10.1007/s10639-025-13487-8 Wang,X.,Pang,H.,Wallace,M.P.,Wang,Q.,&Chen,W.(2024).LearnersâperceivedAIpresencesinAI-supportedlanguage learning:AstudyofAIasahumanizedagentfromcommunityofinquiry.ComputerAssistedLanguageLearning,37(4), 814-840.https://doi.org/10.1080/09588221.2022.2056203 Xu,R.,Sun,Y.,Ren,M.,Guo,S.,Pan,R.,Lin,H.,...Han,X.(2024).AIforsocialscienceandsocialscienceofAI:Asurvey. InformationProcessingandManagement,61(3),103665.https://doi.org/10.1016/j.ipm.2024.103665 Zhang,J.,Yan,Y.,Yan,J.,Zheng,Z.,Piao,J.,Jin,D.,&Li,Y.(2025,July).AParallelizedFrameworkforSimulating Large-ScaleLLMAgentswithRealisticEnvironmentsandInteractions.InProceedingsofthe63rdAnnualMeetingofthe AssociationforComputationalLinguistics(Volume6:IndustryTrack)(p.1339-1349). https://doi.org/10.18653/v1/2025.acl-industry.94 Zhang,L.,Wu,H.,Duan,T.,&Du,H.(2025).Automaticdeductivecodingindiscourseanalysis:anapplicationoflarge languagemodelsinlearninganalytics.SageOpen,15(4),21582440251390054. https://doi.org/10.1177/21582440251390054 Zhang,Z.,Zhang-Li,D.,Yu,J.,Gong,L.,Zhou,J.,Hao,Z.,...&Li,J.(2025,April).Simulatingclassroomeducationwith llm-empoweredagents.InProceedingsofthe2025ConferenceoftheNationsoftheAmericasChapteroftheAssociation forComputationalLinguistics:HumanLanguageTechnologies(Volume1:LongPapers)(p.10364-10379). https://doi.org/10.18653/v1/2025.naacl-long.520