Paper deep dive
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation
Yearim Kim, Njun Baek, Nojun Kwak
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/21/2026, 3:42:19 AM
Summary
This paper introduces PedaCo, a dual-gatekeeping video authoring pipeline designed to ensure pedagogical rigor in AI-generated educational content. The system employs two layers of structured refusal: Layer 1 involves human educators iteratively reviewing and reshaping AI-generated scripts based on Mayer's Cognitive Theory of Multimedia Learning (CTML) principles, while Layer 2 uses automated metrics to evaluate the final video for instructional coherence and narrative-visual synchronization. Evaluation with 23 educators and automated metrics on 7 topics demonstrates that both layers independently improve instructional quality, suggesting that 'principled resistance' enhances rather than hinders AI content creation.
Entities (11)
Relation Signals (10)
Injun Baek â affiliatedwith â Seoul National University
confidence 99% ¡ INJUN BAEK â , Seoul National University, Samsung Electronics, Republic of Korea
Nojun Kwak â affiliatedwith â Seoul National University
confidence 99% ¡ NOJUN KWAK â , Seoul National University, Republic of Korea
Yearim Kim â affiliatedwith â Seoul National University
confidence 99% ¡ YEARIM KIM â , Seoul National University, Republic of Korea
Injun Baek â affiliatedwith â Samsung Electronics
confidence 95% ¡ INJUN BAEK â , Seoul National University, Samsung Electronics, Republic of Korea
PedaCo â hascomponent â Layer 2: Automated Metric
confidence 95% ¡ The second layer performs a post-synthesis evaluation of the finalized video through a composite metric... We automate the assessment of five dimensions: coherence, redundancy, temporal contiguity, modality, and image quality.
PedaCo â hascomponent â Layer 1: Human Educator Review
confidence 95% ¡ The first layer intervenes before any video is rendered... The educator begins by inputting learning content... An LLM then generates an initial script... The human educator then decides what to accept, what to revise manually, and what to regenerate.
PedaCo â usestheory â Mayer's Cognitive Theory of Multimedia Learning
confidence 95% ¡ we developed the PedaCo (Pedagogical Co- creation), a human-AI collaborative system that operationalizes this resistance. Rooted in Mayerâs Cognitive Theory of Multimedia Learning (CTML)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative-visual synchronization. While neither layer is exhaustive, their synergy ensures that principled resistance--the act of deferring AI output until it meets rigorous standards--becomes a catalyst for higher quality. Evaluation combining a study with 23 educators across 3 topics and automated metrics across 7 topics drawn from established science and philosophy curricula shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners.
Tags
Links
- Source: https://arxiv.org/abs/2608.19812v1
- Canonical: https://arxiv.org/abs/2608.19812v1
Trouble viewing inline? Open PDF directly â
Full Text
15,855 characters extracted from source content.
Expand or collapse full text
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation YEARIM KIM â , Seoul National University, Republic of Korea INJUN BAEK â , Seoul National University, Samsung Electronics, Republic of Korea NOJUN KWAK â , Seoul National University, Republic of Korea To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on multimedia learning theory, while the second employs automated metrics to flag violations in instructional coherence and narrative-visual synchronization. While neither layer is exhaustive, their synergy ensures that principled resistanceâthe act of deferring AI output until it meets rigorous standardsâbecomes a catalyst for higher quality. Evaluation combining a study with 23 educators across 3 topics and automated metrics across 7 topics drawn from established science and philosophy curricula shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners. CCS Concepts:⢠Human-centered computingâ User studies;⢠Applied computingâ Interactive learning environments. Additional Key Words and Phrases: Human-AI Collaboration, Educational AI, Generative AI, Multimedia Learning ACM Reference Format: Yearim Kim, Injun Baek, and Nojun Kwak. 2026. When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation. In Proceedings of Understanding and Engaging Critical Resistance to AI in Education (CHI â26 Workshop). ACM, New York, NY, USA, 5 pages. 1 Introduction While modern AI models [1,2] synthesizes professional-looking educational videos in a minute, surface-level polish does not guarantee pedagogical rigor. Current video generation pipelines often prioritize visual appeal over instructional essentials, such as precise temporal alignment of narration or strategic sequencing of prerequisite concepts. Consequently, outputs are frequently optimized for looking good, rather than teaching well. This gap matters as educators are increasingly expected to adopt AI-generated materials with minimal intervention. The pressure toward seamless, friction-free adoption treats any slowdown as inefficiencyâa stance that risks reducing the educatorâs role from professional decision-maker to passive consumer. We argue, however, that pedagogical friction is not a hurdle to be eliminated but a site of professional accountability. Moments of deliberate hesitationâwhether an educator questioning a scriptâs logical flow or an algorithm flagging a narrative truncationâare precisely where instructional quality is forged. â Both authors contributed equally to this research. â Corresponding author Authorsâ Contact Information: Yearim Kim, yerim1656@snu.ac.kr, Seoul National University, Seoul, Republic of Korea; Injun Baek, jjune1416@snu.ac.kr, Seoul National University, and Samsung Electronics, Republic of Korea; Nojun Kwak, Seoul National University, Seoul, Republic of Korea, nojunk@snu.ac.kr. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Š 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM1 arXiv:2608.19812v1 [cs.AI] 20 Aug 2026 2Kim et al. Table 1. Mayerâs 12 CTML principles [3] for reducing extraneous, managing essential, and fostering generative processing. CoherenceSignalingRedundancySpatial ContiguityTemporal ContiguitySegmenting Pre-trainingModalityMultimediaPersonalizationVoiceImage We conceptualize this approach as principled resistance: deliberate, theory-grounded pushback against AI outputs that fail pedagogical standards. Rooted in Mayerâs Cognitive Theory of Multimedia Learning (CTML) [3]âa framework of 12 empirically validated principles for effective multimedia instruction (Table 1)âwe developed the PedaCo (Pedagogical Co- creation), a human-AI collaborative system that operationalizes this resistance. PedaCo integrates two complementary gatekeeping mechanisms: reviewing scripts through CTML-informed criteria, and an automated metric evaluating finished videos against computationally measurable CTML dimensions. This paper presents the design rationale behind this dual approach, summarizes converging evidence from both human and computational evaluations, and poses open questions about how educational AI systems should balance human agency with automated safeguards. 2 Two Layers of Resistance In our framework, principled resistance takes three concrete forms: rejecting (requesting regeneration), revising (manual editing), and overriding (vetoing automated flags). These are not ad hoc reactions but norm-driven decisions grounded in CTML principles. The framework rests on a simple observation: educators and algorithms are good at catching different kinds of problems: while human educators excel at identifying nuanced pedagogical mismatchesâsuch as content being too advanced for a target audienceâalgorithms provide precise, high-resolution verification of structural integrity, such as identifying temporal misalignments between narration and visuals. Building both checkpoints into the same pipeline creates overlapping coverage that neither could achieve alone. 2.1 Layer 1: Human Educator Review at the Script Stage The first layer intervenes before any video is rendered (Figure 1 left). The educator begins by inputting learning content and configuring which CTML principles the system should enforce. An LLM [4] then generates an initial script, which passes through a structured review cycle. An AI reviewerâitself prompted with CTML principlesâproduces feedback organized by principle, identifying potential violations rather than definitive judgments (e.g., âScene 3 introduces technical terms without prior explanation, which may conflict with the Pre-training principleâ). The human educator then decides what to accept, what to revise manually, and what to regenerate. This review loop can be repeated until the educator is satisfied with the script. This design choice to review at the script level is deliberate. Textual revisions are computationally and laboriously efficient, whereas pedagogical errors baked into a rendered visual narrative are nearly impossible to correct post- synthesis. By placing the human checkpoint at this intermediate stage, we make pedagogical critique economically viable. This ensures the system remains advisory, preserving the educatorâs professional authority to say ânoâ to AI suggestions based on their specific curricular context and pedagogical style. 2.2 Layer 2: Automated Metric After Video Synthesis The second layer performs a post-synthesis evaluation of the finalized video through a composite metric (Figure 1 right). We automate the assessment of five dimensions: coherence, redundancy, temporal contiguity, modality, and image quality. The educator reviews the principle-level scores and decides whether to accept the video or return to the script stage for targeted revision. This boundary between human and automated resistance is a strategic design decision. While dimensions like temporal synchronization are amenable to reliable computational measurement, othersâsuch as Manuscript submitted to ACM When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation3 a b c d e f g h i Fig. 1. The Dual Gatekeeping Interface. The left panel (Layer 1: Script Level) generates an initial script (í) from learning content (í) and generation principles (í), then scaffolds the educatorâs revision by providing AI critiques and revised drafts (í) based on review constraints (í). The right panel (Layer 2: Video Level) visualizes invisible pedagogical quality via automated metrics (â,í), allowing users to assess the alignment between the final video (í) and the learning content (í ). Personalization (tone)ârequire a deep understanding of curricular structures and learner psychology that currently remains the sole domain of the human expert. By automating only where algorithmic feasibility aligns with pedagogical necessity, we create a robust safety net that prevents technical regressions without marginalizing human judgment. 3 Evidence from Two Evaluations We conducted a multi-method evaluation to assess the efficacy of the PedaCo framework, combining a human-centric study with educators (3.1) and an algorithmic assessment via automated metrics (3.2). Together, these evaluations provide converging evidence for the value of dual-layered resistance. 3.1 What Educators Found We conducted a within-subject study with 23 educatorsâwho were priorly briefed on the CTML principlesâdirectly used our system to generate and evaluate videos. Guided by the system, participants simulated with the pre-generated videos by inputting raw learning content and iteratively refining the AI-generated scripts with the systemâs CTML feedback. Then, they compared these videos outcomes with baseline videos generated without CTML guidelines. The topics include three different cognitive demands: causal reasoning, abstract concepts, and procedural knowledge. Participants rated each condition on 13 items covering all 12 CTML principles and overall instructional validity. The review-based approach yielded statistically significant improvements across every principle (í< .05, Wilcoxon signed-rank). The mean rating rose from 3.07 to 3.86 on a 5-point scale (+0.79,í< .01), with the most pronounced gains observed in content organization: prerequisite sequencing (+0.86), irrelevant material removal (+0.84), and overall instructional validity (+0.96)ove. Nearly all effects were large (í ⼠.64, computed así/ â í), with one exception (redundancy, í= .42, medium), and consistent across gender and experience level. Notably, educators did not perceive the review process as slowing them down. They rated production efficiency at 4.26/5 and the validity of the CTML-based guidance at 4.04/5 (with remarkably low variance,ííˇ=0.62). One participant captured the tension well: the iterative process was âquite challengingâ but ultimately for producing âa robust and effective learning toolâ (P23). Another explicitly requested âseparate functions where teachers can additionally review, edit, and modifyâ the AI output (P05)âin other words, more resistance, not less. Manuscript submitted to ACM 4Kim et al. 3.2 What the Metrics Found Independently, we applied our automated metrics to a corpus of 14 videos (7 topicsĂ2 conditions) drawn from established science and philosophy curricula. The videos were generated with identical structure, isolating the CTML- informed generation and review as the sole variable. Two of the five metrics showed significant improvement: temporal contiguity (0.294 vs. 0.273,í= .021) and coherence (0.729 vs. 0.646,í= .011). The remaining three showed no significant differenceâmodality and redundancy scored high in both conditions (near ceiling), and image quality did not differ as both conditions used the same video synthesis model. We retain these metrics as safety nets: current models perform well on these dimensions, but they serve as guardrails against future model regressions or hallucination-induced failures. 3.3 Where the Two Evaluations Agree The most compelling evidence for our dual-layered approach is the high degree of convergence between subjective ratings and objective metrics. Despite being conducted independently with different samples and instruments, both evaluations identified coherence and temporal alignment as the dimensions most significantly enhanced by the PedaCo pipeline. Educators ranked coherence among the top three improvements, matching the statistically significant gains identified by the automated system. This triangulation suggests that the two gatekeeping layers are not merely redundant but provide complementary verification of the same underlying instructional quality. 4 Discussion: Reframing Resistance in ducational AI The PedaCo framework offers a theory-grounded perspective on where the boundaries of Generative AI should be drawn in educational settings. By operationalizing "principled resistance" through CTML, we illustrate how intentional friction can sustain pedagogical integrity without marginalizing the professional authority of educators. Our findings suggest that "productive friction" is most effective when it is: (a) theoretically grounded rather than intuitive; (b) strategically embedded upstream at the script stage; and (c) hybridized across human and computational agents. However, this dual-gatekeeping approach surfaces three emergent tensions for the workshop to consider: ⢠Negotiating Agency: When automated flags and educator judgments diverge, how should the interface balance algorithmic safeguards with human autonomy? â˘Sustainability of Friction: While valued in our short-term study, the long-term viability of iterative review in daily classroom preparation remains unknown. We must identify the threshold where productive friction transitions into "friction fatigue." ⢠Beyond Proxy Metrics: Future research must move beyond theoretical proxies to evaluate the direct causal impact of structured resistance on student learning outcomes. 5 Conclusion In this paper, we have argued that educational AI needs structured ways to say ânot yetââbridging human pedagogical expertise with computational precision. Our PedaCo framework builds principled resistance into the video generation pipeline through educator review at the script stage and automated pedagogical metrics at the video stage. Early evidence suggests both layers improve the same instructional dimensions, and that educators experience this friction as productive rather than burdensome. We offer this as one concrete answer to the workshopâs animating question: resistance to AI in education should not be equated with rejection. Instead, it can mean building systems that are designed to push back, on principled grounds, until the output is genuinely ready to teach. Manuscript submitted to ACM When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation5 Acknowledgments This work was funded by the Korean Government through the grants from IITP (RS-2021-I211343) and KOCCA (RS-2024-00398320). References [1]Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. 2024. Video generation models as world simulators. (2024). https://openai.com/research/video-generation-models-as- world-simulators [2] Google DeepMind. 2024. Veo: Googleâs most capable generative video model. https://deepmind.google/models/veo/. Accessed: 2025-09-29. [3] Richard E. Mayer. 2009. Multimedia Principle. Cambridge University Press, 223â241. [4] Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023). Manuscript submitted to ACM