Paper deep dive
Talking Slide Avatars: Open-Source Multimodal Communication Approach for Teaching
Xinxing Wu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/21/2026, 7:16:53 AM
Summary
This paper presents an open-source, practice-based workflow for creating 'talking slide avatars' to enhance instructor presence in online and asynchronous teaching. The workflow integrates OpenVoice for text-to-speech and voice cloning with Ditto-TalkingHead for audio-driven image synthesis. The study frames these avatars as multimodal communication artifacts that serve as a reusable layer for introductions, transitions, and recaps in slide-based instruction. The author emphasizes a modular design approach—separating script, voice, and image—to allow for efficient revisions and aesthetic continuity, while providing practical guidelines on clip length (10-30 seconds), image selection, and ethical disclosure.
Entities (7)
Relation Signals (5)
Xinxing Wu → affiliatedwith → Kentucky State University
confidence 100% · Xinxing Wu 1 * 1 School of Mathematics and Computer Science, Kentucky State University
OpenVoice → partofworkflow → Talking Slide Avatar production
confidence 100% · The workflow integrates OpenVoice... with Ditto-TalkingHead... for creating talking slide avatars
Ditto-TalkingHead → partofworkflow → Talking Slide Avatar production
confidence 100% · The workflow integrates OpenVoice... with Ditto-TalkingHead
OpenVoice → usedfor → Text-to-speech generation
confidence 100% · The workflow integrates OpenVoice for text-to-speech generation and voice cloning
Ditto-TalkingHead → usedfor → Audio-driven talking-image synthesis
confidence 100% · with Ditto-TalkingHead for audio-driven talking-image synthesis
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose the instructor presence, narrative continuity, and expressive framing that help learners connect with content. Full lecture video can partly restore these qualities, but it is time-consuming to record, revise, and reuse. This study addresses that pedagogical and production challenge by presenting a practice-based analysis of an open-source workflow for creating talking slide avatars for slide-based teaching. The workflow integrates OpenVoice for text-to-speech generation and voice cloning with Ditto-TalkingHead for audio-driven talking-image synthesis, enabling instructors to transform a script and a static portrait into a short narrated video that can be embedded in slide decks or HTML-based lecture materials. Rather than treating this workflow merely as a technical solution, the study frames talking slide avatars as multimodal communication artifacts at the intersection of digital pedagogy, aesthetic education, and art-technology practice. Using a practice-based implementation and analytic reflection approach, the study documents the production pipeline, examines its communicative and aesthetic affordances, and proposes practical guidelines for script length, image selection, pacing, disclosure, accessibility, and ethical use. The study makes three primary contributions: it presents an educator-oriented open-source production model, reframes talking avatars as an educational communication design problem, and proposes a responsible pathway for incorporating generative synthetic media into teaching. It concludes that short, transparent, and carefully designed avatars can humanize slide-based instruction while providing a reusable communicative layer for introductions, transitions, reminders, and recaps across online, hybrid, and asynchronous learning environments.
Tags
Links
- Source: https://arxiv.org/abs/2604.23703v1
- Canonical: https://arxiv.org/abs/2604.23703v1
Trouble viewing inline? Open PDF directly →
Full Text
43,373 characters extracted from source content.
Expand or collapse full text
Talking Slide Avatars: Open-Source Multimodal Communication Approach for Teaching Xinxing Wu 1 * 1 School of Mathematics and Computer Science, Kentucky State University, Frankfort, Kentucky, United States *Corresponding author: Xinxing Wu (xinxingwu@gmail.com; xinxing.wu@kysu.edu) Abstract Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose the instructor presence, narrative continuity, and expressive framing that help learners connect with content. Full lecture video can partly restore these qualities, but it is time-consuming to record, revise, and reuse. This study addresses that pedagogical and production challenge by presenting a practice-based analysis of an open-source workflow for creating talking slide avatars for slide-based teaching. The workflow integrates OpenVoice for text-to-speech generation and voice cloning with Ditto-TalkingHead for audio-driven talking-image synthesis, enabling instructors to transform a script and a static portrait into a short narrated video that can be embedded in slide decks or HTML-based lecture materials. Rather than treating this workflow merely as a technical solution, the study frames talking slide avatars as multimodal communication artifacts at the intersection of digital pedagogy, aesthetic education, and art-technology practice. Using a practice-based implementation and analytic reflection approach, the study documents the production pipeline, examines its communicative and aesthetic affordances, and proposes practical guidelines for script length, image selection, pacing, disclosure, accessibility, and ethical use. The study makes three primary contributions: it presents an educator-oriented open-source production model, reframes talking avatars as an educational communication design problem, and proposes a responsible pathway for incorporating generative synthetic media into teaching. It concludes that short, transparent, and carefully designed avatars can humanize slide-based instruction while providing a reusable communicative layer for introductions, transitions, reminders, and recaps across online, hybrid, and asynchronous learning environments. Keywords: AI avatar; multimodal communication; instructional video; art and technology; higher education; talking head synthesis 1. Introduction Slide-based teaching remains one of the most common approaches for organizing and presenting content in higher education, 1-2 yet the communicative experience of slides is often uneven across learning contexts. In face-to-face classrooms, an instructor’s voice, pacing, emphasis, and body language can animate a slide deck, transforming static bullet points into a coherent and engaging explanation. In online, hybrid, and asynchronous settings, however, this live layer of instructional presence is frequently diminished, absent, or costly to reproduce. 3 Recorded lectures can partially mitigate this limitation, but full video production often demands repeated recording, editing, and revision whenever course content changes. For many instructors, particularly those managing multiple sections or updating materials each semester, this creates a practical tension between pedagogical goals and sustainable content production. 4 At the same time, research on multimedia learning and instructional video has shown that voice, temporal pacing, and social cues matter. Well-designed instructional videos can support engagement and comprehension, especially when they balance verbal explanation with clear visual focus. 5-8 Research on pedagogical agents further suggests that learners often interpret human-like on-screen agents socially rather than purely instrumentally, which means that voice, demeanor, and communicative style can shape how instructional messages are received. 9-12 Recent work on AI-based educational avatars extends this conversation by examining how synthetic speakers, embodied agents, and avatar-based interfaces may alter the experience of teaching and learning. 13-14 This paper focuses on talking slide avatars for slide-based teaching, course lectures, and related forms of course communication, and it combines two open-source projects, OpenVoice and Ditto- TalkingHead, in an educator-friendly workflow. We use open-source components to generate speech from a script, optionally clone a desired voice style, animate a static portrait image, and export a short talking video that can be inserted into slides or course webpages. 15-17 The central claim of this paper is that the significance of this workflow extends beyond its technical function. Talking slide avatars can be understood as communicative and aesthetic artifacts situated at the intersection of digital portraiture, speech performance, instructional design, and mediated presence. Their implementation also invites engagement with broader perspectives from art and technology, communication studies, and aesthetic education. The goal of this paper is therefore to offer an original applied contribution to communication- oriented instructional design. Specifically, it: (1) presents a detailed implementation pipeline; (2) examines the communicative, aesthetic, and pedagogical affordances of the resulting artifacts; (3) proposes practical design guidelines for classroom use; and (4) addresses limitations, ethical considerations, and future directions. In doing so, the paper reframes talking avatars as a reusable layer of course communication that can humanize slide-based instruction without requiring full- scale video production for entire lessons. The remainder of the paper reviews the relevant scholarly context, documents the workflow, analyzes its prototype outcomes, considers broader implications for communication, ethics, and pedagogy, and concludes with limitations and future directions. 2. Conceptual and scholarly context This paper draws on three overlapping areas: multimedia learning, pedagogical agents, and contemporary AI-avatar research. First, multimedia learning research has long emphasized that students learn from coordinated combinations of words and images rather than from either channel alone. Instructional videos extend this principle by adding timing, narration, pacing, and attention guidance. Studies of educational video production have shown that relatively small design choices, such as clip length, segmentation, and conversational delivery, can have measurable implications for engagement. In addition, systematic reviews of instructor presence in instructional videos show a more nuanced picture: visually present instructors do not automatically improve learning outcomes, yet carefully designed social and attentional cues can increase affective engagement and support aspects of learning in specific conditions. Second, research on pedagogical agents suggests that learners often respond to computer-based agents as social actors. Animated or embodied instructional agents may trigger what has been described as social agency, in which learners process the message as if they are in a form of guided interaction rather than simply receiving neutral information. This does not mean that any animated face is educationally beneficial; indeed, agent design effects are often mixed and highly dependent on appearance, voice, context, and task type. What the literature does support, however, is the idea that voice quality, politeness, social cues, and a coherent persona matter in multimedia environments. For a project such as talking slide avatars, this means the design question is not “Can a face move with speech?” but “What kind of communicative relationship does the avatar establish with the learner?” Third, the rapid development of generative AI has made synthetic avatar production more accessible. Recent scholarship has begun to examine AI-based avatars as educational tools, highlighting both their potential and their risks. One empirical study comparing video lectures with human instructors and AI-generated instructors found lower engagement for the AI-generated version but comparable learning performance, suggesting that technical realism and social acceptance still shape reception even when content equivalence is preserved. 18 This is an important finding because it pushes design beyond novelty. The challenge is not just to create an avatar, but to create one whose visual and vocal behavior does not distract, alienate, or undermine trust. Within the broader field of arts and communication, AI-generated media are increasingly recognized as transformative forces in image-making, creativity, and cultural production. Recent scholarship on AI, digital art, and mediated communication suggests that generative systems are not simply instruments of automation; they also reconfigure aesthetic authorship, mediation, and reception. 19–23 The talking slide avatar belongs to this broader landscape. It is neither a conventional lecture recording nor merely a software output. Rather, it is a designed performance artifact: a synthetic micro-presentation in which voice, face, timing, and slide content are intentionally orchestrated to communicate meaning. This artistic-communicative framing is particularly relevant in educational settings, where the design of explanation is inseparable from the design of attention, engagement, and trust. The present study contributes to this conversation by offering a concrete, reusable, and comparatively low-cost open-source approach to producing such artifacts. Unlike proprietary avatar-generation platforms that obscure production logic behind closed interfaces, the workflow developed here renders the process visible and editable through scripts, notebooks, and documented procedures. Such transparency matters both academically and pedagogically, because it allows the artifact to be analyzed, adapted, critiqued, and repurposed not only as a teaching resource, but also as a case of communication-oriented design. 3. Materials and methods This study adopts a practice-based implementation and analytic reflection approach. Rather than evaluating learner outcomes through a formal experimental design, the paper presents and reflects on an open-source workflow developed to support the creation of talking slide avatars for teaching. This approach is appropriate because the aim of the study is not to test learner outcomes experimentally, but to document, interpret, and evaluate the communicative and pedagogical significance of an implementable workflow. The implementation accompanying this paper includes a README, a text-to-speech notebook (TextToVoice.ipynb), a talking-image notebook (ImageSpeaking.ipynb), supporting dependencies, and integrated open-source components related to OpenVoice and Ditto-TalkingHead. The purpose of this paper is not to analyze a repository as an object in itself, but to document the workflow, clarify how it translates a written lecture script into a speaking slide artifact, and examine the communicative design possibilities that such a workflow makes available. As summarized in Figure 1, the workflow can be summarized in four stages. First, the instructor prepares a short script for a slide segment, such as a weekly introduction, an assignment explanation, or a concept summary. Second, the script is converted to speech using OpenVoice. In the implementation provided with this paper, the workflow initializes the BaseSpeakerTTS model and the ToneColorConverter, selects an English base speaker checkpoint, and, when desired, extracts speaker embeddings from a reference audio clip to approximate a preferred vocal style. The text is then rendered into a temporary waveform and converted into a cloned or style-adjusted output audio file. Third, the generated audio is combined with a portrait image in Ditto- TalkingHead. In the implementation, this stage uses a Colab-based setup, downloads model checkpoints, installs dependencies, and runs an inference process to synthesize an MP4 talking- avatar clip from the audio-image pair. Fourth, the resulting short video is embedded into lecture slides or HTML-based course pages, where it functions as a compact speaking element rather than as a full-length lecture recording. Several implementation characteristics are especially important. The notebooks are structured for Google Colab and Google Drive, which reduces local installation demands for instructors who may not wish to maintain a dedicated machine-learning environment. The workflow also separates speech generation from facial animation, making the process modular. This modularity is pedagogically useful because instructors can revise scripts, change a voice reference, or replace the avatar image without redesigning the entire presentation. In addition, the implementation recommends short clips, clear front-facing images, and moderate speech pace, all of which align with established guidance on instructional video segmentation and attention management. Short script for: weekly goals assignment prompts concept summaries OpenVoice converts text to speech and can use a reference voice for tone, pacing, and style cues. Ditto - TalkingHead uses speech audio + a portrait image to generate a lip - synced talking clip. The output video is placed into PowerPoint, Google Slides, or HTML slides as a micro - explanation layer. Design logic The instructional value of the workflow comes from segmentation, embodiment, and reuse. Rather than replacing full lectures, short avatar clips are best used for openings, transitions, reminders, recaps, and orientation cues that humanize otherwise static slide sequences. text audio + image mp4 Slides Script writing Voice generation Talking image Slide embedding Open - source workflow for talking slide avatar production Figure 1. Workflow for talking slide avatar production. The analytic reflection in this paper is organized around three dimensions. The first is technical reproducibility: whether the workflow is documented clearly enough to be followed, adapted, and reused by educators. The second is communicative function: what kinds of instructional messages the talking-avatar format is particularly well suited to deliver. The third is aesthetic-pedagogical design: how the synthetic combination of image, speech, timing, and slide context shapes the tone, presence, and meaning of the communication. Because this paper does not report human-subject data, survey data, or classroom experiments, it does not claim measured learning gains. Instead, it offers an original implementation account together with a theoretically grounded discussion of why this workflow matters for educational communication. The technical implementation is intentionally concrete. In the text-to-speech stage, the workflow mounts cloud storage, clones the OpenVoice repository, installs required dependencies, loads checkpoints for the English base speaker and converter, generates a temporary waveform from the instructional script, and then applies tone-color conversion using a reference speaker clip. In the talking-image stage, the workflow mounts storage again, clones Ditto-TalkingHead, downloads checkpoints, installs PyTorch and supporting packages, and runs an offline inference command that takes a source image, an audio waveform, and an output path as its core inputs. The separation of these stages is more than a coding convenience; it creates a pipeline that can be documented, taught, maintained, and revised in modular form. This modularity also has methodological and practical value. Speech quality can be revised without changing the portrait image. Visual presentation can be revised without rewriting the instructional script. Embedding decisions can be adjusted without rerunning earlier media- generation stages once the MP4 artifact is satisfactory. For educational practitioners, this separation reduces the fragility that often discourages adoption of AI media tools. A workflow that breaks under minor revision is difficult to sustain pedagogically; a workflow organized into separable stages is far more likely to remain reusable over time. Table 1 below summarizes the key components of the workflow and their roles in the overall implementation. Table 1. Implementation components. Component Role in workflow Implementation in This Study Script Defines the message to be spoken Short lecture text prepared by the instructor for introductions, assignment explanations, or concept summaries. OpenVoice Text-to-speech generation and voice-style conversion BaseSpeakerTTS generates source audio; ToneColorConverter and speaker embedding extraction allow style transfer from a reference clip. Reference audio Shapes the vocal color or identity of the output A sample MP3 is used to extract speaker embeddings for cloned or style-adjusted speech. Ditto- TalkingHead Audio-driven talking image synthesis The inference pipeline combines a source image and generated WAV audio to produce an MP4 talking- avatar clip. Portrait image Provides the visual anchor for embodiment A clear, front-facing image is recommended to improve lip-sync plausibility and visual stability. Slide or HTML page Presentation context for the avatar The final MP4 can be embedded into PowerPoint, Google Slides, or HTML-based course slides. 4. Prototype outcomes and design analysis The implementation presented in this paper defines a compact production model for talking slide avatars. Its most important outcome is not any single generated video, but the establishment of a repeatable instructional workflow. This is significant because many educators do not require a fully autonomous conversational avatar. Rather, they need a lightweight and sustainable method for adding a brief, teacher-like speaking presence to otherwise static slide sequences. The workflow developed here addresses precisely that instructional niche. 4.1 Instructional use cases From a communication-design perspective, the workflow appears particularly well suited to four scenarios. The first is slide opening and orientation. A short avatar clip can welcome students, introduce the agenda, or frame expectations before the main content begins. The second is transition management. In longer slide decks, the avatar can signal movement from one conceptual block to another, reducing the abruptness that often characterizes asynchronous slide viewing. The third is task explanation. Assignment instructions, weekly reminders, and logistical updates are often too minor to justify full video production, yet they still benefit from a direct voice and visible face. The fourth is emphasis and recap. A short speaking avatar at the end of a section can reinforce a key takeaway and provide a more memorable closing rhythm than a final static bullet list alone. Figure 2. Example of a talking slide avatar embedded in a slide-based lecture interface. The figure illustrates how the avatar functions not as a full lecture substitute, but as a compact communication layer within the slide environment. A major practical strength of the system is its modularity. Because the script, reference voice, portrait image, and embedding context remain separable, instructors can make small revisions efficiently. For example, an instructor may preserve the same avatar image across an entire course while changing only the script from week to week, thereby creating a stable visual persona. Alternatively, one may maintain a recurring opening style while adapting the spoken content for different announcements or modules. This flexibility supports what may be described as aesthetic continuity: the repeated use of a recognizable visual-voice motif that helps students interpret these clips as part of a coherent instructional environment rather than as isolated novelties. The implementation also highlights an important artistic dimension. The avatar clip functions less as a neutral narration layer than as a micro-performance. The generated voice is timed, stylized, and embodied through a portrait image. When embedded within slides, the clip transforms a conventional presentation into a mixed-media object composed of text, image, recorded sound, and synthetic animation. In this sense, the instructor is not only a content author, but also a designer of mediated presence. At the same time, the implementation reveals several design boundaries. The first is duration. The accompanying workflow recommends short clips of approximately 10 to 30 seconds, and this recommendation is well justified. Short duration helps preserve novelty, reduces rendering cost, minimizes the visibility of animation artifacts, and aligns with established guidance favoring segmented instructional video over long, uninterrupted presentation streams. The second boundary is image quality. The avatar performs most effectively when driven by a clear, front-facing image with stable lighting and minimal occlusion. This requirement is both technical and aesthetic: weaker source images not only reduce lip-sync plausibility, but may also diminish perceived trustworthiness and visual harmony with the surrounding slide design. A third boundary concerns realism. Prior work on AI-generated instructors suggests that imperfect facial animation can lower engagement even when instructional performance remains broadly comparable. That finding helps explain why the talking slide avatar is best positioned as a supplementary communicative layer rather than a complete substitute for live teaching or full recorded lecture presence. The format is most persuasive when assigned specific, bounded functions such as orientation, transition, emphasis, or recap. It is less convincing when expected to sustain extended explanation in a highly human-realistic style. Taken together, these outcomes suggest that the value of the prototype lies not in replacing teaching, but in enriching slide-based instruction with a scalable and reusable layer of voice, face, and timing. Its strongest contribution is therefore design-oriented: it offers a practical model for introducing mediated instructor presence into slide environments in ways that are lightweight, adaptable, and aesthetically coherent. 4.2 Modularity and aesthetic continuity An additional outcome of the study is the way the avatar redistributes attention across the slide surface. In a conventional deck, the student’s visual pathway is often determined by bullet order or by whatever element is visually most dominant. The avatar clip introduces a temporal focal point. For a brief interval, it can orient the learner toward a key phrase, a section title, or a transition statement before the rest of the slide is read. This is not simply decorative motion. It is a form of signaling through embodied timing. When the avatar appears at the right moment, it can narrow interpretive ambiguity and indicate where the communicative center of the slide lies. Table 2. Design guidelines for classroom deployment. Design issue Recommendation Rationale Clip length Keep most clips within 10-30 seconds. Short clips preserve attention, reduce rendering time, and align with segmented instructional video design. Message type Use avatars for openings, transitions, reminders, and recaps rather than dense full lectures. The format is strongest for bounded communication tasks, not long continuous exposition. Image selection Use front-facing images with clear lighting and minimal occlusion. Better source images improve both technical performance and perceived trust. Design issue Recommendation Rationale Voice design Prefer natural pacing and moderate speech speed; avoid overly dramatic tone shifts. A stable and conversational voice reduces distraction and supports message clarity. Disclosure Label clips as AI-assisted or AI- generated where appropriate. Transparency helps protect trust and distinguishes legitimate educational media from deceptive synthetic content. Accessibility Provide captions or transcripts whenever possible. Students benefit from multimodal access and review options. Aesthetic consistency Reuse a small set of coherent avatar styles across a course. Consistency helps the avatar function as a recognizable communicative motif rather than a novelty. Frequency of use Insert selectively at key slide moments. Overuse may create redundancy, fatigue, or unnecessary cognitive load. Consent and identity Use only authorized images and voices, or clearly non-personal avatar assets. Ethical deployment requires control over likeness, authorship, and permissions. The workflow is also notable for how it balances sameness and variation. Across a semester, an instructor may wish students to encounter a stable instructional persona without producing a monotonous media experience. The workflow makes that possible by allowing different combinations of changed and unchanged elements. The portrait can remain constant while the script changes weekly; the voice style can remain stable while the image changes to suit a theme; and the same avatar can appear in different slide positions or visual frames depending on the lecture segment. These choices amount to a grammar of synthetic course presence. They enable the instructor to determine how much continuity and how much variation course communication should display. This grammar also has aesthetic consequences. A recurring avatar at the beginning of each weekly module can function almost like a title card or signature motif. In design terms, it becomes part of the course identity system. Students may begin to associate a certain voice timbre, portrait style, and pacing pattern with the start of a topic or the framing of an important reminder. Such patterned repetition is not foreign to art and media practice; it resembles the use of recurring motifs in broadcast design, animation, and interactive media. In an educational context, that recurrence can make course navigation feel more coherent and intentional. 4.3 Rhetorical consequences and deployment guidelines The talking slide avatar also encourages a different relationship to script writing. Traditional slide text is often compressed, noun-heavy, and written to be read silently. By contrast, avatar clips demand spoken language. Sentences must be pronounceable, paced, and listenable. This encourages instructors to translate dense slide prose into more conversational explanations. That shift may indirectly improve communication quality because the script is forced to function as speech rather than as overpacked text. In this sense, the workflow may also serve as a rhetorical discipline, pushing instructors toward clearer oral-style framing. To translate these observations into practical terms, Table 2 presents design guidelines for classroom deployment derived from the present analysis and the supporting literature. These design observations also point beyond implementation itself toward broader questions of communication, ethics, pedagogy, and mediated authorship. 5. Broader implications for communication, ethics, and pedagogy Beyond its immediate instructional use, the workflow presented in this paper has broader implications for communication, ethics, and pedagogy. Its significance lies not only in enabling the production of talking slide avatars, but also in showing how teaching materials are increasingly becoming hybrid media artifacts. In this workflow, instructional communication is shaped through the coordinated design of script, voice, portrait image, timing, animation, and slide context. The resulting avatar is therefore not merely a software-generated output, but a compact communicative performance situated at the intersection of education, media design, and synthetic representation. In this sense, the workflow participates in wider transformations of authorship, mediation, and aesthetic production associated with contemporary AI-generated media. 24-30 5.1 Openness and pedagogical reflexivity A second implication concerns openness and pedagogical reflexivity. Many commercial avatar platforms prioritize convenience and visual polish, but they often conceal the intermediate production process behind closed interfaces. By contrast, the workflow presented here remains comparatively transparent. Scripts can be revised directly, voice references can be inspected and changed, checkpoints and dependencies are explicit, and the final output is produced through documented steps. This transparency matters educationally because it supports adaptation, critique, and learning. The workflow can therefore function not only as a teaching aid, but also as an object of study for students in computing, digital media, communication, and art-and-technology contexts who wish to examine how synthetic media are assembled, interpreted, and deployed. A third implication involves teaching sustainability and accessibility. Short avatar clips can reduce the need to repeatedly record routine course messages, especially in online and hybrid environments where instructors frequently update reminders, weekly introductions, or brief explanations. The same workflow may also support future accessibility enhancements, including multilingual delivery, transcript generation, captioning, and more flexible voice-style adaptation. At the same time, such possibilities should be approached cautiously. Guidance on generative AI in education emphasizes human-centered design, accountability, transparency, and inclusion, all of which are directly relevant to synthetic instructional media. 22 For talking avatars, responsible use therefore requires clear disclosure that the clip is AI-assisted, careful permission practices for voice and image sources, and continued instructor accountability for the pedagogical message being communicated. 5.2 Trust, disclosure, and realism Trust is especially important because synthetic instructional media can approach the communicative territory of deepfakes and related forms of deceptive representation. In educational settings, the distinction between acceptable synthetic assistance and problematic deception depends not only on the model itself, but also on context, disclosure, authorship, and intent. A transparently labeled avatar used by an instructor to present a scripted course message is ethically distinct from an unlabeled synthetic representation that implies a naturally recorded or live performance. The workflow is therefore strongest when it presents the avatar as a clearly designed media object rather than as a hidden simulation of unmediated presence. Questions of realism introduce a further communicative and ethical implication. Research on avatar realism and humanoid interfaces suggests that highly human-like agents may become less effective when facial expression, timing, or motion appear almost realistic but not fully convincing, thereby producing distraction or unease. 23,29 For course communication, the design goal should not necessarily be maximal realism. In many situations, a modestly stylized, clearly artificial, and visually coherent avatar may be more effective than one that strives for perfect human substitution. This reinforces the argument that the talking slide avatar should be understood as a communication artifact whose success depends on aesthetic fit, rhetorical function, and contextual appropriateness rather than on raw model fidelity alone. 5.3 Authorship and broader applications From a pedagogical standpoint, the workflow is especially promising when used selectively. It can strengthen course openings, improve continuity in asynchronous modules, provide more personable assignment guidance, and reduce the abruptness of purely text-based slide transitions. However, overuse may weaken its effectiveness by creating redundancy, slowing slide navigation, or turning a useful communicative cue into decorative excess. This aligns with broader work on multimedia learning and instructional video, which suggests that signaling, segmentation, pacing, and purposeful social cues are generally more beneficial than simply increasing media density. 30- 31 The most effective role for the talking slide avatar is therefore not to dominate a lesson, but to punctuate and frame it. A further implication concerns authorship. In conventional slide-based teaching, authorship is often divided between the instructor as content author and the software as a relatively neutral delivery medium. In a talking-avatar workflow, that distinction becomes less stable. The final artifact emerges through the combined action of instructor decisions, scripted language, source imagery, voice-reference choices, and model behavior. This does not erase human authorship, but it does make authorship more distributed and collaborative. For scholarship in art, media, and communication, that hybrid authorship is itself a meaningful subject of inquiry, particularly as generative systems continue to reshape creative practice, mediation, and aesthetic reception. 24-30 Finally, the workflow may have value beyond conventional lectures. It could be adapted for museum micro-explanations, student project showcases, digital exhibitions, conference posters with embedded narration, or multilingual public-facing educational materials. Because the artifact is essentially a short-narrated speaking image, it can travel across educational and cultural communication settings with relatively little modification. That portability strengthens the case for understanding the workflow not merely as an instructional utility, but as part of a broader art-and- technology and synthetic-media discourse. 24-30 6. Limitations and future directions This study has several limitations. o First, it is an implementation-and-analysis paper rather than a controlled classroom experiment, and it therefore does not provide direct evidence of student learning outcomes, satisfaction, engagement, or retention associated with the workflow. o Second, the discussion is based on one implementation configuration centered on OpenVoice and Ditto-TalkingHead; future tools may alter the balance among output quality, production speed, controllability, and ease of use. o Third, the paper focuses specifically on short, portrait-based avatar clips and does not extend to conversational agents, interactive avatar systems, or fully dynamic video lecturers. Future research should advance in at least four directions. o First, empirical evaluation is needed. Comparisons across static slides, audio-only narration, human-recorded talking-head clips, and AI-generated talking slide avatars would help identify the conditions under which this format is most effective, as well as the points at which it may become distracting or pedagogically limited. o Second, accessibility should be strengthened through more explicit integration of automatic captioning, transcript export, language switching, and playback controls. o Third, further visual experimentation is warranted. Researchers in art, communication, and design could investigate how stylization, framing, gaze direction, and graphic integration influence audience reception and meaning-making. o Fourth, governance will become increasingly important. As avatar production becomes more accessible, institutions will need practical policies regarding consent, disclosure, archival transparency, and acceptable use in formal teaching materials. 7. Conclusion The implementation presented in this paper demonstrates that talking slide avatars can be produced through an open-source workflow that is feasible, modular, and pedagogically meaningful. By combining script writing, synthetic speech generation, talking-image synthesis, and slide integration, the workflow introduces a new layer of course communication that occupies a productive middle ground between static slides and full lecture video. Its value lies not only in technical efficiency, but also in its capacity to add timing, tone, and embodied presence to slide- based instruction. For scholarship in arts and communication, this work is significant because it shows how teaching materials are increasingly becoming hybrid media artifacts shaped by animation, voice design, portrait aesthetics, and synthetic performance. For educators, it offers a practical model for producing short, reusable clips that can welcome, orient, explain, and recap. For future research, it opens a pathway for deeper investigation into trust, realism, design, and ethics in synthetic educational media. Talking slide avatars should not be understood as replacements for teachers. Rather, when used transparently and thoughtfully, they function as compact communicative performances that can make digital teaching environments more coherent, more expressive, and more humanly legible. References 1.Baker JP, Goodboy AK, Bowman ND and Wright A. Does teaching with PowerPoint increase students' learning? A meta-analysis. Computers & Education. 2018; 126: 376-387. doi: 10.1016/j.compedu.2018.08.003 2.Chávez HD, Ramón CP and Castelló TA. Patterns of PowerPoint use in higher education: a comparison between the natural, medical, and social sciences. Innovative Higher Education. 2020: 45(1): 65-80. doi: 10.1007/s10755-019- 09488-4 3.Li W and Wang W. The impact of teaching presence on students’ online learning experience: Evidence from 334 Chinese universities during the pandemic. Frontiers in psychology. 2024; 15: 1291341. doi: 10.3389/fpsyg.2024.1291341 4.Polat H. Instructors’ presence in instructional videos: A systematic review. Education and information technologies. 2023; 28(7): 8537-8569. doi: 10.1007/s10639-022-11532-4 5. Lawson AP, Mayer RE, Adamo-Villani N, Benes B, Lei X and Cheng J. The positivity principle: Do positive instructors improve learning from video lectures?. Educational Technology Research and Development. 2021; 69(6): 3101-3129. doi: 10.1007/s11423-021-10057-w 6. Guo PJ, Kim J, Rubin R. How video production affects student engagement: An empirical study of MOOC videos. In: Proceedings of the First ACM Conference on Learning at Scale Conference. Association for Computing Machinery. 2014; 41-50. doi:10.1145/2556325.2566239 7. Polat H, Taş N, Kaban A, Kayaduman H, Battal A. Human or humanoid animated pedagogical avatars in video lectures: The impact of the knowledge type on learning outcomes. International Journal of Human–Computer Interaction. 2025; 41(14): 8912-8927. doi:10.1080/10447318.2024.2415762 8.Anttonen, R, Kristian K, Eija R, Carita K. Storifying instructional videos on online credibility evaluation: Examining engagement and learning. Computers in Human Behavior. 2024; 161: 108385. doi: 10.1016/j.chb.2024.108385 9. Dai L, Jung M, Postma M, Louwerse M. A systematic review of pedagogical agent research: Similarities, differences and unexplored aspects. Computers & Education. 2022;190: 104607. doi:10.1016/j.compedu.2022.104607 10. Atkinson RK. Fostering social agency in multimedia learning: Examining the impact of an animated agent's voice. Contemporary Educational Psychology. 2005; 44(2): 117-137. doi: 10.1016/j.cedpsych.2004.07.001 11. Wang N, Johnson WL, Mayer RE, Rizzo P, Shaw E, Collins H. The politeness effect: Pedagogical agents and learning outcomes. International Journal of Human-Computer Studies. 2008; 66(2): 98-112. doi: 10.1016/j.ijhcs.2007.09.003 12. Nass C, Steuer J, Tauber ER. Computers are social actors. In Proceedings of the SIGCHI conference on Human factors in computing systems. ACM; 1994: 72-78. doi: 10.1145/191666.191703 13.Wu X. Singing syllabi with virtual avatars: enhancing student engagement through AI-generated music and digital embodiment. arXiv. 2025. arXiv: 2508.11872. 14. Fink MC, Robinson SA, Ertl B. AI-based avatars are changing the way we learn and teach: Benefits and challenges. Frontiers in Education. 2024; 9: 1416307. doi:10.3389/feduc.2024.1416307 15. Qin Z, Zhao W, Yu X, Sun X. OpenVoice: Versatile instant voice cloning. arXiv. 2023. arXiv:2312.01479. 16. Li T, Zheng R, Yang M, Chen J, Yang M. Ditto: Motion-space diffusion for controllable realtime talking head synthesis. arXiv. 2024. arXiv: 2411.19509. 17. Sun A, Zhang X, Ling T, Wang J, Cheng N, Xiao J. Pre-Avatar: An automatic presentation generation framework leveraging talking avatar. In 2022 IEEE 34th International Conference on Tools with Artificial Intelligence (ICTAI). IEEE; 2022:1002-1006. doi: 10.1109/ICTAI56018.2022.00153 18. Arkün-Kocadere S, Çağlar-Özhan Ş. Video lectures with AI-generated instructors: Low video engagement, same performance as human instructors. The International Review of Research in Open and Distributed Learning. 2024; 25(3): 350-369. doi:10.19173/irrodl.v25i3.7815 19. Duester E, Zhang R. Digital and AI transformation in the contemporary art industry in China. Arts & communication. 2025;3(2):3822. doi:10.36922/ac.3822 20. Ramos-Vallecillo N, Murillo-Ligorred V. The phenomenon of artificial intelligence-generated images in university teacher training and its impact on developing critical thinking. Arts & communication. 2025;3(3):5047. doi:10.36922/ac.5047 21. Zhao B, Zhan D, Zhang C, Su M. Computer-aided digital media art creation based on artificial intelligence. Neural Computing and Applications. 2023; 35(35): 24565-24574. doi:10.1007/s00521-023-08584-z 22. Holmes W and Miao F. Guidance for generative AI in education and research. Unesco Publishing. 2023. 23. Mori M, MacDorman KF, Kageki N. The uncanny valley [from the field]. IEEE Robotics & Automation Magazine. 2012;19(2):98-100. doi:10.1109/MRA.2012.2192811 24. Bender S. Generative-AI, the media industries, and the disappearance of human creative labour. Media Practice and Education. 2025; 26(2): 200-217. doi: 10.1080/25741136.2024.2355597 25. Zhou E, Dokyun L. Generative artificial intelligence, human creativity, and art. PNAS nexus. 2024; 3(3): pgae052. doi: 10.1093/pnasnexus/pgae052 26. Bomba F, Antonella DA. Agency and authorship in AI art: Transformational practices for epistemic troubles. International Journal of Human-Computer Studies. 2025; 205: 103652. doi: 10.1016/j.ijhcs.2025.103652 27. Egon K, Eugene R, Rosinski J. AI in art and creativity: exploring the boundaries of human-machine collaboration. OSF Preprints. 2023; 20: 1-11. 28.Hsu TWL. Online Art Therapy: Reimagining Body, Place, Object and Relations in the Digital Era. Doctoral dissertation, Goldsmiths, University of London. 2024. 29.Tao, Z, Liu Y, Qiu J, Li S. Impact of virtual avatar appearance realism on perceptual interaction experience: a network meta-analysis. Frontiers in Psychology. 2025; 16: 1624975. 30.Katalin F. Exploring AI media. Definitions, conceptual model, research agenda. Journal of Media Business Studies. 2024; 21(4): 340-363. 31.Richard EM. Evidence-based principles for how to design effective instructional videos. Journal of Applied Research in Memory and Cognition. 2021; 10(2): 229-240. Acknowledgments The author acknowledges the open-source communities behind OpenVoice and Ditto-TalkingHead, whose publicly released code, models, and documentation made the educator-oriented workflow presented in this paper possible. Availability of data The implementation accompanying this study is publicly available in the VirtualAssistant repository: https://github.com/xinxingwu-uk/VirtualAssistant . A public demo is also available at https://xinxingwu-uk.github.io/projects/demo3/slides.html