Paper deep dive
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild
Haotian Qi, Gabriel Skantze
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/20/2026, 8:48:40 AM
Summary
MuVAP (Multimodal Multiparty Voice Activity Projection) is a causal multimodal framework designed for speaker-aware turn-taking prediction in unconstrained multiparty settings. It addresses the limitations of traditional Voice Activity Projection (VAP) and Active Speaker Detection (ASD) by using a single monaural audio stream and a single camera view. The framework introduces 'Role-Relative Projection' to map N-speaker interactions into a fixed 'current versus next' floor-holder state, ensuring social scalability. To support this, the authors introduce the Audio-Visual Conversation Corpus (AVCC), a 31-hour unedited, single-camera dataset of natural multiparty conversations. The model architecture combines a VAP backbone for prosodic/linguistic cues and a causal ASD backbone for visual grounding, enabling both GlobalVAP (joint state) and SpeakerVAP (speaker-specific) predictions.
Entities (9)
Relation Signals (4)
MuVAP â evaluateson â Audio-Visual Conversation Corpus
confidence 100% · we introduce the Audio-Visual Conversation Corpus... Evaluations demonstrate that MuVAP outperforms strong baselines
MuVAP â extends â Voice Activity Projection
confidence 100% · We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection
GlobalVAP â isatypeof â MuVAP_prediction
confidence 100% · The model predicts two probability distributions at each time step t: GlobalVAP (GVAP)...
MuVAP â uses â Role-Relative Projection
confidence 100% · To address the combinatorial complexity... we propose Role-Relative Projection
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single-camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks across two- and three-speaker settings.
Tags
Links
- Source: https://arxiv.org/abs/2606.16731v1
- Canonical: https://arxiv.org/abs/2606.16731v1
Trouble viewing inline? Open PDF directly â
Full Text
54,908 characters extracted from source content.
Expand or collapse full text
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild Haotian Qi ID , Gabriel Skantze ID Department of Speech Music and Hearing, KTH Stockholm, Sweden haotianq@kth.se, skantze@kth.se Abstract Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their ap- plicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework that extends Voice Activity Projection by grounding acoustic predictions in face tracks, enabling speaker-aware turn-taking predictions from a monaural audio stream and a single camera view. To address the combinatorial complexity of modeling multiple speakers, we propose Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state. Because existing audiovisual datasets contain disruptive editing cuts that break causal tracking, we introduce the Audio-Visual Conversation Corpus, a 31-hour dataset of unedited, single- camera multiparty conversations. Evaluations demonstrate that MuVAP outperforms strong baselines on Shift-Hold and next- speaker prediction tasks across two- and three-speaker settings. Index Terms: Turn-taking, Multimodal Interaction, Active Speaker Detection, Multiparty Dialogue, Next Speaker Predic- tion 1. Introduction Turn-taking is a fundamental aspect of conversation. Since peo- ple cannot easily speak and listen simultaneously, they must coordinate their turns through a complex exchange of cues [1, 2, 3]. For example, a syntactically or semantically incom- plete phrase may signal a turn hold, whereas a complete phrase may be turn-yielding [4]. A filled pause is a strong cue for turn-holding [5, 6]. When syntax and semantics are ambiguous, prosody and visual cues such as lip motion and facial dynamics can also be informative [7]. These cues allow humans to take turns with very small gaps (around 200 ms [8]), while avoiding large overlaps or interruptions [9]. Traditional conversational systems have mostly relied on si- lence thresholds to trigger system responses [10]. This reactive approach leads to unnatural delays and interruptions, motivat- ing a shift towards continuous, predictive turn-taking. Notably, Voice Activity Projection (VAP) [11] utilizes acoustic features and formulates turn-taking as a continuous projection of speech activity into the near future (e.g., 2 seconds). By forecasting the joint activity of speakers, VAP can predict complex phe- nomena, including turn shifts, backchannels, and interruptions, in a zero-shot fashion [11] or by training a linear probe on the learned embeddings [12]. Recent work has shown how the VAP model can be inte- grated into conversational systems, including human-robot in- teraction (HRI), to improve turn-taking and backchannel predic- 1 https://github.com/Haotian-Qi/MuVAP Figure 1: Visualization demo 1 of multiparty turn-taking predic- tion. The top panel shows the video stream with tracking of speakers S0 (red), S1 (green), and S2 (yellow), with the sin- gle channel audio stream below. The bottom panel shows the ground truth speech activity timelines, with the red vertical line marking the current time step t, leading to a turn shift from S1 to S0. The left panel shows the GlobalVAP (GVAP) shift/hold predictions along with individual SpeakerVAP (SVAP) predic- tions. tions [12, 13]. However, for HRI applications, traditional VAP suffers from two critical limitations. First, it is predominantly designed for two-party interactions, whereas HRI frequently in- volves multiparty settings. Second, it often relies exclusively on the speech channel. In multiparty dialogue, the visual channel is indispensable, since the model must determine not only if a turn is ending, but who among the group is the next speaker. We propose MuVAP (Multimodal Multiparty Voice Activ- ity Projection), a framework that anchors voice activity projec- tions to faces detected in the video stream. Unlike prior ap- proaches that rely on separate audio channels or assume fixed dyadic interactions, MuVAP operates in unconstrained multi- party settings without requiring channel separation. This de- sign enables the model to jointly address two complementary tasks. First, it predicts the temporal onset of forthcoming speech (when) by leveraging audiovisual and prosodic cues. Second, it infers the identity of the current and likely next speaker (who) by grounding its predictions in the visible participants within the scene. To address the combinatorial complexity inherent in multi- party interactions, we introduce Role-Relative Projection (Fig- ure 3). Rather than modeling all possible speakerâlistener per- mutations, we adopt a role-centric abstraction inspired by hu- man conversational behavior: interlocutors primarily monitor arXiv:2606.16731v1 [cs.SD] 15 Jun 2026 AVA-AS: (Isolated speaker segment, less concurrent data with scripted expression MSDWild: (Editorial cuts remove natural pause and turn-taking cues) Jump Cut AVCC: (Ours: Preserves natural conversational flow with visible predictive cues) Mutual Silence Breath in Gaze Shift Pre-turn cue Shift Cue Removed Hold Cue Removed No Data Jump Cut Figure 2: Comparison of dataset continuity and structure. AVA- ActiveSpeaker (Top) is limited to many isolated speaker seg- ments rather than full scene, often relying on scripted cues that lack natural turn-taking dynamics. MSDWild (Middle) captures unconstrained scenes but suffers from editorial and jump cuts that disrupt the visual timeline needed for prediction. In con- trast, AVCC (Bottom) preserves the unedited, continuous flow of conversation. the current floor-holder and the most probable next speaker. By projecting group dynamics onto these two relative roles, Mu- VAP attains social scalability â enabling a single, fixed architec- ture to generalize to an arbitrary number of participants without retraining or architectural modification. In contrast to previous studies [14, 15] that rely on data from laboratory settings, our work leverages unconstrained, real-world video. While some existing datasets used for Active Speaker Detection (ASD), such as MSDWild [16], cover more naturalistic settings, they typically contain editing gaps, which make them unsuitable for directly training a causally grounded turn-taking framework like MuVAP. To address this, we introduce the Audio-Visual Conver- sation Corpus (AVCC) by collecting and annotating approxi- mately 31 hours of unedited videos of multiparty dialogue from the web, specifically filtered for static single-camera perspec- tives to preserve genuine social dynamics. Furthermore, to leverage existing high-resource data, we adopt a modular train- ing strategy that utilizes multiple datasets as shown in Table 1: (a) a large-scale telephone corpus (1960 h) to learn acoustic turn-taking signals, (b) standard ASD datasets (140 h) to pre- train the ASD module, and (c) the AVCC dataset to optimize the temporal turn-taking dynamics on multimodal, continuous, unedited sequences of conversation. While previous studies have examined multiparty turn- taking prediction, ASD, and multimodal interaction modeling, explicit real-time next speaker identity forecasting in multiparty settings under a strictly single-view, single-channel setting re- mains largely unexplored. We address this gap by proposing a causal multimodal model that continuously anticipates which participant will take the conversational floor, without exploit- ing spatial audio cues, microphone arrays, or multi-view visual geometry. 2. Related Work 2.1. Active Speaker Detection vs. Turn-Taking We distinguish our work from multimodal Active Speaker De- tection (ASD), which aims to identify the active speaker among multiple faces in a video stream. Current ASD architectures, such as TalkNet [17] and LoCoNet [18], achieve impressive pre- cision on the AVA-ActiveSpeaker (AVA-AS) benchmark [19] by correlating lip motion with audio streams. While sometimes conflated with turn-taking, ASD differs strictly in scope and temporality. Standard ASD models focus on identifying the current speaker state.However, they often overlook the underly- ing conversational context, such as whether a speaker is mid- utterance, providing a brief backchannel, or yielding the floor during silence. In contrast, MuVAP models interaction dy- namics. This approach captures the subtle, predictive signals that reveal a participantâs intent to either maintain, yield, or seize the floor. We also enforce strict causality throughout the architecture. While many ASD benchmarks improve detection accuracy by looking at future frames, MuVAP is designed for the reality of live interaction. Every prediction is made using only the history available at that point. 2.2. Voice Activity Projection (VAP) Voice Activity Projection (VAP) [11] is a self-supervised frame- work designed to predict future conversational patterns directly from raw acoustic inputs. Unlike traditional turn-taking ap- proaches that rely on reactive silence thresholds or specific event labels like pauses or overlaps, VAP formulates turn-taking as a continuous forecasting problem. It projects the joint voice activity of two speakers into a near-future window (typically 2000 ms). This enables Next Speaker Prediction (NSP) tasks within strict dyadic settings, where a Shift indicates that the lis- tener will take the turn. To make this objective tractable, the future window is grouped into activity bins, and the joint activity bins are mapped to a finite codebook of discrete labels. By optimizing a cross- entropy objective over these future states, VAP implicitly learns to recognize complex coordination cues, such as prosodic shifts, backchannels, and filled pauses, without requiring manual event annotations [20, 11]. However, the standard VAP formulation faces a critical bot- tleneck in multiparty settings due to a combinatorial explosion, as seen in Figure 3. Because VAP models the joint state of all speakers simultaneously, the label space grows exponentially with the number of speakers (N ). While recent extensions have successfully adapted VAP for fixed triadic settings [15], this joint-modeling approach cannot scale to arbitrary or dynamic group sizes; it also suffers from coarse label resolution. The codebook would require retraining for every possible value of N , which makes standard VAP unsuitable for dynamic settings where the number of participants fluctuates. 2.3. Multimodal Next Speaker Prediction Next Speaker Prediction (NSP) is a classification task in which the system must predict which speaker, within a multiparty con- versation, will take the next turn. Traditional frameworks rely on explicit pre-utterance indicators like gaze transition patterns [21, 22, 23], head pose dynamics [14], and mouth movement [24] to identify floor acquisition. Speakers use these visual sig- nals alongside prosodic contours to navigate turn-taking [25]. While sentence-level NSP [26] and event-based visual frame- DatasetModality DurationSourceLang. Unedited Module FisherAudio1958h7mTelephoneEnglishâVAP AVA-ASAV38h5mMovieMultiâASD MSDWildAV80h18mVlogs/WildMultiâASD WASDAV30.0hWildMultiâASD AVCC (Ours)AV30h52mWildEnglishâMuVAP Table 1: Comparison of datasets used in this work. AV = Audio-visual. Unedited means continu- ous, uncut conversations essential for long-term turn-taking modeling. AVCC Subset Duration 2-speaker17h31m 3-speaker13h21m Training22h28m Validation8h24m Total30h52m Table 2: Distribution of the col- lected AVCC dataset by speaker count and partition. works [27, 28] achieve high classification accuracy, they treat turn-taking as isolated events, missing the continuous alignment of acoustic and facial dynamics. A major hurdle for existing NSP models is their reliance on controlled or high-resource environments. Many state-of-the- art models [29, 30, 22, 28] require a multi-channel microphone array or a dedicated camera for every speaker to successfully isolate signals. Such hardware dependencies are impractical for real-world Human-Robot Interaction (HRI) because they can- not easily adapt to unconstrained settings. Conversely, existing single-channel monaural systems that lack this hardware sup- port often suffer from identity drift, especially when speakers have similar vocal timbres. MuVAP distinguishes itself by operating under a strict single-camera and single-audio-stream constraint. Unlike pre- vious works that depend on multi-channel inputs, MuVAP avoids explicit pose extraction and treats the visual modality primarily as a separation anchor. By maintaining spatially dis- tinct visual tracks within a single video stream, we anchor the global acoustic history to the correct speaker identity. This ar- chitecture allows the model to predict short periods of future speech activity for each speaker simultaneously. Consequently, the system can identify the next speaker at turn transitions us- ing only the âin the wildâ data available from a single monaural source and one camera view. 3. The AVCC Dataset Currently, publicly available ASD datasets [31, 16] suffer from domain gaps that limit their utility for turn-taking prediction in interactive settings (such as HRI). As illustrated in Figure 2, these sources contain editing artifacts like jump cuts that disrupt the temporal flow of conversation. This prevents models from learning the true causal history required to predict turn-taking. Conversely, egocentric datasets [32] introduce viewpoint bias in which the camera wearerâs actions influence sensory input. Similarly, while multiparty datasets like AMI [33] capture con- tinuous dialogue, they rely on overhead cameras and dedicated lenses for each participant. This distributed array captures a view that is fundamentally disconnected from the singular vi- sual stream a human or robot relies on. The Audio-Visual Conversation Corpus 1 (AVCC) dataset was collected according to three primary criteria to ensure suit- ability for causal turn-taking modeling. First, we selected only continuous, unedited sequences to preserve the natural tempo- ral flow of conversation. This allows the model to observe the full duration of mutual silences and pre-turn behaviors. Sec- ond, we restricted the data to a static third-party perspective. This constraint prevents the model from relying on camera mo- tion or editorial cues, forcing it to focus on participant-level signals. Third, we prioritized spontaneous social interactions to capture realistic phenomena like overlaps, interruptions, and Table 3: Conversational dynamics of Fisher (Telephone) vs. AVCC (Multimodal) with 2 or 3 speakers. Numbers are median seconds with inter-quartile ranges, reported for Gaps (silence between turns), Pauses (silence within turns) and Overlaps (be- tween turns). MetricFisherAVCC (2) AVCC (3) Gaps0.320.390.43 [0.16â0.68][0.15â0.88][0.15â1.00] Pauses0.400.500.61 [0.24â0.76][0.27â1.01][0.36â1.13] Overlaps0.560.440.50 [0.24â1.48][0.17â1.21][0.19â1.23] hesitation cues among familiar friends. We collected the data from YouTube and Twitch livestreams of unscripted multiparty conversations. These formats were chosen because they consist of unscripted dialogue where natural turn-taking and conversa- tional fillers are frequent. In total, the dataset comprises 30 hours and 52 minutes of recordings, spanning both two-speaker and three-speaker in- teraction settings. The distribution of the number of speak- ers, as well as training/validation splits are shown in Table 2. Table 3 also shows some descriptive turn-taking statistics for the dataset, compared to the Fisher telephone dataset. As can be seen, AVCC exhibits wider inter-quartile ranges, confirming that unconstrained interaction in the wild is less predictable than dyadic telephone speech, even in 2-speaker settings. We utilized the InsightFace [34] framework with Reti- naFace [35] backbone for automated face detection and identity tracking. We localized faces using SCRFD [36] with a confi- dence threshold of 0.62. To eliminate background artifacts, we applied a relative area filter that dropped any detections smaller than 10% of the largest face within a given frame. To initialize stable speaker identities for the dataset, we sampled 50 random frames from each video to establish a set of global anchor em- beddings. These anchors, in combination with frame-by-frame embeddings generated by ArcFace [37], were used to maintain temporal continuity. We calculated a joint cost matrix by combining the cosine distance of these embeddings with the Euclidean distance of bounding box centroids. Speaker assignment was resolved us- ing the Hungarian algorithm [38]. This approach ensures that speaker tracks remain stable across all segments of a video. We performed a manual audit of the resulting tracks, checking for identity swaps or missed detections by identifying sudden changes in the distance of bounding box centroids. The syn- chronized audio and resulting face crops were processed by our Speaker Based Role Relative A B C Frame-level Voice Activity Labels 0.6s1.4s1.4s0.6s t HistoryFuture Figure 3: The difference between speaker based projection [15] and (Ours) Role-Relative Projection for 3-speaker conversa- tions. pre-trained ASD model, as shown in Figure 4, with a VAD ob- jective to generate initial voice activity labels for each speaker. All annotations were then manually refined using the VIA anno- tator to ensure high-fidelity ground truth labels, similar to those shown in Figure 1. 4. Model We propose MuVAP (Multimodal Multiparty Voice Activity Projection), a causal framework designed to forecast turn-taking dynamics for an arbitrary number of speakers. Given a single- channel audio waveform A âR T and a set of face tracks V =v 1 , . . . , v N for N detected speakers, the model predicts two probability distributions at each time step t: âą GlobalVAP (GVAP): The joint turn-taking state of the âcur- rent floor holderâ versus the ânext floor holderâ. âą SpeakerVAP (SVAP): A speaker-specific projection for each tracked face for next speaker predictions. The overall architecture and the different sub-modules are shown in Figure 4. 4.1. VAP Backbone The goal of the VAP module is to extract prosodic and linguis- tic cues from the audio waveform features [20]. Following [11], we utilize a contrastive predictive coding (CPC) encoder to map raw audio to dense representations, followed by a causal down- sampling convolution (100Hzâ 25Hz) to align with the video frame rate. These audio features then pass to a Transformer [39] block with ALiBi positional embeddings and causal masking. For the original VAP model [11], the prediction target at each incremental step was defined by a window of 2 seconds containing the future voice activity for both speakers. In that setting, the window was discretized into 8 separate bins (4 for each speaker), with each bin corresponding to time bins of in- creasing duration in seconds [0.2, 0.4, 0.6, 0.8]s. Each bin was assigned a value of 1 if more than half of its frames contain ac- tive speech from that speaker. This 2Ă 4 discretization yields an 8-bit binary vector, corresponding to 256 unique classes. 4.1.1. Role-Relative Projection As noted earlier, the original VAP objective does not scale well to multiparty settings, as the number of unique classes grows exponentially with the number of speakers (e.g., 2 4ĂN , where N denotes the number of speakers). Moreover, when only a sin- gle mono audio channel is available, it is not possible to anchor predictions to individual speakers. The original VAP model [11] also used a mono channel, but it made use of individual voice activity labels as input to anchor the predictions. Our model only uses a mono channel, without any such labels as input. To resolve this, we introduce Role-Relative Projection, where the objective is to predict future voice activity rela- tive to past voice activity, within certain time bins, by letting the classes jointly represent both future and past time bins, as shown in Figure 3. This reduces the variable N -speaker state to a fixed âCurrent vs. Nextâ pairwise representation. This re- duces temporal resolution in bin durations, but is based on the simplifying assumption that most turn-taking typically involves two out of the N speakers at any given point in time. For each time step t, we discretize voice activity into a window of four bins: two history bins and two future bins as [1.4, 0.6, 0.6, 1.4]s duration for each speaker, as shown in Fig- ure 3. We then determine the projection labels via a two-stage ranking: 1. Current Holder (S curr ): We rank all N speakers by their to- tal activity in the history bins. The speaker with the highest activity is designated as the current/past floor holder. 2. Next Holder (S next ): We rank the remaining N â 1 speak- ers by their activity in the future bins. The highest-ranked remaining speaker is designated as the primary next speaker. This reduction transforms the complex multiparty dynamic into a fixed pairwise stateS curr , S next . Following the standard VAP encoding method [11], we extract the binary activity for this pair (2 roles,Ă 4 bins) to form an 8-bit vector. While an 8- bit vector theoretically require 2 2Ă4 = 256 states, we make the codebook invariant to the ordering of the two patterns by treat- ing (S curr , S next ) and (S next , S curr ) as the same state, resulting in 136 unique states. Thus, c([0011, 1100]) = c([1100, 0011]) c([0111, 1000]) = c([1000, 0111]) The VAP module is trained to minimize the cross-entropy loss between the predicted probability distribution Ëy t and the ground truth label y t over the sequence length T from the latent embedding Z VAP : L VAP =â 1 T T X t=1 log P(Ëy t = y t | Z VAP )(1) 4.2. Causal ASD Backbone To ground predictions in visual tracks and distinguish audio be- tween participants, we require robust multimodal embeddings that focus on speech separation based on visual grounding. We pre-train a dedicated ASD module with three datasets. We adopt a modified TalkNet architecture [17], replacing non-causal tem- poral convolutions with causal ones, increasing the dilation to [1, 2, 4, 8, 16], and using a pre-trained CPC audio encoder (the same used for the VAP module). Unlike the traditional binary ASD objective that focuses on current frames, we use a similar bin configuration as the VAP module, with history bins of [0.8, 0.6, 0.4, 0.2]s and future bins [0.2, 0.4]s. We use shorter future bins with the goal of still Z ASD Z ASD CPC f a Self Attention Z VAP VAP Head VAP L VAP Concat Gated Type AddFrozen Train T CPC Visual Encoder Cross Attention Cross Attention Self Attention Z ASD ASD Head f a f v ASD L NCE L A L V L ASD Z VAP Linear LN Linear LN Linear LN KV Q Frame Cross Attention GVAP Self Attention Z GVAP Z ASD Z ASD Z ASD GVAP Head SVAP Head Z SVAP MuVAP T L GVAP L SVAP Preprocessing Face Tracking Visual Clips Single-channel waveform Causal Attention Figure 4: Modular architecture of the MuVAP model. The embeddings from the VAP and ASD modules (Z VAP and Z ASD ) are used as input to the main module (right), which makes both GlobalVAP (GVAP) and SpeakerVAP (SVAP) predictions. capturing certain backchannel expressions within edited video datasets, which results in a 6-bin prediction of y asd . Note that since this module makes predictions for each speaker indepen- dently, the activity of the different bins is not predicted jointly, but as independent binary predictions. The goal is to minimize the binary cross-entropy loss for the 6-bin target using the fused embedding Z ASD . Additionally, we utilize the side streams (L a and L v ) for auxiliary tasks. De- spite sharing the underlying encoders, these branches are con- strained to predict the binary voice activity (y t ) for the current frame t as an auxiliary loss. We also utilize the TalkNCE [40] loss, which performs contrastive loss on the features f a and f v to pull the positive pairs closer, denoted as L NCE , to boost the active speakersâ visual representations with the CPC audio em- beddings. The total loss is a weighted sum, where all coeffi- cients follow the configuration in [40]: L total = L ASD + 0.4(L a + L v ) + 0.3L NCE (2) 4.3. MuVAP: Multimodal & Multiparty We treat multiparty turn-taking as a hierarchical coordination problem. The GlobalVAP (GVAP) acts as a âsocial conductorâ, monitoring the global pulse of the conversation to predict when a transition is imminent. This global context then guides the SpeakerVAP (SVAP), which focuses on individual visual tracks to resolve which specific participant will take the next turn. To fuse different modalities, the model takes global audio embeddings Z VAP and a set of N individual speaker embed- dings,Z 1 ASD , . . . , Z N ASD from ASD modules. These inputs are processed through independent LayerNorm and linear projec- tion layers to map them into a shared 256-dimensional space to match the VAP embeddings, ensuring the representations are aligned before the attention blocks. We then use the Z VAP as the Query and Z N ASD as Key and Value, passing them through a frame transformer to allow the model to attend to all speaker tokens within the same temporal frame. We then pass the output through another temporal trans- former to capture temporal dynamics. This yields the Z GVAP embeddings, which capture the global turn-taking dynamics for multiparty settings. We pass the input Z N ASD embeddings through separate Lay- erNorm and projection layers, then use gated addition with the Z GVAP to let the GVAP dynamically modulate the speaker fea- tures. The modulated features are fed into the SVAP prediction head where the target consists of independent binary predictions of future bin activity with durations [0.2, 0.4, 0.6, 0.8]s, which matches the original VAP bin resolution. The training objective is as follows: L MuVAP = L GVAP + 1 N N X n=1 L (n) SVAP 5. Implementation Our pipeline utilizes five distinct datasets: Fisher (Parts 1 & 2), MSDWild, WASD, AVA-ActiveSpeaker, and our newly intro- duced AVCC dataset. Each dataset is mapped to specific mod- ules to facilitate multi-stage training as shown in Table 1. All modules were optimized using a Cosine Annealing learning rate scheduler with a linear warm-up phase covering the first 10% of total training steps. The learning rate reached a peak of 1Ă 10 â3 and decayed to 1Ă 10 â4 over five epochs, utilizing a Weight Decay of 0.01. Experiments were conducted on a single NVIDIA A100 (40GB) GPU. The model consists of 27.7 million total parameters. This is distributed across three modules: the ASD module (20.7 M), the VAP module (4.9 M), and the MuVAP module (2.1 M). For preprocessing, face crops were extracted based on tracked coordinates and missing frames were padded with a grayscale value of 127. Video streams were standardized to 25 fps, and audio was sampled at 16 kHz in mono. Speaker ASpeaker A/B Speaker A/B MaxGap Offset Silent Hold/Shift Active Hold/Shift Solo Duration Offset MaxGap Figure 5: Visualization of the Active and Silent prediction sam- pling methods. The yellow region indicates the required solo- speaker duration (1000 ms) and offset (100 ms) used to prevent label leakage. The maximum gap duration is 3000 ms for both event types. Active events are derived from Silent events by shift- ing the prediction point into the ongoing speech segment while maintaining a fixed 100 ms offset from the speech endpoint. 5.1. VAP Backbone The VAP module was trained on the Fisher Corpus of spoken dialogue over telephone. We partitioned the data by session ID: sessions divisible by 16 were reserved for validation, while the remainder formed the training set. For the logistic regres- sion probe, groups064, 080 were held out as Test set and the rest were used as the Fit set to train on the linear probe. Au- dio samples were segmented into 30-second windows with 10- second overlaps. The temporal Transformer comprises L = 4 layers, h = 4 attention heads, and a hidden dimension of d model = 768. We use ALiBi positional encoding and set the dropout to 0.1. 5.2. ASD Backbone The ASD backbone was trained on a composite of MSDWild, WASD, and the AVA-ActiveSpeaker training set. We modified the standard TalkNet architecture by replacing all non-causal temporal layers with causal convolutions and 3D to 2D Con- volutions. We also increased the dilation to [1, 2, 4, 8, 16] for more visual history context. The cross-attention mechanism has a hidden dimension of 768, while self-attention blocks were ex- panded to 1024 dimensions. Due to high memory requirements, we employed a batch size of 4 during training. 5.3. MuVAP The complete MuVAP model was trained on the AVCC dataset using a batch size of 16. During this phase, the VAP and ASD backbones were frozen to serve as robust feature extractors. Audio (Z VAP ) and visual (Z ASD ) embeddings were projected into a shared 256-dimensional latent space through indepen- dent Linear-LayerNorm blocks before being processed by the Frame transformer and Temporal transformer. Both transformer blocks have a feed-forward expansion factor of 4, with 4 atten- tion heads and 1 layer. 6. Downstream tasks In the spirit of the original VAP model [11], MuVAP is trained to learn generic embeddings capturing turn-taking dynamics. To measure the usefulness of these learned embeddings, we de- fine a set of downstream tasks for evaluation. 6.1. Shift-Hold vs. Next Speaker Prediction We first distinguish between two types of tasks: The first task is Shift-Hold Prediction, where the model has to determine whether the turn is about to Shift, or whether the current speaker will Hold the turn. This binary task does not require spe- cific speaker identification and relies solely on raw role-relative VAP/GVAP predictions. Thus, it can be done on both multi- modal (AVCC) and unimodal (Fisher) scenarios. This down- stream task is performed using a linear probe (a Logistic Re- gression model) trained on top of the Z GVAP embeddings. To prevent data leakage, we strictly separate the data used for the probe. We extract the event prediction points from the AVCC Train partition to fit the linear probe, and we evaluate its per- formance exclusively on the AVCC Validation partition. This split was based on the number of speakers and scenario settings in each video to ensure a balanced distribution of 2-speaker and 3-speaker configurations across all subsets, as seen in Table 4. The second task is Next Speaker Prediction (NSP), where the model has to determine who (among the N speakers) will be the next speaker, in moments of mutual silence. These decision points are the same as those for the Shift-Hold Prediction task, but the decision is not relative to the previous speaker. Note that the set of potential next speakers includes the current speaker (equivalent to a Hold). To solve this task, the raw VAP/GVAP embeddings are not sufficient; the model also has to make use of the individual SVAP predictions. Note that the NSP task in a 2-speaker scenario is not exactly the same as Shift-Hold Predic- tion, since the model has to successfully anchor the prediction in the set of visible speakers, which is not strictly necessary in the Shift-Hold Prediction task. Additionally, we propose a GVAP-conditioned filtering strategy to ensure individual speaker assignments remain con- sistent with the projected global state. By utilizing the GVAP prediction as a conductor prior, we logically constrain the can- didate set. If the GVAP predicts a Shift, the identified previous speaker is eliminated from the pool of potential candidates. If a Hold is predicted, the next speaker is automatically assigned to the previous floor holder. This hierarchical coupling forces the speaker-level predictions to align with the overall rhythm of the group interaction. To identify the previous floor holder, we aggregate SVAP predictions over a 1-second rolling window, effectively using the 0.2s future-bin predictions as a temporally shifted proxy for recent speech activity, and assign the speaker with the highest cumulative probability. To calculate the likelihood of a specific participant taking the floor, we sum the probabilities of the final two future bins from their respective SVAP head. By comparing these scores across all participants, the model identifies the most probable next speaker. 6.2. Silent vs. Active Prediction We then distinguish between two different ways of determining when to make the prediction: First, we identify points of mutual silence (Silent Prediction), where a speaker has just stopped speaking and no other participant is speaking. These predic- tion points are illustrated in the top part of Figure 5. Prediction points are sampled at a +100 ms offset from the end-of-speech. We only include pauses/gaps†3 s that are preceded by at least 1 s of single-speaker speech. A more proactive version of this task is to predict upcom- ing turn dynamics while a speaker is still active. Unlike silent prediction, which occurs after a vocal offset, we define Active Prediction by sampling points within the ongoing speech of the current floor holder. These points are derived directly from the previously identified silence events by moving the prediction point earlier, into the active speech region. To ensure the model has sufficient causal history, we also shift the solo speaker re- Table 4: Split of the AVCC dataset Train and validation parti- tion, used for the Logistic Regression probe. Samples are cat- egorized by Shift-Hold mode (Active/Silent) and the number of speakers present (2 or 3). Split Mode Total2/3-SpkHOLD SHIFT Train Active1774610265/748179239823 Silent1795110865/708699807971 Val Active58913274/261726313260 Silent62053725/248036232582 quirement backwards to ensure that the prediction point is pre- ceded by at least 1s of uninterrupted clean speech from a single participant. For Active-Shift events, the prediction point is placed within the vocalization of an interlocutor whose turn results in a floor transition. In contrast, Active-Hold events are sampled from turns that precede a within-turn pause where the current speaker continues to hold the floor. By anchoring both classes to a fixed offset of 100ms before the speech ends, the task mea- sures the modelâs capacity to decode terminal signals such as specific prosodic contours or stable facial dynamics that indi- cate turn-yielding or turn-holding intent. This setup evaluates if the model can anticipate an upcoming transition point before the actual onset of mutual silence. The distribution of labels for the training and validation of events are shown in Table 4. 7. Results We evaluate Shift-Hold Prediction using Macro-F1 (due to class imbalance), and Next Speaker Prediction (NSP) using accuracy. Results are reported as mean ± 95% confidence intervals over ten runs using random seeds 42â51. We compare MuVAP against two baselines: Majority Class for the binary Shift-Hold task and Random for the multi- class NSP task. To quantify the benefit of our proposed fusion strategy, we also implement a No-Fusion (MLP) baseline. This model processes VAP and ASD embeddings separately through individual LayerNorm and projection layers. These layers re- duce the input dimension from 512 to 256. We apply a GELU activation and 0.1 dropout before passing to the NSP and Shift- Hold prediction heads, respectively. 7.1. Role-Relative Projection Since we introduce a new VAP training objective in this paper â the Role-Relative Projection (see Section 4.1.1) â we first want to compare this with the speaker based projection in the orig- inal VAP model in a two-speaker scenario where we do have two audio channels, namely the Fisher corpus. We train two models based on these two objectives: One with mono audio and Role-Relative Projection, and one with stereo audio and speaker based projection. We then use the Z VAP embeddings to train a logistic regression probe for the Silent Shift-Hold Pre- diction task (as defined in Section 6.1) on the Fisher Fit set and evaluate on the test set (as defined in Section 5.1). The results are shown in Table 5. The stereo model has bet- ter performance than the mono model, which is expected, given that stereo channels provide explicit speaker attribution, which makes it easier to identify speaker shifts and resolve overlap- ping speech. In the mono setting (which is a requirement for our model), the model has to figure out speaker switches based on speaker characteristics in the speech signal. In light of this, we think that the drop in performance (about 2%) is quite mod- est, and validates the viability of the Role-Relative Projection to be used in our model. Table 5: Performance comparison between speaker based and our Role-Relative VAP training objectives on the Fisher dataset. ObjectiveMacro-F1 Majority class.451 Speaker based (stereo) .799 Role-Relative (mono).778 7.2. Shift-Hold Prediction As described in Section 6 above, both the Silent and Active Shift-Hold prediction tasks are evaluated with a logistic regres- sion probe trained on the Z GVAP embeddings. As baselines, we also use the unimodal Z VAP embeddings from the VAP model trained on Fisher (see box 1 in Figure 4). The results for the Silent and Active Shift-Hold Predic- tions are shown in Table 6. Our model outperforms the base- lines across all tasks, indicating that visual information helps for Shift-Hold Prediction and that the gated fusion of the Mu- VAP model contributes. Results are similar between 2 and 3 speaker settings, even for the VAP model, indicating that Shift- Hold Prediction using Role-Relative Projection works well in both these settings, since individual turn shifts typically involve just two of the participants. While the overall performance in mutual silence is lower than in dyadic telephone settings (Table 5), these results demonstrate that MuVAP effectively handles the increased complexity of multiparty interactions. The perfor- mance is somewhat lower in the Active tasks, which is expected, but still indicates that turn shift forecasting is possible. 7.3. Next Speaker Prediction The results for NSP are shown in Table 7. Again, our model outperforms the baseline MLP model across the different tasks and settings. We also show how the results can be further im- proved by not just selecting the speaker with the highest future SVAP predictions, but by also conditioning the prediction us- ing the GVAP predictions to infer turn holds, in combination with SVAP predictions of the previous speaker (+GVAP). How- ever, these predictions of the previous speaker are not perfect, as can be seen in Table 8. We therefore also show the poten- tial performance if the previous speaker detection was perfect (+GVAP+GT), as an upper bound of the performance. As expected, for NSP, 3-speaker settings are harder than 2-speaker settings, since a turn shift might lead to any of the other two partners taking the turn. In many cases, if the current speaker does not select the next speaker, it is up to the other speakers to decide who would like to take the turn, so-called self-selection [2], which is very hard, if not impossible, to pre- dict. 8. Discussion The results reveal a clear division of labor between the modal- ities in our model. Acoustic, linguistic, and prosodic cues pro- vide the global rhythm of coordination, while the visual channel acts as an essential separation anchor during competitive over- lap and active speech. Rather than relying on explicit behavioral Table 6: Performance on the Silent and Active Shift-Hold Pre- diction tasks on the AVCC dataset. Results for the MLP and our model are reported with 95% confidence intervals across 10 seeds. Model2 Spk (F1) 3 Spk (F1) Silent Majority class.367.351 VAP.672.655 MLP.650±.003.654±.003 MuVAP.696±.003.670±.002 Active Majority class.346.367 VAP.622.634 MLP.610±.003.635±.002 MuVAP.641±.005.652±.002 Table 7: Results (accuracy) for the Silent and Active Next Speaker Prediction (NSP) task on the AVCC dataset, with Random and MLP baselines, depending on speaker counts, with 95% confidence intervals across 10 seeds. (+GVAP = conditioning on the GVAP and previous speaker predictions. +GVAP+GT = conditioning on GVAP predictions and ground truth previous speaker labels.) Model2 Spk (acc) 3 Spk (acc) Random.500.333 Silent MLP.617±.002.464±.002 MuVAP.637±.003.477±.001 MuVAP (+GVAP).666±.003.508±.003 MuVAP (+GVAP+GT).702±.003.547±.003 Active MLP.543±.001.429±.001 MuVAP.560±.002.441±.002 MuVAP (+GVAP).605±.005.483±.003 MuVAP (+GVAP+GT).652±.003.516±.002 cues, our ASD modules ground the acoustic signal to individual visual tracks. The performance gain over unimodal baselines in Active Shift-Hold prediction shows that visual grounding helps disentangle overlapping acoustic signals and assign future ac- tivity to the correct participant. While this structural grounding proves effective, previous research demonstrates that explicit behavioral signals like gaze and head pose strongly govern turn allocation [14, 22, 23, 28]. Integrating these specific visual fea- tures on top of our separation tracks provides a clear path for future performance gains. Role-Relative Projection can be seen as mimicking human social attention, to some extent. Multiparty turn-taking operates through a highly local coordination system [41]. Therefore, hu- mans likely do not distribute their cognitive load equally to track every participant. They structurally reduce the interaction to a primary exchange by prioritizing the current floor holder and the next projected speaker. By mapping group dynamics onto these relative roles, the model can more easily scale to different group sizes. This approach bypasses the combinatorial explo- sion of joint modeling and allows the system to handle flexible group sizes without retraining. The contrast between the Fisher telephone corpus and our AVCC dataset, shown in Table 3, highlights the structural dif- ferences between constrained dyadic speech and unconstrained Table 8: Results (accuracy) for inferring the previous speaker, based on SVAP-based predictions, with 95% confidence inter- vals across 10 seeds. Model 2 Spk (acc) 3 Spk (acc) Random.500.333 SilentSVAP.837±.001.760±.001 ActiveSVAP.821±.001.742±.001 multiparty interaction. The performance drop from the Fisher corpus to the AVCC dataset highlights the inherent difficulty of in-the-wild interactions. Telephone conversations enforce a strict alternating rhythm because speakers lack visual feed- back. In contrast, unconstrained video data have longer gaps and pauses with more open floor states and visual backchannels. This domain shift makes prediction significantly harder. Train- ing on unedited streams ensures the model encounters these nat- ural conversational rhythms, which are artificially truncated by jump cuts in many datasets. 8.1. Limitations While the Role-Relative Projection effectively simplifies the multiparty problem, it inherently introduces a severe class im- balance in the training data. Because speakers are strictly or- dered into âCurrentâ and âNextâ based on activity, states rep- resenting a continuation of speech (holds) occur significantly more frequently than states representing a change in speaker (turn shifts). The model also relies on the history bins to desig- nate the current speaker. This dependence creates a delay of up to two seconds before the system recognizes a new floor holder. Additionally, the visual backbone relies on a standard ASD module. This module separates speaker tracks but misses subtle facial expressions. Upgrading to a more detailed visual encoder could improve early turn shift detection. Finally, the architecture supports variable-sized groups, but our evaluation only covers groups of two and three speaking English. Test- ing on larger groups in diverse languages is required to confirm actual scalability. 9. Conclusion We introduced MuVAP, a multimodal framework for predict- ing turn-taking in unconstrained multiparty settings. To address the scaling bottlenecks of joint modeling, we proposed a Role- Relative Projection that compresses complex group dynamics into a scalable pairwise state. Crucially, MuVAP operates on a single monaural audio stream and a single camera view, bypass- ing the traditional requirement for microphone arrays or spatial audio. We demonstrated that visual grounding via ASD tracks can effectively substitute for spatial separation, allowing the model to resolve speaker attribution even during competitive overlaps in a single-channel mix. On the AVCC dataset, Mu- VAP demonstrates consistent improvements in preemptive turn- taking prediction. Our results suggest that combining scalable state reduction with audiovisual grounding is a robust path for deploying responsive conversational agents in real-world social environments using only standard commodity hardware. Future work will transition this framework from offline evaluation to live deployment. Specifically, we plan to integrate MuVAP into a physical robot architecture to evaluate real-time, closed-loop turn-taking performance with human subjects. 10. Acknowledgment This work was supported by the Wallenberg AI, Autonomous Systems and Software Program (WASP) funded by the Knut and Alice Wallenberg Foundation (KAW), and the Swedish Re- search Council project 2020-03812. The computations and data handling were enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Supercomputer Centre. 11. Generative AI Use Disclosure Generative AI tools were used in this paper exclusively for edit- ing and polishing the text to remove grammatical errors. These tools were not used for writing any major parts of the paper. 12. References [1] S. Duncan, âSome signals and rules for taking speaking turns in conversations,â Journal of Personality and Social Psychology, vol. 23, p. 283â292, 08 1972. [2] H. Sacks, E. A. Schegloff, and G. Jefferson, âA simplest system- atics for the organization of turn-taking for conversation,â Lan- guage, vol. 50, no. 4, p. 696â735, 1974. [3] S. Duncan Jr, âOn the structure of speakerâauditor interaction dur- ing speaking turns,â Language in Society, vol. 3, no. 2, p. 161â 180, 1974. [4] C. E. Ford and S. A. Thompson, âInteractional units in conver- sation: Syntactic, intonational, and pragmatic resources for the management of turns,â Studies in interactional sociolinguistics, vol. 13, p. 134â184, 1996. [5] H. H. Clark, Using language. Cambridge University Press, 1996. [6] P. Ball, âListenersâ responses to filled pauses in relation to floor apportionment,â British Journal of Social & Clinical Psychology, vol. 14, p. 423â424, 11 1975. [7] S. Duncan and D. Fiske, Face-to-face interaction: research, meth- ods and theory. Wiley, 1977. [8] T. Stivers, N. J. Enfield, P. Brown, C. Englert, M. Hayashi, T. Heinemann, G. Hoymann, F. Rossano, J. P. De Ruiter, K.-E. Yoon et al., âUniversals and cultural variation in turn-taking in conversation,â Proceedings of the National Academy of Sciences, vol. 106, no. 26, p. 10 587â10 592, 2009. [9] S. C. Levinson and F. Torreira, âTiming in turn-taking and its im- plications for processing models of language,â Frontiers in psy- chology, vol. 6, p. 136034, 2015. [10] G. Skantze, âConversational interaction with social robots,â in Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction, 2021, p. 717. [11] E. Ekstedt and G. Skantze, âVoice Activity Projection: Self- supervised Learning of Turn-taking Events,â in Proc. Interspeech 2022, 2022, p. 5190â5194. [12] K. Inoue, D. Lala, G. Skantze, and T. Kawahara, âYeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,â in Proceedings of the 2025 Confer- ence of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, p. 7171â7181. [13] G. Skantze and B. Irfan, âApplying general turn-taking models to conversational human-robot interaction,â in Proceedings of the ACM/IEEE International Conference on Human-Robot Interac- tion (HRI), 2025. [14] M.-C. Lee, M. Trinh, and Z. Deng, âMultimodal turn analysis and prediction for multi-party conversations,â in Proceedings of the 25th International Conference on Multimodal Interaction, 2023, p. 436â444. [15] M. Elmers, K. Inoue, D. Lala, and T. Kawahara, âTriadic multi- party voice activity projection for turn-taking in spoken dialogue systems,â in Proc. Interspeech 2025, 2025, p. 3015â3019. [16] T. Liu, S. Fan, X. Xiang, H. Song, S. Lin, J. Sun, T. Han, S. Chen, B. Yao, S. Liu, Y. Wu, Y. Qian, and K. Yu, âMSDWild: Multi- modal Speaker Diarization Dataset in the Wild,â in Proc. Inter- speech 2022, 2022, p. 1476â1480. [17] R. Tao, Z. Pan, R. K. Das, X. Qian, M. Z. Shou, and H. Li, âIs someone speaking? exploring long-term temporal features for audio-visual active speaker detection,â in Proceedings of the 29th ACM International Conference on Multimedia, 2021, p. 3927â 3935. [18] X. Wang, F. Cheng, and G. Bertasius, âLoconet: Long-short con- text network for active speaker detection,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, p. 18 462â18 472. [19] J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xi et al., âAVA Active Speaker: An audio-visual dataset for active speaker detection,â in ICASSP 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, p. 4492â4496. [20] E. Ekstedt and G. Skantze, âHow much does prosody help turn- taking? Investigations using voice activity projection models,â in Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. Edinburgh, UK: Association for Computational Linguistics, Sep. 2022, p. 541â551. [Online]. Available: https://aclanthology.org/2022.sigdial-1.51 [21] R. Ishii, K. Otsuka, S. Kumano, M. Matsuda, and J. Yamato, âPre- dicting next speaker and timing from gaze transition patterns in multi-party meetings,â in Proceedings of the 15th ACM on Inter- national conference on multimodal interaction, 2013, p. 79â86. [22] S. Heo, C. Miller, C. Murdock, and M. Proulx, âGaze-enhanced multimodal turn-taking prediction in triadic conversations,â in Proc. Interspeech 2025, 2025, p. 1068â1072. [23] K. Onishi, H. Tanaka, and S. Nakamura, âMultimodal voice activ- ity prediction: Turn-taking events detection in expert-novice con- versation,â in Proceedings of the 11th International Conference on Human-Agent Interaction, 2023, p. 13â21. [24] R. Ishii, K. Otsuka, S. Kumano, R. Higashinaka, and J. Tomita, âPrediction of who will be next speaker and when using mouth- opening pattern in multi-party conversation,â Multimodal Tech- nologies and Interaction, vol. 3, no. 4, p. 70, 2019. [25] V. Petukhova and H. Bunt, âWhoâs next? speaker-selection mech- anisms in multiparty dialogue,â in Proceedings of the Workshop on the Semantics and Pragmatics of Dialogue, 2009, p. 19â26. [26] M.-C. Lee, W. A. Li, and Z. Deng, âA computational study on sentence-based next speaker prediction in multiparty conversa- tions,â in Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents, 2024, p. 1â4. [27] M. Roddy, G. Skantze, and N. Harte, âMultimodal continuous turn-taking prediction using multiscale rnns,â in Proceedings of the 20th ACM International Conference on Multimodal Interac- tion, 2018, p. 186â190. [28] S. O. Russell and N. Harte, âVisual cues enhance predictive turn- taking for two-party human interaction,â in Findings of the As- sociation for Computational Linguistics: ACL 2025, 2025, p. 209â221. [29] M. Cheng, F. Su, C. Li, J. Liu, and M. Li, âMulti-channel sequence-to-sequence neural diarization: Experimental results for the misp 2025 challenge,â in Proc. Interspeech 2025, 2025, p. 1898â1902. [30] M.-C. Lee and Z. Deng, âEnhancing gaze prediction in multi- party conversations via speaker-aware multimodal adaptation,â in Proceedings of the 27th International Conference on Multimodal Interaction, 2025, p. 200â208. [31] O. K Ì op Ì ukl Ì u, M. Taseska, and G. Rigoll, âHow to design a three- stage architecture for audio-visual active speaker detection in the wild,â arXiv preprint arXiv:2106.03932, 2021. [32] K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al., âEgo4d: Around the world in 3,000 hours of egocentric video,â in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, p. 18 995â19 012. [33] W. Kraaij, T. Hain, M. Lincoln, and W. Post, âThe ami meeting corpus,â in Proc. International Conference on Methods and Tech- niques in Behavioral Research, 2005, p. 1â4. [34] J. Deng, J. Guo, X. An, Z. Zhu, and S. Zafeiriou, âMasked face recognition challenge: The insightface track report,â in Proceed- ings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, p. 1437â1444. [35] J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, âReti- naface: Single-shot multi-level face localisation in the wild,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, p. 5203â5212. [36] J. Guo, J. Deng, A. Lattas, and S. Zafeiriou, âSample and com- putation redistribution for efficient face detection,â arXiv preprint arXiv:2105.04714, 2021. [37] J. Deng, J. Guo, N. Xue, and S. Zafeiriou, âArcface: Additive an- gular margin loss for deep face recognition,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2019, p. 4690â4699. [38] H. W. Kuhn, âThe hungarian method for the assignment prob- lem,â Naval research logistics quarterly, vol. 2, no. 1-2, p. 83â 97, 1955. [39] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, âAttention is all you need,â in Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wal- lach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30.Curran Associates, Inc., 2017. [Online]. Avail- able:https://proceedings.neurips.c/paper files/paper/2017/file/ 3f5e243547dee91fbd053c1c4a845a-Paper.pdf [40] C. Jung, S. Lee, K. Nam, K. Rho, Y. J. Kim, Y. Jang, and J. S. Chung, âTalknce: Improving active speaker detection with talk- aware contrastive learning,â in ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, p. 8391â8395. [41] G. Skantze, âTurn-taking in conversational systems and human- robot interaction: A review,â Computer Speech & Language, vol. 67, p. 101178, 2021.