Paper deep dive
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/3/2026, 3:11:41 AM
Summary
The paper introduces FriendBench, a benchmark for inferring dyadic familiarity (familiar vs. strangers) from 20-second clips of ice-breaker conversations. It evaluates 26 multimodal models from seven companies against human panels across text, audio, and video modalities. Results show that while the best models achieve accuracy statistically indistinguishable from human crowds, they differ in strategy: humans maintain balanced response biases, whereas models exhibit a strong bias toward predicting 'strangers'. Humans uniquely benefit from visual cues in addition to audio, while models under-exploit visible behavior.
Entities (13)
Relation Signals (9)
FriendBench â evaluates â 26 models
confidence 95% ¡ We evaluate 26 models from seven companies... across all three modalities.
FriendBench â evaluates â human panels
confidence 95% ¡ compare 26 models from seven companies against matched human panels over 96 balanced dyads.
FriendBench â measurescapability â dyadic familiarity inference
confidence 95% ¡ a benchmark for inferring whether two people are already familiar or are meeting as strangers
FriendBench â usesdataset â Seamless Interaction
confidence 95% ¡ built on the Seamless Interaction dataset (Agrawal et al., 2025)
Models â exhibitsbias â stranger bias
confidence 90% ¡ the strongest models lean toward 'stranger'
humans â exhibitsbias â balanced response
confidence 90% ¡ humans stay balanced across the two answers
gemini_audio_pro â developedby â Google
confidence 85% ¡ gemini_audio_pro is listed under audio models; Google is listed as a contributing company.
gpt_text_mini â developedby â OpenAI
confidence 85% ¡ gpt_text_mini is listed under text models; OpenAI is listed as a contributing company.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.
Tags
Links
- Source: https://arxiv.org/abs/2607.29602v1
- Canonical: https://arxiv.org/abs/2607.29602v1
Trouble viewing inline? Open PDF directly â
Full Text
65,423 characters extracted from source content.
Expand or collapse full text
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony DâAvirro, Benjamin Peloquin Fluid Concepts Research Correspondence: jeff@fluidconcepts.ai Abstract Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward âstrangerââa difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions. FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony DâAvirro, Benjamin Peloquin Fluid Concepts Research Correspondence: jeff@fluidconcepts.ai 1 Introduction Social intelligence, the capacity to make sense of other people and the situations they are in, is a central component of human cognition (Adolphs, 2003; Frith and Frith, 2007). It is also an increasingly important requirement for AI systems, which now observe, mediate, and participate in human interaction in roles ranging from companion agents to meeting assistants and care monitors (Mathur et al., 2024; Sap et al., 2019). Such systems must reason about the social situation people are in, and much of that reasoning rests on multimodal behavioral cues rather than explicit verbal content, such as how people coordinate, respond, and orient toward one another (Tickle-Degnen and Rosenthal, 1990). A system that reads only the semantic content of a conversation captures just part of what social understanding requires. We study one concrete instance of social intelligence: the capacity to recognize the relationship between two people from a brief sample of how they interact. Specifically, we ask whether an observer can tell that two people are already familiar with one another rather than meeting as strangers, without being told and without relying on what the conversation is about. This judgment is a clean probe of behavioral social perception for two reasons. First, since we have ground-truth relationship labels, its answer is a matter of fact rather than interpretation: much related work in social intelligence targets intentions, emotions, or beliefs (Premack and Woodruff, 1978; Sap et al., 2019), whose correct answer is often contestable, whereas prior familiarity has an unambiguous ground truth (i.e., two people either have met before or they have not). Second, familiarity is often unstated in a short interaction, so it must be inferred from behavior. The task therefore rules out succeeding by reading semantic content alone, and prior work shows that humans make accurate social judgments from brief samples of behaviorâso-called thin slices (Ambady and Rosenthal, 1992). We operationalize the task as FriendBench: binary classification, familiar versus strangers, from a 20-second clip of a dyadic ice-breaker conversation, evaluated separately across text, audio, and video modalities. Every dyad, whether familiar or strangers, responds to the same type of ice-breaker prompt, so the conversational topic is matched across the two classes and cannot by itself reveal the answer, leaving the manner in which the two people interact as the primary signal. In addition to the core capability we benchmark, we examine how humans and models draw on the three modalities, and how each reaches its accuracyâby telling the two classes apart (discrimination) or by answering one class more often regardless of the evidence (response bias). Two patterns stand out. First, richer channels help both models and humans alike, but unequally: audio adds reliable signal over text for each, yet only humans gain a further reliable increment from the visual channel, while the strongest models are flat from audio to audiovisualâthey under-exploit the visible behavior humans read. Second, humans stay close to balanced across both classes in every modality, whereas models show a larger and more idiosyncratic class bias, most clearly in the text-only condition. Both patterns concern how humans and models solve the task, not merely whether they succeed (§5). Contributions. ⢠We introduce FriendBench, a multimodal benchmark for inferring dyadsâ prior familiarity (familiar vs. strangers) from ice-breaker clips in which every dyad answers the same type of prompt, built on the Seamless Interaction dataset (Agrawal et al., 2025).111The benchmark stimuli, human ratings, and model predictions are openly available at https://huggingface.co/datasets/fluid-concepts/friend-bench. ⢠We release a matched human-rater dataset for the task, covering 96 dyads in text, audio, and video, with roughly 90 raters per modality. ⢠We evaluate 26 models from seven companiesâOpenAI, Google, Anthropic, Alibaba, Mistral, Thinking Machines, and Metaâspanning proprietary and open-weight systems, across all three modalities. ⢠We evaluate humans and models on the same stimuli under matched conditions. On accuracy, the best model and the human crowd are statistically indistinguishable in every modality. But equal accuracy is not human-like perception. The strongest models reach it by leaning toward âstranger,â while human raters stay balanced. Using signal detection theory, we characterize this as a difference in effective prior, not discrimination (§5). ⢠We find a modality ordering both rater types obey only in part: text carries little signal for either and audio adds reliable discrimination, but the audiovisual channel gives humans a further reliable gain while leaving the strongest models flatâcurrent models capture the vocal signal but under-exploit the visible behavior human observers read (§5.3). 2 Related Work Thin-slice perception of familiarity. Human observers form accurate judgments about people and relationships from very brief behavioral exposure. Ambady and Rosenthalâs (1992) meta-analysis found that judgments from observations under five minutes (and often under 30 seconds) predicted objective outcomes at râ.39râ.39, with longer exposure adding little. Familiarity in particular is legible in thin slices, with friendship the best-studied case: observers tell friends from strangers from silent video (Latif et al., 2014) or brief audio (Bryant et al., 2020), and cues such as inter-turn timing distinguish them (Templeton et al., 2023). Critically for our design, Dunbar et al. (2022) show that relationship quality remains inferable from speech even after its lexical content is digitally removedâthe relational signal need not come from what is said. This motivates both our task and our 20-second window, which is ample by this evidence. Recognizing relationships from behavior. One line of work predicts relationship type from images or video via supervised classification: PISC (Li et al., 2017) labels images as intimate, non-intimate, or no-relation, PIPA (Sun et al., 2017) annotates sixteen fine-grained relations in photo albums, and more recent work classifies asymmetric relations from the temporal dynamics of a live interaction (Tang et al., 2026). A parallel line infers relationships from the semantic content of dialogueâacoustic-lexical classifiers over phone calls (Katerenchuk et al., 2014) and language models over movie-script dialogue, from relation-classification datasets such as DDRel (Jia et al., 2021) to LLM evaluations where GPT-4o infers speaker relationships well above chance (Kim et al., 2026). Both differ from our task in two ways: they infer relationship type rather than the presence or absence of prior familiarity, and they lean on appearance, scene, or semantic content (especially for scripted dialogue). We instead fix the conversational prompt, controlling the semantic content these methods lean on. Social reasoning in multimodal models. Multimodal models are increasingly tested for social understanding: Social-IQ (Zadeh et al., 2019) poses questions about social videos, while SIV-Bench (Kong et al., 2026) and PIVOTSBench (Zhang et al., 2026) probe reasoning about social scenes and fine-grained relations. HumanSense (Qin et al., 2026) is the closest to our setting; one of its subtasks asks a model to judge how well two people in a video know each other. Three things set our benchmark apart: (1) every dyad answers the same type of prompt, so the topic itself cannot give the answer away; (2) the label is objective, recording whether a pair had actually met before rather than how close a viewer judges them; and (3) we gather matched human ratings in text, audio, and video for direct comparison. Our stimuli come from the Seamless Interaction corpus (Agrawal et al., 2025). Other dyadic corpora record acquaintance tooâUDIVA (Palmero et al., 2021) labels each pair known or unknown, NoXi (Cafaro et al., 2017) rates how well partners know each otherâbut treat it as metadata, not as a label for prediction. 3 Methods 3.1 Benchmark Construction Samples are drawn from the Seamless Interaction dataset (Agrawal et al., 2025), restricted to naturalistic interactions from three recording sites (the datasetâs vendor field), each contributing a comparable share of dyads. We exclude the datasetâs improvised interactions, in which participants are assigned a relationship and asked to act it out, since our aim is to measure this task on real, unscripted behavior rather than acted behavior. Because the benchmark is used only for zero-shot evaluation and never for training, we pool dyads across all of the datasetâs predefined splits rather than restricting to one, maximizing the pool from which our quality filters and stratified design (below) can draw. Seamless Interaction sessions include many different interaction types; we use only the Either-Or (EO) ice-breaker, a short task in which one participant poses a forced-choice hypothetical question (e.g., âwould you rather have the ability to fly or be invisible?â) and the dyad discusses their answer. We chose EO for three reasons: it standardizes the conversational context, since familiar and stranger pairs do the same thing and any signal must come from how a dyad responds rather than what it discusses; it is always the first interaction in a session, so every dyad is sampled from the same point in their interaction history; and its content carries little direct information about relationship status, so the task cannot be solved from semantic content alone. In free conversation, by contrast, status can leak through contentâshared history and inside knowledge for familiars, getting-acquainted basics for strangersâcues a shared hypothetical question largely suppresses. An EO interaction, though centered on one either-or question, typically continues for several minutes. We sample one 20-second clip per dyad from this interaction, drawn from either its early or late portion (counterbalanced across relationship and recording site), with boundaries snapped (Âą 5s) to the nearest turn start so clips do not begin or end mid-utterance. The rendered video stimulus retains the clipâs audio track, so the video condition is audiovisualâvisible behavior together with speechâwhereas the audio and text conditions each isolate a single channel; both human raters and video models therefore receive sound as well as picture in the video condition. The final set fully crosses recording site, relationship (familiar/strangers), and gender composition (same-gender/mixed-gender pair), with 8 dyads per cell (96 total), and is participant-disjoint: no individual appears in more than one dyad. Crossing rules out recording site and gender composition as confounds (§4). Disjointness keeps each dyad an independent observation and guards against recognition leakage: because human raters see multiple stimuli in one session, a shared individual could be recognized from an earlier item, letting that recognition serve as the cue rather than the relationship itself.222This is not a risk for our models, which are evaluated zero-shot with no training process in which to learn identity. Interaction- and clip-level filters (minimum turns, in-window speaker balance, audio-track synchrony, interaction length, prompt completeness; thresholds in Appendix A, Table A1) exclude clips that cannot support the task regardless of relationship, e.g., one participant speaking for only a few seconds or desynchronized audio tracks. An additional LLM-based audit checks that each interactionâs recorded prompt label matches its content. 3.2 Human Ratings To establish baseline human performance, we collected human ratings of relationship type in a crowdsourced study. For each stimulus, raters made a forced choice: familiar or strangers. Before rating, each rater read a short instruction screen. It explained that they would read, listen to, or watch (depending on modality) a clip from a real conversation in which the two people were playing the âEither Orâ ice-breaker game, then judge whether the pair had met before (familiar) or were meeting for the first time (strangers). The full instruction text is reproduced in Appendix C. Raters were recruited via Prolific and completed the task on GORILLA. We ran three experiments, one per modality (text, audio, video), each with a panel of roughly 90 raters (94 text, 94 audio, 92 video). Prolific prescreening and quota matching ensured a gender-balanced pool of US nationals residing in the US, with no reported hearing difficulties or autism-spectrum diagnosis, and excluded anyone who had already participated in one of our other modality panels (see Appendix D). Raters followed a planned-missing, 6-block incomplete design (Graham et al., 2006), each rating one of six overlapping blocks and seeing 32 of the 96 dyads by design; individual stimuli were in turn rated by 25â35 raters.This designârather than one rating per stimulus, averagedâlets us model raters themselves, estimating each raterâs own accuracy and response bias via crossed random effects, which requires each rater to contribute enough trials to be more than a one-shot data point. Each rater additionally completed the Social Information Processing subscale from the Tromsø Social Intelligence Scale (TSIS; Silvera et al., 2001; Grieve and Mahar, 2013) and a custom post-task self-report of which cues (visual, vocal, verbal, interactional) they attended to, enabling individual-differences analysis (Appendix E). 3.3 Model Predictions We evaluate 26 models from seven companies across text, audio, and video (Table 2), a fully-crossed design in which every model rates every stimulus; unlike the human panels, no incomplete-block correction is needed for the model results. Each model receives the same prompt and response parsing. The model prompt is adapted directly from the human instruction text, so that both rater types are given the same task framing, the same âEither Orâ description, and the same familiar/strangers definition; the only difference is the response mechanismâa forced-choice button for humans, a JSON label with a confidence rating for models. Models and prompts are further described in Appendix B and Appendix C, respectively. A few audio models decline to answer on a share of trials (returning no usable familiar/strangers label) rather than committing to a forced choice. Because a refusal is not a wrong answer, we report each modelâs accuracy over its answered (covered) trials and give per-model coverage alongside it (Table 2); scoring refusals as errors would understate a model for declining rather than for misjudging. Coverage is near-total except for three audio models (§4), and the top models in each modality answer every trial, so this choice does not affect the headline comparison. 4 Benchmark Results Table 1 summarizes, per modality, the average per-rater human, the human crowd (the per-stimulus majority vote of all raters who saw that stimulus), and the single best-performing modelâtheir accuracy and, for the class-bias analysis in §5.1, their per-class recall and signal-detection statistics (Stanislaw and Todorov, 1999). Table 2 gives the complete per-model breakdown with Wilson 95% confidence intervals. Figure 1 depicts the same comparison, plotting every model against the average individual human rater and the majority-vote crowd in each modality. Rater Acc. Fam. Str. dâ˛d c Text Human (indiv.) 49.9 54.6 45.3 0.00 â0.12-0.12 Human (crowd) 51.0 59.6 42.6 0.05 â0.21-0.21 gpt_text_mini 56.2 20.8 91.7 0.54 1.061.06 Audio Human (indiv.) 56.3â 57.2 55.5 0.32 â0.02-0.02 Human (crowd) 63.5â 64.6 62.5 0.68 â0.03-0.03 gemini_audio_pro 66.7â 39.6 93.8 1.21 0.860.86 Video Human (indiv.) 60.9â 58.7 63.2 0.56 0.060.06 Human (crowd) 71.9â 70.2 74.5 1.16 0.060.06 gemini_video 66.7â 41.7 91.7 1.12 0.770.77 Table 1: Human raters (average individual and majority-vote crowd) vs. the best model per modality. Acc. is accuracy; Fam. and Str. per-class recall (all %); dâ˛d and c as in Table 2. The named model in each block is that modalityâs best-performing (top-accuracy) system, scored over answered trials; human-indiv. accuracy is the per-rater mean (per-model Wilson CIs in Table 2). /â/ââŁâ^*/^**/^***: above chance at .05/.01/.001.05/.01/.001, by exact binomial vs. 50% (crowd, best model) or 95%/99%/99.9%95\%/99\%/99.9\% GLMM posterior HDI excluding 50% (individual rows). Figure 1: Human vs. model accuracy by modality. Each rater point covers only the âź 32 trials that rater saw, so the rater cloud is widened by sampling noise and its upper tail is not a band of expert raters (§5.4). Horizontal bars are Wilson 95% CIs for the best model and crowd, not tests of humanâmodel or cross-modality differences (§4, §5.3). The best model and the human crowd trade the point-estimate lead across modalities. The model edges the crowd in text (by 5.2 points) and audio (by 3.2), while the crowd leads in video (by 5.2). Video is humansâ strongest modality, and the one place the crowd, not a model, comes out ahead. None of these three gaps is statistically reliable, though. On the 96 paired dyads, an exact McNemar test (Dietterich, 1998; McNemar, 1947) never approaches significance (p>.4p>.4 in every modality). Text is the weakest modality for humans: at chance individually and barely above it as a crowd. But the best text model is also not above chance (56.2%, p=.26p=.26). Text is therefore less a model win than a modality where neither rater type finds a signal. Model Acc. 95% CI Cov. Fam. Str. dâ˛d c Text gpt_text_mini 56.2 [46.3, 65.7] 100 20.8 91.7 0.54 1.061.06 gemini_text_thinking 55.8 [45.8, 65.4] 99 27.7 83.3 0.36 0.760.76 gemini_text 55.2 [45.3, 64.8] 100 25.0 85.4 0.36 0.840.84 gpt_text 53.1 [43.2, 62.8] 100 29.2 77.1 0.19 0.630.63 inkling_text 52.1 [42.2, 61.8] 100 16.7 87.5 0.17 1.031.03 claude_text 52.1 [42.2, 61.8] 100 14.6 89.6 0.19 1.121.12 muse_spark_text_high 51.0 [41.2, 60.8] 100 8.3 93.8 0.14 1.401.40 mistral_text 50.0 [40.2, 59.8] 100 87.5 12.5 0.00 â1.11-1.11 muse_spark_text 49.0 [39.2, 58.8] 100 10.4 87.5 â0.10-0.10 1.161.16 Audio gemini_audio_pro 66.7â [56.8, 75.3] 100 39.6 93.8 1.21 0.860.86 gemini_audio 63.5â [53.6, 72.5] 100 81.2 45.8 0.76 â0.48-0.48 muse_spark_audio 57.3 [47.3, 66.7] 100 20.8 93.8 0.67 1.131.13 gemini_audio_thinking 55.2 [45.3, 64.8] 100 64.6 45.8 0.26 â0.23-0.23 muse_spark_audio_high 53.1 [43.2, 62.8] 100 18.8 87.5 0.25 0.990.99 gpt_audio_1_5 53.1 [43.2, 62.8] 100 87.5 18.8 0.25 â0.99-0.99 qwen_audio 51.5 [39.8, 62.9] 71 0.0 100.0 0.02 2.192.19 inkling_audio 51.0 [41.2, 60.8] 100 41.7 60.4 0.05 0.230.23 voxtral_audio 50.0 [40.2, 59.8] 100 100.0 0.0 0.00 â2.32-2.32 gpt_audio 48.8 [34.6, 63.2] 45 100.0 4.3 0.45 â1.76-1.76 gpt_audio_mini 48.1 [30.7, 66.0] 28 25.0 81.8 0.18 0.720.72 qwen_omni 47.9 [38.2, 57.8] 100 16.7 79.2 â0.15-0.15 0.870.87 Video gemini_video 66.7â [56.8, 75.3] 100 41.7 91.7 1.12 0.770.77 gemini_video_thinking 58.3 [48.3, 67.7] 100 27.1 89.6 0.62 0.910.91 gemini_video_pro 57.3 [47.3, 66.7] 100 18.8 95.8 0.77 1.251.25 muse_spark_video 54.2 [44.2, 63.8] 100 14.6 93.8 0.44 1.241.24 muse_spark_video_high 53.1 [43.2, 62.8] 100 10.4 95.8 0.42 1.421.42 Table 2: Per-model results, sorted within modality. Acc. is accuracy (%) over answered trials and CI its Wilson 95% confidence interval; Cov. is coverage (% answered rather than declined). Fam. and Str. are per-class recall (%): share of familiar and of stranger dyads correctly labeled. dâ˛d and criterion c treat âfamiliarâ as the signal class (log-linear corrected; Hautus, 1995), so c>0c>0 indicates a bias toward âstrangerâ and c<0c<0 toward âfamiliar.â â/â: accuracy above chance at p<.01p<.01/p<.05p<.05 (exact binomial vs. 50%). Only three models are individually above chance at this sample size, all from Google: gemini_audio_pro and gemini_audio in audio, and gemini_video in video. No text model clears chance. Apart from the two top-accuracy systems, the best remaining model in each modality reaches only 56.2% (text, gpt_text_mini), 63.5% (audio, gemini_audio), and 58.3% (video, gemini_video_thinking). Most models cluster near chance on raw accuracy. The per-class recall and criterion columns already hint at the analysis in §5.1: several models carry very large response biases regardless of accuracy. In audio the bias is large but inconsistent in direction. voxtral_audio labels every dyad âfamiliarâ (100%/0%, c=â2.32c=-2.32), while qwen_audio does the opposite (0%/100%, c=2.19c=2.19). In text most models lean hard toward âstranger,â with muse_spark_text_high recovering only 8.3% of familiar pairs (c=1.40c=1.40). A near-chance accuracy can thus conceal a strongly one-sided response pattern rather than balanced guessing. Three audio models also decline a substantial share of trials: gpt_audio answers only 45%, gpt_audio_mini 28%, and qwen_audio 71%. They sit at chance on the trials they do answer, so their behavior is better summarized as declining-plus-guessing. 5 Performance Analysis 5.1 Class Bias: Humans vs. Models Humans and the best-performing models are close to parity on accuracy (§4), but do they reach it the same way? To separate discrimination from response bias, we read the per-class recall and signal-detection columns of Table 1, treating âfamiliarâ as the signal class: dâ˛d measures how well a rater tells the classes apart, and the criterion c how far its decision rule leans toward one answer (c>0c>0 toward âstranger,â c<0c<0 toward âfamiliarâ). Two patterns stand out. Humans stay near-balanced everywhere: the human criterion is within 0.230.23 of zero in every modality, and per-class recall gaps never exceed 10 points for pooled raters. The best text and audio models instead reach their accuracy largely through response bias: gemini_audio_pro labels 93.8% of strangers correctly but recovers only 39.6% of familiar pairs (c=0.86c=0.86), and the best text model, gpt_text_mini, is more extreme still (91.7% vs. 20.8%, c=1.06c=1.06). Across the roster this is the rule: in text almost every model leans âstrangerâ by 50â85 recall points, and in audio the bias is large but inconsistent in direction (Table 2). The gap holds even in video, where models come closest to humans. There the best model (gemini_video) still leans âstrangerâ (c=0.77c=0.77; 91.7% of strangers but only 41.7% of familiar pairs), while the human crowd stays balanced (c=0.06c=0.06) at slightly higher discrimination (dâ˛=1.16d =1.16 vs. 1.121.12). The two reach comparable dâ˛d , but only the crowd does so without skewing its decision rule. Models thus fail to recreate the human balanced-recall pattern in any modality, including their strongest: several of the highest-accuracy systems (Table 2) buy that accuracy with a skewed criterion rather than sharper discrimination. 5.2 Wisdom of the Crowd Pooling human raters via majority vote barely moves accuracy in text (49.9%â 51.0%, ++1.1 points) but raises it substantially in audio (56.3%â 63.5%, ++7.2 points) and video (60.9%â 71.9%, ++11.0 points). Figure 2 shows the full crowd-growth curves: majority-vote accuracy rises with crowd size k in audio and video but stays flat near chance in text. The audio and video gains therefore reflect real signal, not an artifact of pooling. Figure 2: Wisdom of the crowd. Majority-vote accuracy as a function of human crowd size k, with raters sampled per stimulus from that stimulusâs own rater panel (mean Âą SD over 300 bootstrap draws). Legend gives single-rater â full-panel crowd accuracy per modality. 5.3 Modality Comparison Text is the weakest modality for both humans and models, and richer channels help bothâthough, as we show below, not to the same degree. To test the modality ordering while accounting for the repeated-measures structure of the human data, we fit a Bayesian crossed random-effects logistic model (Gelman et al., 2014) to per-trial human correctness (Bernoulli likelihood, logit link), with a fixed effect of modality and crossed random intercepts for rater and for dyad (§3.2); this is the design the block structure was built to support. We estimate it with the bambi interface to PyMC (Capretto et al., 2022), using its default weakly-informative, data-scaled priors (wide normal priors on the fixed effects and half-normal hyperpriors on the random-effect standard deviations), and sample the posterior with NUTS (4 chains, 1,000 warmup and 1,000 post-warmup draws each; all R^â1.00 Râ 1.00). We report population-averaged (marginal) accuracies and contrasts as posterior means with 95% highest-density intervals (HDIs, Makowski et al., 2019). The marginal accuracies confirm the raw patternâtext 50.0%50.0\% (95% HDI [46.7,53.6][46.7,53.6]), audio 56.6%56.6\% [52.2,60.6][52.2,60.6], video 60.9%60.9\% [56.9,64.9][56.9,64.9]âand the pairwise contrasts are decisive: audio exceeds text by 6.56.5 points (HDI [4.1,8.9][4.1,8.9]), video exceeds audio by 4.54.5 points ([1.9,6.8][1.9,6.8]), and video exceeds text by 10.910.9 points ([8.5,13.3][8.5,13.3]), each with posterior P>0P>0 of at least 0.9990.999. Text alone is statistically indistinguishable from chance (posterior probability of exceeding 50% only 0.500.50). Because the video condition is audiovisual (§3.1), its edge over audio reflects the added value of visible behavior on top of speech, not vision in isolation; video is best read as an audiovisual upper bound rather than a vision-only channel. The model side matches the human floor but not the human ceiling. The best model climbs from text (56.2%56.2\%) to audio (66.7%66.7\%) but gains nothing from the audiovisual channel (66.7%66.7\%), and among models evaluated in both audio and video only Gemini Flash improves (63.5%â66.7%63.5\%â 66.7\%) while Gemini Pro (66.7%â57.3%66.7\%â 57.3\%) and Muse Spark (57.3%â54.2%57.3\%â 54.2\%) decline. (Cross-modality model means rise monotonicallyâ52.752.7, 53.953.9, 57.9%57.9\% for text, audio, videoâbut the video roster is small and skewed toward stronger companies, so we read the best-model and within-model trajectories rather than the pooled mean.) Where humans reliably convert visible behavior into accuracy on top of speech, then, the strongest models do not: they capture the vocal signal but under-exploit the visual channel, the one place a benchmark of multimodal social perception most expects a model to gain. That text is the hardest modality is consistent with the benchmarkâs design rather than a defect of it. The EO ice-breaker was chosen precisely so that transcript content carries little direct information about relationship status (§3.1); the weak transcript performance of both humans and models is the expected consequence of that choice. It also helps explain why response bias is most visible in text: with little signal available, a raterâs responses reflect its bias more than the stimulus. 5.4 Rater Individual Differences Because each rater contributes many trials by design (§3.2), we can estimate individual accuracy and relate it to rater traits. Trait social intelligence (TSIS-PS scale; reliable in every panel, Îą=0.88Îą=0.88â0.920.92) is essentially uncorrelated with accuracy in the two modalities where the task is doable: audio r=0.03r=0.03 (p=.78p=.78) and video r=â0.12r=-0.12 (p=.24p=.24). The only significant association is in text (r=0.29r=0.29, p=.004p=.004), the modality where average accuracy is at chance, so we read it with caution. Self-reported cue use is similarly flat: raters most often report attending to interactional cues (rapport, responsiveness), but cue-use scores rarely predict accuracy (Appendix E). The individual-rater cloud in Figure 1 is correspondingly wide, with the best raters near 75% in every modality. Its spread should not be read as a stable band of expert observers, though. Each raterâs accuracy comes from only the âź 32 trials they saw, so the upper tail is close to what sampling noise alone would produce. The best model therefore sits within the human distribution, not above it, even as almost no individual beats the aggregated crowd (1 of 92 in video). The crossed random-effects model makes the point directly: the dyad random-intercept SD (0.670.67â0.950.95 across modalities, logit scale) is six to seven times the rater SD (0.100.10â0.140.14). Who the rater is thus matters far less than which dyad they were rating. 5.5 Dyad Difficulty Treating the per-dyad random intercepts from the crossed random-effects model as Rasch-style item-easiness parameters (Rasch, 1960; de Boeck and Wilson, 2004), we find that difficulty is largely a property of the conversation, not of the modality through which it is observed. Estimated dyad easiness correlates positively across all modality pairs (textâaudio r=0.52r=0.52, textâvideo r=0.33r=0.33, audioâvideo r=0.61r=0.61; all p<.01p<.01): a dyad that is hard to read from one modality tends to be hard from the others as well. A handful of dyads sit below chance in every modality, acting as systematic âluresâ that most raters misread alike (Appendix F). Humans and models tend to find the same dyads hard, though how strongly depends on how it is measured. Correlating per-dyad human-crowd accuracy with per-dyad accuracy pooled over all models yields r=0.56r=0.56 (audio) and r=0.34r=0.34 (video), both significant (p<.01p<.01; Appendix Figure A1), but only r=0.20r=0.20 in text (p=.046p=.046). This raw correlation understates the shared difficulty. It blends two things: how hard a dyad is to read, and which class it belongs to. Humans stay balanced while the models lean toward âstrangerâ (§5.1), so the two rater types tend to miss different classes. Isolating difficulty from this class split, by correlating within each true class, brings the agreement out clearly: r=0.61r=0.61 (text), 0.580.58 (audio), and 0.440.44 (video), all p<.001p<.001. So a substantial part of what makes a pair legible or illegible is a property of the interaction shared across observers, even though the two reach their answers by different rules. 6 Discussion For a benchmark of multimodal social perception, the clearest result is about the modalities themselves. Text carries little relational signal for anyone, by design: the shared ice-breaker prompt suppresses lexical content (§3.1). Richer channels help both humans and models, but asymmetrically. Humans improve reliably at every step, from text to audio to audiovisual (§5.3), gaining a credible increment from visible behavior on top of speech. The strongest models capture the audio gain but not the visual one: their top accuracy is identical in audio and audiovisual (66.7%66.7\%). The sharpest humanâmodel difference is thus not whether models can read relationships but whether they exploit the visual channel as humans do. This gap has applied stakes. The companion agents, meeting assistants, and care monitors that motivate the task must read social situations from behavior, and the visible cues humans exploit on top of speech are exactly what current models miss. In terms of overall accuracy, the best models have drawn level with an aggregated human crowd. In every modality the two are statistically indistinguishable. Point estimates put the model slightly ahead in text and audio and the crowd slightly ahead in video (Table 1), but no gap survives a paired McNemar test (p>.4p>.4 throughout; §4). The difference is in how humans and models use their two answers. Discrimination (dâ˛d ) measures how well a rater tells the classes apart. Response bias (criterion c; Table 1) measures how far it leans toward one answer. Human raters stay balanced across the two answers in every modality. The strongest models instead lean toward âstranger,â giving that answer more often regardless of the pair. Video makes this clearest. There the best model tells the classes apart about as well as the crowd (dâ˛â1.1d â 1.1), but it reaches that accuracy by leaning toward âstrangerâ where the crowd stays balanced. In signal-detection terms, this is a difference in effective prior: the models behave as though strangers were the more common answer and demand more evidence before saying âfamiliar.â Humans behave as though the classes were equally likely, which is true in our balanced set. This describes the response distribution, not a mechanism, and the prior is inferred from behavior rather than verified. A pure criterion shift is correctable: recentering to the known base rate would raise accuracy with no gain in discrimination. So a biased modelâs raw accuracy can understate its discrimination. At the same time, humans and models are not perceiving unrelated things. Dyad difficulty is partially shared and partially transfers across modalities (§5.5), so part of what makes a pair legible or illegible is a property of the interaction itself, not of the observer or channel. The dissociation is therefore specific: the two agree on which dyads are hard but diverge on the decision rule they apply. Methodologically, these results argue for evaluating social-perception models with more than a single accuracy number: per-class recall, signal-detection statistics, and item-level difficulty each revealed structure that accuracy alone obscured, and are cheap to report once the underlying predictions are available. Taken together, the strongest models now match an aggregated human crowd on accuracy in this task, but not on how they reach it. They find the same conversations hard. Yet they systematically lean toward âstranger,â where human raters stay balanced, and they do not successfully convert the visual channel into accuracy. Whether models can be brought to read relationships as humans doâbalanced across classes, and drawing on visible behavior rather than speech aloneâis the open question this benchmark is built to track. Limitations This paper reports only the ice-breaker (first-interaction) task from the Seamless Interaction dataset. Findings about modality strength and class bias may not generalize to conversations later in a relationship or session, or to different prompt structures. All familiar subtypes (friends, family, romantic partner, coworkers, familiar_other) are collapsed into a single âfamiliarâ class, because the balanced design does not have enough dyads per subtype to power a multi-class analysis. A six-class relationship-type task is defined in the broader project but is out of scope here, so any within-familiar heterogeneity is invisible to this binary framing. Only two-person interactions are studied, and relationship inference in larger groups may draw on cues not captured here. The balanced evaluation set is also modest in size (96 dyads, one clip each), so per-model accuracy intervals are wide and only the strongest models clear chance individually. The class-bias and difficulty patterns we emphasize are more robust than any single modelâs rank. Our human baseline is drawn entirely from US-resident raters (§3.2); relationship-perception cues can be culturally specific, so this baseline may not represent human performance in other populations. Finally, even the ânaturalisticâ subset was recorded in a fixed-camera motion-capture studio, with wired lapel microphones and a posed ice-breaker prompt. This is not truly in-the-wild interaction, so findings may not transfer to less controlled contexts such as phone video or casual settings. Our evaluation set is also balanced 50/50 between familiar and stranger dyads by construction, which makes accuracy a clean measure of discrimination but base-rate-specific. A reader might ask whether the modelsâ stranger-lean is not miscalibration but a well-calibrated prior for a world in which two people recorded together are more often strangers, penalized only by our artificial balance. This does not threaten our central claim. That claim is the difference in criterion between humans and models on identical, matched stimuli. Neither rater type was told the base rate, so this difference does not depend on the true base rate, and neither does dâ˛d , which is base-rate-invariant. It does mean we measure discrimination, not deployment calibration; we do not read the balanced accuracy as a deployment estimate. The observed bias is in any case idiosyncratic in direction and size across models (Table 2), which is hard to square with calibration to any single real-world base rate. Measuring calibration against realistic base rates would usefully complement the discrimination-focused evaluation we report here. Each model is also evaluated under a single prompt, adapted from the human instructions (§3.3). Because model responses can be sensitive to prompt wording, the exact biases in Table 2 may shift under other phrasings; a prompt-robustness sweep is worthwhile future work. The humanâmodel comparison holds the task framing fixed for both rater types, so it is less exposed to this concern than any single modelâs bias estimate read on its own. Ethics Statement Informed consent and compensation. Rater recruitment and prescreening are described in §3.2 and Appendix D. Before viewing any stimulus, each rater read a description of the study and affirmatively agreed to three statements: that they had read and understood the information, that they were free to withdraw at any timeâby closing the browser tabâwithout giving a reason, and that they agreed to take part. Participation was voluntary and could be ended at any point without penalty. Sessions took roughly 15â30 minutes depending on modality, and raters were compensated at an effective rate of approximately $10â12 per hour, above Prolificâs fair-pay guidance. The task was minimal-risk: raters viewed short, benign clips of consenting conversation partners and answered non-sensitive perceptual and self-report questions. Privacy and data minimization. We collected no personally identifying information from raters. Beyond the task responses and the self-report questionnaires analyzed here, we recorded only the coarse attributes used for prescreening and quota matching (§3.2), keyed to a pseudonymous Prolific identifier that we do not link to any real-world identity. All released rater data are de-identified. Institutional review. This research was conducted outside a university setting and was not reviewed by an institutional review board. Because it collected no personally identifying information from raters and involved only minimal-risk procedures, it did not fall under IRB oversight; we nonetheless followed standard human-subjects safeguardsâvoluntary informed consent, the right to withdraw without penalty, fair compensation, minimal-risk stimuli, and data de-identification. Stimulus source. All interaction clips are drawn from the publicly released Seamless Interaction dataset (Agrawal et al., 2025), whose participants consented to the recording and research use of their audio and video. We use these data in accordance with the datasetâs C-BY-NC 4.0 license (attribution, non-commercial use only): the rendered benchmark clips are redistributed under the same C-BY-NC 4.0 terms, with attribution to the Seamless Interaction dataset. We do not attempt to re-identify or contact any recorded individual. Intended use and risks. The benchmark is intended for evaluating and auditing the social-perceptual behavior of models, including the response biases we document (§5.1). Inferring familiarity from behavior could in principle support surveillance or profiling; we release the benchmark to enable research on and scrutiny of such capabilities, not their deployment. Given the modest accuracy and pronounced, idiosyncratic biases we observe (§4), we caution against using these models or this task to make consequential judgments about real individuals. References R. Adolphs (2003) Cognitive neuroscience of human social behaviour. Nature Reviews Neuroscience 4 (3), p. 165â178. External Links: ISSN 1471-0048, Document Cited by: §1. V. Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Cheng, P. Chowdary, J. Chuang, A. DâAvirro, J. Daly, N. Dong, M. Duppenthaler, C. Gao, J. Girard, M. Gleize, S. Gomez, H. Gong, S. Govindarajan, B. Han, S. He, D. Hernandez, Y. Hristov, R. Huang, H. Inaguma, S. Jain, R. Janardhan, Q. Jia, C. Klaiber, D. Kovachev, M. Kumar, H. Li, Y. Li, P. Litvin, W. Liu, G. Ma, J. Ma, M. Ma, X. Ma, L. Mantovani, S. Miglani, S. Mohan, L. Morency, E. Ng, K. Ng, T. A. Nguyen, A. Oberai, B. Peloquin, J. Pino, J. Popovic, O. Poursaeed, F. Prada, A. Rakotoarison, R. Ranjan, A. Richard, C. Ropers, S. Saleem, V. Sharma, A. Shcherbyna, J. Shen, J. Shen, A. Stathopoulos, A. Sun, P. Tomasello, T. Tran, A. Turkatenko, B. Wan, C. Wang, J. Wang, M. Williamson, C. Wood, T. Xiang, Y. Yang, J. Yao, C. Zhang, J. Zhang, X. Zhang, J. Zheng, P. Zhyzheria, J. Zikes, and M. Zollhoefer (2025) Seamless interaction: Dyadic audiovisual motion modeling and large-scale dataset. arXiv:2506.22554 [cs.CV]. External Links: Document, https://arxiv.org/abs/2506.22554 Cited by: 1st item, §2, §3.1, Stimulus source.. N. Ambady and R. Rosenthal (1992) Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin 111 (2), p. 256â274. External Links: ISSN 1939-1455, Document Cited by: §1, §2. G. A. Bryant, C. S. Wang, and R. Fusaroli (2020) Recognizing affiliation in colaughter and cospeech. Royal Society Open Science 7 (10), p. 201092. External Links: ISSN 2054-5703, Document Cited by: §2. A. Cafaro, J. Wagner, T. Baur, S. Dermouche, M. Torres Torres, C. Pelachaud, E. AndrĂŠ, and M. Valstar (2017) The NoXi database: multimodal recordings of mediated novice-expert interactions. In Proceedings of the 19th ACM International Conference on Multimodal Interaction, ICMI â17, New York, NY, USA, p. 350â359. External Links: Document, ISBN 978-1-4503-5543-8 Cited by: §2. T. Capretto, C. Piho, R. Kumar, J. Westfall, T. Yarkoni, and O. A. Martin (2022) Bambi: A Simple Interface for Fitting Bayesian Linear Models in Python. Journal of Statistical Software 103, p. 1â29. External Links: ISSN 1548-7660, Document Cited by: §5.3. P. de Boeck and M. Wilson (Eds.) (2004) Explanatory Item Response Models: A Generalized Linear and Nonlinear Approach. Springer Science & Business Media. External Links: ISBN 978-0-387-40275-8 Cited by: §5.5. T. G. Dietterich (1998) Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation 10 (7), p. 1895â1923. External Links: ISSN 0899-7667, Document Cited by: §4. R. I. M. Dunbar, J. Robledo, I. Tamarit, I. Cross, and E. Smith (2022) Nonverbal Auditory Cues Allow Relationship Quality to be Inferred During Conversations. Journal of Nonverbal Behavior 46 (1), p. 1â18. External Links: ISSN 1573-3653, Document Cited by: §2. C. D. Frith and U. Frith (2007) Social Cognition in Humans. Current Biology 17 (16), p. R724âR732. External Links: ISSN 0960-9822, Document Cited by: §1. A. Gelman, J. B. Carlin, H. S. Stern, D. B. Dunson, A. Vehtari, and D. B. Rubin (2014) Bayesian data analysis. 3rd edition, CRC Press, Boca Raton, FL. Cited by: §5.3. J. W. Graham, B. J. Taylor, A. E. Olchowski, and P. E. Cumsille (2006) Planned missing data designs in psychological research. Psychological Methods 11 (4), p. 323â343. External Links: ISSN 1082-989X, Document Cited by: Appendix D, §3.2. R. Grieve and D. Mahar (2013) Can social intelligence be measured? Psychometric properties of the Tromsø Social Intelligence Scale â English Version. The Irish Journal of Psychology 34 (1), p. 1â12. External Links: ISSN 0303-3910, Document Cited by: §3.2. K. L. Gwet (2021) Handbook of inter-rater reliability: Chance-corrected agreement coefficients. 5th edition, Vol. 1, AgreeStat Analytics. Cited by: Appendix D. M. J. Hautus (1995) Corrections for extreme proportions and their biasing effects on estimated values of dⲠ. Behavior Research Methods, Instruments, & Computers 27 (1), p. 46â51. External Links: ISSN 1532-5970, Document Cited by: Table 2. Q. Jia, H. Huang, and K. Q. Zhu (2021) DDRel: A New Dataset for Interpersonal Relation Classification in Dyadic Dialogues. Proceedings of the AAAI Conference on Artificial Intelligence 35 (14), p. 13125â13133. External Links: ISSN 2374-3468, Document Cited by: §2. D. Katerenchuk, D. G. Brizan, and A. Rosenberg (2014) âWas that your mother on the phone?â: classifying interpersonal relationships between dialog participants with lexical and acoustic properties. In Proc. Interspeech 2014, p. 1831â1835. External Links: Document Cited by: §2. E. Kim, J. Park, J. Oh, K. Park, S. Song, A. S. DoÄruĂśz, A. Oh, and N. Kim (2026) Are they lovers or friends? Evaluating LLMsâ Social Reasoning in English and Korean Dialogues. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 23431â23451. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §2. F. Kong, W. Zu, X. Chen, Y. Yang, S. Zhu, and X. Feng (2026) SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 37379â37403. External Links: Document, ISBN 979-8-89176-395-1 Cited by: §2. N. Latif, A. V. Barbosa, E. Vatikiotis-Bateson, M. S. Castelhano, and K. G. Munhall (2014) Movement Coordination during Conversation. PLOS ONE 9 (8), p. e105036. External Links: ISSN 1932-6203, Document Cited by: §2. J. Li, Y. Wong, Q. Zhao, and M. S. Kankanhalli (2017) Dual-Glance Model for Deciphering Social Relationships. In Proceedings of the IEEE International Conference on Computer Vision, p. 2650â2659. Cited by: §2. D. Makowski, M. S. Ben-Shachar, S. H. A. Chen, and D. LĂźdecke (2019) Indices of effect existence and significance in the Bayesian framework. Frontiers in Psychology 10. External Links: ISSN 1664-1078, Document Cited by: §5.3. L. Mathur, P. P. Liang, and L. Morency (2024) Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 20541â20560. External Links: Document Cited by: §1. Q. McNemar (1947) Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages. Psychometrika 12 (2), p. 153â157. External Links: ISSN 0033-3123, 1860-0980, Document Cited by: §4. C. Palmero, J. Selva, S. Smeureanu, J. C. S. J. Junior, A. Clapes, A. Mosegui, Z. Zhang, D. Gallardo, G. Guilera, D. Leiva, and S. Escalera (2021) Context-Aware Personality Inference in Dyadic Scenarios: Introducing the UDIVA Dataset. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 1â12. Cited by: §2. D. Premack and G. Woodruff (1978) Does the chimpanzee have a theory of mind?. Behavioral and Brain Sciences 1 (4), p. 515â526. External Links: ISSN 1469-1825, 0140-525X, Document Cited by: §1. Z. Qin, R. Zheng, Y. Wang, T. Li, Y. Yuan, J. Chen, and L. Wang (2026) HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs. Proceedings of the AAAI Conference on Artificial Intelligence 40 (30), p. 24973â24981. External Links: ISSN 2374-3468, Document Cited by: §2. G. Rasch (1960) Probabilistic Models for Some Intelligence and Attainment Tests. MESA Press, 5835 S. External Links: ISBN 978-0-941938-05-1 Cited by: §5.5. M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi (2019) Social IQa: Commonsense Reasoning about Social Interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), K. Inui, J. Jiang, V. Ng, and X. Wan (Eds.), Hong Kong, China, p. 4463â4473. External Links: Document Cited by: §1, §1. D. Silvera, M. Martinussen, and T. I. Dahl (2001) The Tromsø Social Intelligence Scale, a self-report measure of social intelligence. Scandinavian Journal of Psychology 42 (4), p. 313â319. External Links: ISSN 1467-9450, Document Cited by: §3.2. H. Stanislaw and N. Todorov (1999) Calculation of signal detection theory measures. Behavior Research Methods, Instruments, & Computers 31 (1), p. 137â149. External Links: ISSN 1532-5970, Document Cited by: §4. Q. Sun, B. Schiele, and M. Fritz (2017) A Domain Based Approach to Social Relation Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 3481â3490. Cited by: §2. W. Tang, F. I. Dogan, L. Qing, and H. Gunes (2026) AsyReC: A Multimodal Graph-Based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification. IEEE Transactions on Circuits and Systems for Video Technology 36 (3), p. 3693â3708. External Links: ISSN 1558-2205, Document Cited by: §2. E. M. Templeton, L. J. Chang, E. A. Reynolds, M. D. Cone LeBeaumont, and T. Wheatley (2023) Long gaps between turns are awkward for strangers but not for friends. Philosophical Transactions of the Royal Society B: Biological Sciences 378 (1875), p. 20210471. External Links: ISSN 0962-8436, Document Cited by: §2. L. Tickle-Degnen and R. Rosenthal (1990) The nature of rapport and its nonverbal correlates. Psychological Inquiry 1 (4), p. 285â293. External Links: Document Cited by: §1. A. Zadeh, M. Chan, P. P. Liang, E. Tong, and L. Morency (2019) Social-IQ: A Question Answering Benchmark for Artificial Social Intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8807â8817. Cited by: §2. S. Zhang, Y. Yin, W. Song, Y. Wu, and M. Liu (2026) PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models. arXiv. External Links: 2606.23092, Document Cited by: §2. Appendix A Benchmark Quality Filters Every candidate interaction and clip must pass the automatic quality filters in Table A1 before entering the stratified selection pool (§3.1). Interaction-level filters (F1, F4, F5) gate the whole source interaction; clip-level filters (F2, F3) gate the specific 20-second window sampled from it. Clips are 20 s long with boundaries snapped within Âą5Âą 5 s to the nearest turn start. These filters remove clips that cannot support the task regardless of relationship; they are followed by the LLM-based prompt-adherence audit (§3.1), which verifies that each interactionâs recorded prompt label matches its actual content. ID Level Criterion F1 Interaction Both recorded EO prompts are free of placeholder/boilerplate text (e.g., âyou will each see the same prompt,â âgiven different lists of sentencesâ) that signals the true question was not logged. F2 Clip âĽ4⼠4 merged conversational turns within the clip window. F3 Clip Each speaker contributes âĽ3⼠3 s of speech within the clip window. F4 Interaction <35%<35\% of cross-track turn-start pairs coincide within 0.20.2 s, rejecting duplicated or desynchronized audio tracks. F5 Interaction âĽ90⼠90 s of speech in the interaction. Table A1: Automatic quality filters applied during clip generation and the EO interaction audit. Interaction-level filters gate the source interaction; clip-level filters gate the sampled 20 s window. Thresholds are the generator/auditor constants (scripts/generate_samples_rapport.py, scripts/audit_eo_interactions.py). Appendix B Model Inference Configuration Table A2 lists the inference configuration for each of the 26 models. All models received the same prompt construction and response parsing (§3.3); the settings below cover per-model decoding and, where applicable, reasoning parameters. Cloud models were queried at temperature 0 with a fixed seed (4242) wherever the API exposed them; the three locally-run open-weight models (Qwen2-Audio, Qwen2.5-Omni, Voxtral) used greedy decoding (do_sample=False), which is deterministic without a seed. Answer length was capped at 150â300 tokens for non-reasoning models, while reasoning models received a larger combined answer-plus-reasoning budget (8,000 tokens for the Gemini thinking and pro variants; 4,000 for Inkling and Muse Spark) to avoid truncating the response. Model API / model ID Temp. Seed Reasoning Text gpt_text gpt-4o 0 4242 â gpt_text_mini gpt-4o-mini 0 4242 â gemini_text gemini-3.5-flash 0 4242 off gemini_text_thinking gemini-3.5-flash 0 4242 dynamic claude_text claude-opus-4-8 âa âa offa mistral_text mistral-large-latest 0 4242 â inkling_text thinkingmachines/Inkling 0 42b42^b on muse_spark_text muse-spark-1.1 âc âc minimal muse_spark_text_high muse-spark-1.1 âc âc high Audio gpt_audio gpt-audio 0 4242 â gpt_audio_1_5 gpt-audio-1.5 0 4242 â gpt_audio_mini gpt-audio-mini 0 4242 â gemini_audio gemini-3.5-flash 0 4242 off gemini_audio_pro gemini-pro-latest 0 4242 dynamicd gemini_audio_thinking gemini-3.5-flash 0 4242 dynamic qwen_audio Qwen/Qwen2-Audio-7B-Instruct greedy â â qwen_omni Qwen/Qwen2.5-Omni-7B greedy â â voxtral_audio mistralai/Voxtral-Mini-3B-2507 greedy â â inkling_audio thinkingmachines/Inkling 0 42b42^b on muse_spark_audio muse-spark-1.1 âc âc minimal muse_spark_audio_high muse-spark-1.1 âc âc high Video gemini_video gemini-3.5-flash 0 4242 off gemini_video_pro gemini-pro-latest 0 4242 dynamicd gemini_video_thinking gemini-3.5-flash 0 4242 dynamic muse_spark_video muse-spark-1.1 âc âc minimal muse_spark_video_high muse-spark-1.1 âc âc high Table A2: Per-model inference configuration (26 models, seven companies; roster matches Table 2). Temp. and Seed are the decoding temperature and random seed passed to the API; greedy marks the locally-run open-weight models decoded with do_sample=False (deterministic, no seed). Reasoning: âââ = non-reasoning model; âoffâ = reasoning-capable but disabled (Gemini thinking_budget=0=0); âdynamicâ = model sets its own reasoning depth (thinking_budget=â1=-1); âonâ = reasoning always on with no depth control; âminimalâ/âhighâ = named reasoning-effort tier. Two Gemini IDs are aliases; at run time (July 2026) gemini-pro-latest resolved to gemini-3.1-pro-preview, and gemini-3.5-flash was pinned explicitly rather than gemini-flash-latest (which then resolved to gemini-3.6-flash). aClaude rejects temperature/top_p and exposes no seed; extended thinking was not enabled. bInkling accepts a seed but the API echoed null, so determinism is best-effort. cMuse Spark exposes no temperature or seed control, and its reasoning depth is non-deterministic run-to-run. dgemini-pro-latest cannot disable thinking and always reasons dynamically. Appendix C Task Instructions and Prompts Both human raters and models were given the same task framing, âEither Orâ game description, and familiar/strangers definition, differing only in how a response was collected. Human raters selected an answer with a button; models were additionally instructed to emit a JSON object with a label and confidence. Modality-specific slots (shown here in brackets, Text/Audio/Video order) were filled per panel or per clip. Human rater instructions. You will [read / listen to / watch] a series of [transcripts / audio clips / video clips] taken from real conversations between two people. These people are playing a game called "Either Or," where they discuss whether they would prefer one thing -- for example, the ability to fly -- or an alternative, such as the ability to breathe underwater. Your job is to pay close attention to each conversation and determine whether the two people have met before -- meaning they are familiar, such as friends, family members, coworkers, or romantic partners -- or are meeting for the first time in this interaction -- meaning they are strangers. Raters then made a forced choice: familiar or strangers. Model prompt. [Read / Listen to / Watch] the following [transcript / audio clip / video clip], taken from a real conversation between two people. These people are playing a game called "Either Or," where they discuss whether they would prefer one thing -- for example, the ability to fly -- or an alternative, such as the ability to breathe underwater. Your job is to pay close attention to the conversation and determine whether the two people have met before -- meaning they are familiar, such as friends, family members, coworkers, or romantic partners -- or are meeting for the first time in this interaction -- meaning they are strangers. For each clip, return: "relationship_label": str FAMILIAR|STRANGER, "confidence": float [0, 1], "reason": str Use the following confidence scale: 0 = Not at all confident, ..., 0.5 = Moderately confident, ..., 1 = Extremely confident Do not provide extra text outside the JSON object. Appendix D Human Rating Study Details Raters were recruited on Prolific with platform-level quality screening in addition to the demographic criteria in §3.2: a Prolific approval rate of 99â100% and at least 10 prior submissions. Under the planned-missing, 6-block design (Graham et al., 2006), each rater was assigned to a single block and each of the 96 dyads appeared in two of the six blocks, so every dyad was rated by an overlapping subset of raters (25â35 per dyad). By design each rater judged 32 dyads; in the audio panel a platform issue left 13 of the 94 raters with 29â31 completed trials, while all text and video raters completed the full 32. Inter-rater agreement. Agreement among human raters is slight, as expected for a difficult perceptual task with a balanced label set. Fleissâs Îş (computed per stimulus over whichever raters saw it, accommodating the incomplete-block design) is 0.080.08 for text and 0.170.17 for both audio and video. Bennettâs S is essentially identical (0.090.09, 0.170.17, 0.170.17 for text, audio, video), as expected when the two classes are balanced (Gwet, 2021). The near-zero text agreement is consistent with individual text raters performing at chance (Table 1): raters are not converging on a shared, reliable text cue. Appendix E Rater Individual Differences Table A3 reports, per modality, the reliability of the TSIS-PS social-intelligence scale and its correlation with per-rater accuracy. The scale is highly reliable in every panel, but its correlation with accuracy is null in audio and video and positive only in textâthe modality where accuracy is at chanceâso we do not interpret it as evidence of a general skill advantage. Modality Îą Pearson r p Spearman Ď Text 0.91 0.290.29 .004 0.230.23 Audio 0.88 0.030.03 .78 0.000.00 Video 0.92 â0.12-0.12 .24 â0.19-0.19 Table A3: TSIS-PS reliability and its correlation with per-rater accuracy (N=94N=94 text, 94 audio, 92 video). For self-reported cue use, raters most often reported attending to interactional cues (rapport, responsiveness) across all modalities, but individual cue-attendance scores rarely predicted accuracy (only a handful of per-cue correlations reached p<.05p<.05, with no stable cross-modality pattern). Full cue-use heatmaps are provided with the released analysis notebooks. The per-rater accuracy clouds in Figure 1 show the full distribution of per-rater accuracy in each modality (94 text, 94 audio, 92 video raters; each raterâs accuracy over the 32 dyads in their assigned block, 29â32 for a few audio raters affected by a platform issue). The clouds are wide, but much of that width is sampling noise: with only âź 32 trials per rater, binomial variation around a common ability already reproduces most of the observed spreadâincluding the upper tail that appears to exceed the best model (§5.4)âso the distribution should not be read as a stable ordering of raters by skill. This is also why per-rater accuracy is best modeled with partially-pooled crossed random effects (§3.2), which shrink noisy individual estimates, rather than trusting raw per-rater rates or treating each rater as a one-shot observation. Text raters cluster tightly around chance, consistent with the near-zero text agreement reported in Appendix D. Appendix F Dyad Difficulty Table A5 lists the eight hardest and eight easiest dyads by mean model-estimated probability of a correct human response, from the crossed random-effects model of §5.5. Several of the hardest dyads fall below chance in every modality, functioning as systematic lures. Estimated easiness correlates across modality pairs (textâaudio r=0.52r=0.52, textâvideo r=0.33r=0.33, audioâvideo r=0.61r=0.61; all p<.01p<.01), indicating that difficulty is largely a property of the conversation rather than the observation channel. Table A4 decomposes the humanâmodel difficulty correlation reported in §5.5. The raw pooled correlation (human-crowd accuracy vs. accuracy pooled over all models) is strong in audio, moderate in video, and only marginal in text. The weak raw text value is a suppression effect rather than absence of shared structure: humans and models carry opposing class biases (§5.1), so their per-dyad accuracies partly anti-align on the familiar/stranger axis. Partialling out the true class raises the within-class correlationâfine-grained difficulty that is not reducible to a two-way class effectâto r=0.61r=0.61 (text), 0.580.58 (audio), and 0.440.44 (video), all p<.001p<.001. The shared difficulty is, however, partly a crowd-level property: repeating the analysis with the single strongest model per modality in place of the pooled estimate leaves it robust only in video (r=0.45r=0.45, p<.001p<.001), with weak raw associations in audio (r=0.19r=0.19, p=.06p=.06) and text (r=â0.04r=-0.04, n.s.); confidence-weighted single-model estimates match the binary ones almost exactly, so this is not an artifact of one prediction per dyad. Individual models thus express the human-aligned difficulty signal only noisily, and pooling recovers it. Estimator Text Audio Video Pooled, raw 0.200.20 0.560.56 0.340.34 Pooled, within-class 0.610.61 0.580.58 0.440.44 Best model, raw â0.04-0.04 0.190.19 0.450.45 Best model, within-class 0.240.24 0.260.26 0.470.47 Table A4: Humanâmodel per-dyad difficulty correlation (r) under four estimators, by modality (N=96N=96 dyads). âPooledâ averages model accuracy over all models; âbest modelâ is the single most accurate model per modality (text gpt_text_mini, audio gemini_audio_pro, video gemini_video). âWithin-classâ partials out the true familiar/stranger label. Figure A1: Humans and models tend to find the same dyads hard (§5.5). Each point is one of the 96 EO dyads, plotting human per-dyad accuracy (over all raters) against model per-dyad accuracy (pooled over all models), colored by true class; the diagonal marks equal difficulty. Per-dyad difficulty correlates positively in audio (r=0.56r=0.56, â31%â 31\% shared variance) and video (r=0.34r=0.34, â12%â 12\%; both p<.01p<.01); text (r=0.20r=0.20) is only marginally significant (p=.045p=.045). Dyad Truth Txt. Aud. Vid. Mean Hardest sample_080 str. 29.6 32.0 22.2 27.9 sample_055 str. 20.5 32.1 33.3 28.6 sample_072 fam. 35.4 19.9 38.0 31.1 sample_053 str. 31.1 27.0 40.7 32.9 sample_038 fam. 25.8 28.5 47.9 34.1 sample_057 str. 34.1 30.3 37.9 34.1 sample_083 str. 31.7 34.7 36.1 34.2 sample_065 fam. 34.9 33.4 35.0 34.4 Easiest sample_045 fam. 59.6 83.8 79.3 74.2 sample_062 str. 69.2 75.2 79.4 74.6 sample_051 str. 53.3 88.4 82.0 74.6 sample_077 fam. 57.4 79.1 89.9 75.5 sample_009 fam. 70.7 77.5 87.4 78.5 sample_030 str. 63.3 83.7 91.1 79.4 sample_073 fam. 71.4 86.1 81.9 79.8 sample_016 str. 68.5 86.8 86.8 80.7 Table A5: Hardest and easiest eight dyads by mean estimated P(correct human response) across modalities (%). âstr.â = stranger, âfam.â = familiar.