Paper deep dive
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/22/2026, 1:59:33 AM
Summary
This paper investigates the 'delegation asymmetry' in agentic recommender systems, specifically within online dating platforms. Using large-scale surveys (N=2,894 and N=2,617), the authors develop a latent-variable measurement model to distinguish between 'send receptivity' (willingness to delegate conversation to an agent) and 'receive receptivity' (willingness to engage with agent-mediated communication). The study finds that while these constructs are highly correlated (rho=0.92), they are distinct, with users significantly more willing to deploy their own agents than to engage with others' agents. This asymmetry limits the viability of fully autonomous agent-mediated matching, with only 4-13% of random dyads combining deployment and engagement. The paper proposes a 'receptivity audit' methodology and demonstrates that routing contacts based on receive receptivity can triple per-contact engagement.
Entities (8)
Relation Signals (5)
Autonomous LLM Agents → usedin → Online Dating
confidence 96% · online dating is its natural frontier... autonomous LLM agents that hold conversations on the user’s behalf
Send Receptivity → correlatedwith → Receive Receptivity
confidence 95% · highly correlated (rho=0.92) but separable
Delegation Asymmetry → causedby → Difference in Receptivity Thresholds
confidence 93% · deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32)
Receive Receptivity → improves → Per-Contact Engagement
confidence 90% · routing agent contacts on receive receptivity triples per-contact engagement
Reciprocity Requirement → reduces → Interaction Volume
confidence 88% · a reciprocity requirement cuts interaction volume by half or more
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.
Tags
Links
- Source: https://arxiv.org/abs/2608.18058v1
- Canonical: https://arxiv.org/abs/2608.18058v1
Trouble viewing inline? Open PDF directly →
Full Text
69,703 characters extracted from source content.
Expand or collapse full text
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online DatingConference: The 20th ACM International Conference on Web Search and Data Mining; February 15–19, 2027; Hong KongWSDM ’27, February 15–19, 2027, Hong KongCCS: Information systems Recommender systemsCCS: Human-centered computing Empirical studies in collaborative and social computing Daria Leshchikova Affiliation: Fleamily, Inc., Delaware, USA , Valentina V. Kuskova Affiliation: Lucy Family Institute for Data & Society University of Notre Dame, Notre Dame, IN, USA , Dmitry Zaytsev Affiliation: Lucy Family Institute for Data & Society University of Notre Dame, Notre Dame, IN, USA and Valerii Klimov Affiliation: Fleamily, Inc., Delaware, USA 2027; © acmlicensed Abstract. Autonomous LLM agents that converse on a user’s behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N = 2,894 on generative profile features; N = 2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs. They are highly correlated (ρ=0.92ρ=0.92) but separable (ΔBIC=52 =52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one’s own agent requires far lower receptivity (threshold θ=−0.38θ=-0.38) than engaging a counterpart’s agent (θ=+0.32θ=+0.32; full engagement +1.39+1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4–13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement halves interaction volume by excluding two-thirds of would-be deployers, while routing agent contacts on receive receptivity triples per-contact engagement: a lift that survives out-of-sample validation in which the routing score is estimated without the target item (AUC 0.88, 3.1× quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking. Keywords: LLM agents, agentic recommender systems, online dating, reciprocal recommendation, item response theory, user acceptance 1. Introduction Somewhere tonight, a dating-app user will get a warm, well-crafted opening message from a match: composed, sent, and if things go well, followed up on by a language model while its principal sleeps. The scenario is no longer speculative. Recent survey evidence suggests that more than a quarter of U.S. singles have already used AI somewhere in their dating lives, and a majority of dating-app users believe they have, at some point, exchanged messages with an algorithm rather than a person (Ante 2026). A cottage industry of “AI wingman” tools writes openers and replies on demand; users who receive such messages report feeling deceived when they find out (Ante 2026), and merely suspecting AI authorship is enough to depress trust (Jakesch et al. 2019; Kadoma et al. 2025). What began as Cyrano de Bergerac with a keyboard - one party quietly borrowing better words - is being industrialized. Platforms are now productizing the next step. The first wave of generative features was assistance: AI-written profile summaries, suggested openers, tone polish, with the user remaining the author of record. The emerging wave is delegation: autonomous LLM agents that hold conversations on the user’s behalf are initiating contact, exchanging messages, screening counterparts, even negotiating common ground with the other side’s agent before the humans ever speak. The research community is actively building this machinery (Peng et al. 2025; Zhang et al. 2024; Wang et al. 2024; Ye et al. 2025), and online dating is its natural frontier: communication is not a feature of the product but the product itself, interaction volume is high, and the stakes are as personal as software gets. There is, however, a condition for this market that the building has run ahead of. Agent-mediated matching is a two-sided proposition: every delegated message needs a willing receiver on the other end. A market in which many users deploy agents but few will engage with one is not a matching market at all; it is automated outreach landing on closed doors, at scale, in the most trust-sensitive consumer domain there is. Yet the two sides of this condition have never been measured together. A large literature measures acceptance of AI tools by their users (Davis 1989; Venkatesh et al. 2003), and a growing one documents receiver-side penalties when AI involvement is suspected (Jakesch et al. 2019; Hancock et al. 2020); what is missing is a joint, within-person measurement of willingness to send and willingness to receive agent communication on a common scale - the quantity that determines whether both sides of a delegation market exist, and the quantity this paper provides. We measure both sides. Working with two large surveys fielded to active users of a major dating platform (N = 2,894 on generative profile features; N = 2,617 on autonomous conversational agents, in two languages), we build a latent-variable measurement model of agent receptivity: a graded response model with latent regression in which seven attitudinal items load on a send dimension (willingness to delegate one’s own communication) and a receive dimension (willingness to engage with agent-mediated communication from others). Four findings emerge. First, model comparison shows send and receive receptivity are distinct constructs: correlated at ρ=0.92ρ=0.92 but decisively separable (ΔBIC=51.8 =51.8), with partial measurement invariance across the two languages. Second, the model quantifies a delegation asymmetry: on the common scale, deploying one’s own agent requires receptivity θ=−0.38θ=-0.38 while engaging a counterpart’s agent requires θ=+0.32θ=+0.32, a displacement of 0.71 SD that survives every recoding we test. With that, population deployment propensity exceeds engagement propensity roughly threefold. Third, receptivity is associated with unmet need: it concentrates among users whose matches fail to convert to conversations and who report frustration about stalled matches. It is lower among women and long-tenured users, and the gender gap is a trait difference, not an artifact of differential item functioning. Fourth, under a random-pairing stated-preference counterfactual, only 4–13% of directed dyads combine deployment with engagement, with a pronounced gender-directional imbalance; counterfactuals quantify the design levers available to a platform. This paper makes four contributions. The first is a construct and a measurement design. We introduce send receptivity and receive receptivity as distinct constructs and measure them jointly, within person, on a common latent scale. Prior work compares AI users and AI receivers across separate studies and samples, confounding population differences with role differences. The within-person design makes the send–receive asymmetry observable without confounding it with between-sample population differences. The second is receptivity audit, a general methodology. We package the pipeline, from short ordinal instrument → latent measurement model → model-implied per-user deploy/engage endorsement propensities → out-of-sample validation → two-sided market counterfactuals as a reusable, pre-deployment procedure (Section 3). It tells a platform whether both sides of an agent-mediated feature’s market exist before the agents are built, and hands the matching layer a pre-deployment routing signal, validated out of sample. Nothing in the procedure is specific to dating. Its methodological content is the bridge: item-level psychometric parameters used directly as market primitives for agentic-system design. The third is empirical findings at platform scale. On a major dating platform we establish that send and receive receptivity are distinct (ρ=0.92ρ=0.92, ΔBIC=51.8 =51.8), quantify the delegation asymmetry (0.71 SD, CI [0.65,0.77][0.65,0.77], robust to recodings and partial measurement invariance), show receptivity concentrates among users reporting unmet matching need and is gendered, and identify a stable asymmetric-delegator segment (≈ 25%). The asymmetry is visible without any model: 40.7% of respondents individually rate sending above receiving; 2.1% the reverse. The fourth is market consequences and quantified design levers. Under a random-pairing counterfactual on the fitted stated-preference model, we price the design space: baseline viability of 4–13% of directed dyads, a gender-directional imbalance, the cost of a reciprocity norm (−-65% of deployers), and a receptivity-aware routing frontier whose 3× per-contact engagement gain survives leave-target-item-out cross-validation, with receive-side consent emerging as a first-class design primitive. 2. Related Work Online dating as a matching market. Online dating is now the dominant channel through which couples meet in several countries (Rosenfeld et al. 2019), and its market structure is well documented: users sort on observable attributes with strong revealed preferences (Hitsch et al. 2010). As users direct most of their attention upward toward more desirable counterparts (Bruch and Newman 2018), most first messages go unanswered: reply behavior is scarce, skewed, and predictable from user and dyad features (Xia et al. 2014). The domain’s relationship psychology is surveyed by Finkel et al. (Finkel et al. 2012). This attention-scarcity structure is the backdrop for our results: automated delegation promises to multiply first-message volume precisely where reply willingness is already the binding resource, and our receive-side measurements quantify how much colder the reception gets when the message is agent-authored. Reciprocal recommendation. Algorithmically, dating recommenders differ from item recommenders in requiring mutual interest. The reciprocal-recommendation line runs from RECON’s harmonic combination of directional preferences (Pizzato et al. 2010; Pizzato et al. 2013) through behavioral models on large platforms (Xia et al. 2015), latent-factor formulations (Neve and Palomares 2019), and the survey of Palomares et al. (Palomares et al. 2021). Recent work treats matching platforms explicitly as two-sided markets with exposure trade-offs: fairness-aware reciprocal recommendation (Tomita and Yokoyama 2024), welfare-based balancing of match rates and fairness, and off-policy evaluation for matching interventions (Hayashi et al. 2025); agent-based models reproduce engagement cycles in swiping apps from simple behavioral rules (Cela et al. 2025). Users, for their part, hold rich folk theories about dating algorithms and counter-strategize against perceived platform interests (Alizadeh et al. 2024). This literature optimizes who is shown to whom, taking the communication medium as given; agent mediation makes the medium itself a market variable, adding a second reciprocity layer: both sides must accept not only each other but the mode of contact. Our send/receive distinction formalizes this second layer, and our routing counterfactual (§7) is a reciprocal-recommendation intervention on it. AI-mediated communication. Hancock et al. (Hancock et al. 2020) define AI-mediated communication and its research agenda. Receiver-side penalties are well documented: profiles suspected of AI authorship are trusted less (Jakesch et al. 2019), suspicion of LLM use falls unevenly across writers (Kadoma et al. 2025), and smart replies measurably alter language and relationships (Hohenstein et al. 2023). Closest to our setting, Ante (Ante 2026) interviews both LLM-assisted senders and receivers in online dating, finding an “authenticity paradox” among senders and “digital betrayal” among receivers. These qualitative asymmetries motivate our contribution: measuring send- and receive-willingness jointly, within person, on one latent scale, at population size. LLM agents and agentic recommenders. LLM agents that plan, remember, and act (Park et al. 2023) are moving into recommendation: as autonomous user simulacra for collaborative filtering (Zhang et al. 2024), as planning agents over recommender tools (Wang et al. 2024), as user-side shields that renegotiate the user–platform relationship (Xu et al. 2025), and as multi-agent matching prototypes for dating specifically (Ye et al. 2025); Peng et al. (Peng et al. 2025) survey the space. This literature builds the supply side of agent-mediated interaction. Our results measure the demand side and show the receiving half of that demand is the binding constraint. Measuring attitudes toward AI. Stated-preference measurement of technology adoption descends from TAM and UTAUT (Davis 1989; Venkatesh et al. 2003). For AI specifically, experimental work established that people both discount algorithmic judgment after seeing it err (Dietvorst et al. 2015) and, in other conditions, prefer it to human judgment (Logg et al. 2019). This is a polarity that our rejector/enthusiast class structure recovers at population scale. A family of dedicated self-report instruments has followed: the two-factor GAAIS (Schepman and Rodway 2020; Schepman and Rodway 2023), the brief AIAS-4 (Grassini 2023), and the unidimensional ATTARI-12 (Stein et al. 2024). These scales share three properties that limit them for our question: they measure attitudes toward AI in general rather than toward a concrete interpersonal deployment; they are sum-scored composites, so response-category thresholds are not modeled quantities; and they measure the respondent as a user of AI, with no construct for being on the receiving end of someone else’s AI. Our instrument and model address all three: concrete agent scenarios, threshold-level modeling, and joint within-person measurement of send and receive dispositions. Psychometric machinery. We use item response theory (Embretson and Reise 2000): the graded response model for ordinal items (Samejima 1969), estimated by marginal maximum likelihood (Bock and Aitkin 1981), extended with latent regression on covariates and formal invariance testing in the factorial-invariance tradition (Meredith 1993; Millsap 2011). The same model family aggregates ordinal expert ratings into latent indices in cross-national measurement programs (Pemstein et al. 2018), where, in the same way as here, the scientific claims live in threshold locations and cross-group comparability rather than in raw scores. The payoff for us is direct: the delegation asymmetry is a statement about where two thresholds sit on one latent scale, which requires thresholds to be first-class parameters rather than artifacts of scale construction, and the partial-invariance analysis (§6) is the standard machinery for licensing our cross-language pooling. 3. The Receptivity Audit Agent-mediated communication features pose a question that arrives before any system exists to log behavior from: do both sides of the market for this feature exist? Answering it post-launch is expensive: the failure mode is deployed agents generating unwanted contact at scale in a trust-sensitive domain. Answering it with off-the-shelf attitude scales is not possible, because existing instruments measure attitudes toward AI in general, score respondents as users only, and collapse responses into sums that cannot locate thresholds (§2). We therefore formalize the procedure this paper instantiates as a general, four-stage receptivity audit: (1) Instrument. A short battery of ordinal items presenting the concrete feature to its intended population, with items deliberately spanning both roles: scenarios in which the respondent deploys the capability, and scenarios in which the respondent is on the receiving end of someone else’s deployment. Within-person coverage of both roles is the non-negotiable design element; it is what separates role effects from population effects. (2) Measurement model. A confirmatory multidimensional graded response model with latent regression: items load on send and receive factors, latent traits regress on user covariates, and dimensionality itself is a model-comparison question rather than an assumption. The fitted model yields three things no sum score provides: threshold locations on a common scale (the asymmetry is a statement about thresholds), formal invariance tests licensing pooling across languages or segments, and per-user factor scores with principled uncertainty. (3) Propensities, validated out of sample. Model-implied endorsement propensities computed from the fitted item response functions at each user’s factor scores, such as the probability of endorsing deployment, and the probability of endorsing engagement when contacted, under transparent strict/soft operationalizations of intermediate response categories, treated as scenario bounds rather than calibrated behavior. Before the propensities feed any policy, the score behind them is validated predictively: the target item is held out, the score is re-estimated from the remaining items and covariates under respondent-level cross-validation, and out-of-sample discrimination and lift are reported. This is what licenses using the score as a routing signal. (4) Market counterfactuals. A dyadic simulation over the empirical population that converts propensities into market quantities such as interaction volume, engagement per contact, directional imbalances, and prices design levers - eligibility rules (who may deploy), routing rules (who receives), and disclosure or consent mechanics. Because the levers are expressed in the model’s own parameters, each counterfactual is a closed-form computation, not a new experiment. Two remarks on scope. First, the audit is deliberately parsimonious: every component we use is established machinery (Samejima 1969; Bock and Aitkin 1981; Meredith 1993); the contribution we claim is the construct and the bridge to market primitives, not a new estimator. Second, nothing above mentions dating. The audit applies wherever one side delegates communication to an agent and the other side must accept the medium: recruiting outreach, sales prospecting, agent-mediated customer contact, professional networking. Sections 4–7 instantiate the four stages on one such market; §8 returns to the general case. 4. Data and Instruments Two survey instruments were fielded to active users of a large dating platform, Fledge.Love as voluntary, self-administered online questionnaires. Respondents were recruited through an in-app prompt shown to active users. The agents instrument (B) was fielded first, from November 12 to December 15, 2025, as parallel Russian- and English-language forms collected concurrently. The generative-features instrument (A) followed from March 22 to April 1, 2026, in Russian. The two samples are not linked at the respondent level. Instrument A: generative features (N = 2,894). Respondents were shown three generative-AI feature concepts in sequence, each with a short description: (1) an AI profile summary: a generated synopsis of the user’s personality-test results displayed on their profile; (2) AI conversation tips: generated suggestions for messaging a specific match; and (3) an AI couple description: a generated compatibility narrative for a user–match pair. Each concept was rated on a three-level interest scale (not interesting / interesting / very interesting), with an open “other” field that a minority of respondents used for free-text commentary instead of a rating (n = 473 comments across items; these respondents are excluded from the corresponding item’s denominator and their comments reviewed for context but not analyzed further). Respondents then chose the best of the three concepts (with a “none” option). They further answered whether such features would increase their interest in browsing profiles (four levels from definitely yes to engagement would not increase), alongside gender, age band, platform-usage frequency, and everyday AI usage. The sample (N = 2,894) is 59.7% male; 25.3% aged 18–24, 40.1% aged 25–34, 25.5% aged 35–44, and 9.1% 45+. It skews toward engaged users: 80.1% report opening the app several times a week, and toward AI-familiar ones: 47.2% report using AI services regularly and a further 29.4% occasionally. A contact field completed by 1,037 respondents volunteering for follow-up interviews is excluded from all analysis and from the deposit. Instrument B: autonomous agents (N = 2,617; RU and EN). The second instrument presented a concrete agent feature through a sequence of annotated interface mockups: a conversational agent that messages the user’s matches on their behalf, with user-configurable activity level (initiates vs. replies only) and tone. Items were interleaved with the mockups so that each scenario was rated immediately after being shown. The seven attitudinal items, their roles, and their response options are given in Table 1; English wording is shown, with the Russian form a professional parallel translation. The design deliberately spans both roles of the delegation relation: Y1–Y3 and Y7 place the respondent as the agent’s principal (seeing the concept, configuring the agent, delegating one’s own conversations, valuing the product), while Y4–Y6 place the respondent as the counterpart (receiving messages from someone else’s agent, observing agent-to-agent pre-conversation about oneself, sharing a group chat with agents). This within-person role coverage is stage one of the audit (§3). The instrument also captured six covariates on ordinal scales: gender; age band (18–24 / 25–34 / 35–44 / 45+); platform tenure (first month / several months / over a year); perceived match volume (very few / not many / enough / too many); match-to-conversation conversion (very few / less than half / more than half / almost all); and negative affect about stalled matches (none / neutral / yes). Two survey questions on conversation-decay attributions and coping strategy are not used in the measurement model. Coding. Responses are coded so that higher categories indicate greater receptivity. Two items required ordering judgments: Y1 offers two negative options (not interesting, didn’t seem trustworthy) that we collapse into a single lowest category, and Y4 similarly offers two rejections (would react negatively, weird - only want a real person); Y7’s 1–10 rating is binned into five ordered levels. Section 6 reports sensitivity analyses in which the collapsed categories are instead treated as distinct ordered levels; the headline quantities are stable (asymmetry gap 0.70–0.72 across codings). The full codebook, including both language forms, ships with the released pipeline. Complete cases across items and covariates: N = 2,499 (95.5% of responses), of which 212 are from the EN form. The released dataset. The deposit contains one record per response, in two files mirroring the instruments. The Instrument B file (2,617 records: 2,385 RU, 232 EN) carries: a language flag; the seven ordinal items of Table 1 both as verbatim categorical responses and under the analysis coding; the six ordinal covariates (gender, age band, tenure, match volume, match-to-conversation conversion, negative affect about stalled matches); and two auxiliary categorical fields not used in the measurement model (attributed reason for conversation decay; coping strategy for stalled matches). The Instrument A file (2,894 records) carries: the three concept ratings on the three-level scale (with free-text “other” responses replaced by a nonresponse flag); the best-of-three choice; the browsing-interest item; and gender, age band, usage frequency, and AI-usage frequency. An auxiliary file provides, for each of the 2,499 complete cases, the model-derived quantities used in Sections 6–7: EAP factor scores θ^send,θ^recv θ^send, θ^recv and the four deploy/engage propensities under strict and soft operationalizations. Removed relative to the raw collection: contact information, all free-text (473 comments in Instrument A, reviewed for context only and not released), and exact submission timestamps, which are coarsened to the ISO week; demographic cross-classifications were audited so that no released cell isolates fewer than k=5k=5 respondents (binding only in the small EN subsample). A codebook documents both language forms verbatim, every coding map, and the provenance of each derived field. The released files are the analysis dataset of record: the k-anonymity recoding (which alters the age code of three EN respondents) was applied before the final model fits, so every quantity in this paper reproduces exactly from the deposit. Human-subjects review. Both surveys were designed, fielded, and collected by the platform for product research; the authors received the response data as an existing dataset and conducted secondary analysis only. The analysis protocol was reviewed by the authors’ institutional review board and determined not to constitute human subjects research, as secondary analysis of de-identified data (protocol 26-08-10287). Availability. Data, codebook, and the full analysis pipeline (harmonization, model fitting, robustness battery, simulation, and table generation) are archived at DOI 10.5281/zenodo.21971273 . Every number, table, and figure in this paper is reproducible from the deposit via the released scripts. Table 1. Instrument B: agent-receptivity items. English form shown; higher codes indicate greater receptivity. P = respondent as principal (send dimension), C = respondent as counterpart (receive dimension). Item Role Stem (abridged) Response options (code) Y1 P First emotional reaction to the agent concept very interesting (2); neutral (1); didn’t seem trustworthy (0); not interesting (0) Y2 P Attitude to configuring the agent’s activity and tone useful – saves me time (2); too complicated – can’t be bothered (1); risky – might write something inappropriate (0) Y3 P “How would you feel if the agent talked with a potential partner while you are away?” I want to try it (2); I might try it (1); I don’t want an agent responding for me (0) Y4 C “How would you feel if someone’s AI agent responded to you?” would engage in conversation with the agent (2); might engage (1); weird – only want a real person (0); would react negatively (0) Y5 C Your agent chats with another agent to find common interests before you join super idea (3); funny – would watch their conversation (2); doubtful – why are we needed then? (1); strange and frightening (0) Y6 C Humans and agents converse together in a group chat super idea (2); funny – would participate (1); doubtful / strange (0) Y7 P Overall value of the feature for staying active and not losing matches, 1–10 binned: 1–2 (0); 3–4 (1); 5–6 (2); 7–8 (3); 9–10 (4) Table 2. Sample descriptives, agent-survey complete cases (N = 2,499). Female 35.4% Age 18–24 21.1% Age 25–34 40.8% Age 35–44 29.3% Age 45+ 8.8% Tenure: <<1 month 24.6% Tenure: Several months 44.5% Tenure: >>1 year 30.9% EN-language instrument 8.5% 5. A Measurement Model of Agent Receptivity We model the seven ordinal items with a graded response model (Samejima 1969) with latent regression. Person i’s response to item j with KjK_j ordered categories follows (1) Pr(Yij≥k∣i)=σ(aj(θi,d(j)−bjk)), (Y_ij≥ k θ_i)=σ\! (a_j( _i,d(j)-b_jk) ), with discrimination aj>0a_j>0 and ordered thresholds bj1<⋯<bj,Kj−1b_j1<…<b_j,K_j-1. In the two-dimensional specification, d(j)d(j) assigns items to a send factor (Y1, Y2, Y3, Y7) or a receive factor (Y4, Y5, Y6); the latent traits follow (2) i=B⊤i+i,i∼(,[1ρ1]), θ_i=B x_i+ _i, _i \! (0, bmatrix1&ρ\\ ρ&1 bmatrix ), where ix_i collects centered covariates (gender, age, tenure, match volume, conversion, negative affect, language). Why measurement models outperform sum scores. We treat receptivity as a latent variable in the structural tradition (Bollen 1989; Bollen 2002; Borsboom et al. 2003): the items are fallible ordinal indicators of unobserved dispositions, and the quantities of scientific interest are parameters of the measurement model, not arithmetic on raw responses. Sum scoring is not a model-free alternative: it is itself a latent variable model with strict implicit constraints, notably equal item weights (McNeish and Wolf 2020), and those constraints are untenable here on three grounds. The items have heterogeneous category counts (three to five levels), so equal weighting is arbitrary by construction. The scientific claim at stake, that engaging a counterpart’s agent requires more receptivity than deploying one’s own, is a claim about where response-category thresholds sit on a common latent scale, and thresholds simply do not exist as quantities in a composite score. And pooling a two-language sample requires the formal invariance machinery of the factorial-invariance tradition (Meredith 1993; Millsap 2011), including the partial-invariance fallback we end up needing, which is native to latent variable models and unavailable to composites. The graded response model, a generalized latent variable model for ordinal indicators (Skrondal and Rabe-Hesketh 2004), delivers all three properties at the cost of distributional assumptions that we subject to model comparison rather than assume. Estimation and inference. We estimate by marginal maximum likelihood (Bock and Aitkin 1981) with 13213^2-node Gauss–Hermite quadrature over the correlated latent space, optimizing with gradient-based quasi-Newton methods under automatic differentiation. Identification fixes both trait variances to one and centers all covariates, so thresholds absorb location and the latent regression is a pure slope structure. Dimensionality (1D vs. 2D) and the latent correlation are selected by BIC and likelihood-ratio test; per-user factor scores are expected a posteriori (EAP) values under the fitted model; uncertainty for all reported quantities comes from a 200-replicate nonparametric bootstrap over respondents (replicates refit at reduced quadrature for tractability; percentile intervals are recentered on the full-precision estimates); and differential item functioning is tested per item by likelihood ratio with Bonferroni correction, with a partial-invariance refit verifying that substantive conclusions survive freeing flagged items. Code for the full pipeline is released for reuse in the deposit described in §4. 6. Results: The Structure of Receptivity 6.1. Send and receive are distinct constructs Table 3. Model comparison. The two-dimensional model is preferred. Model −logL- L Params BIC 1D 13,707.4 31 27,657.2 2D 13,650.1 39 27,605.4 The two-dimensional model dominates (ΔBIC=51.8 =51.8; LRT =114.5=114.5, 88 df, p<10−15p<10^-15), with latent correlation ρ=0.92ρ=0.92: willingness to send and willingness to receive agent communication are strongly related but not interchangeable. Table 4 reports the item parameters. All seven items discriminate strongly (a between 1.7 and 4.2); the single most informative item on the send dimension is willingness to deploy one’s own agent (Y3, a=4.2a=4.2), and the receive-side items discriminate comparably well. The weakest item is the overall 1–10 product rating (Y7, a=1.7a=1.7), whose thresholds sit far to the right of every other send-side item: even users well above average in receptivity stop short of rating the product concept highly. Curiosity about agents and perceived product value are related but partially decoupled, which is a caution against reading feature-level enthusiasm off a single product rating, in either direction. Table 4. Item parameters of the 2D graded response model: discrimination a and ordered thresholds bkb_k (the latent-trait location at which the probability of responding in category k or higher reaches 50%). Item content and response categories are given in Table 1; dashes mark items with fewer categories; S = send, R = receive dimension. Item Dim. a b1b_1 b2b_2 b3b_3 b4b_4 Y1 S 2.97 −0.23-0.23 0.360.36 — — Y2 S 3.44 −0.33-0.33 −0.12-0.12 — — Y3 S 4.20 −0.38-0.38 0.380.38 — — Y4 R 2.82 0.320.32 1.391.39 — — Y5 R 3.28 −1.64-1.64 −0.36-0.36 0.990.99 — Y6 R 3.36 −0.09-0.09 1.171.17 — — Y7 S 1.72 0.480.48 0.940.94 1.531.53 2.092.09 6.2. The delegation asymmetry Figure 1. Item response curves for deploying one’s own agent (Y3) and engaging a counterpart’s agent (Y4). The engagement curve is displaced ≈ 0.7 SD to the right. On the anchored latent scale, endorsing deployment of one’s own agent requires θ=−0.38θ=-0.38, while endorsing engagement with a counterpart’s agent requires θ=+0.32θ=+0.32 (full engagement +1.39+1.39), a displacement of 0.71 SD (95% nonparametric bootstrap CI [0.65,0.77][0.65,0.77], 200 replicates; Figure 1). The displacement is stable across the alternative category orderings we test (0.700.70–0.720.72), and the latent correlation is estimated precisely (ρ=0.92ρ=0.92, CI [0.90,0.94][0.90,0.94]). Mean model-implied deployment endorsement is 0.38 (strict) / 0.50 (soft) against engagement endorsement of 0.12 / 0.26. The asymmetry is not a modeling artifact. Y3 and Y4 differ in wording and response options as well as role, so the threshold displacement could in principle absorb formulation differences. Three checks show the asymmetry is not an artifact of dimensional specification or coding, and is visible directly within respondents. • Model-free paired evidence: on the collapsed three-category codings, 40.7% of respondents individually rate deploying their own agent strictly above engaging a counterpart’s, against 2.1% in the opposite direction: a 19:1 imbalance (Wilcoxon signed-rank p<10−189p<10^-189; sign test p<10−232p<10^-232). The modal cell of the paired table is joint rejection (36.1%), and full send-endorsement coexists with full receive-endorsement in only 11.0% of respondents. • Dimensionality: refitting with all seven items on a single latent dimension, so both thresholds live on literally the same scale, yields a displacement of 0.680.68 SD, within a few hundredths of the 2D estimate. • Coding: the displacement is stable (0.700.70–0.720.72) across the alternative category orderings of §4. These checks cannot, however, separate a pure role effect from residual scenario-framing differences between the two stems (“while you are away” vs. “responded to you”); a paired vignette experiment randomizing role against fixed wording is the clean follow-up. 6.3. Receptivity is associated with unmet need Table 5. Latent regression of send and receive receptivity on centered covariates (trait SD = 1). Brackets: 95% nonparametric bootstrap CIs (200 replicates over respondents), recentered on the full-precision estimates. Covariate βsend _send βrecv _recv Female −0.38-0.38 [-0.57, -0.23] −0.30-0.30 [-0.47, -0.18] Age band +0.07+0.07 [0.00, 0.15] −0.01-0.01 [-0.07, 0.06] Platform tenure −0.14-0.14 [-0.24, 0.00] −0.11-0.11 [-0.19, 0.04] Match volume +0.09+0.09 [-0.04, 0.17] +0.08+0.08 [-0.02, 0.14] Match→ rate −0.17-0.17 [-0.27, -0.11] −0.16-0.16 [-0.26, -0.09] Negative affect (stalled matches) +0.14+0.14 [0.07, 0.26] +0.14+0.14 [0.06, 0.23] Language: EN +0.01+0.01 [-0.14, 0.16] +0.07+0.07 [-0.09, 0.22] Table 5 reports the latent regression. Women are substantially less receptive on both dimensions (send β=−0.38β=-0.38, CI [−0.57,−0.23][-0.57,-0.23]; receive β=−0.30β=-0.30, CI [−0.47,−0.18][-0.47,-0.18]; the largest effects in the model), and receptivity declines modestly with platform tenure, though the tenure intervals are marginal (Table 5). The two most diagnostic effects concern need: users whose matches already convert to conversations are less receptive (β≈−0.17β≈-0.17 on both dimensions), while users who report negative affect about stalled matches are more receptive (β≈+0.14β≈+0.14). Receptivity concentrates where the friction is: it is highest among users reporting that the human-authored funnel is failing them. These are cross-sectional associations: the design does not license a causal reading, but the pattern is exactly what a need-based account predicts. This has a double edge. It is consistent with agents addressing a real, self-identified pain point rather than a manufactured one; it equally implies that the users most likely to adopt are those in the most emotionally loaded position, a targeting question we return to in the ethics discussion. The language coefficient is near zero on both dimensions, consistent with the pooling of the two samples. 6.4. Measurement invariance Per-item likelihood-ratio DIF tests (Bonferroni-corrected) reject full cross-language invariance for two items (Y4: LRT =15.4=15.4, p=.001p=.001; Y5: LRT =41.3=41.3, p<10−7p<10^-7); the remaining five are invariant, and a partial-invariance model freeing Y4/Y5 for the EN group leaves all headline quantities essentially unchanged (gap =0.69=0.69, ρ=0.92ρ=0.92, gender effects within 0.0050.005 of the constrained fit). By gender, only the agent-to-agent item shows DIF (Y5: LRT =16.6=16.6); the two items defining the asymmetry (Y3, Y4) function equivalently for men and women, so the gender gap reflects trait differences rather than differential item functioning. Where the invariance breaks, and what it may mean. The non-invariance is not randomly located: both flagged items are receive-side, while every send-side item, including deployment of one’s own agent, functions equivalently across the two language forms. The freed parameters show the direction. For both items the lowest threshold sits substantially higher in the EN form (Y4: b1=0.47b_1=0.47 vs. 0.310.31; Y5: b1=−0.94b_1=-0.94 vs. −1.73-1.73): at equal latent receptivity, EN-form respondents more often select the categorical-rejection options (“weird – only want a real person”; “strange and frightening”). Simultaneously, the EN group’s latent receive-side mean rises to +0.20+0.20 under partial invariance: two offsetting effects that the fully constrained model conflated into a near-zero language coefficient (CI [−0.09,0.22][-0.09,0.22]), and that a sum-score analysis would never have separated. One reading is substantive: delegating is a largely private utility judgment (time saved, control retained), whereas accepting agent-mediated contact is a normative one, where a judgment about politeness, authenticity, and what counts as strange, and norms are exactly what varies across linguistic–cultural contexts, consistent with the socially constructed character of authenticity judgments in AI-mediated communication (Hancock et al. 2020; Ante 2026). We state this as a hypothesis rather than a finding: with 232 EN respondents self-selecting the English form on a single platform, language is confounded with unmeasured population composition, and translation artifacts in the rejection options’ emotional valence cannot be separated from genuine norm differences. What the analysis does establish is methodological: receive-side receptivity is where measurement travels least well, so cross-locale deployments of the audit must invariance-test per market, and calibrate receive-side routing thresholds per locale rather than assume a single calibration. Cross-cultural replication with balanced samples is the natural extension, and the receive side is where variation should be sought. 6.5. User segments A latent class analysis over the seven items: BIC-selected at four classes (27,632, vs. 28,031 at K=3K=3 and 27,633 at K=5K=5), with relative entropy 0.830.83 and all 20 random restarts recovering the identical solution (mean ARI =1.00=1.00), gives the population structure behind the continuous traits. The four- and five-class solutions are nearly tied on BIC (Δ≈1 ≈ 1); the choice is immaterial for our claims, because under K=5K=5 the asymmetric-delegator class persists essentially unchanged (98.9% of its members map to a K=5K=5 class with the same profile and share, the extra class instead splitting the ambivalent group), as do the rejector and enthusiast classes. Rejectors (≈ 31%) are negative on every item; women are overrepresented here (36% of women vs. 29% of men). Enthusiasts (≈ 19%) are positive throughout, including toward receiving agents. Ambivalents (≈ 25%) sit near the population mean with lukewarm responses across the board. The remaining class (≈ 26%) is the structurally consequential one: asymmetric delegators, who overwhelmingly want their own agent (98% respond “want to try” or “might try”), enjoy the spectacle of agent-to-agent conversation, yet only 10% would fully engage an agent that contacts them, and they rate overall product value low. This is the send/receive gap embodied as a population segment, roughly a quarter of the market wants to emit agent traffic it would not itself accept, and it is the segment a reciprocity policy prices out (§7). 6.6. Passive features versus active delegation The companion instrument (N = 2,894) shows how specific this resistance is to delegation. Among respondents using the closed response options (a minority left free-text comments instead, tabulated separately), passive generative features draw broad interest: an AI-written couple description reaches 80.1% top-two-box interest and an AI profile summary 78.6% (and 70.6% say such features would raise their browsing interest). The one assistive feature that touches live conversation - AI communication tips - drops to 62.4%, with 37.6% explicitly uninterested; it is also the only feature men rate higher than women (65.1% vs. 58.3% top-two-box, against a female advantage on both passive features), consistent with the unmet-need pattern of §6. Reception of generative AI in this population is therefore not monolithic aversion; it is graded by how deeply the feature intrudes into interpersonal communication, with autonomous delegation at the far end of the gradient. 6.7. The routing signal predicts out of sample The routing counterfactual of §7 ranks receivers by receive receptivity and asks how engagement changes when agent traffic is routed to the most receptive. If receptivity were scored using the engagement item itself, that exercise would be partly circular: Y4 would help construct the score that then “predicts” Y4. We therefore validate the signal with the target item held out entirely. Under five-fold respondent-level cross-validation, we refit the two-dimensional model on the training folds excluding Y4, score each held-out respondent’s receive receptivity by EAP from the six non-target items and covariates only, calibrate a logistic link from score to Y4 on the training folds, and evaluate on the held-out folds. Pooled out-of-fold performance: AUC =0.89=0.89 for any engagement (Y4 ≥1≥ 1) and 0.880.88 for full engagement (Y4 =2=2); Brier scores 0.1320.132 and 0.0790.079 against reference (base-rate) values of 0.2400.240 and 0.1070.107. Ranked into out-of-fold quartiles, actual full-engagement rates run 0.0%0.0\% / 2.4%2.4\% / 8.2%8.2\% / 37.9%37.9\% from least to most receptive quartile (any-engagement: 0.3%0.3\% / 18.4%18.4\% / 59.1%59.1\% / 81.4%81.4\%). The top-quartile lift over the population base rate is 3.1×3.1× (strict) - essentially the in-sample routing gain, now earned without the target item. A short instrument, minus the very item that defines engagement, ranks receivers well enough to triple per-contact engagement; this supports treating receive receptivity as a pre-deployment routing signal (validation against behavior under live agent contact remains future work). Table 6. Routing-score baselines on identical folds (full engagement, out of fold). The logistic baselines are trained on the target labels; the latent score never uses Y4. Score AUC Top-quartile lift Covariates only (logistic) 0.60 1.5× Raw Y5+Y6Y5+Y6 sum score 0.85 3.0× Logistic on six non-target items 0.89 3.2× Latent (EAP, leave-Y4-out) 0.88 3.1× Against simpler scores. Table 6 benchmarks the latent score on the identical folds. Covariates alone barely beat chance (AUC 0.59): receive receptivity is not reducible to demographics or platform experience. A raw Y5+Y6 sum score reaches 0.85; a plain logistic on the six non-target items reaches 0.89. The latent score sits at 0.88, which is predictive parity with the supervised classifier. One asymmetry is worth noting: the logistic baselines are trained on the target labels in every training fold, while the latent score is estimated without ever observing Y4 (the logistic link fitted afterwards is monotone, so it does not affect ranking metrics). We claim parity, not superiority, over supervised classifiers, while additionally delivering what prediction-only scores cannot: threshold locations (the asymmetry itself), formal invariance testing, calibrated uncertainty, and a construct that transfers across instruments and markets. 7. Stated-Preference Market Counterfactuals The measurement model yields, for every respondent, model-implied endorsement propensities: pidepp^dep_i, the probability of endorsing deployment of one’s own agent (from item Y3 at θ^isend θ^send_i), and piengp^eng_i, the probability of endorsing engagement with a counterpart’s agent (from Y4 at θ^irecv θ^recv_i). These are stated-preference quantities, not calibrated behavior; we report a strict operationalization (top response category only) and a soft one (half-credit for the intermediate “might try” / “might engage” category) and read the pair as scenario bounds. Population means already reveal the imbalance: deployment 0.38/0.50 (strict/soft) against engagement 0.12/0.26. We then compute a random-pairing counterfactual over directed dyads: sender i initiates an agent contact with probability pidepp^dep_i, receiver j engages with probability pjengp^eng_j, representing an independence baseline, not an equilibrium matching model. Two outcomes matter to a platform. Under a sender-eligibility rule S and a receiver-routing rule R (subsets of the population; the unconstrained market has S and R equal to everyone), the viable dyad rate measures volume, (3) V(S,R)=i[pidepi∈S]j[pjengj∈R],V(S,R)\;=\;E_i\! [p^dep_i1\i∈ S\ ]\,E_j\! [p^eng_j1\j∈ R\ ], and engagement per agent contact measures quality, Q(R)=j∈R[pjeng]Q(R)=E_j∈ R[\,p^eng_j\,]. Every design lever in this section is a choice of (S,R)(S,R), so its price is a closed-form computation on fitted quantities in stage four of the audit. Table 7 reports both outcomes under four regimes. Table 7. Random-pairing stated-preference counterfactuals. Vol. = viable dyad rate; Qual. = engagement per agent contact. Strict Soft Scenario Vol. Qual. Vol. Qual. Baseline (unconstrained) 0.044 0.116 0.128 0.256 Male → female contacts 0.035 0.084 0.114 0.213 Female → male contacts 0.042 0.134 0.122 0.279 Reciprocity gate 0.020 0.116 0.046 0.256 Routing: top 10% receivers 0.025 0.648 0.041 0.809 Routing: top 25% receivers 0.037 0.394 0.080 0.640 Routing: top 50% receivers 0.043 0.227 0.117 0.469 Routing: top 75% receivers 0.044 0.154 0.127 0.338 Baseline. Unconstrained, only 4.4% (strict) to 12.8% (soft) of directed dyads yield an engaged agent interaction, and engagement per contact is 11.6%/25.6%: most modeled agent contacts do not result in endorsed engagement. Stated enthusiasm for having an agent coexists with a market in which most of what agents would produce is, by the recipients’ own statements, unwanted. Direction matters. Because women are less receptive on both dimensions, the market is directionally imbalanced: male-deployed agents contacting women achieve 8.4% engagement per contact (strict) against 13.4% for female-deployed agents contacting men. The heavier prospective traffic (male-to-female) faces the colder reception: the same structural pattern that plagues human-authored first messages, reproduced and potentially amplified by automation. Reciprocity gate. Requiring symmetric willingness, when a user may deploy an agent only if their own soft engagement propensity is at least 0.50.5, excludes 65% of likely deployers and halves interaction volume (0.044 → 0.020 strict). This prices the free-rider structure directly: about a quarter of users want to send agents but will not receive them, and a reciprocity norm removes exactly this traffic. Figure 2. Receptivity-aware routing: engagement per agent contact (soft operationalization) as a function of receiver coverage q. The dotted line marks the unrouted baseline. Receptivity-aware routing. Routing agent contacts only to receivers in the top receptivity quartile raises engagement per contact from 11.6% to 39.4% (strict; 25.6% → 64.0% soft), or a 3.4× quality gain, at the cost of restricting coverage to 25% of receivers. This is not an artifact of scoring receivers with the engagement item itself: with Y4 held out and receptivity re-estimated under cross-validation, the top-quartile lift is 3.1×3.1× (§6.7). Figure 2 traces the full quality–coverage frontier, which a platform can tune like any exposure policy. Routing on receive receptivity is a matching-layer intervention invisible to senders; it converts the measurement model into a pre-deployment policy input. Two caveats bound these numbers. The simulation assumes random dyad formation, ignoring matching structure and homophily in receptivity. It treats propensities as static, whereas receivers plausibly update after good or bad agent encounters, reflecting that an equilibrium version is future work. The qualitative conclusions: receive-side scarcity, directional imbalance, and the volume/quality lever structure, do not depend on these simplifications. 8. Design Implications for Agentic Recommenders Receive-side consent is a first-class primitive. Current agent designs treat deployment as the consent event: the sender opts in. Our results locate the scarce resource on the other side. A receiving-consent surface, or an explicit setting for whether, and from whom, agent-mediated contact is acceptable, is not a compliance nicety but the mechanism that determines market quality; without it, three quarters of agent contacts land on users who did not want them (Table 7). Disclosure follows directly: undisclosed agents convert unwilling receivers into deceived ones, the “digital betrayal” documented qualitatively by (Ante 2026), and our receive-side estimates should be read as an upper bound on tolerance for undisclosed contact. Receptivity is a routable signal. Because receive receptivity is measurable from a handful of items and predicts held-out engagement responses with AUC 0.880.88 even when the engagement item itself is excluded from scoring (§6.7), it can enter the matching layer like any exposure feature (in production, behavioral proxies would play the items’ role). The quality–coverage frontier (Figure 2) is the tuning knob: routing agent traffic to the top receptivity quartile more than triples per-contact engagement while leaving human-authored contact untouched for everyone else. This is the cheapest intervention in our counterfactual set and requires no visible product change for senders. Reciprocity is a values choice with a known price. A symmetric-willingness rule (deploy only if you would engage) halves agent-interaction volume and excludes two thirds of would-be deployers - almost exactly the asymmetric-delegator segment. Whether to charge that price is a platform-values decision, not a technical one; our contribution is that the price is now quantified. Beyond dating: auditing delegated-communication markets. The audit transfers wherever delegation meets a human receiver. Recruiting is the nearest analogue: agent-authored candidate outreach is already deployed, candidate-side receptivity is unmeasured, and the market shares dating’s structure of concentrated attention and scarce replies. Sales prospecting, agent-mediated customer contact, and professional networking follow the same pattern. In each, the audit’s outputs map onto the same levers: deployment eligibility, receive-side routing (invariance-tested and calibrated per locale, per §6), and disclosure mechanics. The same failure mode looms: send-side enthusiasm masking receive-side scarcity. We conjecture the delegation asymmetry itself generalizes, because its plausible mechanisms (retained control over one’s own agent, none over others’; effort saved when sending, authenticity lost when receiving) are not dating-specific; testing that conjecture requires only re-fielding the instrument, which is the point of packaging the procedure. Sequence the rollout by intrusion depth. The passive-to-active gradient (§6) suggests a staged path: profile-level generative features first (75%+ interest, low relational intrusion), conversation assistance with strong user control second, autonomous delegation last and opt-in on both sides. Deploying the far end of the gradient first, to the most frustrated users, is the demand-efficient and trust-corrosive order. 9. Limitations Our estimates are stated preferences elicited from concept mockups, not behavior under deployed agents; the technology-acceptance literature suggests stated and revealed adoption correlate but diverge, and the direction of divergence for receiving agents is unknown. Both samples are self-selected respondents from a single platform, and the EN subsample is small (n = 232) and self-selected, so language is confounded with population composition. Our cross-lingual claims are therefore limited to the partial-invariance form we report, and the cultural interpretation of the receive-side DIF remains a hypothesis. Ordinal codings required judgment calls for response options without a canonical order; the sensitivity analyses show the headline quantities are robust to the defensible alternatives, but cannot rule out orderings we did not consider. The two instruments were fielded to non-linked respondents, so the passive-versus-active contrast is between-sample. The market simulation assumes random dyad formation and static propensities, abstracting from matching structure, homophily in receptivity, and the plausible dynamics in which receivers update after agent encounters. Finally, attitudes toward agent mediation are likely nonstationary as the technology normalizes; our estimates are a 2026 snapshot of a moving object. 10. Conclusion Agent-mediated matching is being built sender-first, but it will succeed or fail receiver-first. Measuring both sides of delegation within person, we find distinct constructs separated by a robust 0.7 SD asymmetry, a quarter of users wanting to send agent traffic they would not fully accept, and, under a random-pairing counterfactual, only 4–13% of directed dyads combining deployment with engagement unless receive-side consent becomes a design primitive. The measurement model that produces these numbers is cheap with only seven ordinal items, and its pre-deployment outputs are directly usable, from a validated routing signal to the quantified price of a reciprocity norm. As LLM agents enter interpersonal products, we offer this as a template for measuring readiness and a caution against mistaking enthusiasm for delegation as enthusiasm for its receipt. Ethical Considerations Data handling and review. Both surveys were voluntary, fielded by the platform to its users for product research, and provided to the authors as existing data; this study is a secondary analysis of de-identified data, determined by the authors’ institutional review board not to constitute human subjects research (protocol 26-08-10287). Analyses are reported in aggregate: no individual-level data, no free-text excerpts attributable to individuals, and no demographic cells small enough to risk re-identification (a concern for the small EN subsample in particular). Contact information volunteered by interview candidates was excluded from analysis. The public deposit (§4) contains only the anonymized structured responses: direct identifiers and free-text are removed, timestamps are coarsened, and demographic cells were audited so that no released cross-classification isolates fewer than k=5k=5 respondents; raw responses are not released. Dual use of receptivity scoring. A model that scores users’ receptivity to AI contact can protect users (routing agent traffic away from those who do not want it) or exploit them (segmenting users for maximal AI exposure regardless of preference). The unmet-need association sharpens this: the most receptive users are those most frustrated with their current experience, so demand-efficient targeting concentrates an experimental technology on users in an emotionally vulnerable position. We consider receive-side routing defensible precisely because it acts on the receiving user’s own preference; deployment-side targeting of frustrated users optimizes the platform’s adoption curve against the user’s state, and we caution against it. Gender differences. We report robust gender differences in receptivity, verified to reflect trait differences rather than differential item functioning. These describe aggregate distributions and must not be used to gate features by gender or to justify differential treatment of individuals; their design-relevant content is directional (the heaviest prospective agent traffic faces the least receptive audience), not individual. Deception and disclosure. Undisclosed agent communication converts our receive-side estimates from consent measurements into deception measurements: users who would decline agent contact cannot decline what they cannot detect. Our results therefore argue for mandatory disclosure of agent authorship, and we note that prior work finds trust penalties even for suspected AI authorship, so disclosure is also in platforms’ long-run interest. Agent-mediated intimacy. Delegating early romantic communication to agents raises questions beyond any platform: whose words form the basis of a relationship, and what happens at the transition to unassisted interaction. Our data show users themselves articulate these concerns unprompted. We present measurement and market analysis to inform this debate, not to settle it. References (1) Alizadeh et al. (2024) Fatemeh Alizadeh, Dennis Lawo, Gunnar Stevens, Douglas Zytko, and Motahhare Eslami. 2024. When the “Matchmaker” Does Not Have Your Interest at Heart: Perceived Algorithmic Harms, Folk Theories, and Users’ Counter-Strategies on Tinder. Proceedings of the ACM on Human-Computer Interaction 8, CSCW2 (2024), 1–29. doi:10.1145/3689710 Ante (2026) Lennart Ante. 2026. The Cyrano Effect: LLM-Assisted Impression Management and Authenticity in Online Dating. Telematics and Informatics 108 (2026), 102422. doi:10.1016/j.tele.2026.102422 Bock and Aitkin (1981) R. Darrell Bock and Murray Aitkin. 1981. Marginal Maximum Likelihood Estimation of Item Parameters: Application of an EM Algorithm. Psychometrika 46, 4 (1981), 443–459. Bollen (1989) Kenneth A. Bollen. 1989. Structural Equations with Latent Variables. Wiley, New York. Bollen (2002) Kenneth A. Bollen. 2002. Latent Variables in Psychology and the Social Sciences. Annual Review of Psychology 53 (2002), 605–634. Borsboom et al. (2003) Denny Borsboom, Gideon J. Mellenbergh, and Jaap van Heerden. 2003. The Theoretical Status of Latent Variables. Psychological Review 110, 2 (2003), 203–219. Bruch and Newman (2018) Elizabeth E. Bruch and M. E. J. Newman. 2018. Aspirational Pursuit of Mates in Online Dating Markets. Science Advances 4, 8 (2018), eaap9815. Cela et al. (2025) Hektor Cela, Mária Vogrin, Thomas Schmickl, and Guilherme Wood. 2025. Emotional Dynamics and Engagement Cycles in Swiping Dating Apps: An Agent-Based Modeling Approach. Computers in Human Behavior Reports (2025), 100775. Davis (1989) Fred D. Davis. 1989. Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology. MIS Quarterly 13, 3 (1989), 319–340. Dietvorst et al. (2015) Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2015. Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Journal of Experimental Psychology: General 144, 1 (2015), 114–126. Embretson and Reise (2000) Susan E. Embretson and Steven P. Reise. 2000. Item Response Theory for Psychologists. Lawrence Erlbaum Associates. Finkel et al. (2012) Eli J. Finkel, Paul W. Eastwick, Benjamin R. Karney, Harry T. Reis, and Susan Sprecher. 2012. Online Dating: A Critical Analysis From the Perspective of Psychological Science. Psychological Science in the Public Interest 13, 1 (2012), 3–66. Grassini (2023) Simone Grassini. 2023. Development and Validation of the AI Attitude Scale (AIAS-4): A Brief Measure of General Attitude toward Artificial Intelligence. Frontiers in Psychology 14 (2023), 1191628. Hancock et al. (2020) Jeffrey T. Hancock, Mor Naaman, and Karen Levy. 2020. AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations. Journal of Computer-Mediated Communication 25, 1 (2020), 89–100. Hayashi et al. (2025) Yudai Hayashi, Shuhei Goda, and Yuta Saito. 2025. Off-Policy Evaluation and Learning for Matching Markets. In Proceedings of the 19th ACM Conference on Recommender Systems (RecSys). 340–349. Hitsch et al. (2010) Günter J. Hitsch, Ali Hortaçsu, and Dan Ariely. 2010. Matching and Sorting in Online Dating. American Economic Review 100, 1 (2010), 130–163. Hohenstein et al. (2023) Jess Hohenstein, René F. Kizilcec, Dominic DiFranzo, Zhila Aghajari, Hannah Mieczkowski, Karen Levy, Mor Naaman, Jeffrey Hancock, and Malte F. Jung. 2023. Artificial Intelligence in Communication Impacts Language and Social Relationships. Scientific Reports 13 (2023), 5487. doi:10.1038/s41598-023-30938-9 Jakesch et al. (2019) Maurice Jakesch, Megan French, Xiao Ma, Jeffrey T. Hancock, and Mor Naaman. 2019. AI-Mediated Communication: How the Perception that Profile Text was Written by AI Affects Trustworthiness. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13. doi:10.1145/3290605.3300469 Kadoma et al. (2025) Kowe Kadoma, Danaë Metaxa, and Mor Naaman. 2025. Generative AI and Perceptual Harms: Who’s Suspected of Using LLMs?. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–17. doi:10.1145/3706598.3713897 Logg et al. (2019) Jennifer M. Logg, Julia A. Minson, and Don A. Moore. 2019. Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organizational Behavior and Human Decision Processes 151 (2019), 90–103. McNeish and Wolf (2020) Daniel McNeish and Melissa Gordon Wolf. 2020. Thinking Twice about Sum Scores. Behavior Research Methods 52, 6 (2020), 2287–2305. doi:10.3758/s13428-020-01398-0 Meredith (1993) William Meredith. 1993. Measurement Invariance, Factor Analysis and Factorial Invariance. Psychometrika 58, 4 (1993), 525–543. Millsap (2011) Roger E. Millsap. 2011. Statistical Approaches to Measurement Invariance. Routledge. Neve and Palomares (2019) James Neve and Iván Palomares. 2019. Latent Factor Models and Aggregation Operators for Collaborative Filtering in Reciprocal Recommender Systems. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys). 219–227. Palomares et al. (2021) Iván Palomares, Carlos Porcel, Luiz Pizzato, Ido Guy, and Enrique Herrera-Viedma. 2021. Reciprocal Recommender Systems: Analysis of State-of-Art Literature, Challenges and Opportunities Towards Social Recommendation. Information Fusion 69 (2021), 103–127. Park et al. (2023) Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST). Pemstein et al. (2018) Daniel Pemstein, Kyle L. Marquardt, Eitan Tzelgov, Yi-ting Wang, Juraj Medzihorsky, Joshua Krusell, Farhad Miri, and Johannes von Römer. 2018. The V-Dem Measurement Model: Latent Variable Analysis for Cross-National and Cross-Temporal Expert-Coded Data. Working Paper. Varieties of Democracy Institute. Peng et al. (2025) Qiyao Peng, Hongtao Liu, Hua Huang, Jian Yang, Qing Yang, and Minglai Shao. 2025. A Survey on LLM-powered Agents for Recommender Systems. In Findings of the Association for Computational Linguistics: EMNLP 2025. 11574–11583. doi:10.18653/v1/2025.findings-emnlp.620 Pizzato et al. (2013) Luiz Pizzato, Tomasz Rej, Joshua Akehurst, Irena Koprinska, Kalina Yacef, and Judy Kay. 2013. Recommending People to People: The Nature of Reciprocal Recommenders with a Case Study in Online Dating. User Modeling and User-Adapted Interaction 23, 5 (2013), 447–488. Pizzato et al. (2010) Luiz Pizzato, Tomasz Rej, Thomas Chung, Irena Koprinska, and Judy Kay. 2010. RECON: A Reciprocal Recommender for Online Dating. In Proceedings of the Fourth ACM Conference on Recommender Systems (RecSys). 207–214. Rosenfeld et al. (2019) Michael J. Rosenfeld, Reuben J. Thomas, and Sonia Hausen. 2019. Disintermediating Your Friends: How Online Dating in the United States Displaces Other Ways of Meeting. Proceedings of the National Academy of Sciences 116, 36 (2019), 17753–17758. Samejima (1969) Fumiko Samejima. 1969. Estimation of Latent Ability Using a Response Pattern of Graded Scores. Psychometrika Monograph Supplement 34, 4, Pt. 2 (1969). Schepman and Rodway (2020) Astrid Schepman and Paul Rodway. 2020. Initial Validation of the General Attitudes towards Artificial Intelligence Scale. Computers in Human Behavior Reports 1 (2020), 100014. Schepman and Rodway (2023) Astrid Schepman and Paul Rodway. 2023. The General Attitudes towards Artificial Intelligence Scale (GAAIS): Confirmatory Validation and Associations with Personality, Corporate Distrust, and General Trust. International Journal of Human–Computer Interaction 39 (2023). doi:10.1080/10447318.2022.2085400 Skrondal and Rabe-Hesketh (2004) Anders Skrondal and Sophia Rabe-Hesketh. 2004. Generalized Latent Variable Modeling: Multilevel, Longitudinal, and Structural Equation Models. Chapman & Hall/CRC. Stein et al. (2024) Jan-Philipp Stein, Tanja Messingschlager, Timo Gnambs, Fabian Hutmacher, and Markus Appel. 2024. Attitudes towards AI: Measurement and Associations with Personality. Scientific Reports 14 (2024), 2909. Tomita and Yokoyama (2024) Yoji Tomita and Tomohiko Yokoyama. 2024. Fair Reciprocal Recommendation in Matching Markets. In Proceedings of the 18th ACM Conference on Recommender Systems (RecSys). Venkatesh et al. (2003) Viswanath Venkatesh, Michael G. Morris, Gordon B. Davis, and Fred D. Davis. 2003. User Acceptance of Information Technology: Toward a Unified View. MIS Quarterly 27, 3 (2003), 425–478. Wang et al. (2024) Yancheng Wang, Ziyan Jiang, Zheng Chen, Fan Yang, Yingxue Zhou, Eunah Cho, Xing Fan, Yanbin Lu, Xiaojiang Huang, and Yingzhen Yang. 2024. RecMind: Large Language Model Powered Agent for Recommendation. In Findings of the Association for Computational Linguistics: NAACL 2024. 4351–4364. doi:10.18653/v1/2024.findings-naacl.271 Xia et al. (2014) Peng Xia, Hua Jiang, Xiaodong Wang, Cindy Chen, and Benyuan Liu. 2014. Predicting User Replying Behavior on a Large Online Dating Site. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). 545–554. doi:10.1609/icwsm.v8i1.14516 Xia et al. (2015) Peng Xia, Benyuan Liu, Yizhou Sun, and Cindy Chen. 2015. Reciprocal Recommendation System for Online Dating. In Proceedings of the IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). 234–241. Xu et al. (2025) Wujiang Xu, Yunxiao Shi, Zujie Liang, Xuying Ning, Kai Mei, Kun Wang, Xi Zhu, Min Xu, and Yongfeng Zhang. 2025. iAgent: LLM Agent as a Shield between User and Recommender Systems. In Findings of the Association for Computational Linguistics: ACL 2025. 18056–18084. Ye et al. (2025) Wanghao Ye, Sihan Chen, Yiting Wang, Shwai He, Bowei Tian, Guoheng Sun, et al. 2025. CogniPair: From LLM Chatbots to Conscious AI Agents — GNWT-Based Multi-Agent Digital Twins for Social Pairing — Dating & Hiring Applications. arXiv:2506.03543. Zhang et al. (2024) Junjie Zhang, Yupeng Hou, Ruobing Xie, Wenqi Sun, Julian McAuley, Wayne Xin Zhao, Leyu Lin, and Ji-Rong Wen. 2024. AgentCF: Collaborative Learning with Autonomous Language Agents for Recommender Systems. In Proceedings of the ACM Web Conference 2024 (W). 3679–3689.