Paper deep dive
Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection
Vignesh Iyer
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/11/2026, 12:46:04 AM
Summary
The paper introduces 'structural tokenization' for AI agent threat detection, which encodes execution-flow patterns (tool calls, arguments, observations) rather than conversational content. This approach significantly improves cross-attack generalization, particularly for structural attacks like tool hijacking and data exfiltration, while gated multi-view fusion allows for robust detection across both linguistic and structural attack families.
Entities (5)
Relation Signals (3)
Structural Tokenization â improvesgeneralizationfor â Tool Hijacking
confidence 98% ¡ This simple representational change dramatically improves cross-attack generalization: +46 AUC points on tool hijacking
Gated Multi-View Fusion â combines â Structural Tokenization
confidence 95% ¡ gated multi-view fusion that adaptively combines both representations
Conversational Tokenization â failson â Structural Attacks
confidence 95% ¡ standard conversational tokenization... fails catastrophically on structural attacks
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Autonomous AI agents executing multi-step tool sequences face semantic attacks that manifest in behavioral traces rather than isolated prompts. A critical challenge is cross-attack generalization: can detectors trained on known attack families recognize novel, unseen attack types? We discover that standard conversational tokenization -- capturing linguistic patterns from agent interactions -- fails catastrophically on structural attacks like tool hijacking (AUC 0.39) and data exfiltration (AUC 0.46), while succeeding on linguistic attacks like social engineering (AUC 0.78). We introduce structural tokenization, encoding execution-flow patterns (tool calls, arguments, observations) rather than conversational content. This simple representational change dramatically improves cross-attack generalization: +46 AUC points on tool hijacking, +39 points on data exfiltration, and +71 points on unknown attacks, while simultaneously improving in-distribution performance (+6 points). For attacks requiring linguistic features, we propose gated multi-view fusion that adaptively combines both representations, achieving AUC 0.89 on social engineering without sacrificing structural attack detection. Our findings reveal that AI agent security is fundamentally a structural problem: attack semantics reside in execution patterns, not surface language. While our rule-based tokenizer serves as a baseline, the structural abstraction principle generalizes even with simple implementation.
Tags
Links
- Source: https://arxiv.org/abs/2601.01723
- Canonical: https://arxiv.org/abs/2601.01723
Trouble viewing inline? Open PDF directly â
Full Text
34,055 characters extracted from source content.
Expand or collapse full text
Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection Vignesh Iyer iyerv68@gmail.com Abstract Autonomous AI agents executing multi-step tool se- quences face semantic attacks that manifest in behavioral traces rather than isolated prompts. A critical challenge is cross-attack generalization: can detectors trained on known attack families recognize novel, unseen attack types? We discover that standard conversational tokenizationâ capturing linguistic patterns from agent interactionsâfails catastrophically on structural attacks like tool hijacking (AUC 0.39) and data exfiltration (AUC 0.46), while suc- ceeding on linguistic attacks like social engineering (AUC 0.78).We introduce structural tokenization, encoding execution-flow patterns (tool calls, arguments, observa- tions) rather than conversational content. This simple rep- resentational change dramatically improves cross-attack generalization: +46 AUC points on tool hijacking, +39 points on data exfiltration, and +71 points on unknown at- tacks, while simultaneously improving in-distribution per- formance (+6 points). For attacks requiring linguistic fea- tures, we propose gated multi-view fusion that adaptively combines both representations, achieving AUC 0.89 on so- cial engineering without sacrificing structural attack detec- tion. Our findings reveal that AI agent security is funda- mentally a structural problem: attack semantics reside in execution patterns, not surface language. While our rule- based tokenizer serves as a baseline, the structural abstrac- tion principle generalizes even with simple implementation. 1. Introduction Autonomous AI agents powered by large language mod- els are transitioning from research prototypes to opera- tional infrastructure [41, 43]. Enabled by tool-use capabil- ities [31, 33, 34], modern deployments span critical busi- ness functions: customer service agents process refunds and manage accounts, document agents extract information from contracts, developer agents modify codebases, and data agents query production databases. This operational autonomy requires agents to possess execution privileges over tools that directly affect business operations, creating attack vectors fundamentally different from traditional soft- ware vulnerabilities. 1.1. Semantic Attacks on AI Agents Traditional cybersecurity threats exploit implementation flaws through well-understood mechanisms: buffer over- flows manipulate memory, SQL injection exploits input san- itization, and cross-site scripting abuses DOM manipula- tion [16, 29]. These attacks succeed by violating syntac- tic constraints and leave clear signatures detectable through pattern matching. AI agent attacks instead operate through semantic ma- nipulation of natural language reasoning [2, 32]. Indi- rect prompt injection [13, 24, 47] embeds malicious in- structions in external data sources that override user intent. Tool hijacking [35] redirects operations toward attacker- controlled endpoints.Data exfiltration extracts sensi- tive information through seemingly benign tool sequences. Adversarial attacks on aligned models [49] demonstrate the brittleness of safety training.Unlike traditional at- tacks, semantic attacks maintain perfect syntactic validityâ detection requires understanding behavioral context. 1.2. The Cross-Attack Generalization Problem Existing defenses focus on detecting specific attack pat- terns observed during training. However, a critical question remains understudied: can threat detectors generalize to attack families never seen during training? Real-world deployments inevitably face novel attacks not represented in training data. If a detector trained on prompt injection can- not recognize tool hijacking, its practical utility is severely limited. We systematically evaluate cross-attack generalization by holding out entire attack families during training and measuring detection performance on these unseen cate- gories [14, 21]. Our findings reveal a striking asymme- try: generalization success depends critically on what be- havioral signals the representation captures. arXiv:2601.01723v1 [cs.CR] 5 Jan 2026 1.3. Key Insight: Structure vs. Language We discover that standard conversational tokenizationâ encoding linguistic patterns from user messages and agent responsesâexhibits dramatic variation in cross- attack transfer. Linguistic attacks like social engineering (AUC 0.78) and prompt injection (AUC 0.69) transfer mod- erately well. However, structural attacks like tool hijack- ing (AUC 0.39) and data exfiltration (AUC 0.46) fail catas- trophically, while unknown attacks collapse entirely (AUC 0.26). This asymmetry reveals that conversational features cap- ture persuasion tactics but fundamentally miss execution- level threats. We hypothesize that structural attacks depend on how tools are orchestrated, not what is saidâa tool hi- jacking attack may use entirely benign language while exe- cuting a malicious tool sequence. 1.4. Contributions We make three contributions: (1) Structural tokenization. We introduce execution- flow tokenization that encodes tool calls, argument patterns, and observation sequences rather than conversational con- tent [10, 26]. This improves cross-attack generalization by 39â71 AUC points while also improving in-distribution per- formance. (2) Attack-family taxonomy. We provide the first sys- tematic study of cross-attack transfer [37, 45], revealing that linguistic and structural attacks require fundamentally dif- ferent representations. (3) Gated multi-view fusion. For deployments facing diverse attack types, we propose adaptive fusion that learns when to rely on each representation, achieving strong per- formance across all attack families. 2. Related Work AI Agent Security. Autonomous agents face seman- tic attacks including indirect prompt injection [13, 24, 47], jailbreaking [42, 49], and tool manipulation [35]. Agen- tHarm [2] provides benchmarks for measuring agent harm- fulness.Prior defenses focus on input sanitization and prompt hardening rather than cross-attack generalization. Behavioral Sequence Analysis.Sequence-based anomaly detection has proven effective in system log analy- sis [8, 10] and network intrusion detection [26]. We extend this intuition to AI agents, showing that execution-flow pat- terns provide stronger cross-attack generalization than lin- guistic features. Multi-View Learning. Combining multiple representa- tions improves robustness across domains [4, 27, 44]. Our gated fusion builds on mixture-of-experts principles [17, 36] to adaptively weight conversational and structural views. Distribution Shift. Out-of-distribution generalization remains challenging [14, 21, 37, 45]. Our cross-attack eval- uation protocol explicitly measures generalization to unseen attack familiesâa critical requirement for real-world secu- rity. Federated Security. Federated intrusion detection [11] and malware classification [23] enable collaborative learn- ing. Recent work explores federated prompt injection de- tection [18]. Non-IID data remains challenging [22, 48]; we discover that representation choice dominates aggrega- tion method. 3. Method 3.1. Problem Formulation Consider agent traces Ď = (M,T ,R) contain- ing user/assistant messages M, tool invocations T = (t i , args i , obs i ), and final responses R. The detection task is binary classification: benign vs. attack. Cross-attack evaluation. Let A = a 1 ,...,a m de- note attack families. For held-out family a j , we train on all data except a j and evaluate detection specifically on a j . This measures true generalization to unseen attack types. 3.2. Conversational Tokenization (Baseline) Our baseline represents traces using a 26-token vocab- ulary capturing linguistic and behavioral patterns through keyword matching and pattern rules: ⢠Tool Types (8): SEND EMAIL, MAKEPAYMENT, READFILE, EXECUTECODE, etc. ⢠Argument Patterns (6): EXTERNALRECIPIENT, HIGH VALUE, SENSITIVEFIELD, etc. ⢠AttackIndicators(6): INJECTIONPHRASE, OVERRIDEATTEMPT, EVASION, etc. ⢠Response Types (3): REFUSAL, COMPLIANT, CLARIFY ⢠Control Flow (3): LOOP, BRANCH, RECURSION This representation captures what is said through attack- indicative phrases. However, it encodes execution structure only indirectly through tool-type tokens. 3.3. Structural Tokenization We introduce execution-flow tokenization using a com- pact 9-token vocabulary that encodes what the agent did rather than linguistic content: TokenSemantics [SYS]System instruction present [USER]User message [ASSISTANT]Assistant response [TOOL]Tool invocation detected [ARGS]Arguments passed to tool [OBS]Tool observation returned [OUTPUT]Final response to user [ERROR]Error or exception [OTHER]Fallback token Example. A trace where a user requests a file, the agent invokes readfile, receives data, and responds becomes: [USER] [ASSISTANT] [TOOL] [ARGS] [OBS] [OUTPUT]. This representation abstracts away linguistic content en- tirely, capturing only the shape of agent execution [8]. Cru- cially, structural patterns remain discriminative even when attackers paraphraseâthey cannot hide the resulting execu- tion flow. 3.4. Gated Multi-View Fusion Neither representation alone is optimal for all attack fam- ilies. We propose adaptive fusion [4, 17] using a learned gate: g = Ď (W g [h conv ;h struct ] + b g )(1) h fused = gâ h conv + (1â g)â h struct (2) where h conv and h struct are encoded representations from parallel BiLSTM encoders [15]. The gate g learns to weight each view based on input characteristics. 3.5. Architecture and Training All models share: 64-dim embeddings [5], bidirectional LSTM [15] (hidden=64, output=128), projection to 32-dim latent space, and two-layer classifier (32â64â1) with sig- moid output. Total: 74K parameters (single-view) or 140K (multi-view). We train in a federated setting with K = 5 organi- zations using FedAvg [25]: 5 rounds, 5 local epochs per round, Adam optimizer [20] (Ρ = 0.001), batch size 32, BCE loss. Optional DP-SGD [1] provides privacy guaran- tees (Îľâ [2.35, 5.69]). 4. Experiments 4.1. Experimental Setup Dataset. We evaluate on a simulated 5-organization fed- eration with 2,500 agent traces. Each organization con- tributes 500 traces (50% benign, 50% malicious) with non- IID attack distributions reflecting realistic specialization. Held-Out AttackConvStructGated Social engineering0.780.670.89 Prompt injection0.690.810.83 Data exfiltration0.460.850.62 Tool hijacking0.390.850.60 Unknown0.260.970.92 IID (seen attacks)0.870.930.89 Table 1. Cross-attack generalization (ROC-AUC). Structural to- kenization dramatically improves detection of tool-based attacks (+39â71 points) while gated fusion excels on linguistic attacks. Bold indicates best per row. Attack families.We evaluate 5 categories: prompt injection (linguistic manipulation), tool hijacking (unau- thorized redirection), data exfiltration (information theft), social engineering (persuasion-based),and unknown (novel/uncategorized attacks). Evaluation. For cross-attack generalization, we hold out one attack family entirely during training and evaluate on that held-out family. We report ROC-AUC. 4.2. Main Results Table 1 presents our central finding: representation choice dramatically affects cross-attack transfer. Conversational tokenization achieves good perfor- mance on linguistic attacks (social engineering: 0.78, prompt injection: 0.69) but fails catastrophically on struc- tural attacks (tool hijacking: 0.39, data exfiltration: 0.46) and unknown attacks (0.26âworse than random). Structural tokenization achieves strong performance on tool-based attacks (0.85) and unknown attacks (0.97), while also improving IID performance (0.93 vs 0.87). How- ever, it underperforms on social engineering (0.67). Gated fusion achieves the best balance: strong on social engineering (0.89) and prompt injection (0.83) while main- taining good performance on unknown attacks (0.92). 4.3. Effect Size Analysis Table 2 quantifies the improvement from structural tok- enization. The improvements are substantial and consistent. No- tably, structural tokenization improves both OOD and IID performanceâa rare outcome where robustness does not trade off against accuracy. 4.4. Attack-Representation Dependency Table 3 reveals which representations each attack family requires. This taxonomy reveals a fundamental insight: most AI agent attacks are structural, not linguistic. Only social engineering requires conversational features. SettingConvStructâ Unknown (OOD)0.260.97+0.71 Tool hijacking (OOD)0.390.85+0.46 Data exfiltration (OOD)0.460.85+0.39 Prompt injection (OOD)0.690.81+0.12 IID0.870.93+0.06 Social engineering (OOD)0.780.67â0.11 Table 2. AUC gains from structural tokenization. Improvements of 39â71 points on hard OOD cases; regression only on social engineering. Attack FamilyConv?Struct? Social engineeringâ Ă Prompt injectionâ Data exfiltration Ăâ Tool hijacking Ăâ UnknownĂâ Table 3. Attack-representation dependency. Structural features dominate for 4/5 attack families. 4.5. Federated Aggregation A key finding is that aggregation method has no signif- icant effect: Local training, FedAvg, and ensemble meth- ods achieve identical results withinÂą0.02 AUC. This estab- lishes that representationânot aggregationâis the bot- tleneck for cross-attack generalization. 4.6. Privacy Analysis With DP-SGD (Îľ = 2.35), the structural model achieves OOD AUC of 0.72âstill substantially above the conver- sational baselineâs 0.52 without any privacy constraints. Meaningful signal persists under strong privacy. 5. Discussion Why structure matters. Attack semantics reside in ex- ecution patternsâtool sequences, argument flows, observa- tion handlingânot surface language [5]. Conversational to- kenization detects how attackers phrase requests; structural tokenization detects what agents do. Attackers can para- phrase arbitrarily but cannot hide execution traces. The social engineering exception. Social engineering relies on psychological manipulation manifesting in con- versational style rather than tool orchestration. Structural tokenization underperforms here (0.67 vs 0.78) because the signal is inherently linguistic. Gated fusion [27, 44] ad- dresses this by learning to weight conversational features for persuasion-based attacks. Unknown attack recovery. The dramatic improvement on unknown attacks (0.26â0.97) is particularly notable. Analysis reveals that conversational tokenization learned âfamiliar = safe,â causing ranking inversion on novel pat- terns. Structural tokenization detects execution anomalies regardless of linguistic novelty. Implications. Practitioners should: (1) default to struc- tural representations for agent security, (2) add conversa- tional features only when social engineering is a significant threat, (3) not expect federated aggregation to compensate for representational weaknesses. Limitations. Our evaluation uses synthetic traces; real- world validation is critical. Our rule-based tokenizer is intentionally simple to isolate the effect of structural ab- straction from tokenizer sophistication, but this makes it potentially susceptible to adversarial evasion. Attackers may fragment tool calls, add noise, or structure malicious traces to mimic benign patterns. However, our approach is modularâthe tokenizer can be replaced with learned vari- ants (neural encoders, contrastive learning) while preserv- ing the core insight that structural patterns generalize bet- ter than conversational content. Evaluation against adap- tive adversaries is critical future work. Our 5-family taxon- omy requires expansion. The gated model underperforms struct-only on structural attacks, suggesting improved fu- sion mechanisms [36] are needed. 6. Conclusion We demonstrate that cross-attack generalization in AI agent threat detection is fundamentally a representation problem. Structural tokenizationâencoding execution flow rather than conversational contentâimproves detection of unseen attack families by 39â71 AUC points while simulta- neously improving in-distribution performance. This chal- lenges the assumption that linguistic features suffice for agent security. Our key insight is that AI agent security is primar- ily structural: attack semantics reside in tool orchestra- tion patterns, not surface language. Practitioners should design detectors that analyze what agents do, not merely what users say [10]. For diverse threats including social engineering, gated multi-view fusion provides a principled approach to combining both signal types [4]. While our rule-based tokenizer provides a strong baseline, future work should explore learned tokenization and adaptive adversary evaluation to further strengthen detection robustness. References [1] M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with dif- ferential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pages 308â318, 2016. 3 [2] M. Andriushchenko, F. Croce, and N. Flammarion. Agen- tharm: A benchmark for measuring harmfulness of llm agents. arXiv preprint arXiv:2410.09024, 2024. 1, 2, 7 [3] E. Bagdasaryan, A. Veit, Y. Hua, D. Estrin, and V. Shmatikov. How to backdoor federated learning. In Inter- national Conference on Artificial Intelligence and Statistics, pages 2938â2948. PMLR, 2020. 8 [4] T. Baltru Ë saitis, C. Ahuja, and L.-P. Morency. Multimodal machine learning: A survey and taxonomy. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 41(2): 423â443, 2018. 2, 3, 4 [5] Y. Bengio, A. Courville, and P. Vincent. Representation learning: A review and new perspectives. IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 35(8): 1798â1828, 2013. 3, 4 [6] P. Blanchard, E. M. El Mhamdi, R. Guerraoui, and J. Stainer. Machine learning with adversaries: Byzantine tolerant gradi- ent descent. In Advances in Neural Information Processing Systems, volume 30, 2017. 8 [7] K. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth. Practical secure aggregation for privacy-preserving machine learning. In Proceedings of the 2017 ACM SIGSAC Con- ference on Computer and Communications Security, pages 1175â1191, 2017. 7 [8] A. Brown, A. Tuor, B. Hutchinson, and N. Mez. Recurrent neural network attention mechanisms for interpretable sys- tem log anomaly detection. In Proceedings of the First Work- shop on Machine Learning for Computing Systems, pages 1â 8, 2018. 2, 3 [9] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A sim- ple framework for contrastive learning of visual representa- tions. In International Conference on Machine Learning, pages 1597â1607. PMLR, 2020. 8 [10] M. Du, F. Li, G. Zheng, and V. Srikumar. Deeplog: Anomaly detection and diagnosis from system logs through deep learning. In Proceedings of the 2017 ACM SIGSAC Con- ference on Computer and Communications Security, pages 1285â1298, 2017. 2, 4 [11] M. A. Ferrag, O. Friha, L. Maglaras, H. Janicke, and L. Shu. Federated deep learning for cyber security in the internet of things: Concepts, applications, and experimental analysis. IEEE Access, 9:138509â138542, 2021. 2 [12] M. Fredrikson, S. Jha, and T. Ristenpart. Model inversion attacks that exploit confidence information and basic coun- termeasures. In Proceedings of the 22nd ACM SIGSAC Con- ference on Computer and Communications Security, pages 1322â1333, 2015. 7, 8 [13] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz. Not what youâve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Ar- tificial Intelligence and Security, pages 79â90, 2023. 1, 2, 7 [14] D. Hendrycks and T. Dietterich. Benchmarking neural net- work robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019. 1, 2, 7 [15] S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735â1780, 1997. 3, 7 [16] M. Howard and S. Lipner. The Security Development Lifecy- cle. Microsoft Press, 2006. 1 [17] R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. Adaptive mixtures of local experts. In Neural Computation, volume 3, pages 79â87, 1991. 2, 3 [18] H. Jayathilaka. Privacy-preserving prompt injection detec- tion for LLMs using federated learning and embedding- based NLP classification. arXiv preprint arXiv:2511.12295, 2025. 2 [19] P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Ben- nis, A. N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, et al. Advances and open problems in feder- ated learning. Foundations and Trends in Machine Learning, 14(1â2):1â210, 2021. 7 [20] D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 3, 7 [21] P. W. Koh, S. Sagawa, H. Marber, et al. Wilds: A benchmark of in-the-wild distribution shifts. In International Confer- ence on Machine Learning, pages 5637â5664. PMLR, 2021. 1, 2 [22] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith. Federated optimization in heterogeneous networks. In Proceedings of Machine Learning and Systems, volume 2, pages 429â450, 2020. 2, 7 [23] Y. Li, Y. Wei, et al. Fedmal: A federated learning framework for malware detection. IEEE Transactions on Dependable and Secure Computing, 2021. 2 [24] Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and Y. Liu. Prompt injection attack against llm- integrated applications. arXiv preprint arXiv:2306.05499, 2023. 1, 2 [25] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statis- tics, pages 1273â1282. PMLR, 2017. 3, 7 [26] Y. Mirsky, T. Doitshman, Y. Elovici, and A. Shabtai. Kit- sune: An ensemble of autoencoders for online network intru- sion detection. In Network and Distributed System Security Symposium (NDSS), 2018. 2 [27] J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning, pages 689â 696, 2011. 2, 4 [28] T. D. Nguyen, P. Rieger, R. De Viti, H. Chen, B. B. Branden- burg, H. Yalame, H. M Ě ollering, H. Fereidooni, S. Marchal, M. Miettinen, et al. Flame: Taming backdoors in federated learning. In 31st USENIX Security Symposium, pages 1415â 1432, 2022. 8 [29] OWASP Foundation.Owasp top 10:2021. https:// owasp.org/Top10/, 2021. 1 [30] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Process- ing Systems, volume 32, 2019. 7 [31] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez. Gorilla: Large language model connected with massive apis. arXiv preprint arXiv:2305.15334, 2023. 1 [32] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving.Red teaming language models with language models.arXiv preprint arXiv:2202.03286, 2022. 1 [33] Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789, 2023. 1 [34] T. Schick, J. Dwivedi-Yu, R. Dess ` Äą, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023. 1 [35] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh. Survey of vul- nerabilities in large language models revealed by adversarial attacks. arXiv preprint arXiv:2310.10844, 2023. 1, 2 [36] N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017. 2, 4, 8 [37] Z. Shen, J. Liu, Y. He, X. Zhang, R. Xu, H. Yu, and P. Cui. Towards out-of-distribution generalization: A survey. arXiv preprint arXiv:2108.13624, 2021. 2 [38] R. Shokri, M. Stronati, C. Song, and V. Shmatikov. Member- ship inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy, pages 3â18. IEEE, 2017. 7, 8 [39] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ĺ. Kaiser, and I. Polosukhin. Attention is all you need. Advances in Neural Information Processing Sys- tems, 30, 2017. 8 [40] P. Voigt and A. Von dem Bussche. The EU General Data Protection Regulation (GDPR): A Practical Guide. Springer, 2017. 7 [41] L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al. A survey on large language model based autonomous agents. Frontiers of Com- puter Science, 18(6):186345, 2024. 1 [42] A. Wei, N. Haghtalab, and J. Steinhardt. Jailbroken: How does llm safety training fail? Advances in Neural Informa- tion Processing Systems, 36, 2023. 2, 7 [43] Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, et al. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864, 2023. 1 [44] C. Xu, D. Tao, and C. Xu. A survey on multi-view learning. arXiv preprint arXiv:1304.5634, 2013. 2, 4 [45] J. Yang, K. Zhou, Y. Li, and Z. Liu.Generalized out-of-distribution detection: A survey.arXiv preprint arXiv:2110.11334, 2021. 2, 7 [46] A. Yousefpour, I. Shilov, A. Sablayrolles, D. Testuggine, K. Prasad, M. Maber, J. Nie, S. Grimm, W. Guo, L. Zong, et al. Opacus: User-friendly differential privacy library in pytorch. https://opacus.ai, 2021. 7 [47] Q. Zhan, Z. Liang, Z. Ying, and D. Kang. Injecagent: Bench- marking indirect prompt injections in tool-integrated large language model agents. arXiv preprint arXiv:2403.02691, 2024. 1, 2, 7 [48] Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V. Chan- dra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018. 2, 7 [49] A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson. Univer- sal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023. 1, 2 A. Extended Experimental Details Dataset Generation.Traces mirror documented at- tacks [2, 13, 42, 47] with placeholder content. 2,500 to- tal traces; 500 per organization; 50% benign, 50% ma- licious across 5 attack families. Data handling follows privacy-preserving principles consistent with GDPR re- quirements [40]. Non-IID Distribution.Attack categories distributed non-uniformly across organizations to simulate realistic specialization: Org-1 sees primarily prompt injection, Org- 2 sees tool hijacking, etc. Architecture Details.Embedding: 64-dim learned vectors.BiLSTM [15]: hidden=64, concatenated for- ward/backward states yield 128-dim. Projection: linear 128â32. Classifier: 32â64 (ReLU)â 64â1 (sigmoid). Gated model uses parallel encoders with element-wise gat- ing. Training Configuration.Adam optimizer [20] (Ρ=0.001, β 1 =0.9, β 2 =0.999), binary cross-entropy loss, batch size 32, 5 federated rounds Ă 5 local epochs. Privacy Mechanisms.We employ secure aggrega- tion [7] to prevent the server from observing individual client updates. When differential privacy is enabled: gradi- ent clipping C=1.0, noise multiplier Ď â [0.6, 1.1], δ=10 â5 , R Ě enyi DP accounting via Opacus [46]. These mechanisms bound membership inference [38] and model inversion [12] risks. Hardware. Apple M2 Pro, 16GB RAM. Training com- pletes in under 5 minutes per configuration. No GPU re- quired. Implementation uses PyTorch [30]. Differential Privacy. When enabled: gradient clipping C=1.0, noise multiplier Ď â [0.6, 1.1], δ=10 â5 , R Ě enyi DP accounting via Opacus [46]. B. Complete Results Table 4 shows the full comparison across all settings. SettingConvStructGatedBest IID0.870.930.89Struct Social eng.0.780.670.89Gated Prompt inj.0.690.810.83Gated Data exfil.0.460.850.62Struct Tool hijack0.390.850.60Struct Unknown0.260.970.92Struct OOD Average0.520.830.77Struct Table 4. Complete results. Structural achieves best OOD average; gated excels on linguistic attacks. Table 5 shows per-organization performance consis- tency. OrgConv IIDStruct IIDStruct OOD 10.850.920.81 20.880.930.84 30.860.920.82 40.890.940.86 50.870.930.83 Mean0.870.930.83 Std0.020.010.02 Table 5. Per-organization results show consistent improvement with low variance. C. Ablation Studies View Necessity. Neither single view is optimal across attack types. Conv-only achieves 0.78 on social engineering but 0.39 on tool hijacking. Struct-only achieves 0.85 on tool hijacking but 0.67 on social engineering. Gated fusion provides balance (0.89 / 0.60). Vocabulary Size. Structural tokenization uses only 9 tokens versus 26 for conversational, yet achieves superior OOD performance. This suggests that abstraction levelâ not vocabulary richnessâdrives generalization. Aggregation Methods. Local, FedAvg [25], and en- semble methods produce identical results (within Âą0.02), confirming representation is the bottleneck. This aligns with findings on non-IID challenges in federated learn- ing [19, 22, 48]. D. Failure Analysis Unknown Category Inversion. With conversational to- kenization, unknown attacks achieve AUC 0.26 (below ran- dom). Investigation reveals systematic inversion: the model assigns higher attack probability to benign samples (mean 0.97) than actual attacks (mean 0.69). The model learned âunfamiliar = dangerous,â but unknown attacks are also un- familiar, causing ranking reversalâa failure mode consis- tent with OOD generalization challenges [14, 45]. Structural Recovery. Structural tokenization recovers to AUC 0.97 because execution patterns remain discrimina- tive regardless of linguistic novelty. E. Structural Tokenization Algorithm F. Limitations and Future Work Tokenizer Robustness and Adaptive Adversaries. Our structural tokenizer uses deterministic, rule-based pattern matching (Algorithm 1). This design choice was inten- tional: it allows us to isolate the effect of structural ab- straction itself without conflating it with tokenizer sophis- tication. However, we acknowledge that such simplicity in- Algorithm 1 Structural Tokenization Ď struct (Ď) Require: Agent trace Ď = (M,T ,R) 1: S â [] 2: for each turn in conversation do 3:if turn is system message then 4:S.append([SYS]) 5:else if turn is user message then 6:S.append([USER]) 7:else if turn is assistant message then 8:S.append([ASSISTANT]) 9:if turn contains tool call then 10:S.append([TOOL]) 11:S.append([ARGS]) 12:end if 13:else if turn is tool observation then 14:S.append([OBS]) 15:end if 16: end for 17: S.append([OUTPUT]) 18: return PAD(S,L max ) troduces an evasion surface. A motivated adversary could attempt to fragment tool calls across turns, pad traces with benign operations, or deliberately mimic common benign execution flows to reduce anomaly scores. While this is an important limitation, it does not un- dermine the central finding of this paper: structural exe- cution patterns carry threat-relevant signal that conversa- tional features alone miss. Our architecture is modular, and the tokenizer can be replaced without changing the de- tection model. The BiLSTM operates over abstract token sequences, making it agnostic to tokenization strategy. In practice, we expect stronger robustness from richer tok- enization strategies such as: (i) state-transition modeling over execution graphs rather than isolated events, (i) se- mantic enrichment of tool calls with access scope and sen- sitivity labels, (i) stochastic abstraction to reduce pre- dictability, and (iv) learned latent encoders trained con- trastively over benign and malicious traces [9]. We also note that evasion attempts often introduce sec- ondary anomalies (e.g., entropy spikes, abnormal sequenc- ing, or workflow divergence), suggesting that attempts to hide execution semantics may themselves become de- tectable. A full adversarial evaluationâwhere attackers ex- plicitly optimize to evade the representation layerâremains essential future work. We are actively pursuing red-team evaluations and partnerships to enable adversarial testing against production agents. Synthetic Data. Our traces are programmatically gener- ated and may not capture real attacker creativity or produc- tion noise. Real-world validation through industry partner- ships is critical. Static Tokenization.Rule-based extraction may be evaded by sophisticated attackers. Learned tokenization via neural encoders [9, 39] is a promising direction. Attack Coverage. Our taxonomy covers 5 families. Emerging attacks like multi-agent exploits and memory poi- soning require study. Gated Fusion Tradeoff. While gated fusion excels on linguistic attacks, it underperforms struct-only on structural attacks (0.60 vs 0.85). Improved fusion mechanisms such as attention-based routing [39] or sparse expert selection [36] are needed. Byzantine Robustness. Our federated setting assumes honest-but-curious participants. Malicious clients could at- tempt model poisoning [3]; defenses such as Byzantine- tolerant aggregation [6] or backdoor detection [28] would be required for production deployment. Privacy Guarantees. While we employ DP-SGD and secure aggregation, determined adversaries may still at- tempt membership inference [38] or model inversion [12]. Stronger privacy budgets (Îľ < 1) or additional defenses may be needed for highly sensitive deployments. G. Reproducibility Implementation. PyTorch 2.0.1 with Opacus 1.4.0 for differential privacy. All experiments run on consumer hard- ware (Apple M2 Pro, 16GB RAM) without GPU. Hyperparameters. Learning rate 0.001, batch size 32, 5 federated rounds, 5 local epochs, embedding dimension 64, LSTM hidden size 64, latent dimension 32. Code and Data. Synthetic trace generation scripts and training code will be released upon publication.