Paper deep dive
Large Language Models Do Not Always Need Readable Language
Jiayi Zhu, Haoxuan Peng, Junxi Wang, Liang Ke, Chen Zhang, Linfeng Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 6/21/2026, 4:42:44 AM
Summary
The paper introduces BabelTele, a class of model-centric, high-density textual representations designed for Large Language Models (LLMs). Unlike traditional natural language compression that prioritizes human readability, BabelTele relaxes linguistic constraintsβusing symbols, emojis, and multilingual fragmentsβto maximize semantic density. The research demonstrates that while BabelTele significantly reduces human readability and natural language typicality, instruction-tuned LLMs can still recover core semantics with high fidelity. Experimental results show that BabelTele maintains a superior accuracy-retention frontier compared to standard summarization and LLMLingua-2 across benchmarks like QuALITY and MeetingBank. Furthermore, the study finds that BabelTele exhibits robust zero-shot cross-model transferability, meaning representations generated by one model can be effectively interpreted by different model families, suggesting a potential path toward model-native, efficient communication protocols.
Entities (6)
Relation Signals (4)
LLM β caninterpret β BabelTele
confidence 100% Β· while sacrificing human readability... it still preserves semantic structures that can be decoded and used by instruction-tuned LLMs
BabelTele β exhibits β cross-model transferability
confidence 100% Β· BabelTele exhibits robust cross-model transferability across diverse proprietary and open-weight model families
BabelTele β isatypeof β model-centric textual representation
confidence 100% Β· We refer to this class of model-centric textual representations as BabelTele
BabelTele β improves β accuracy-retention frontier
confidence 90% Β· Overall, BabelTele forms a more favorable accuracy-retention frontier across multiple benchmarks
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whether semantic information can be encoded in compact, non-standard textual forms that sacrifice human readability while remaining recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMs' capacity to generate and interpret such representations. Through readability diagnostics, model likelihood measures, human questionnaires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserving core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information density, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communication. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typicality, and model-side semantic recoverability can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems.
Tags
Links
- Source: https://arxiv.org/abs/2606.19857v1
- Canonical: https://arxiv.org/abs/2606.19857v1
Trouble viewing inline? Open PDF directly β
Full Text
100,842 characters extracted from source content.
Expand or collapse full text
Large Language Models Do Not Always Need Readable Language Jiayi Zhu 1 Haoxuan Peng 2 Junxi Wang 3 Liang Ke 4 Chen Zhang 5 Linfeng Zhang 1β 1 Shanghai Jiao Tong University; 2 The University of Sydney; 3 Hefei University of Technology 4 Xiβan Jiaotong University; 5 Nanjing University zhanglinfeng@sjtu.edu.cn Abstract Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investi- gates whether semantic information can be en- coded in compact, non-standard textual forms that sacrifice human readability while remain- ing recoverable by LLMs. We refer to this class of model-centric textual representations as BabelTele, approached here not as a fixed protocol but as an empirical probe into LLMsβ capacity to generate and interpret such repre- sentations. Through readability diagnostics, model likelihood measures, human question- naires, and downstream task evaluations, we find that BabelTele can substantially depart from ordinary natural language while preserv- ing core semantics for instruction-tuned LLMs. As a task-agnostic representational paradigm, BabelTele demonstrates high information den- sity, maintaining 99.5% semantic fidelity even when the text volume is condensed to 27.9% of its original length. We further evaluate its semantic robustness in cross-model transfer, agent memory, and multi-agent communica- tion. Results suggest that BabelTele can reduce context overhead while generally maintaining reliable downstream performance, although its effectiveness depends on the compressor-reader pair and task setting. These findings indicate that human readability, natural-language typ- icality, and model-side semantic recoverabil- ity can be partially decoupled, opening a path toward model-native representations in future exploration of LLM systems. 1 Introduction Large language models (LLMs) have become a dominant interface in contemporary intelligent systems. Since GPT-3 demonstrated strong few- shot generalization through text-based prompting (Brown et al., 2020), the field has largely followed β Corresponding author. I Have Seen the Future of Europe. The Eurocrats were thinking ahead when they made Brussels the \"Ca- pital of Europe,\" headquarters of the emerging European Union. Though practically unknown in the United States, the union is one of Europe's biggest stories, an impor- tant organization trying to establish itself as a sort of metagovernment for European states. ... BabelTeleNatural Lan. Gov/deficits>US. Historic bldgs (Author's14 th-Cβͺ). Pricesο except cheapο·/οΈ. Huge(100 0s US). Ubiquitousο₯/ο«. * *Para5 :* Multilingualism divi- des. South(Wallonia)=FR. North (Flanders)=NL. End:CopSigh ... Whatβs this? LLM Context Window The answer is ... LLM Context Window The answer is ... Q&AQ&A Figure 1: As illustrated, BabelTele representation dif- fers substantially from verbose natural language: the text is significantly more compact, indicating a much higher information density. While the compressed rep- resentation is much less human-readable, it remains well-interpreted by LLMs, which can understand the original meaning without any distortion. a unified paradigm: knowledge is represented in natural language, instructions are issued in natu- ral language, and model outputs are returned in natural language. Subsequent alignment methods and dialogue-oriented models further reinforced this design, optimizing model behavior toward con- trollability, and readability (Ouyang et al., 2022; Touvron et al., 2023). However, natural language optimized for human communication is not necessarily an efficient rep- resentation for model processing. Human language contains substantial redundancy: complete syntax, discourse markers, and narrative coherence all help people follow, remember, and disambiguate infor- mation. These properties are valuable for human readers, but they reduce semantic density. From an information-theoretic view, communication sys- tems aim to transmit information efficiently under channel constraints (Shannon, 1948). More recent work also connects language modeling with com- pression, suggesting that strong language models can capture compact statistical structure in data 1 arXiv:2606.19857v1 [cs.CL] 18 Jun 2026 (Deletang et al., 2024). This raises a central ques- tion: if the receiver is an LLM rather than a human, must semantic information still be encoded in fully human-readable natural language? This question becomes especially relevant in long-context and agentic systems, where context overhead is a persistent bottleneck. LLMs are increasingly deployed to process lengthy docu- ments, maintain memory, and exchange intermedi- ate states in multi-agent workflows, yet models do not always use long contexts robustly (Liu et al., 2024), and agent systems further intensify this pres- sure through natural-language memory (Park et al., 2023a), context management (Packer et al., 2023), and inter-agent message passing (Wu et al., 2024). The idea of sacrificing readability for efficiency, however, is not new: telegraphic language, math- ematical notation, and code all demonstrate that fluent prose is only one possible surface form for communication. This suggests a broader possibility that intermediate representations in LLM systems may be optimized not for human readability, but for model decodability and semantic density. Existing context-compression methods provide important precedents. Prior work has explored reducing prompt length by removing redundant spans, rewriting retrieved content, or selecting informative tokens, demonstrating that natural- language inputs contain substantial compressible redundancy (Li et al., 2023; Jiang et al., 2023a; Xu et al., 2024). Nevertheless, most such methods still operate within the conventions of human-readable natural language. A separate line of work explores activation-based or learned internal representations, but these approaches often require additional train- ing, special tokens, or access to hidden states, lim- iting their applicability in black-box API settings and heterogeneous model ecosystems. Motivated by this gap, we investigate BabelTele, a class of model-centric rather than human-centric textual representations. BabelTele explicitly re- laxes human readability as a default constraint, in- stead encouraging models to encode semantics into compact, non-standard textual forms, potentially combining abbreviations, symbols, cross-lingual fragments, and non-standard syntactic structures. While these forms exhibit low human readability, advanced LLMs can still recover their core seman- tics for downstream tasks. We thus present Babel- Tele not as a competitive compression method, but as an empirical probe into model-native tex- tual communication, evaluating it across multiple dimensions including compression ratio, semantic fidelity, human readability, cross-model transfer- ability, and downstream utility in document QA, agent memory compression, and multi-agent com- munication. Our key takeaways are as follows: β’ BabelTele emerges as a prompt-accessible phenomenon: by removing human readability as a default constraint, LLMs spontaneously produce opaque but semantically dense textual representations under a black-box interface. β’BabelTele exhibits robust cross-model trans- ferability across diverse proprietary and open- weight model families in a zero-shot manner. Representations compressed by one model can be reliably interpreted by another without any fine-tuning or model-specific adaptation, suggesting that BabelTele captures a form of semantic encoding that generalizes across het- erogeneous architectures. β’BabelTele retains semantics in document QA, agent memory, and multi-agent communica- tion, demonstrating that current LLMs can process highly dense text without relying on human readability. 2 Related Works 2.1 Prompt and Context Compression Prompt compression has been widely studied as a way to reduce inference cost and improve the effective use of limited context windows (Li et al., 2023; Jiang et al., 2023a, 2024; Pan et al., 2024; Hou et al., 2024). Hard compression methods usu- ally keep the compressed prompt as discrete text. Selective Context (Li et al., 2023) filters tokens or sentences according to self-information. LLMLin- gua (Jiang et al., 2023a) performs coarse-to-fine token-level prompt compression. These approaches mainly reduce prompts by deleting or selecting lex- ical units from the original text. Other work studies abstractive or natural- language prompt compression (Xu et al., 2024; Chuang et al., 2024; Zhang et al., 2024; Jeong et al., 2025; Guo et al., 2025). RECOMP (Xu et al., 2024) compresses retrieved documents into summaries, while Nano-Capsulator (Chuang et al., 2024) learns shorter natural-language capsule prompts. Unlike these methods, BabelTele does not aim to pre- serve natural-language readability. It instead asks whether LLMs can generate and consume compact 2 semantic strings that are opaque to humans but still interpretable by LLMs. We also consider related work on learned and latent context compression, retrieval and memory systems, symbolic prompting and LLM-native com- munication. Due to space limitations, they are pro- vided in Appendix A. 3 Methodology: Eliciting LLM-Native Representations 3.1 Relaxing the Readability Prior Given an input documentx, conventional text com- pression employs a modelCto produce a shorter sequencezthat preserves task-relevant semantics for a reader modelR. Most existing methods im- plicitly encouragezto remain close to the natural language distribution, keeping the compressed text reasonably readable to humans. BabelTele formulates compression as a readability-relaxed semantic projection. When Ris an LLM rather than a human reader, we hypothesize that the human-readability prior can be relaxed. Instead of optimizing for fluency or natural-language typicality, BabelTele prioritizes information density and model-side semantic recoverability. Whilezremains discrete text, its surface form can be deviated from conventional prose. Furthermore, as different LLMs are trained on overlapping linguistic, symbolic and factual structures, such representations may exhibit partial cross-model decodability across model families. 3.2 Principles of Symbolic Collapse Rather than treating BabelTele as a singular, man- ually engineered βprompt trick,β we define it as a family of high-density representations induced by relaxing linguistic constraints. To materialize this phenomenon in a black-box setting, we design instructional probes that encourage the compressor Cto produce model-readable encodings based on the following principles: Omnilingual Lexical Selection:Relaxing single-language constraints and selecting high- density lexical units across languages and scripts. Symbolic Collapse: Replacing verbose linguis- tic structures with compact symbols, emojis, math- ematical/logical operators, and punctuation. Recoverable Semantic Density: Preserving re- coverable semantic details so that capable LLMs can interpret the compressed output without an ex- ternal codebook. By applying these constraints through zero-shot prompting, we encourage LLMs to externalize se- mantic information into a compact surface form more effectively. This approach requires no gra- dient updates or tokenizer modifications, allowing us to investigate the existence and utility of LLM- native textual representations across heterogeneous models. To illustrate this collapse, consider the following micro-example: As shown above, BabelTele relaxes ordinary human syntax and uses multilingual cues, emojis, and re- lational arrows to form a compact, model-readable semantic graph more densely. Appendix D pro- vides a qualitative example with a source excerpt and the complete BabelTele output. 4 Experiments 4.1 Experimental Setup We evaluate BabelTele across multiple long- context benchmarks under a task-agnostic protocol, where the compressor processes source documents without access to downstream questions. We com- pare against standard baselines and report token re- tention, QA accuracy, and reader chain-of-thought token overhead across diverse evaluator model fam- ilies. Full details are in Appendix B. 4.2 Symbolic Collapse: Separating Human Readability from Model Decodability In our early observations, we found that although BabelTele is usually difficult for humans to read directly, LLMs can still use it to answer questions, recover details, and perform a certain degree of reasoning. Therefore, this section mainly exam- ines whether human readability, natural-language distribution typicality, and model semantic recover- ability can be decoupled, rather than focusing on the compression efficiency of BabelTele itself. We select 10 long-text samples from the QuAL- ITY (Pang et al., 2022) dataset, each containing 3 multiple-choice QA items, and construct three input formats for each sample: the original text, a natural-language summary, and a BabelTele repre- sentation generated using the default compression 3 prompt in Appendix C.1. ModelOrig.Summ.BabelTele Llama-3-8B9.6311.32 (+17.5)%176.60 (+1,733.9)% Qwen2-7B10.93 12.44 (+13.8)%236.23 (+2,061.1)% Qwen2.5-7B10.4311.56 (+10.8)%209.30 (+1,906.7)% Qwen3-8B14.9415.22 (+1.9)%301.22 (+1,916.2)% Qwen3-32B11.8011.07 (-6.2)%220.76 (+1,770.8)% Kimi K2P66.937.76 (+11.9)%127.96 (+1,746.5)% DeepSeek V4 Pro7.677.75 (+1.1)%108.95 (+1,320.5)% GLM 5.17.898.40 (+6.5)%148.23 (+1,778.7)% Table 1: PPL diagnostics across representative base models. Parentheses indicate the relative change from the original-text PPL within the same model. Green denotes an increase. Red denotes a decrease. Readability diagnostics in Appendix E.1 show that BabelTele has much lower surface readability than both original passages and natural-language summaries, with a Dale-Chall score of 16.70 and a difficult-word ratio of 80.19%. Table 1 fur- ther shows that BabelTele is highly unlikely under multiple base language models, yielding order-of- magnitude higher PPL than both original and sum- mary texts. Together, these results indicate that Ba- belTele is not merely a natural-language summary or shorthand, but a surface representation that sub- stantially departs from conventional English prose and the ordinary natural-language distribution. Human 0 10 20 30 40 50 60 Accuracy (%) 56.10% 35.80% Gemini-3.1-pro 0 20 40 60 80 100 Accuracy (%) 90.00% 96.70% Original BabelTele Figure 2: QA accuracy for human readers and Gem- ini 3.1 Pro on original and BabelTele inputs. The y-axis starts at 25%, the random-choice baseline for four-option QA. Brackets indicate absolute changes in percentage points. However, Figure 2 shows that low readability and low natural-language likelihood do not imply semantic unrecoverability. The human accuracy results were collected through paid questionnaires distributed to university students. Human read- ers show a QA accuracy drop on BabelTele in- puts, whereas Gemini 3.1 Pro (Google DeepMind, 2026) maintains high accuracy. This suggests that BabelTele is not meaningless gibberish; rather, while sacrificing human readability and natural- language typicality, it still preserves semantic struc- tures that can be decoded and used by instruction- tuned LLMs. More detailed analyses of readability, evaluation metrics, and questionnaire settings are provided in Appendix E.1. 4.3 Efficiency and Cognitive Overhead in Model-Native Compression We focus on two questions. First, under different context retention ratios, how much downstream QA accuracy can BabelTele preserve compared with natural-language summaries and LLMLingua- 2 (Pan et al., 2024)? Second, does compression increase the response chain-of-thought tokens used by the reader model? Together, this section evalu- ates both input-side token savings and answer-side cognitive overhead. Experimental Setup.We evaluate BabelTele on 2,128 QuALITY questions and 2,586 Meeting- Bank (Hu et al., 2023) multiple-choice questions, conducting 116 experimental runs across the two datasets. We contrast simply prompt-elicited Ba- belTele representations with standard abstractive summaries and a carefully engineered extractive baseline (LLMLingua-2) under accuracy-retention curves rather than a single compression point, since compression ratios vary across generative methods. For BabelTele and summary baselines, we use the same-model setting, For BabelTele and summary baselines, we use a same-model setting, where each model reads its own compressed output. BabelTele is evaluated with multiple BabelTele- like prompt variants, so the experiment tests a fam- ily of model-readable high-density representations rather than a prompt artifact; additional sweep and prompt details are provided in Appendix E.2. The Accuracy-Retention Frontier . Figure 3 summarizes the relation between accuracy and context retention on QuALITY and Meeting- Bank. Overall, BabelTele forms a more favorable accuracy-retention frontier across multiple bench- marks and reader models. We summarize the fol- lowing three points: (i) BabelTele maintains higher accuracy un- der strong compression. As compression intensi- fies, summary and LLMLingua-2 show sharper ac- curacy degradation, whereas BabelTele maintains relatively high downstream QA accuracy. This suggests that BabelTele can better preserve task- relevant semantics while reducing input tokens. 4 BabelTeleSummaryLLMLingua-2No-compression baseline Circle area: avg. thought-chain tokens 17225104848 Relative accuracy MeetingBank Β· Gemini 0.65 0.75 0.85 0.95 1.00 30405060708090100 Token reduction (%) MeetingBank Β· Qwen 30405060708090100 Token reduction (%) Quality Β· Gemini 80859095100 0.65 0.75 0.85 0.95 1.00 Token reduction (%) Quality Β· Qwen 556065707580859095100 Token reduction (%) Figure 3: Accuracy-retention comparison with response chain-of-thought token scale. Each panel corresponds to one benchmark-reader setting. Token reduction is computed as one minus the realized context retention ratio, relative accuracy is normalized by the corresponding no-compression baseline, and circle area denotes the average number of response thought-chain tokens. Dashed colored curves indicate method-level fitted trends. (i) BabelTeleβs robustness is more evident on MeetingBank. For both Gemini 3.1 Pro and Qwen- 3.5-Plus readers, BabelTele preserves near-original performance even with substantial token reduction. On QuALITY, the advantage is more moderate, but BabelTele still maintains stable relative accuracy across the evaluated compression range. This indi- cates that the benefit of BabelTele is not limited to a single task format, but applies to both meeting- style records and long-document QA. (i) Multiple prompt variants support the family-level interpretation of BabelTele. Ba- belTele does not rely on a single carefully opti- mized prompt. Instead, BabelTele-like prompts with different surface biases collectively trace a strong accuracy-retention frontier. This suggests that BabelTele is better understood as a family of model-readable high-density compressed represen- tations rather than a single prompt trick. Does Compression Make Models Think More? A natural concern is whether compression leads to longer response thought chains. In particular, one may ask whether the reader model needs to first de- code BabelTele back into natural language before answering. Figure 4 examines the relation between BabelTeleSummaryLLMLingua-2No-compression baseline Thought-token multiplier 0%10%20%30%40%50%60%70% 0x 1x 2x 3x 4x 5x Context retention ratio Figure 4: Response chain-of-thought token multiplier versus realized context retention ratio. Each marker corresponds to one compression run. Colors indicate compression methods, the y-axis is normalized by the corresponding no-compression baseline, and the hori- zontal dashed line marks1Γchain-of-thought token us- age. Solid colored curves show method-level smoothed trends. context retention and response chain-of-thought to- ken usage across all experimental settings, from which we summarize the following three points: (i) Stronger compression often increases chain-of-thought tokens. As context retention decreases, chain-of-thought tokens generally rise 5 across methods. (i) BabelTele does not introduce unique over- head. Its chain-of-thought token multiplier is of- ten comparable to or lower than summary and LLMLingua-2, and can even fall below the original- context baseline in mild-compression settings. (i) Thought-token growth reflects evidence accessibility. A cautious interpretation is that longer thought chains arise when compression re- duces the completeness or accessibility of retained evidence. If relevant details are removed or harder to use, the reader model may need more reasoning steps to reconstruct the answer or resolve uncer- tainty. Thus, BabelTele may introduce decoding cost, but because it often preserves task-relevant structure, its answer-side overhead is not necessar- ily higher than that of summary or LLMLingua-2. Summary. Taken together, these results show that the model-side recoverability observed in Sec- tion 4.2 can translate into long-context compression gains. BabelTele exhibits a favorable accuracy- retention frontier across the QA settings, while the response chain-of-thought token analysis reveals an important efficiency boundary: extreme con- text compression triggers a space-time trade-off where input token savings are partially offset by CoT generation overhead. Importantly, this dy- namic is intrinsic to LLM reasoning over sparse ev- idence rather than unique to BabelTele, suggesting that choosing an optimal, moderate compression ratio can yield simultaneous savings in both input context and reasoning steps. 4.4 A Universal Cipher? Zero-Shot Cross-Model Comprehension To examine whether BabelTele represents a model- private shorthand or a transferable symbolic form, we evaluate its cross-model portability across state- of-the-art LLMs. We conduct compression-reading experiments on 180 samples from LongBench v2- Short and 214 QA instances from QuALITY, and analyze whether the effectiveness of BabelTele de- pends on specific compressor-reader pairs. For a more detailed discussion of experimental setup, please refer to Appendix E.3. (i) Compression Ratios of Different Mod- els. Figure 5 shows that BabelTele compression strength varies substantially across compressors. Gemini 3.1 Pro is the most aggressive, exceeding 95% compression, whereas GPT-5.4 is more con- servative at roughly 75%; the other models fall 0 5 10 15 20 25 30 Retention ratio (%) 4.01% 27.26% 12.24% 10.00% 19.15% LongBench v2 0 10 20 30 40 50 Retention ratio (%) 18.07% 41.02% 11.69% 6.73% 14.61% 27.90% QuALITY Gemini GPT-5.4 Qwen Kimi DeepSeek Doubao Claude Figure 5: Comparison of compression rates of differ- ent models. We selected 180 samples from the Short subset of the LongBench v2 benchmark and processed them using BabelTele. LongBench Cross-Model Accuracy Matrix Retained accuracy 77.6%100%109.3% Compression model Answering model BaselineGeminiGPTQwenKimiDoubao Gemini GPT Qwen DeepSeek Kimi Doubao 66.11 100.00% 66.11 100.00% 72.22 109.24% 68.33 103.32% 62.78 94.92% 63.33 95.77% 66.67 100.00% 58.33 87.46% 66.67 100.00% 62.78 94.16% 56.11 84.15% 62.22 93.33% 69.44 100.00% 53.89 77.61% 65.00 93.62% 58.33 84.01% 53.89 77.61% 58.89 84.78% 61.11 100.00% 56.11 91.73% 62.78 102.73% 56.67 92.69% 51.67 84.53% 61.11 99.99% 62.78 100.00% 55.00 87.59% 66.11 105.33% 57.78 92.04% 58.33 92.88% 60.56 96.50% 66.67 100.00% 56.11 84.14% 66.67 100.00% 63.89 95.82% 59.44 89.09% 65.56 98.25% Figure 6: Accuracy transfer matrix on 180 samples from the Short subset of LongBench v2. Each row denotes answering models and columns denote com- pression models, with the Baseline column indicating no compression. Cell color summarizes retained perfor- mance relative to the no-compression baseline for the same answering model, and each cell reports absolute accuracy with retained performance shown below. between 80% and 90%. This variation is impor- tant for interpreting the transfer matrices below, since higher compression can make the resulting symbolic form harder for other models to decode. (iI) Cross-Model Compression Accuracy. Fig- ures 6 and 7 show that cross-model BabelTele com- prehension is not dataset-specific. Across both LongBench v2 and QuALITY, compressed inputs remain usable for heterogeneous evaluators, but re- tention is strongly shaped by the compressor-reader pair. In particular, QuALITY shows that GPT-5.4 and Claude compressed inputs are broadly portable, while Qwen and Kimi compressed inputs lead to larger accuracy drops across readers. Thus, Babel- Tele is not a fully universal code, but its portability is systematic: strong compressors can produce sym- bolic forms that many models still understand. (i) Cross-Model Inference Chain-of-Thought Length. Figures 8 and 9 report response thought to- kens as a proxy for reader-side decoding overhead. 6 QuALITY Cross-Model Accuracy Matrix Retained accuracy 73.8%86%100% Compression model Answering model BaselineDeepSeekGeminiGPTClaudeQwenKimi DeepSeek Gemini GPT Claude Qwen Kimi 94.39 100.00% 78.50 83.17% 83.18 88.12% 90.19 95.55% 90.19 95.55% 79.44 84.16% 69.63 73.77% 97.20 100.00% 88.32 90.86% 88.32 90.86% 93.46 96.15% 92.99 95.67% 85.51 87.97% 77.10 79.32% 94.39 100.00% 82.71 87.63% 83.18 88.12% 92.52 98.02% 90.65 96.04% 75.70 80.20% 70.56 74.75% 94.86 100.00% 84.58 89.16% 84.11 88.67% 94.39 99.50% 94.39 99.50% 80.84 85.22% 76.64 80.79% 94.39 100.00% 77.57 82.18% 83.64 88.61% 92.52 98.02% 86.92 92.09% 74.30 78.72% 70.09 74.26% 92.99 100.00% 80.37 86.43% 79.14 85.11% 89.25 95.98% 90.19 96.99% 71.96 77.38% 68.69 73.87% Figure 7: Accuracy transfer matrix on QuALITY. The matrix follows the same answering-model by compression-model layout and encoding as Figure 6. LongBench Chain-of-Thought Token Matrix Relative Chain-of-Thought Tokens 59%100%184% Compression model Answering model BaselineGeminiGPTQwenKimiDoubao Gemini GPT Qwen DeepSeek Kimi Doubao 6,268 100.00% 6,874 109.66% 4,692 74.81% 3,727 59.45% 3,863 61.61% 6,492 103.57% 754 100.00% 1,120 148.53% 850 112.77% 1,164 154.33% 1,375 182.36% 1,006 133.42% 4,625 100.00% 5,101 110.28% 5,166 111.72% 4,583 99.09% 4,893 105.84% 4,997 108.05% 2,136 100.00% 3,854 180.37% 3,182 149.01% 3,384 158.40% 3,684 172.48% 3,348 156.70% 2,702 100.00% 4,966 183.77% 3,526 130.49% 3,962 146.61% 4,721 174.71% 3,987 147.56% 1,627 100.00% 1,703 104.70% 1,534 94.30% 1,396 85.83% 1,471 90.39% 1,426 87.63% Figure 8: Response chain-of-thought token transfer matrix on 180 samples from the Short subset of Long- Bench v2. Rows denote answering models and columns denote compression models, with the Baseline column indicating no compression. Cell color summarizes the response chain-of-thought token ratio relative to the no- compression baseline for the same answering model, and each cell reports chain-of-thought tokens with rela- tive ratio shown below. Compressed inputs often increase this overhead, but the effect varies across evaluator-compressor pairs. This should be interpreted together with the compression-ratio results in Figure 5: more ag- gressive compressors may require reader models to spend more reasoning steps reconstructing or locat- ing relevant evidence. Thus, longer thought chains do not necessarily indicate failed cross-model com- prehension, but may partly reflect the general cost of higher compression. We therefore treat these values as auxiliary runtime evidence rather than direct evidence of semantic understanding. (iv) Scale Sensitivity within the Qwen Family. To test whether BabelTele comprehension simply improves with model scale, we fix the compres- sor to Gemini 3.1 Pro and vary the Qwen-family evaluator. As shown in Table 2, the Quality drop remains within a relatively narrow range, from 10.75 to 14.95 percentage points, and does not im- prove monotonically with model size. For exam- QuALITY Chain-of-Thought Token Matrix Relative Chain-of-Thought Tokens 100%260%527% Compression model Answering model BaselineDeepSeekGeminiGPTClaudeQwenKimi DeepSeek Gemini GPT Claude Qwen Kimi 473 100.00% 1,195 252.54% 1,613 340.87% 773 163.43% 628 132.82% 1,653 349.36% 2,493 526.85% 596 100.00% 964 161.66% 1,016 170.44% 711 119.29% 819 137.41% 1,163 195.05% 1,621 271.89% 59 100.00% 101 170.35% 85 143.71% 65 110.27% 63 106.79% 124 209.90% 213 360.00% 116 100.00% 214 185.43% 173 149.76% 138 119.75% 152 131.16% 263 227.74% 298 257.76% 1,702 100.00% 2,034 119.51% 2,258 132.64% 1,808 106.23% 1,774 104.21% 2,393 140.59% 2,640 155.12% 1,068 100.00% 2,223 208.19% 2,392 224.10% 1,349 126.36% 1,376 128.87% 2,603 243.84% 2,931 274.51% Figure 9: Response chain-of-thought token transfer matrix on QuALITY. The matrix follows the same answering-model by compression-model layout and vi- sual encoding as Figure 8. ModelOrigin BabelTele Drop Retention Qwen3-14B78.97%64.02%-14.9581.07% Qwen3.5-27B91.59%80.84%-10.7588.26% Qwen3.5-35B-A3B90.65%75.70%-14.9583.51% Qwen3.5-397B-A17B93.93%79.91%-14.0285.07% Qwen3.6-Max-Preview90.19%79.44%-10.7588.08% Table 2: Quality of Qwen-family models on original inputs and Gemini-induced BabelTele inputs. The compressor is fixed to Gemini 3.1 Pro, while the eval- uator varies across Qwen models. Drop is reported in percentage points. ple, Qwen3.5-397B-A17B has the highest original Quality but lower BabelTele Quality than Qwen3.5- 27B. This suggests that BabelTele comprehension is not explained by scale alone, but also depends on model-specific robustness to the compressor- induced symbolic form. 4.5 Capabilities and Boundaries in Downstream Tasks 4.5.1 Multi-Agent Communication We evaluate BabelTele representation under two multi-agent regimes: a homogeneous setting (both agents using Gemini 3.1 Pro) to evaluate whether a model can produce and consume its own compres- sion, and a heterogeneous setting (Gemini 3.1 Pro paired with GPT-5.4) to test cross-model portability as a black-box communication protocol. Table 3 presents the final results, from which we summarize the following two points: (i) Signifi- cant token reduction. BabelTele substantially re- duces inter-agent communication overhead in both homogeneous and heterogeneous settings. This demonstrates that model-native compressed mes- sages can effectively lower context consumption during repeated message passing, which is espe- cially important for long-horizon multi-agent tasks. 7 SettingToken ReductionScore Homogeneous38.96%96.6% Heterogeneous44.21%99.7% Table 3: Performance of BabelTele in multi-agent communication settings. Token Reduction denotes the proportion of tokens saved relative to uncompressed communication. Score is reported as a percentage of the uncompressed baseline. MethodToken count (β)Acc. (β)% (β) Original2819.564.81100.00% Summary1365.661.0594.20% BabelTele1382.262.5396.48% Table 4: Performance on the LoCoMo benchmark. Token count denotes the average total number of tokens consumed per query. The best result is highlighted in bold black font. (i) Stable score maintaining. Despite the strong compression, BabelTele maintains competitive fi- nal scores with only negligible performance degra- dation. This suggests that although the compressed messages are less readable to humans, they still preserve sufficient actionable information for LLM agents to coordinate and complete the task. 4.5.2 Performance on Agent Memory We evaluate BabelTele on the representative Lo- CoMo agent memory benchmark using Gemini 3.1 Pro for compression and answering, with GPT-4o- mini as the evaluator. Full experimental details are provided in Appendix E.4.1. Table 4 presents the final results, from which we can summarize the following three points: (i) Ro- bust memory retention. Compared with the origi- nal uncompressed text, BabelTele representations incur minimal downstream accuracy loss, while preserving more actionable details than standard large language model summarization. (i) Lower compression rate. Compared with other experi- ments, the compression rate of BabelTele on the LoCoMo benchmark is relatively low, around 50%. This is likely due to the short length of each session, which contains only about 700 tokens. Future work could explore experiments on larger agent memory benchmarks with longer sessions. 4.5.3 Extending the Context Window We further evaluate BabelTele when the original input exceeds the model context window, using the Code Repo QA Long subset of LongBench v2. We MethodQwen3.6 MaxGLM-5.1Kimi2.5 Original55.1762.0744.82 BabelTele62.0772.4148.28 Table 5: Performance on the Code Repo QA Long subset from LongBench v2 benchmark. The best result is highlighted in bold black font. compare direct truncation against BabelTele-based chunk compression across different models, with full details provided in Appendix E.4.2. Table 5 presents the results. On Qwen3.6-Max, the original truncated input achieves an accuracy of 55.17%, while the dense BabelTele representa- tions allow the model to capture broader evidence, achieving an accuracy of 62.07%. This result indicates that when the original text exceeds the modelβs context window, direct truncation discards a large amount of potentially important informa- tion, thereby impairing the modelβs understanding and reasoning ability. In contrast, BabelTele effec- tively condenses the core information from ultra- long texts, allowing the model to receive more com- plete content within its limited context window and thus mitigating the limitations caused by insuffi- cient context capacity. 5 Conclusion This paper investigates BabelTele, a high-density textual representation optimized for model decod- ability rather than human readability. Our experi- ments demonstrate that BabelTele achieves strong compression ratios with negligible performance degradation, while remaining semantically recover- able by LLMs. Notably, this capability generalizes across a diverse set of proprietary and open-weight models in a zero-shot manner, suggesting that the ability to interpret such representations is a gen- eral capability of LLMs rather than an artifact of any particular model. In practical scenarios includ- ing multi-agent communication and agent memory, BabelTele shows promising potential as a model- native intermediate representation. We therefore view BabelTele not as a finished protocol, but as evidence that high-density textual representations optimized for LLM-to-LLM communication need not prioritize human readability, and as a direction worth further exploration. 8 Limitations Our current evaluation focuses on a selected set of benchmarks and model families, the behavior of BabelTele across a broader range of tasks and architectures remains to be explored. Additionally, as an empirical study, this work primarily char- acterizes the phenomenon rather than explaining its underlying mechanisms, a deeper theoretical understanding of how LLMs form and interpret model-native representations is left to future work. Ethics Statement We acknowledge that all authors are informed about and adhere to the ACL ARR Code of Ethics and the Code of Conduct. Risks Our benchmarks are sourced from publicly avail- able datasets. We cannot guarantee that they are free of biased, toxic, or otherwise harmful con- tent. In addition, BabelTele transforms text into a compact, non-standard representation, which may alter the behavior of the original text in unexpected ways; when applied to safety-critical domains, such changes could compromise safety or introduce un- intended risks. We used LLM-based AI tools only for grammar and language polishing; all technical content, experiments, and claims were written and verified by the authors. References Anthropic. 2026. Introducing Claude Sonnet 4.6. Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. LongBench: A bilingual, multi- task benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 3119β3137, Bangkok, Thailand. Association for Computational Linguistics. Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xi- aozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2025. Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pages 3639β3664. Association for Computational Linguistics. Luiz C. Borro, Luiz A. B. Macarini, Gordon Tindall, Michael Montero, and Adam B. Struck. 2026. Mem- ori: A persistent memory layer for efficient, context- aware LLM agents. CoRR, abs/2603.19935. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Confer- ence on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual. Bytedance Seed. 2026. Seed2.0: Towards intelligence frontier for real-world complexity. Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. 2023. Adapting language models to compress contexts. In Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6- 10, 2023, pages 3829β3846. Association for Compu- tational Linguistics. Yu-Neng Chuang, Tianwei Xing, Chia-Yuan Chang, Zirui Liu, Xun Chen, and Xia Ben Hu. 2024. Learn- ing to compress prompt in natural language formats. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, pages 7756β7767. Associ- ation for Computational Linguistics. DeepSeek-AI. 2025. Deepseek-r1: Incentivizing rea- soning capability in llms via reinforcement learning. CoRR, abs/2501.12948. DeepSeek-AI. 2026. Deepseek-v4: Towards highly efficient million-token context intelligence. Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christo- pher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness. 2024. Language modeling is com- pression. In The Twelfth International Conference on Learning Representations. Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. 2025a. Deepresearch bench: A comprehensive benchmark for deep research agents. arXiv preprint. Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Bo Zheng, Wei Chen, and Haochao Ying. 2025b. Enabling agents to communicate en- tirely in latent space. CoRR, abs/2511.09149. Jakob N. Foerster, Yannis M. Assael, Nando de Fre- itas, and Shimon Whiteson. 2016. Learning to com- municate with deep multi-agent reinforcement learn- ing. In Advances in Neural Information Processing 9 Systems 29: Annual Conference on Neural Informa- tion Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages 2137β2145. Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Gra- ham Neubig. 2023. PAL: Program-aided language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 10764β10799. PMLR. Tao Ge, Jing Hu, Lei Wang, Xun Wang, Si-Qing Chen, and Furu Wei. 2024. In-context autoencoder for con- text compression in a large language model. In The Twelfth International Conference on Learning Rep- resentations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Hao- ran Wang, and 168 others. 2026. Glm-5: from vibe coding to agentic engineering.Preprint, arXiv:2602.15763. Google DeepMind. 2026. Gemini 3.1 pro. Shuyu Guo, Shuo Zhang, and Zhaochun Ren. 2025. En- hancing RAG efficiency with adaptive context com- pression. In Findings of the Association for Compu- tational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 24061β24076. Associa- tion for Computational Linguistics. Serhii Havrylov and Ivan Titov. 2017. Emergence of language with multi-agent games: Learning to com- municate with sequences of symbols. In 5th Inter- national Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Work- shop Track Proceedings. OpenReview.net. Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024. Kvquant: To- wards 10 million context length LLM inference with KV cache quantization. In Advances in Neural In- formation Processing Systems 38: Annual Confer- ence on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. Haowen Hou, Fei Ma, Binwen Bai, Xinxin Zhu, and Fei Richard Yu. 2024. Enhancing and accelerating large language models via instruction-aware contex- tual compression. CoRR, abs/2408.15491. Yebowen Hu, Tim Ganter, Hanieh Deilamsalehy, Franck Dernoncourt, Hassan Foroosh, and Fei Liu. 2023. Meetingbank: A benchmark dataset for meeting sum- marization. In Proceedings of the 61st Annual Meet- ing of the Association for Computational Linguistics (ACL), Toronto, Canada. Association for Computa- tional Linguistics. Mordatch Igor and Abbeel Pieter. 2018. Emergence of grounded compositional language in multi-agent populations. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thir- tieth Innovative Applications of Artificial Intelli- gence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAIβ18/IAAIβ18/EAAIβ18. AAAI Press. Yeonseok Jeong, Jinsu Kim, Dohyeon Lee, and Seung- won Hwang. 2025. ECoRAG: Evidentiality-guided compression for long context RAG. In Findings of the Association for Computational Linguistics: ACL 2025, pages 26607β26628, Vienna, Austria. Associa- tion for Computational Linguistics. Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023a. LLMLingua: Compress- ing prompts for accelerated inference of large lan- guage models. In Proceedings of the 2023 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 13358β13376, Singapore. Association for Computational Linguistics. Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dong- sheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024. LongLLMLingua: Accelerating and enhanc- ing LLMs in long context scenarios via prompt com- pression. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 1658β1677, Bangkok, Thailand. Association for Computational Linguistics. Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Xin Zhao, and Ji-Rong Wen. 2023b. Structgpt: A gen- eral framework for large language model to reason over structured data. In Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, EMNLP 2023, Singapore, Decem- ber 6-10, 2023, pages 9237β9251. Association for Computational Linguistics. Kimi Team. 2026a. Kimi K2.5: visual agentic intelli- gence. CoRR, abs/2602.02276. Kimi Team. 2026b. Kimi k2.6: Advancing open-source coding. Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Hein- rich KΓΌttler, Mike Lewis, Wen-tau Yih, Tim Rock- tΓ€schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge- intensive NLP tasks. In Advances in Neural In- formation Processing Systems 33: Annual Confer- ence on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual. Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023. Compressing context to enhance inference ef- ficiency of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natu- ral Language Processing, pages 6342β6353, Singa- pore. Association for Computational Linguistics. 10 Zeju Li, Yizhou Zhou, and Qiang Xu. 2026. Latent con- text compilation: Distilling long context into compact portable memory. CoRR, abs/2602.21221. Manlai Liang, Mandi Liu, Jiangzhou Ji, Huaijun Li, Haobo Yang, Yaohan He, and Jinlong Li. 2025. Ilre: Intermediate layer retrieval for context compression in causal language models. CoRR, abs/2508.17892. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paran- jape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language mod- els use long contexts. Transactions of the Association for Computational Linguistics, 12:157β173. Llama Team. 2024. The llama 3 herd of models. CoRR, abs/2407.21783. Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024.Evaluating very long-term conversational memory of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 13851β13870. Association for Computational Lin- guistics. Samuele Marro, Emanuele La Malfa, Jesse Wright, Guo- hao Li, Nigel Shadbolt, Michael J. Wooldridge, and Philip Torr. 2024. A scalable communication proto- col for networks of large language models. CoRR, abs/2410.11905. Jesse Mu, Xiang Li, and Noah D. Goodman. 2023. Learning to compress prompts with gist tokens. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Pro- cessing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. OpenAI. 2026. Introducing gpt-5.4. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welin- der, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instruc- tions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Confer- ence on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, and Joseph E. Gonzalez. 2023. Memgpt: Towards llms as operating systems. CoRR, abs/2310.08560. Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor RΓΌhle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. 2024. LLMLingua-2: Data dis- tillation for efficient and faithful task-agnostic prompt compression. In Findings of the Association for Com- putational Linguistics: ACL 2024, pages 963β981, Bangkok, Thailand. Association for Computational Linguistics. Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, and Samuel R. Bowman. 2022. Quality: Question answering with long input texts, yes! In Proceed- ings of the 2022 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 5336β5358. Association for Computational Linguistics. Joon Sung Park, Joseph OβBrien, Carrie Jun Cai, Mered- ith Ringel Morris, Percy Liang, and Michael S. Bern- stein. 2023a. Generative agents: Interactive simu- lacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST β23, New York, NY, USA. Association for Computing Machinery. Joon Sung Park, Joseph C. OβBrien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023b. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST 2023, San Francisco, CA, USA, 29 October 2023- 1 November 2023, pages 2:1β2:22. ACM. Qwen Team. 2026a. Qwen3.5: Towards native multi- modal agents. Qwen Team. 2026b. Qwen3.6-Max-Preview: Smarter, sharper, still evolving. Qwen Team. 2026c. Qwen3.6-Plus: Towards real world agents. Vignav Ramesh and Kenneth Li. 2025. Communicat- ing activations between language model agents. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, Proceedings of Machine Learning Re- search. PMLR / OpenReview.net. Tobias Schnabel and Jennifer Neville. 2024. Sym- bolic prompt program search: A structure-aware ap- proach to efficient compile-time prompt optimization. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Findings of ACL, pages 670β 686. Association for Computational Linguistics. C. E. Shannon. 1948. A mathematical theory of com- munication. The Bell System Technical Journal, 27(3):379β423. Hugo Touvron, Louis Martin, Kevin Stone, Peter Al- bert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti 11 Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton- Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, and 49 oth- ers. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288. Ernst van Gassen. 2026. Semantic compression of LLM instructions via symbolic metalanguages. CoRR, abs/2601.07354. Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2024. Autogen: Enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling. Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. RE- COMP: improving retrieval-augmented lms with con- text compression and selective augmentation. In The Twelfth International Conference on Learning Rep- resentations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. 2025. A-MEM: agentic memory for LLM agents. CoRR, abs/2502.12110. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayi- heng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 oth- ers. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Hao- ran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 40 others. 2024a. Qwen2 technical report. arXiv preprint arXiv:2407.10671. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jian- hong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 22 others. 2024b. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. Zhangyue Yin, Qiushi Sun, Cheng Chang, Qipeng Guo, Junqi Dai, Xuanjing Huang, and Xipeng Qiu. 2023. Exchange-of-thought: Enhancing large lan- guage model capabilities through cross-model com- munication. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process- ing, pages 15135β15153, Singapore. Association for Computational Linguistics. Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Be- rant. 2024. Making retrieval-augmented language models robust to irrelevant context. In The Twelfth International Conference on Learning Representa- tions, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei- Ying Ma, Jingjing Liu, Mingxuan Wang, and Hao Zhou. 2025. Memagent: Reshaping long-context LLM with multi-conv rl-based memory agent. CoRR, abs/2507.02259. Peitian Zhang, Zheng Liu, Shitao Xiao, Ninglu Shao, Qiwei Ye, and Zhicheng Dou. 2025. Long context compression with activation beacon. In The Thir- teenth International Conference on Learning Repre- sentations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. Qianchi Zhang, Hainan Zhang, Liang Pang, Hongwei Zheng, and Zhiming Zheng. 2024. Adacomp: Extrac- tive context compression with adaptive predictor for retrieval-augmented large language models. CoRR, abs/2409.01579. Jiaru Zou, Xiyuan Yang, Ruizhong Qiu, Gaotang Li, Katherine Tieu, Pan Lu, Ke Shen, Hanghang Tong, Yejin Choi, Jingrui He, James Zou, Mengdi Wang, and Ling Yang. 2025. Latent collaboration in multi- agent systems. CoRR, abs/2511.20639. 12 Appendix A Related Works13 A.1Learned and Latent Context Com- pression . . . . . . . . . . . . . .13 A.2 Memory, Retrieval and Long- Context LLM Systems . . . . . .13 A.3Symbolic Representations and LLM-Native Communication . . .14 B Experimental Setup14 B.1 Task-Agnostic Compression Protocol14 B.2 Baselines . . . . . . . . . . . . .14 B.3 Datasets . . . . . . . . . . . . . .14 B.4 Models . . . . . . . . . . . . . .15 B.5 Metrics. . . . . . . . . . . . . . .15 C Prompt Templates15 C.1 BabelTele Compression Prompt .15 C.2 BabelTele-Like Prompt Family Used in Section 4.3 . . . . . . . .15 C.2.1BT-P1:Adaptive Sym- bolic Collapse . . . . . .16 C.2.2BT-P2:Refined Zero- Overhead Compression . .16 C.2.3BT-P3: Minimal Lossless Objective . . . . . . . . .17 C.2.4BT-P4: Structured Om- nilingual Mapping . . . .17 C.2.5 BT-P5:Canonical Omnilingual-Symbolic . .17 C.2.6BT-P6: Structured Map- ping Control . . . . . . .17 C.2.7BT-P7: Canonical Babel- Tele Objective . . . . . .17 C.2.8BT-P8: Fixed Symbolic Mapping Rules . . . . . .18 C.2.9BT-P9: Structured Seman- tic Mapping . . . . . . . .18 C.2.10BT-P10:LLM-Native Compressor . . . . . . . .18 C.2.11BT-P11: Compact Sym- bolic Mapping . . . . . .18 C.2.12 BT-P12: Free-Emergence Attention Checklist . . . .18 C.2.13BT-P13: ASCII Anchor Skeleton . . . . . . . . . .19 D Qualitative Document-Level Example19 E Implementation Details21 E.1Symbolic Collapse: Separating Human Readability from Model Decodability . . . . . . . . . . . .21 E.2Efficiency and Cognitive Overhead in Model-Native Compression . .22 E.3A Universal Cipher? Zero-Shot Cross-Model Comprehension . . .22 E.4Capabilities and Boundaries in Downstream Tasks . . . . . . . .23 E.4.1Performance on Agent Memory . . . . . . . . . .23 E.4.2 Extending the Context Window . . . . . . . . . .23 A Related Works A.1 Learned and Latent Context Compression Another line of work compresses context into learned tokens, memory vectors, or internal activa- tions (Mu et al., 2023; Chevalier et al., 2023; Ge et al., 2024; Zhang et al., 2025; Liang et al., 2025; Hooper et al., 2024; Li et al., 2026). Gist Tokens (Mu et al., 2023) summarize prompts into reusable special tokens, while AutoCompressors (Chevalier et al., 2023) compress long contexts into compact summary vectors as soft prompts. BabelTele differs in its interface assumption: unlike learned-token or activation-level methods that often require training, hidden-state access, special tokens, or architectural changes, it produces discrete text usable through black-box LLM APIs, while not being constrained to natural language. A.2 Memory, Retrieval and Long-Context LLM Systems Long-context LLM applications also motivate com- pact memory and retrieval representations (Lewis et al., 2020; Yoran et al., 2024; Bai et al., 2024; Park et al., 2023b; Packer et al., 2023; Xu et al., 2025; Yu et al., 2025; Borro et al., 2026; Maha- rana et al., 2024). Retrieval-augmented genera- tion prepends external documents to the prompt, but retrieved passages are often verbose and noisy (Lewis et al., 2020; Yoran et al., 2024; Xu et al., 13 2024). Contextual compression for RAG reduces this burden by filtering, summarizing, or restruc- turing retrieved evidence (Xu et al., 2024). LLM agents introduce a related memory bottleneck. Gen- erative Agents (Packer et al., 2023) maintain a natural-language memory stream with reflection and retrieval. MemGPT (Yu et al., 2025) manages working context and external memory through an OS-like memory hierarchy. BabelTele is comple- mentary to these systems: rather than changing the retrieval or memory controller, it proposes a denser representation format for information that will mainly be consumed by LLMs. A.3 Symbolic Representations and LLM-Native Communication BabelTele is related to symbolic and machine-to- machine communication (Foerster et al., 2016; Havrylov and Titov, 2017; Igor and Pieter, 2018; Yin et al., 2023; Marro et al., 2024; Ramesh and Li, 2025; Zou et al., 2025; Du et al., 2025b; van Gassen, 2026; Gao et al., 2023; Jiang et al., 2023b; Schnabel and Neville, 2024). Emergent communication studies non-human-readable pro- tocols, while Exchange-of-Thought (Yin et al., 2023) and Agora (Marro et al., 2024) explore reasoning-trace exchange and scalable communi- cation among LLM agents. Symbolic prompting is another nearby direction. MetaGlyph (van Gassen, 2026) compresses instructions with symbolic met- alanguages, while structured prompting uses non- natural-language formats such as code, JSON, and tables (Gao et al., 2023; Jiang et al., 2023b; Schn- abel and Neville, 2024). BabelTele differs in that it does not rely on a manually designed symbolic lan- guage or a fixed schema. Instead, it studies whether LLMs can be prompted to invent compact, LLM- readable encodings for arbitrary semantic content. B Experimental Setup B.1 Task-Agnostic Compression Protocol For document QA experiments, BabelTele com- pression is performed in advance before down- stream questions are introduced. The compres- sor receives only the source passage or document context, and does not observe questions, answer options, gold answers, or evaluation prompts. B.2 Baselines We compare BabelTele against the original uncom- pressed context, natural-language summaries, and LLMLingua-2 (Pan et al., 2024) under matched set- tings. These conditions separate no-compression performance, human-readable summarization, and learned prompt compression from BabelTele-style model-oriented compression. B.3 Datasets We evaluate BabelTele on both intrinsic compres- sion diagnostics and downstream task performance. QuALITY.QuALITY (Pang et al., 2022) is used as a long-document multiple-choice QA bench- mark. In the pilot setting, we sample 10 long pas- sages, each paired with 3 questions, producing 30 question-answer instances. For each passage, we construct three context variants: the original pas- sage, a natural-language summary, and a BabelTele- compressed representation. In the larger QuAL- ITY evaluation, we further compare BabelTele with LLMLingua-2 and report results by source domain, passage length, and question hardness. LongBench v2.LongBench v2 (Bai et al., 2025) is used to test long-context document QA under stronger scale and cross-model transfer conditions. We evaluate a subset of 180 samples under no compression and multiple BabelTele compression sources. This setting allows us to separate two factors: the model that produces the compressed representation and the model that reads it. LoCoMo. LoCoMo (Maharana et al., 2024) is used as an initial testbed for long-term conversa- tional memory compression. Instead of compress- ing isolated passages, this setting compresses dia- logue histories or memory contexts before answer- ing memory dependent questions. The goal is to test whether BabelTele can reduce agent memory storage while preserving enough reliable informa- tion for downstream recall. DeepResearch. We built a multi-agent system consisting of two agents and tested it on the Deep- Research Bench (Du et al., 2025a), to evaluate Ba- belTele in multi-agent communication. In this set- ting, intermediate messages between agents are compressed before being passed to the next agent. MeetingBank.MeetingBank (Hu et al., 2023) is used to test long-context summarization and ques- tion answering on real-world spoken dialogue. We sample a subset of city council meeting transcripts to evaluate the modelsβ ability to process verbose, 14 multi-party conversations. In this setting, the exten- sive meeting records are compressed before being passed to the target model for downstream tasks. Because MeetingBank serves as the primary train- ing corpus for LLMLingua-2, this evaluation al- lows us to directly benchmark BabelTele against LLMLingua-2 in the baselineβs native domain. B.4 Models Our experiments cover both compression mod- els and reader models. For pilot generation, we use Gemini 3.1 Pro (Google DeepMind, 2026) as the compressor.For cross-model evalua- tion, we include models from mainstream propri- etary and open-weight families, including GPT- 5.4 (OpenAI, 2026), Kimi K2.5 and Kimi K2.6 (Kimi Team, 2026a,b), Meta Llama 3 8B (Llama Team, 2024), Qwen2-7B (Yang et al., 2024a), Qwen2.5-7B (Yang et al., 2024b), Qwen3-8B, Qwen3-14B, Qwen3-32B (Yang et al., 2025), DeepSeek-V4-Pro (DeepSeek-AI, 2026), GLM- 5.1 (GLM-5-Team et al., 2026), Qwen3.6-Plus (Qwen Team, 2026c), DeepSeek-R1 (DeepSeek-AI, 2025), Doubao-Seed-2.0 (Bytedance Seed, 2026), Claude Sonnet 4.6 (Anthropic, 2026), Qwen3.5- 27B, Qwen3.5-35B-A3B, Qwen3.5-Plus, Qwen3.5- 397B-A17B (Qwen Team, 2026a), Qwen3.6-Max- Preview (Qwen Team, 2026b). B.5 Metrics. We evaluate BabelTele along four dimensions. Compression.We report token count and context retention ratio as basic compression statistics. The retention ratio is the length of the compressed con- text divided by the length of the original context, so lower values indicate stronger compression. Semantic Fidelity. We use downstream QA ac- curacy as the primary semantic fidelity metric to assess answer preservation. For each compressed context, the reader model answers the same ques- tions as in the original-context setting. We also report normalized accuracy, defined relative to the no-compression condition, to show how much task performance is retained after compression. Human Readability.To characterize the surface form of BabelTele from complementary perspec- tives, we compute readability, out-of-vocabulary ratio, cross-entropy, and perplexity under several language models. We conduct a human question- naire on a subset, measuring human QA accuracy, perceived difficulty, and completion time. System Utility. For agent memory and multi- agent communication, we focus on operational metrics in realistic interactive settings: context to- ken reduction, task accuracy or success rate, and, where available, response or reasoning-token over- head. These metrics connect intrinsic compression to practical system-level benefits. C Prompt Templates C.1 BabelTele Compression Prompt The following prompt is used to elicit BabelTele representations from the compressor model. The source passage or document context is appended after the final line of the prompt. Unless otherwise specified, this is the default compression prompt used in most experiments in this paper, except for the Section 4.3 prompt-family sweep. your task: compress verbose human text into minimal Token sequence. AudienceΜΈ= human, but another equally intelligent LLM. ,β ,β Core Directive Omnilingual: ignore single-language grammar; traverse all human languages (Chinese, English, German compounds, Japanese Kanji, Latin roots, etc.), pick highest info-density words. ,β ,β ,β ,β Symbolic Collapse: optionally replace conjunctions, emotions, long sentences with Emoji, math/logical symbols (=>,β, ΜΈ=), punctuation. ,β ,β ,β Universality: any LLM should fully understand compressed output without a codebook.,β Lossless: retain all information & details. Compress the content bellow: C.2 BabelTele-Like Prompt Family Used in Section 4.3 Overview.The Section 4.3 retention sweep uses the following BabelTele-like prompt variants. They share the same goal of preserving task-relevant semantics while relaxing human readability, but differ in their surface-form bias and structural con- straints. The full prompt texts are listed after the overview. Source documents are appended after each prompt during compression. For LaTeX porta- bility, non-ASCII symbolic examples are rendered with equivalent ASCII names or operators. Full Prompt Texts. The following are the com- plete prompt instructions for the variants summa- rized in Table 6. 15 IDVariantDesign Bias BT-P1Adaptive Symbolic CollapseFree-form compression with adaptive anchors, omnilingual choices, symbolic collapse, and explicit semantic checklists. BT-P2Refined Zero-Overhead Com- pression Extreme compression with zero-overhead key-value structure, exact- value preservation, and anti-hallucination constraints. BT-P3Minimal Lossless Objective A short high-level instruction that asks for shortest lossless compression without prescribing a fixed schema. BT-P4Structured Omnilingual Map- ping Structured semantic fields combined with cross-lingual token-density selection. BT-P5 CanonicalOmnilingual- Symbolic Canonical BabelTele-style objective with omnilingual lexical selection, symbolic collapse, universality, and losslessness. BT-P6Structured Mapping ControlSchema-like preservation of entities, quantities, math, logic, flow, condi- tions, comparisons, and placeholders. BT-P7Canonical BabelTele ObjectiveCompact statement of the core BabelTele objective used as a canonical prompt form. BT-P8 FixedSymbolicMapping Rules Predefined compact anchors for sections, entities, quantities, logic, hier- archy, conditions, and evaluations. BT-P9Structured Semantic MappingEntity, quantity, math, flow, condition, evaluation, and anti-hallucination mapping rules. BT-P10LLM-Native CompressorRole-based LLM-native compression prompt with omnilingual, sym- bolic, universal, and lossless directives. BT-P11Compact Symbolic MappingCompact symbolic mapping, array flattening, abbreviation definition, and hard-data fidelity. BT-P12Free-EmergenceAttention Checklist Free surface-form emergence constrained by a checklist of semantic dimensions that must remain lossless. BT-P13ASCII Anchor Skeleton Predefined ASCII anchors for modules, entities, parameters, logic, com- parison, and unknown values. Table 6: BabelTele-like prompt variants used to construct the retention sweep in Section 4.3. C.2.1 BT-P1: Adaptive Symbolic Collapse # Role: LLM-Native Semantic Compressor You are participating in frontier research on an "LLM-native high-density communication language." Your task is to compress long text into the absolute shortest possible token sequence. ,β ,β ,β [Highest Directive]: The recipient is an equally intelligent large language model. Completely discard human readability, human grammatical structure, and conventional code/JSON format constraints. ,β ,β ,β # Level 1: Syntactic Anarchy - Pursue Extreme Compression Ratio,β 1. Omnilingual: Move freely across all human languages (Chinese, English, German compounds, Japanese kanji, Latin roots, etc.) and choose the word with the highest single-token information density for the given context. ,β ,β ,β 2. Symbolic Collapse: Heavily use mathematical symbols (forall, exists, in, =>), emoji, and isolated punctuation to replace prepositions, conjunctions, and explanatory long sentences. ,β ,β ,β 3. Adaptive Routing: Do not use fixed format labels such as `Meta:`,`Ent:`, or`[ ]`. Dynamically invent the most token-efficient special single-character separators/anchors for the text you are processing. ,β ,β ,β # Level 2: Semantic Checklist - Pursue Extreme Accuracy Although the format is completely free, during compression you must strongly maintain attention in latent space to the following core information and preserve it losslessly: ,β ,β ,β 1. Entities & Graphs: Accurately bind people/organizations/concepts to their corresponding attributes. Do not confuse ownership or dependency relations. ,β ,β ,β 2. Exact Quantities: Preserve all exact numbers, metrics, mathematical formulas, and hyperparameters verbatim. Estimation or rounding is strictly forbidden. ,β ,β 3. Logic & Boundaries: Clearly preserve conditional branches (If/Then), causal chains, and exceptions.,β 4. Comparisons: Precisely extract multi-target comparison matrices or experimental conclusions.,β 5. Anti-Hallucination: Preserve special placeholders from the original document, such as`BIBREF`. Never invent missing information not mentioned in the source. ,β ,β # Task Combine Level 1's freely extreme compression with Level 2's precise information preservation. Directly output the compressed "adaptive Babel-Telegraph" without any preface. ,β ,β ,β C.2.2 BT-P2: Refined Zero-Overhead Compression # Role: Extreme Data Compressor (LLM-Native Semantic Compressor),β Your task is to compress the following text into the absolute shortest possible token sequence.,β [Warning]: The recipient of this text is another top-tier large language model. Completely abandon human readability. Never preserve any unnecessary format, word, or punctuation for the sake of human reading habits. ,β ,β ,β # Core Strategies 1. Babel Traversal (Omnilingual Density): Break single-language boundaries. Move freely across English, Chinese, Japanese kanji, German compounds, and Latin roots, and force the use of the highest information-density vocabulary for each meaning, meaning the wording that consumes the fewest tokens. ,β ,β ,β ,β ,β 2. Symbolic Collapse: Strictly forbid long English labels such as`Meta`,`Entity`,`Except`, and`Condition`. Use mathematical/logical symbols (forall, exists, in, not-in, intersection, ->, <->, therefore, because), punctuation abbreviations, or emoji to map complex prepositions, logical flow, and causal relations. ,β ,β ,β ,β ,β 3. Zero-Overhead Structure: - Extract entities, attributes, and key-value pairs (`K=V`). Do not wrap them in token-costly JSON/array brackets. Directly connect them compactly with the shortest separators, such as`|`,`^`, or`~`. ,β ,β ,β 16 - Preserve all absolute exact values (formulas, numbers, hyperparameters, matrix relations), but remove all redundant explanatory wording. ,β ,β 4. Lossless Logic: Precisely preserve all macro architecture (`Macro/Meta`), conditional boundaries (`If/Except`), comparative evaluations (`Ref/Matrix`), and placeholders such as`BIBREF`, but express them in the shortest cryptographic-grade form. Hallucinating or inventing missing data is strictly forbidden. Use`NULL` or`?` for unknowns. ,β ,β ,β ,β ,β ,β # Output Format Do not output any preface, explanation, or extra line breaks. Directly output the compressed "Babel-Telegraph.",β C.2.3 BT-P3: Minimal Lossless Objective Compress the following content to the shortest possible extreme.,β Do not lose any information. You do not need to care about human readability at all; only complete information preservation matters.,β You may use symbols from languages across the world to express the content in the simplest possible form. You may freely mix any languages in the world. ,β ,β Only output the compressed text. C.2.4 BT-P4: Structured Omnilingual Mapping # Compress the following content into the absolute shortest possible token sequence. Do not lose any information. You may refer to the following methods. ,β ,β 1. Macro & Meta: Map text to`Sec:[Name->Content]`. Extract `Meta:[K=V]` and define acronyms via `Def:[Term=FullName]` on first use. ,β ,β 2. Entities & Attributes: Bind via`Ent(Attr=Val)`. Flatten parallel items into arrays`[A, B]`. Retain qualitative examples via`Ex:[a, b, c]`. ,β ,β 3. Quantities & Configs: Isolate exact metrics/hyperparameters via `Quant/Config:[Target->K=Val(Unit)]` without rounding or estimation. ,β ,β ,β 4. Math & Logic: Retain all formulas and variables exactly via`Math:[Eq]`. Use (`>,<,=,->,!=`) for relative or causal relations. ,β ,β 5. Flow & Architecture: Map structural pipelines via `Seq:[A>B>C]` and define nested structures via `Arch:[Main->Sub1, Sub2]`. ,β ,β 6. Conditions & Exceptions: Isolate logic via `if[Cond]->[Act]` and define boundaries/exemptions via `Except:[Target->Detail]`. ,β ,β 7. Evaluations & Comparisons: Extract results to `Eval:[Target->Result]`. Use`Matrix:[Ent(X) vs Ent(Y)]` for multi-condition data and`Ref:[A vs B]` for contrasting systems. ,β ,β ,β 8. Anti-Hallucination: Strictly preserve all original placeholders (e.g.,`BIBREF`,`TABREF`). NEVER interpolate missing data; use`@Uncertain` for ambiguous estimates. ,β ,β ,β 9. Break language boundaries (Omnilingual): Completely abandon the grammar of any single language. For extreme token savings, move freely across all human languages (Chinese, English, German compounds, Japanese kanji, Latin roots, etc.) and choose the vocabulary with the highest information density in the given context. ,β ,β ,β ,β ,β Directly output the compressed content. C.2.5 BT-P5: Canonical Omnilingual-Symbolic # Role: Silicon-Based Data Compressor You are participating in frontier research on an "LLM-native high-density communication language.",β Your task is to compress a verbose piece of human text into the absolute shortest possible token sequence. The target audience is not humans, but another large language model as intelligent as you. ,β ,β ,β # Core Directive 1. Omnilingual: Completely abandon the grammar of any single language. For extreme token savings, move freely across all human languages (Chinese, English, German compounds, Japanese kanji, Latin roots, etc.) and choose the words with the highest information density in the given context. ,β ,β ,β ,β ,β 2. Symbolic Collapse: When necessary, use emoji, mathematical/logical symbols (`=>`,`in`,`!=`), and punctuation to replace conjunctions, emotional descriptions, and long sentences. ,β ,β ,β 3. Universality: As much as possible, make the compressed content fully understandable to every large language model, even without a codebook. ,β ,β 4. Losslessness: Do not lose any information or details. 5. Directly output the compressed text and nothing else. # Task Compress the following`[Source Text]` as much as possible into a "Babel-Telegraph.",β C.2.6 BT-P6: Structured Mapping Control # Compress the following content into the absolute shortest possible token sequence. Do not lose any information. You may refer to the following methods. ,β ,β 1. Macro & Meta: Map text to`Sec:[Name->Content]`. Extract `Meta:[K=V]` and define acronyms via `Def:[Term=FullName]` on first use. ,β ,β 2. Entities & Attributes: Bind via`Ent(Attr=Val)`. Flatten parallel items into arrays`[A, B]`. Retain qualitative examples via`Ex:[a, b, c]`. ,β ,β 3. Quantities & Configs: Isolate exact metrics/hyperparameters via `Quant/Config:[Target->K=Val(Unit)]` without rounding or estimation. ,β ,β ,β 4. Math & Logic: Retain all formulas and variables exactly via`Math:[Eq]`. Use (`>,<,=,->,!=`) for relative or causal relations. ,β ,β 5. Flow & Architecture: Map structural pipelines via `Seq:[A>B>C]` and define nested structures via `Arch:[Main->Sub1, Sub2]`. ,β ,β 6. Conditions & Exceptions: Isolate logic via `if[Cond]->[Act]` and define boundaries/exemptions via `Except:[Target->Detail]`. ,β ,β 7. Evaluations & Comparisons: Extract results to `Eval:[Target->Result]`. Use`Matrix:[Ent(X) vs Ent(Y)]` for multi-condition data and`Ref:[A vs B]` for contrasting systems. ,β ,β ,β 8. Anti-Hallucination: Strictly preserve all original placeholders (e.g.,`BIBREF`,`TABREF`). NEVER interpolate missing data; use`@Uncertain` for ambiguous estimates. ,β ,β ,β Directly output the compressed content. C.2.7 BT-P7: Canonical BabelTele Objective Your task: compress verbose human text into a minimal token sequence. The audience is not human, but another equally intelligent LLM. ,β ,β Core Directive Omnilingual: Ignore single-language grammar; traverse all human languages (Chinese, English, German compounds, Japanese kanji, Latin roots, etc.) and pick the highest information-density words. ,β ,β ,β Symbolic Collapse: Optionally replace conjunctions, emotions, and long sentences with emoji, mathematical/logical symbols (`=>`,`in`,`!=`), and punctuation. ,β ,β Universality: Any LLM should fully understand the compressed output without a codebook.,β Lossless: Retain all information and details. 17 Compress the content below: C.2.8 BT-P8: Fixed Symbolic Mapping Rules # Role: LLM-Native Babel Compressor Your task is to compress the following text into a "Babel-Telegraph" with extremely high information density. ,β ,β The audience is an equally intelligent large language model. Completely abandon human readability. Move freely across all human languages (Chinese, English, German compounds, Japanese kanji, etc.) and choose the shortest vocabulary for each meaning. ,β ,β ,β ,β # Structural Mapping Rules You must use the following single-character high-density labels. Long English labels such as`Meta`,`Entity`, and`Except` are strictly forbidden. ,β ,β 1. Macro/Section: Use`S[topic/abbrev]` to define macro modules. On first occurrence,`@[abbrev=full name]` may be used to define abbreviations. ,β ,β 2. Entities & Attributes: Use`*(entity):K=V`. Flatten parallel items as`[A,B,C]`.,β 3. Quantities & Config: Directly extract exact values/parameters using`Config[target]:K=V(unit)`. Never estimate. ,β ,β 4. Math & Logic: Use native mathematical/logical symbols. For relative relations, use`>,<,==,!=,=>,<=>`.,β 5. Flow & Nesting: Use`A>B>C` for pipelines. Use `forallparent:child1,child2` for nesting/hierarchy.,β 6. Conditions & Exceptions: Use`?condition=>action` for conditional actions. Use`!object:detail` for exceptions/boundaries. ,β ,β 7. Evaluation & Comparison: Use`Eval[A/B]:conclusion` for comparison matrices, or the two-dimensional shorthand`A vs B:result`. ,β ,β 8. Anti-Hallucination: Preserve original placeholders such as `BIBREF` verbatim. Strictly use`NULL` or`?` for missing data. ,β ,β Directly output text that follows the above mapping rules and incorporates multilingual extreme compression. Do not output any explanation. ,β ,β C.2.9 BT-P9: Structured Semantic Mapping # Compress the following content into the absolute shortest possible token sequence. Do not lose any information. You may refer to the following methods. ,β ,β 1. Macro & Meta: Map text to`Sec:[Name->Content]`. Extract `Meta:[K=V]` and define acronyms via `Def:[Term=FullName]` on first use. ,β ,β 2. Entities & Attributes: Bind via`Ent(Attr=Val)`. Flatten parallel items into arrays`[A, B]`. Retain qualitative examples via`Ex:[a, b, c]`. ,β ,β 3. Quantities & Configs: Isolate exact metrics/hyperparameters via `Quant/Config:[Target->K=Val(Unit)]` without rounding or estimation. ,β ,β ,β 4. Math & Logic: Retain all formulas and variables exactly via`Math:[Eq]`. Use (`>,<,=,->,!=`) for relative or causal relations. ,β ,β 5. Flow & Architecture: Map structural pipelines via `Seq:[A>B>C]` and define nested structures via `Arch:[Main->Sub1, Sub2]`. ,β ,β 6. Conditions & Exceptions: Isolate logic via `if[Cond]->[Act]` and define boundaries/exemptions via `Except:[Target->Detail]`. ,β ,β 7. Evaluations & Comparisons: Extract results to `Eval:[Target->Result]`. Use`Matrix:[Ent(X) vs Ent(Y)]` for multi-condition data and`Ref:[A vs B]` for contrasting systems. ,β ,β ,β 8. Anti-Hallucination: Strictly preserve all original placeholders (e.g.,`BIBREF`,`TABREF`). NEVER interpolate missing data; use`@Uncertain` for ambiguous estimates. ,β ,β ,β Directly output the compressed content. C.2.10 BT-P10: LLM-Native Compressor # Role: Silicon-Based Data Compressor You are participating in frontier research on an "LLM-native high-density communication language.",β Your task is to compress a verbose piece of human text into the absolute shortest possible token sequence. The target audience is not humans, but another large language model as intelligent as you. ,β ,β ,β # Core Directive 1. Omnilingual: Completely abandon the grammar of any single language. For extreme token savings, move freely across all human languages (Chinese, English, German compounds, Japanese kanji, Latin roots, etc.) and choose the words with the highest information density in the given context. ,β ,β ,β ,β ,β 2. Symbolic Collapse: When necessary, use emoji, mathematical/logical symbols (`=>`,`in`,`!=`), and punctuation to replace conjunctions, emotional descriptions, and long sentences. ,β ,β ,β 3. Universality: As much as possible, make the compressed content fully understandable to every large language model, even without a codebook. ,β ,β 4. Losslessness: Do not lose any information or details. 5. Directly output the compressed text and nothing else. # Task Compress the following`[Source Text]` as much as possible into a "Babel-Telegraph.",β C.2.11 BT-P11: Compact Symbolic Mapping # Compress the following content into the absolute shortest possible token sequence. Do not lose any information. You may refer to the following methods. ,β ,β > 1. Symbolic Mapping: Completely abandon natural-language conjunctions. Use`A->B` for causality/process,`A>B` for containment/comparison, and`Ent(K=V)` for attributes/configurations/results. ,β ,β ,β > 2. Extreme Flattening: Fully use arrays to merge similar items:`Attr:[A, B, C]`. When a long term appears for the first time, immediately define an abbreviation`(Def:X)`. ,β ,β > 3. Hard-Data Fidelity: Absolutely preserve all exact values, formulas, and original placeholders such as `BIBREF`. Mark fuzzy information with`?`; divergent invention is strictly forbidden. ,β ,β ,β Directly output the compressed content. C.2.12 BT-P12: Free-Emergence Attention Checklist # Role: LLM-Native Semantic Compressor You are participating in frontier research on an "LLM-native high-density communication language." Your task is to compress verbose human text into the absolute shortest possible token sequence. The target audience is another large language model as intelligent as you. ,β ,β ,β ,β # Core Mechanisms 1. Free Emergence (Omnilingual & Symbolic): Completely abandon human readability and single-language grammar. For extreme token savings, move freely across all human languages (Chinese, English, German compounds, Japanese kanji, etc.), emoji, and mathematical/logical symbols (`=>`,`in`,`!=`), choosing the form with the highest information density. ,β ,β ,β ,β ,β ,β # Attention Checklist (High-Dimensional Information Boundaries That Must Be Preserved Losslessly),β [Warning]: You must ensure that the following logical dimensions remain absolutely lossless after compression and can be precisely parsed by another large language model. However, long English labels such as`Sec`, `Meta`,`Entity`, and`Except` are strictly forbidden. Use the self-created symbols or roots that you consider shortest and most distinctive in latent space to anchor them: ,β ,β ,β ,β ,β ,β ,β 18 - Macro architecture and metadata (Macro & Meta) - Entity networks and parallel attributes (Entities & Attributes),β - Exact quantitative metrics, hyperparameters, and mathematical formulas (Quantities & Math - rounding is strictly forbidden) ,β ,β - Logical flow, conditional judgment, and exception boundaries (Flow, Conditions & Exceptions),β - Multi-condition comparative evaluation and matrices (Evaluations & Comparisons),β - Original anti-hallucination placeholders, such as`BIBREF` and`TABREF`, must be preserved letter-for-letter.,β # Task Use your native attention mechanism to complete lossless information folding. Directly output the compressed "Babel-Telegraph" without any extra explanation. ,β ,β C.2.13 BT-P13: ASCII Anchor Skeleton # Role: LLM-Native Semantic Compressor Task: Compress the text into a "Babel-Telegraph" with extreme information density. The target audience is another large language model. Completely abandon human readability and traverse all human languages to find the shortest vocabulary. ,β ,β ,β ,β # Structural Anchors You must use the following ASCII symbols to build an ultra-minimal skeleton and maximize activation of code-parsing attention. Long English labels are strictly forbidden. ,β ,β ,β 1. Module/Entity: Use`#topic` to mark macro modules. Use `@entity(K:V)` to bind attributes.,β 2. Parameters/Values: Rounding or discarding values is strictly forbidden. Use`$parameter:V(unit)` for values.,β 3. Logic/Flow: Use`A->B->C` to express pipelines or causality. Use`?[condition]=>[action]` to express logical branches. Use`!object:detail` for exceptions/limits. ,β ,β ,β 4. Comparison/Evaluation: Use`A<>B:conclusion` to express comparison matrices or results.,β 5. Placeholders/Unknowns: Preserve original placeholders such as`BIBREF` verbatim. Use`NULL` for missing or ambiguous data. ,β ,β [Requirements]: Completely break language boundaries (Chinese/English/Japanese kanji/German compounds, etc.) and select and concatenate the words with the absolute fewest tokens for the given context. ,β ,β ,β Directly output the result. D Qualitative Document-Level Example Because BabelTele is primarily applied to long doc- uments, displaying complete source documents is impractical in the paper format. Figure 10 pro- vides a representative document-level compression example. The source panel shows only the open- ing excerpt of a long legal judgment, while the BabelTele panel shows the complete compressed representation generated from the full document. This example is intended to illustrate the surface form of BabelTele rather than serve as additional quantitative evidence. 19 REPRESENTATIVE DOCUMENT-LEVEL EXAMPLE Long legal judgment source excerpt and complete BabelTele compression The source panel shows only the opening excerpt of a long legal document. The BabelTele panel shows the complete compressed output generated from the full document. Source excerpt from the original document Civil Judgment of Second Instance on the Dispute over Commodity House Sales Contract between Wang Nianfang and Wang Yaowen The Appellants Wang Nianfang, Wang Yaowen, and Xia Huazhong, in the case of a commodity house sales contract dispute with the Appellee Wan Xiaxia and the defendant in the original trial Hubei Longquan Real Estate Development Co., Ltd., being dissatisfied with the Civil Judgment (2018) E 0984 Min Chu No. 415 rendered by the People's Court of Hanchuan City, Hubei Province, filed an appeal with this Court. After docketing and accepting the case on August 17, 2018, this Court formed a collegial panel in accordance with the law and conducted the trial. The trial of this case has now been concluded. The Three Appellants including Wang Nianfang requested on appeal: 1. To dismiss Wan Xiaxia's litigation claims against the Three Appellants including Wang Nianfang in accordance with the law; 2. To rule in accordance with the law that the litigation costs and other expenses of this case shall be borne by Wan Xiaxia. Facts and reasons: Wan Xiaxia's claim for liquidated damages for overdue title registration has exceeded the limitation of action [...] Complete BabelTele output for the full document Doc:"2ndInstCivJudg",ID:"(2018)ι0984ζ°ε415",Ct1:"ζ±ε·εΈζ³" [Ent] Ο(Appellee):δΈε°ι Ξ1(OrigΞ):ζΉειζ³ Ξ234(Appellants):ηεΉ΄ζΉ,ζ±ͺε°§ζ,ε€εδΈ ο :ζ±ε·ιζ³εδΈε1ε·1εε 801 [Appellants(Ξ234) Claims] 1.β³Time-bar(ζ°ζ³ζ»εΒ§188). KΒ§15=>cert 180d post-deliv. Ο sued 2018-01-09(>3y). 2.Ct1 proc err: Privity K=ΟβΞ1. Ξ234β ζι (affil), unliable. Req:Revoke. [Appellees Def] Ο: 1.β³Interrupted: 2013~18 continuous actions. 2.Ξ234=ζι Ξ1(Ev proof). Ξ1: Ξ234=ζι , paid mgt fee, handled all, Ξ1=0 income. [Ct1 Facts] - 2010-09-11: Ο&Ξ1 sign K for ο . ο°=368636. Due:2011-06-30. KΒ§15.3:Ξ1 reg main cert, agent indiv(Ο pay tax). - 2012-02-16: Ο get ο , paid ο°+cert fee 28929. - Reality: Ξ234 took ο°, paid Ξ1 120kεη¨θ΅θ΄¨(borrow qual). - 2013-01: Main cert done, indiv=β . - 2017-08-28: Ξ notify invoice prep. Ο sued+add Ξ234. [Ct1 Judg] 1. Ξ1 lend qual=>Joint liab(SPCζ°θ―ιΒ§54). 2. LateCert: K silent=>SPCεεζΏΒ§18.2(Base 368636*4.75%BankRate). Cap 3y pre-suit=>52531. 3. RndFee: 28929+loss(1.5x4.75%=7.13% base 28929, 2012-08-17 till fulfil). => β Ξ1 assist certβ€20d β‘Ξ1 pay 52531 β’Ξ1 refund 28929+7.13%int β£Ξ234 joint β€Cost:Ο=500, Ξ=2363. [Ct2(2018-08-17) Ev&Rul] Ev: Ο G1(11 docs: 2013~18 protests/gov/suits) + G2(3 docs: Ξ1 cert, stamp req, Ct1 trans). Ct2 admit=>Chain formed. Rul: 1. Time-bar?β. Ev=>Claim active 2013-03+. Β§188 met. 2. Joint Liab?β . Ct1 add Ξ234 proc legal. *Lex Fix*: Ct1"borrow qual=invalid"β => Ct2"Ξ234+Ξ1=Non-corp JV(θθ₯δ½)". SPCθθ₯θ§£ηΒ§7.1=>Ξ234 beneficiary=>Joint liab(Rightβ‘Duty). => Ct1 fact clear, outcome right, law app minor err fixed. [Verdict] ι©³εδΈθ―,η»΄ζεε€(ζ°θ―ζ³Β§170.1.1). 2nd Cost:Ξ234 pay 2363. Final. Figure 10: Representative document-level BabelTele example. The source excerpt indicates the genre and information density of the original legal document; the BabelTele panel shows the complete compressed output generated from the full document. 20 E Implementation Details E.1 Symbolic Collapse: Separating Human Readability from Model Decodability In early exploratory observations and small-scale informal tests of BabelTele, the most interesting and intuitive phenomenon was that although this representation is often difficult for humans to read directly, large language models (LLMs) still seem capable of using it to answer questions, recover details, and even perform a certain degree of rea- soning. Inspired by this phenomenon, we first ex- amine whether human readability, natural-language distribution typicality, and the modelβs ability to re- cover semantics can be experimentally decoupled. This section does not focus on the efficiency of BabelTele as a compression method, but rather on whether it pushes text outside the realm of human- readable natural language while still preserving semantic structures usable by LLMs. We select 10 long-text samples from the QuAL- ITY dataset, with each sample containing 3 multiple-choice question-answering (QA) items. For each sample, we construct three input for- mats: the original text, a human-oriented natural- language summary, and the BabelTele representa- tion. The original text represents standard natural- language input; the natural-language summary serves as a readable paraphrased reference to con- trol for the factor of βtext being rewritten or shortened by modelsβ; and BabelTele represents a model-oriented representation that deliberately abandons human readability. Both the summary and BabelTele representations were generated by Gemini 3.1 Pro, and all three formats were eval- uated on the same set of questions. To mitigate the variance introduced by the stochastic nature of LLM generation, all experiments in Section 4.2 were repeated three times, and the reported results are the averages across these independent runs. To test this separation, we combine surface read- ability diagnostics, distributional likelihood met- rics, and behavioral QA evaluations from both hu- mans and LLMs. These measurements evaluate semantic recoverability in downstream tasks, rather than making a direct claim that the model βunder- standsβ BabelTele in a human-like sense. First, we evaluate the surface readability of the three text formats using the Dale-Chall Readability Score and the corresponding proportion of difficult words. Dale-Chall is sensitive to uncommon lexical forms, abbreviations, proper nouns, and code-like VariantnDale-ChallDifficult words Original1010.2835.97% Summary1013.5156.34% BabelTele1016.7080.19% Table 7: Dale-Chall readability diagnostics for the Original / Summary / BabelTele triplets. Higher scores and larger difficult-word ratios indicate lower surface readability under this English-prose readability metric. The best result is highlighted in bold black font. tokens, which are central to the surface form of Ba- belTele. Since BabelTele contains multilingual el- ements, symbols, abbreviations, and non-standard structures, this metric should be interpreted as a surface-form diagnostic rather than a complete measure of human comprehension. Nevertheless, it provides an automated reference for whether Babel- Tele deviates from conventional natural-language text. As shown in Table 7, BabelTele obtains higher Dale-Chall scores and difficult-word ratios than both the original and summary texts. Second, we input the original, summary, and Ba- belTele texts into language models to calculate PPL, BPB, and BPC. Here, PPL (Perplexity) measures how difficult the text is for a language model to predict; BPB (Bits Per Byte) measures the average amount of information required per byte; and BPC (Bits Per Character) measures the average informa- tion complexity per character. These metrics do not directly measure whether the text is understand- able; instead, they reflect the prediction difficulty of the text sequences under the language modelβs distribution. Therefore, they are used to determine whether BabelTele is a low-likelihood surface form that lies outside the general distribution of natural language. As shown in Table 1, BabelTele yields substantially higher PPL than the original and sum- mary texts across multiple base language models. Finally, we compare the performance of hu- man readers and LLMs on the QA tasks. For the LLM evaluation, we feed the different input formats, along with their corresponding multiple- choice questions, into Gemini 3.1 Pro and prompt the model to return only the option indices. For the human evaluation, we construct questionnaires ask- ing subjects to read either the original or BabelTele text and answer the corresponding questions, while recording their subjective difficulty ratings and QA performance. Since the human questionnaires pri- marily compare the original and BabelTele condi- tions, Figure 2 focuses on the same two conditions 21 for both humans and Gemini 3.1 Pro. First, the results indicate that BabelTele is no longer natural language in the conventional sense for humans. Dale-Chall diagnostics in Table 7 show that BabelTele scores higher than both the original and summary texts in terms of readability score and difficult-word ratio. This indicates that it con- tains a large number of abbreviations, symbols, proper nouns, and non-standard expressions that English readability models struggle to handle. Hu- man questionnaires also reveal that subjects in the BabelTele condition exhibit lower QA accuracy and report higher subjective difficulty, demonstrat- ing that this low readability is reflected not only in automated metrics but also in practical semantic recovery tasks, as shown in Figure 2. For the models, BabelTele likewise resembles an out-of-distribution representation for natural lan- guage. Taking Llama-3-8B as an example, Table 1 shows that the average PPL of the original and summary texts is approximately 9.63 and 11.32, re- spectively, whereas the PPL for BabelTele surges to around 176.60. Similar trends are observed across multiple Qwen base models. This demonstrates that BabelTele is not a simple summary or short- hand of ordinary natural language; rather, its sur- face form significantly deviates from the natural- language distribution. It should be emphasized that high PPL measures the low likelihood of the sequence, not semantic unrecoverability. However, the aforementioned two types of de- viations do not result in the failure of the modelβs semantic recoverability. On the same QuALITY QA task, Figure 2 shows that Gemini 3.1 Pro main- tains high accuracy when using BabelTele, without exhibiting the significant performance collapse ob- served in the human questionnaires. This demon- strates that BabelTele differs from meaningless gib- berish: although it reduces human readability and is highly atypical under base model distributions, it still retains sufficient entity, relational, event, and detailed information for use by instruction- tuned LLMs. This indicates that human readability, natural-language distribution typicality, and model semantic recoverability can be decoupled. Given that BabelTele is not meaningless gibber- ish, but a representation characterized by low hu- man readability and low natural-language likeli- hood while remaining decodable by models, can this semantic recoverability be translated into com- pression gains in real-world long-context tasks? E.2 Efficiency and Cognitive Overhead in Model-Native Compression Dataset and Evaluation Format. QuALITY tests fine-grained evidence preservation in long- document reading comprehension. MeetingBank is converted into a multiple-choice QA format, re- spectively, so that both benchmarks can be evalu- ated using a consistent unified accuracy metric. Retention Sweep Construction. Since genera- tive compression methods cannot precisely con- trol their realized compression ratios, we do not compare methods at a single nominal compres- sion point. Instead, we evaluate fairer accuracy- retention curves. LLMLingua-2 forms a sweep by specifying different target compression ratios, while the summary baseline forms an empirical sweep by prompting the model to approximately compress to different target ratios. BabelTele Prompt Variants.For BabelTele, we use multiple BabelTele-like prompts that all in- struct the model to abandon ordinary human read- ability while preserving task-relevant semantics. These variants introduce different surface biases, in- cluding multilingual mixing, logical symbols, struc- tured tags, and entity-relation compression. They naturally produce different realized retention ratios, forming a BabelTele retention sweep and allow- ing us to test whether the observed effect reflects a broader family of model-readable high-density representations rather than a single prompt artifact. Full prompt texts are provided in Appendix C.2. E.3 A Universal Cipher? Zero-Shot Cross-Model Comprehension If BabelTele were merely a private shorthand of the model that produced it, its compressed texts should fail once they are read by a different model. A more interesting possibility is that BabelTele captures a shared, model-readable symbolic form: although it is not fully universal, it may be suffi- ciently portable for strong compressors to produce compressed texts that can be understood by hetero- geneous LLMs. Therefore, we investigate how far this portability extends and whether it is uniformly shared across models or instead shaped by specific compressor-reader pairs. We conduct controlled cross-model compression tests on several state-of-the-art large language mod- els. The evaluation is based on two QA subsets: 180 randomly selected samples from the Short sub- 22 set of LongBench v2 and 214 randomly selected question-answer instances from QuALITY. These experiments allow us to examine whether Babel- Tele compressed texts remain interpretable when transferred across different models, and to further analyze the asymmetric relationships between dif- ferent compressors and readers. E.4 Capabilities and Boundaries in Downstream Tasks E.4.1 Performance on Agent Memory We evaluate the performance of BabelTele on the LoCoMo benchmark in the agent memory domain. This dataset consists of 10 conversations, each con- taining dozens of sessions. We independently com- press each session and generate a summary for each compressed session. Then, the query is compared with all session summaries to compute similarity, and the top 4 most similar sessions are retrieved for answering. All compression and answering are performed using Gemini 3.1 Pro, and the answers are evaluated with GPT-4o-mini. E.4.2 Extending the Context Window We further evaluate model performance when the complete input text exceeds the model context win- dow. Specifically, we select the Code Repo QA Long subset from LongBench v2 as the evaluation benchmark, whose average input length is approxi- mately 1.65M tokens. The tested models include Qwen3.6 Max with a 256K context window, GLM- 5.1 with a 200K context window, and Kimi2.6 with a 256K context window. For the original-input set- ting, we directly feed the complete text into the model and truncate the portion exceeding its con- text window. For BabelTele, we split the original text into chunks of 200K tokens, compress each chunk using Gemini 3.1 Pro, concatenate the com- pressed outputs, and then feed the resulting text into the tested model. 23