Paper deep dive
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
Feiyang Ren, Shengtao Wen, Lingbing Guo, Yu Tian, Yuanning Cui, Xiang Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 87%
Last extracted: 8/22/2026, 2:34:42 AM
Summary
The paper introduces ESTP (Entropy-and-Semantic Token Pooling), a lightweight framework for predicting LLM output lengths to optimize serving efficiency. It addresses the limitations of entropy-only methods (like EGTP) by combining token-wise entropy with attention-based semantic importance scores derived from self-attention weights. ESTP demonstrates superior prediction accuracy and reduced padding ratios on the ForeLen benchmark across multiple LLM backbones.
Entities (19)
Relation Signals (21)
ESTP → combines → Entropy
confidence 95% · ESTP ... addresses this issue by combining entropy with attention-based importance scores.
ESTP → combines → attention-based importance scores
confidence 95% · ESTP ... addresses this issue by combining entropy with attention-based importance scores.
ESTP → evaluatedon → ForeLen
confidence 95% · On the ForeLen benchmark, ESTP outperforms baseline methods
ESTP → outperforms → EGTP
confidence 95% · On the ForeLen benchmark, ESTP outperforms baseline methods... Compared with EGTP, Our method reduces the Avg MAE by over 9 points
Attention weights → captures → semantic importance
confidence 90% · this allows ESTP to capture both uncertainty and semantic importance
Entropy → captures → Uncertainty
confidence 90% · this allows ESTP to capture both uncertainty and semantic importance
attention-based importance scores → derivedfrom → self-attention weights
confidence 90% · These scores are derived directly from the self-attention weights computed during the LLM prefill phase
ESTP → improves → Throughput
confidence 90% · When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduces the overhead. This advantage is especially pronounced in long-context reasoning and reinforcement learning applications. Existing approaches, such as entropy-guided token pooling, use token-wise entropy as their primary signal, but they tend to ignore differences in semantic content across tokens. So, important tokens are often underweighted, and tokens carrying little information receive disproportionate emphasis. This hurts the reliability of length prediction. We introduce ESTP (Entropy-and-Semantic Token Pooling), a lightweight framework that addresses this issue by combining entropy with attention-based importance scores. These scores are derived directly from the self-attention weights computed during the LLM prefill phase, and this allows ESTP to capture both uncertainty and semantic importance with minimal additional computation. Since the framework reuses prefill activations, it adds almost no extra memory overhead and introduces only minimal latency. On the ForeLen benchmark, ESTP outperforms baseline methods, achieves better prediction accuracy and lower error rates in most scenarios. When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput and reduce the padding ratio. Our results offer a practical and effective building block for length-aware LLM serving systems.
Tags
Links
- Source: https://arxiv.org/abs/2608.15592v1
- Canonical: https://arxiv.org/abs/2608.15592v1
Trouble viewing inline? Open PDF directly →
Full Text
49,514 characters extracted from source content.
Expand or collapse full text
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction Feiyang Ren 1∗ , Shengtao Wen 1∗ , Lingbing Guo 2 , Yu Tian 3 , Yuanning Cui 4 , Xiang Chen 1† 1 MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics 2 Nanjing University 3 Tsinghua University 4 Nanjing University of Information Science and Technology feiyang_ren, xiang_chen@nuaa.edu.cn Abstract Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput. Predicting output lengths in advance makes it possible to adopt length-aware scheduling, and this reduces the overhead. This advantage is especially pronounced in long-context reasoning and reinforcement learning applica- tions. Existing approaches, such as entropy-guided token pool- ing, use token-wise entropy as their primary signal, but they tend to ignore differences in semantic content across tokens. So, important tokens are often underweighted, and tokens carrying little information receive disproportionate emphasis. This hurts the reliability of length prediction. We introduce ESTP(Entropy-and-Semantic Token Pooling), a lightweight framework that addresses this issue by combining entropy with attention-based importance scores. These scores are de- rived directly from the self-attention weights computed during the LLM prefill phase, and this allows ESTP to capture both uncertainty and semantic importance with minimal additional computation. Since the framework reuses prefill activations, it adds almost no extra memory overhead and introduces only minimal latency. On the ForeLen benchmark, ESTP outper- forms baseline methods, achieves better prediction accuracy and lower error rates in most scenarios. When integrated with a length-aware scheduler in end-to-end system tests, it further helps improve overall throughput and reduce the padding ra- tio. Our results offer a practical and effective building block for length-aware LLM serving systems. Introduction Efficient serving of Large Language Models (LLMs) remains a critical system-level challenge. Although techniques such as continuous batching, PagedAttention (Kwon et al. 2023), PowerAttention (Chen et al. 2025a), and SARATHI (Agrawal et al. 2023) have improved throughput, the barrel effect, which wastes computation by padding shorter sequences to the longest sequence in a batch (Deshmukh, Raparti, and Hsu ∗ These authors contributed equally. † Corresponding author. Copyright © 2027, Association for the Advancement of Artificial Intelligence (w.aaai.org). All rights reserved. Prompt: John has 6 children. Assuming that the gender of each child is determined independently and with equal likelihood of male and female, what is the probability that John has more sons than daughters or more daughters than sons? High entropy tokens (e.g, transitional /reasoning-boundary/filler words) are heavily weighted. John has 6 children . Assuming that the gender ... more daughters than sons ? 1 2 3 4 5 6 7 8 9 ... n-1 n n+1 n+2 Entropy-only Weights (EGTP) Semantics-only Weights (Attention-Based) Semantic-critical tokens (entities, numbers, constraints, key instructions) are emphasized. ESTP Weights (Entropy+Semantics) ESTP balancees uncertainty and semantic importance, highlighting both critical content and necessary connector. (a) Balanced Token Weighting with ESTP 0% 5% 10% 15% 20% 25% 30% 35% 40% O ... U ... Quantitative Statistics (512 samples) P r o p o r t i o n ( % ) 23.43% Over-weighted by EGTP Under-weighted by EGTP A substantial portion of tokens are misaligned when relying on entropy alone. u23.43% of tokens in the Top 40% high-entropy group have low semantic importance. u21.48% of tokens in the Bottom 40% low-entropy group have high semantic importance. 21.48% (b) Quantitative Results of Entropy-Semantic Misalignment Figure 1: Limitations of Entropy-Only Token Impor- tance Assessment. (a) Comparison of Token-Level Weight- ing Strategies: Entropy-Only, Semantics-Only, and ESTP. (b) Quantitative Analysis of Token Misalignment Between En- tropy and Semantic Importance. 2025), remains a fundamental bottleneck. This overhead is especially severe in workloads with high length variance, including long-context reasoning (Wang et al. 2025b; Chen et al. 2025b) and dynamic Reinforcement Learning (RL) arXiv:2608.15592v1 [cs.AI] 16 Aug 2026 sampling (Jiang et al. 2025; Wang et al. 2025a). Accurate pre-generation output length prediction enables length-aware scheduling, reducing padding waste and im- proving hardware utilization. Existing prediction meth- ods range from costly external auxiliary models, such as DistilBERT-based classifiers (Jin et al. 2023; Qiu et al. 2024), to LLM-intrinsic approaches such as TRAIL (Shahout et al. 2025a). Recently, EGTP (Xie et al. 2026) achieved state-of- the-art performance by reusing LLM on-the-fly activations and token-level entropy for weighted pooling with negligi- ble overhead. However, as illustrated in Figure 1a, relying solely on entropy introduces a systematic blind spot: entropy captures generation uncertainty rather than semantic impor- tance. Core content tokens in a prompt, including key enti- ties, constraints, and instructions, are often processed with high confidence (low entropy), causing them to be under- weighted. In contrast, high-entropy transitional or filler to- kens may be over-amplified, leading to substantial semantic signal loss during prediction. We quantitatively verify this semantic misalignment through a controlled diagnostic on 512 RL-sampled in- stances. For each token, we compute both its entropy and its semantic importance, and we use the self-attention weights from the LLM prefill phase as a proxy for the latter. The results show a stark divergence between the two measures. As illustrated in Figure 1b, when tokens are split at the 60th percentile, a substantial fraction of high-entropy tokens carry little semantic importance, and many low-entropy tokens are semantically critical. This confirms that entropy alone does not fully capture token relevance, and this makes the predic- tor vulnerable in complex reasoning scenarios where output length is tightly governed by prompt semantics. To recover these lost semantics, we propose ESTP (Entropy-and-Semantic Token Pooling), a lightweight pre- diction framework shown in Figure 3. For each token, ESTP computes an attention-based semantic weight, linearly com- bines it with the token’s entropy, and applies a temperature- scaled softmax to produce the final pooling weights. The re- sulting aggregated representation captures both uncertainty signals and semantic content. Following EGTP, we employ a soft-label distribution prediction head to handle the heavy- tailed nature of length distributions. We discretize target lengths into bins and optimize a combined Cross-Entropy and MSE loss, which provides distance-aware regression with strong robustness to outliers. Because ESTP directly reuses hidden states from the prefill phase, it adds only marginal VRAM overhead and incurs virtually no inference latency. We evaluate ESTP on ForeLen (Xie et al. 2026), a custom dataset spanning three challenging scenarios (long-sequence generation, complex reasoning, and dynamic RL sampling), using four LLM backbones (Qwen2.5-3B/7B and Llama3.2- 1B/3B). Our main contributions are summarized as follows. • We empirically verify a systematic mismatch between en- tropy and semantic importance. Tokens that carry critical semantic content often exhibit high prediction confidence, and this leads to insufficient weight assignment. This find- ing challenges the common assumption that uncertainty alone can serve as a reliable indicator of token relevance. • We propose ESTP. It fuses entropy with attention-based semantic importance and a soft-label regression head. By balancing uncertainty and semantic importance, it recov- ers the semantic signal that entropy underweights, and it does this at negligible additional inference cost. • Extensive evaluations show that ESTP offers a practi- cal improvement for length prediction. Across a range of models and most scenarios on ForeLen benchmark, our method generally reduces prediction error (MAE), im- proves binning accuracy, and further enhances throughput while reducing the padding ratio on end-to-end tasks. Related Work LLM Serving Optimization Existing LLM serving optimizations, including continuous batching (Kwon et al. 2023), batch prompting (Qiu and Sri- vastava 2025), and PagedAttention (Kwon et al. 2023), im- prove throughput by dynamically managing requests. Recent studies have further advanced the field through improved scheduling strategies, memory management, and parallel computing techniques. For example, SARATHI (Agrawal et al. 2023) improves decoding throughput with chunked prefill and decode piggybacking. However, these methods do not address the “barrel effect”, which refers to the computa- tional waste caused by padding shorter sequences to match the longest sequence in a batch (Deshmukh, Raparti, and Hsu 2025). Our work takes a complementary approach by reusing hidden states generated internally by LLMs. This enables ac- curate output length prediction and length-aware scheduling, reducing padding overhead while complementing existing system-level optimizations. LLM Output Length Prediction Conventional approaches to LLM output length prediction often rely on external auxiliary models, such as DistilBERT- based classifiers (Jin et al. 2023; Qiu et al. 2024; Hu et al. 2024). These models only process input prompts and require substantial computational resources and training time. To overcome these limitations, TRAIL (Shahout et al. 2025a) introduced a prediction method based on the target LLM it- self, achieving low overhead and accurate predictions. More recently, inspired by studies connecting model internals, such as token entropy, with generation characteristics (Li et al. 2025), EGTP (Entropy-Guided Token Pooling) (Xie et al. 2026) directly reused LLM on-the-fly activations and token entropy for accurate static prediction with negligi- ble overhead. This approach significantly reduces training costs. However, while this lightweight paradigm enables ef- ficient inference and accurate prediction, relying solely on entropy weighting can lead to semantic information loss. Our framework addresses this limitation by combining entropy with token-level feature weights to construct an improved weighted pooling mechanism. Preliminaries On the Inherent Limitations of EGTP Mechanisms. During LLM inference, token entropy reflects generation Joint Distribution of Attention Scores and Importance Attention Scores Importance Frequency 0.30.40.5 0.6 0.70.8 15 30 45 60 75 90 105 120 0.25 0 0.30 0.35 0.40 0.45 0.50 0.55 0.60 0.65 0.70 Regression Line(r=0.424) Figure 2: Displaying the joint distribution of Attentions Scores and Token Importance, with the regression line con- firming a significant positive correlation (r=0.424). CategoryReasoning RL Average Entity34.233.834.0 Key instructions17.510.013.8 Constraints25.037.531.2 Nonsemantic23.318.721.0 Table 1: Category proportion among high-attention tokens. uncertainty (Xu and Lu 2025; Nguyen, Payani, and Mirza- soleiman 2025; Xu et al. 2026; Wu, Zhang, and Gao 2026) and has been shown to correlate positively with output length prediction, which motivates entropy-weighted pooling for lightweight length prediction (Xie et al. 2026). But relying only on entropy has a key limitation. High-entropy tokens reflect model uncertainty, and this mechanism dilutes the contribution of semantically core tokens. So, critical infor- mation is lost, irrelevant tokens are over-weighted, and pre- diction accuracy ultimately suffers. Motivation. Length prediction needs an understanding of semantic structure. The model generally need to find key entities and core instructions from the query. These signals often show up in tokens that have low entropy but high atten- tion weights. In Transformers, attention weights indicate how much a token contributes to the overall contextual representa- tion (Huangyw et al. 2025), such as the use of attention value vectors to capture sentence-level semantics (Zhang et al. 2026) and the temperature-controlled Softmax sharpening approach that directs model focus toward the most relevant tokens (Ram, Xia, and Soatto 2025). And prior work has shown them to be a reliable proxy for semantic importance in various settings (Zhang et al. 2025; Chen et al. 2024; Xing et al. 2024; Han et al. 2025). We sampled 120 instances from the Reasoning and RL scenarios to verify the connection between attention and semantics. We manually annotated se- mantically important tokens. As shown in the table 1, about 80% of high-attention tokens are identified as semantically important. This provides empirical support for the alignment between high attention weights and semantic importance. To better understand the role of token-level attention, we use a gradient-based attribution method (Sundararajan, Taly, and Yan 2017; Bacha and George 2025; Yang et al. 2025) to measure each token’s contribution to the final prediction. As shown in the figure 2, we observe a clear positive correlation between attention scores and token importance. This result confirms that high-attention tokens are especially informa- tive for length prediction and provides empirical support for the design of ESTP. We conduct experiments on 512 RL sam- ples to verify the limitations of entropy-only weighting. We compute token entropy and attention-based semantic scores from LLM activations. We classify tokens into high- and low-entropy groups and high- and low-semantics groups. We calculate the misalignment ratio (Figure 1b). Over 20% of the top 40% high-entropy tokens have low semantic importance. This shows that entropy-only pooling overweights meaning- less tokens and underweights core prompt information. Ab- lation studies confirm that semantic features contribute more than entropy alone. Also, over 20% of the bottom 40% low- entropy tokens carry high semantic value. This shows that entropy-only downplays the contribution of semantically im- portant tokens. To address this issue, we introduce seman- tic features as a complementary signal. ESTP fuses entropy and semantic weights to balance uncertainty and importance. This achieves more accurate LLM output length prediction. Method Attention-based semantics Importance Modeling We propose a semantic attention-based weighting strategy to calculate the semantic importance of each token with strict handling of padding tokens and self-similarity to ensure the rationality and accuracy of the weights. We compute token importance by averaging the attention values along the col- umn dimension of the attention matrix A∈ R N×N . Here, A denotes the multi-head self-attention matrix averaged across all heads, extracted from the final layer of LLM. Specifi- cally, the importance of the j-th token is determined by the average amount of attention it receives from all other valid (non-padding) tokens: S j = 1 N N X i=1 A ij ,(1) where N is the input sequence length, A ij denotes the atten- tion weight from token i to token j. Semantics-Entropy Joint Weight Fusion For token entropy computation, we follow the calculation scheme introduced in EGTP. We first derive the entropy value H i of each individual token. Concretely, given the hidden state h i mapped from input token x i , entropy is quantified based on the next-token probability distribution P(v|x <i ) across the entire vocabulary V : H i =− X v∈V P(v|x <i )logP(v|x <i ),(2) 푺 풋 = ퟏ 푵 풊=ퟏ 푵 푨 풊풋 푯 풊 =− 풗∈푽 풑(풗)풍풐품풑(풗) 푻 풊 =흀 ퟏ 푺 풊 +흀 ퟐ 푯 풊 풆 풊 = 풆풙풑(푻 풊 ) σ 풋=ퟏ 풏 풆풙풑(푻 풋 ) 0 0.2 0.4 0.6 0.8 1 123 ... K Probability Length Bins Diverse Scenarios Data Question: John has 6 children. Assuming that the gender of each child is determined independently and with equal likelihood of male and female, what is the probability that John has more sons than daughters or more daughters than sons? (Math) Input Embedding + Position Encoding LLM ( Pre-fill ) Block 1 ... Block 2 Block N Token Weight e: 풉 ′ = 풊=ퟏ 풏 풉 풊 ∙풆 풊 푯 ′ 흐ℝ 풅 Pooled Representation Predicted Length 푳= 풌=ퟏ 푲 풌∙풑 풌 Softmax probs1 probs2 ... probs n Hidden_states 푯=[풉 ퟏ ,풉 ퟐ ...,풉 풏 ]훜ℝ 풏×풅 Predicted Length Length Distribution(Soft-label) 1234N ... ... 1 2 3 4 N ... ... 1234N ... ... 123 4 n ... ... Semantic Weights Entropy Pooling Weighted Pooling Joint Pooling logits1 logits2 ... logits N token1 token2 ... token N Self-Attention scores1 scores2 ... scores n Entropy and Semantic Pooling (ESTP) Figure 3: Overview of the ESTP framework. The framework consists of three main components: (1) Construction of Hidden States, which extracts hidden state representations from diverse scenarios including reasoning, long-sequence, and RL tasks; (2) Semantic and Entropy Pooling, which computes attention-based semantic weights and aggregates entropy values to identify high-entropy frequent tokens (e.g., logical connectives, suppositional words, inference steps), yielding combined weightsT i ; and (3) Lightweight Prediction, which applies softmax-normalized weights to hidden states and feeds the weighted representation into an MLP predictor to estimate the output length with both soft and hard losses. Then, we weight and combine the semantics S i and the en- tropy H i for each token to obtain the final T i . T i = λ 1 S i + λ 2 H i (3) Subsequently, these values are leveraged to derive attention weights. We adopt the softmax function to normalize the values into weight distribution e i , where a temperature coef- ficient α is introduced to adjust the concentration degree of the distribution: e i = exp(T i ) P n j=1 exp(T j ) (4) Ultimately, the aggregated representation h is obtained by weighted summation of hidden states with semantic-entropy: h = n X i=1 e i h i (5) Soft Label–Guided Distribution Regression for Sequence Length To address the limitation that the standard MSE loss is highly sensitive to outliers and heavy-tailed distributions, we adopt the prediction head design from EGTP (Xie et al. 2026). The design employs a joint loss function that combines cross- entropy loss and mean squared error loss, leading to more accurate prediction performance. Using the feature repre- sentations described above, we design a dedicated prediction head for sequence length estimation. We first convert the continuous length target y into a soft probability distribu- tion p as the supervised ground truth. The continuous length space is partitioned intoK predefined bins, and the probabil- ity assigned to bin j decreases monotonically with growing distance from the ground-truth bin i. This is computed as: p j = exp(−|j− i|) P K k=1 exp(−|k− i|) (6) Subsequently, taking the feature vector h derived from ESTP as input, the model produces two parallel outputs. One is aK- dimensional probabilistic classification prediction ˆp obtained via softmax projection. The other is the ultimate regression result ˆy, calculated as the expectation of the predicted distri- bution. Let c i represent the center value of the i-th bin, the calculation formula is defined as: ˆy = K X i=1 ˆp i · c i (7) Ultimately, the model is optimized via a combined loss func- tion integrating cross-entropy and mean squared error losses, ModelScenario Prediction Method SSJF-RegSSJF-MCLTR-CTRAILEGTPESTP Qwen2.5-3B LongSeq246.76318.71164.64147.9297.5096.85 Reasoning322.77242.26287.77132.20143.11120.8 RL120.37206.4296.65 159.1599.7872.56 Avg229.97255.80183.22146.42112.4596.74 Qwen2.5-7B LongSeq206.11346.96169.28134.1881.7285.34 Reasoning299.45284.53256.72124.19135.32120.33 RL123.56222.6599.70155.5195.24 74.62 Avg209.71284.71175.33137.96104.1093.43 Llama3.2-1B LongSeq442.27199.40402.98145.3582.6481.30 Reasoning320.31203.96299.10148.28138.04132.63 RL88.43 192.75122.37161.78108.5375.61 Avg283.67198.30274.82151.80107.5396.51 Llama3.2-3B LongSeq431.21210.73395.66143.62122.88115.57 Reasoning359.37200.60334.44177.16100.4099.21 RL108.65 207.20123.48152.85114.3193.61 Avg299.74206.18284.53157.88111.93102.80 Table 2: Comparison of MAE across different length prediction methods (lower is better). Best results are highlighted in bold, and second-best results are underlined. ESTP consistently achieves lower average prediction error across all backbone LLMs. with hyperparameter λ 3 balancing the two components: L = λ 3 L CE (p, ˆp) + (1− λ 3 )L MSE (y, ˆy)(8) L CE matches ˆp to p, providing stable gradients and improved training.L MSE minimizes the gap between predicted length ˆy and ground truth y for precise length prediction. Experiments Experimental Setup Dataset. To evaluate the performance of our proposed ESTP in complex and realistic scenarios, we conduct ex- tensive experiments on ForeLen (Xie et al. 2026), custom- built dataset constructed specifically for assessing predic- tor capabilities under challenging condictions. ForeLen en- compasses three primary scenarios: (1) Long-Sequence and Complex Reasoning Generation, prompts for this scenario are sourced from three well-established benchmark datasets: LongBench (Bai et al. 2023), ZeroSCROLLS (Shaham et al. 2023), and IFEval (Zhou et al. 2023). (2) Dynamic RL Sampling, prompts are selected from six widely used math and code reasoning datasets: CRUXEval (Gu et al. 2024), GSM8K (Cobbe et al. 2021), LiveCodeBench (Jain et al. 2025), MATH (Hendrycks et al. 2021b), MBPP (Austin et al. 2021), and MMLU-STEM (Hendrycks et al. 2021a). Metrics. To quantify prediction deviation, we use MAE as the main metric and also compare RMSE across methods. We bin output lengths by the RL dataset distribution and compute per-bin accuracy. We measure inference efficiency by average time and GPU memory on fixed samples. To evaluate the end- to-end system performance, we use Throughput, JCT(Job Completion Time) and Padding Ratio(detailed calculation of experimental metrics in the appendix). Models and Baselines. To evaluate the effectiveness of our approach, we conduct extensive experiments on four models: Qwen2.5-3B/7B (Yang et al. 2024) and Llama3.2- 1B/3B (Team 2024). For each architecture, we compare six configurations: SSJF-Reg (Qiu et al. 2024), SSJF-MC (Qiu et al. 2024), TRAIL (Shahout et al. 2025b), LTR-C (Fu et al. 2024), EGTP (Xie et al. 2026), and our proposed ESTP. Experimental Settings. We use AdamW (Kingma and Ba 2015; Loshchilov and Hutter 2019) as the optimizer, and fix the global random seed to 42 to ensure experimental re- producibility. Two hyperparameters λ 1 and λ 2 are utilized to balance entropy weights and semantic feature weights. Meanwhile, through weight fusion analysis (sweeping λ 1 :λ 2 from 1:9 to 9:1), we observe λ 1 drives the predictions much more than λ 2 , confirming that semantic features are the key contributor here. The coefficient λ 3 adjusts the ratio between cross-entropy loss and Mean Squared Error loss. K denotes the number of bins into which the continuous output length space is partitioned. We provide more detailed hyperparam- eter analysis and experimental settings in the appendix. Main Results For Q1: How Does Mean Absolute Error Perform Over- all in Length Prediction Tasks? We evaluate the Mean Absolute Error across four widely used LLMs, as presented in Table 2. Across all evaluated models, ESTP consistently achieves the lowest average MAE and outperforms the EGTP on most scenarios. Compared with EGTP, Our method re- duces the Avg MAE by over 9 points, especially over 20 points in RL scenarios. We also performed the statistical sig- nificance analysis(detailed statistical results in the appendix), in which the p-value in the RL scenario is well below 1%, fully demonstrating its substantial improvement on dynamic ModelInput Length Prediction Method SSJF-RegSSJF-MCLTR-CTRAILEGTPESTP Qwen2.5-3B Short50.3759.7027.5658.8222.0359.83 Medium80.7970.8762.1566.4286.0184.32 Long50.7548.2133.2050.4851.0051.57 Overall67.6761.6060.4661.1562.3672.54 Qwen2.5-7B Short72.4965.2548.3062.5356.6674.13 Medium74.7671.0966.7752.1974.8575.27 Long58.8757.3128.7063.2147.6559.82 Overall70.6469.3760.2258.3363.6173.26 Table 3: Accuracy comparison across different input length bins in RL scenarios(higher is better). Best results are highlighted in bold, and second-best results are underlined. ESTP achieves the highest overall accuracy on the Qwen2.5-3B/7B while maintaining competitive performance across Short (<200 tokens), Medium (200–500 tokens), and Long (>500 tokens). sampling tasks in RL training. Overall, our method yields broad gains across a variety of models and most scenarios, further improves LLM length prediction performance. RMSE 100 140 180 220 260 300 340 380 111 101 257 210 372 244 147 154 192 300 324 335 113 125 134 148 262 119 124 127 194 219 319 233 Figure 4: Comparison of RMSE across different length pre- diction methods on the Qwen2.5-7B. For Q2: How Does Root Mean Square Error Penalize Large Deviations in Length Prediction? To further eval- uate the robustness of our proposed ESTP method, we com- pare the RMSE values of all approaches using the Qwen2.5- 7B model, which allows us to examine the sensitivity of each method to large prediction errors. As shown in Figure 4, ESTP achieves the best performance in both RL and Rea- soning scenarios, and achieves the lowest average RMSE. These results demonstrate that our approach provides more accurate and reliable predictions than the baseline methods. We report additional results on other models in the appendix. For Q3: How Does Prediction Accuracy Vary Across Dif- ferent Length Buckets? To better analyze accuracy across different length bins, we evaluate prediction performance within three intervals (Table 3).ESTP achieves strong overall accuracy, exceeding 70% on Qwen2.5 models, and substan- tially outperforms TRAIL and EGTP, two state-of-the-art ap- proaches that also leverage the internal activations of LLMs. The SSJF-series methods achieve competitive performance in terms of overall accuracy, but they rely on an external aux- iliary BERT model. In contrast, ESTP introduces negligible inference latency and memory overhead, making it more ef- ficient and practical than SSJF variants. We report additional results on other models in the appendix. Pooling Method LongSeq Reasoning RL Average ESTP(Ours)96.85120.8 72.56 96.74 Average Pooling 163.32 133.74 92.44 129.83 Entropy Pooling 97.50143.11 99.78 112.45 Semantic Pooling 100.10 124.84 74.28 99.74 Table 4: Ablation Study on the Qwen2.5-3B. Ablation Analysis To investigate different pooling strategies, we evaluate four variants: average pooling, entropy-only pooling, semantic- only pooling, and our ESTP joint pooling (Table 4). ESTP achieves the best overall performance with the lowest average MAE, and it achieves the best results in each of the three sce- narios. Compared with average pooling, our method gives large performance gains across all three scenarios. Com- pared with entropy-only EGTP, ESTP combines entropy and semantic importance, and it shows outstanding gains on the RL dataset. Overall, both entropy and semantics are essen- tial to LLM length prediction. The combination of these two aspects leads to better prediction performance. We addition- ally verified the impact of features from different layers on prediction performance (detailed in the appendix), and found using features from the final layer shows better performance than those from intermediate layers(including 16 and 20). ScenarioMetrics Prediction Method SSJF-RegSSJF-MCLTR-CTRAILEGTPESTP Reasoning Padding Ratio↓1.261.211.011.221.180.37 Throughput↑17.8019.6121.016.0322.5123.46 Avg. JCT↓11.7810.538.5511.58.33 7.69 LongSeq Padding Ratio↓1.341.251.151.231.13 1.01 Throughput↑14.2520.1428.6721.0330.74 48.45 Avg. JCT↓16.2510.637.312.116.073.29 Table 5: End-to-End System Performance Comparison on Llama3.2-3B. Best results are highlighted in bold, and second-best results are underlined. ESTP achieves the best performance among all methods. 10 0 Inference timeMemory consumption 10 1 10 2 10 3 Figure 5: Average time(per sample) and memory(per batch of 32 samples) overhead for predicting in LongSeq scenarios. End-to-End System Performance Comparison We integrated ESTP and baseline predictors with the Short- est Job First (SJF) scheduler in an end-to-end system backed by the vLLM serving engine. As shown in the table 5, ESTP achieves the best results across Reasoning and LongSeq sce- narios. In the Reasoning setting, ESTP gives a much lower Padding Ratio than EGTP, decreased from 1.18 to 0.37. In the LongSeq, our method significantly improves throughput and reduces average JCT compared to EGTP, the strongest base- line. In summary, our ESTP is a practical and effective build- ing block for length-aware LLM serving systems, facilitating more efficient resource allocation and request scheduling and helping to improve overall system throughput. Efficiency Analysis We compare the time and memory cost of ESTP with base- lines in Figure 5. On LongSeq, ESTP has much lower infer- ence latency than methods that use external auxiliary models like SSJF (BERT) and LTR-C (OPT-350M), since these mod- els bring significant computational overhead and longer re- sponse time. For memory consumption, our ESTP based on internal state features also use far less memory than those re- lying on external auxiliary models. Overall, ESTP achieves a user_prompt:Write a function with the following signature: def check_str(string).Write a function to check whether the given string is starting with a vowel or not using regex. Baseline(EGTP) : Only Entropy Weighted Pooling EFTP(Ours) : Feature and Entropy Pooled Write a function with the following signature: def check_str(string).Write a function to check whether the given string is starting with a vowel or not using regex. High Weights Write a function with the following signature: def check_str(string).Write a function to check whether the given string is starting with a vowel or not using regex. Low Weights ESTP(Ours) : Semantic and Entropy Joint Weighted Pooling High WeightsLow Weights Figure 6: Qualitative comparison demonstrates the token weighting distributions of EGTP and ESTP on specific sam- ples. Red tokens indicate high-weight tokens, and the corre- sponding prediction results are provided for reference. favorable trade-off between average inference time and mem- ory footprint, and performs well on both dimensions. Case Study We visualize token-level weights on real examples to show the advantage of ourESTP method, which combines entropy- based signals with semantic features. The results illustrate that our model does a better job at capturing the core meaning of the input prompts. In Figure 6, we see that our method puts more weight on important words and entities, which helps extract key information more precisely. In contrast, EGTP amplifies the weights of tokens that have little semantics, while underweighting those carry important information. For final length prediction, ESTP with combined entropy and attention-based semantic information achieves significantly lower prediction errors than EGTP only with entropy. Conclusion and Future Outlook In this paper, we first point out the major drawback of entropy- only guided pooling: it does not reflect the semantic sig- nificance of individual tokens well. To address this prob- lem, we present ESTP, a lightweight framework that com- bines attention-based semantic representations with token- level entropy and uses soft-label regression to cope with the heavy-tailed nature of generation lengths. Experiments on several tasks from the ForeLen benchmark show that our ap- proach achieves better results than prior methods, and it also brings substantial improvements in end-to-end system per- formance. While ESTP demonstrates strong empirical per- formance across multiple open-source LLMs, its attention- based proxy is correlational, not causal, and its closed-source applicability is limited by inaccessible activations. Future work pursues causal attribution for semantic–positional dis- entanglement, black-box adaptation without internal activa- tions, and multi-feature fusion for scenario-wise robustness. References Agrawal, A.; Panwar, A.; Mohan, J.; Kwatra, N.; Gulavani, B. S.; and Ramjee, R. 2023. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. CoRR, abs/2308.16369. Austin, J.; Odena, A.; Nye, M. I.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C. J.; Terry, M.; Le, Q. V.; and Sutton, C. 2021. Program Synthesis with Large Language Models. CoRR, abs/2108.07732. Bacha, A.; and George, T. 2025. Training Feature Attribution for Vision Models. CoRR, abs/2510.09135. Bai, Y.; Lv, X.; Zhang, J.; Lyu, H.; Tang, J.; Huang, Z.; Du, Z.; Liu, X.; Zeng, A.; Hou, L.; Dong, Y.; Tang, J.; and Li, J. 2023. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. CoRR, abs/2308.14508. Chen, L.; Xu, D.; An, C.; Wang, X.; Zhang, Y.; Chen, J.; Liang, Z.; Wei, F.; Liang, J.; Xiao, Y.; and Wang, W. 2025a. PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention. CoRR, abs/2503.03588. Chen, L.; Zhao, H.; Liu, T.; Bai, S.; Lin, J.; Zhou, C.; and Chang, B. 2024. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision- Language Models. In Leonardis, A.; Ricci, E.; Roth, S.; Rus- sakovsky, O.; Sattler, T.; and Varol, G., eds., Computer Vi- sion - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part LXXXI, volume 15139 of Lecture Notes in Computer Science, 19– 35. Springer. Chen, Q.; Qin, L.; Liu, J.; Peng, D.; Guan, J.; Wang, P.; Hu, M.; Zhou, Y.; Gao, T.; and Che, W. 2025b. Towards Reason- ing Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models. CoRR, abs/2503.09567. Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. CoRR, abs/2110.14168. Deshmukh, A.; Raparti, V. Y.; and Hsu, S. 2025. Zen- Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs. CoRR, abs/2508.17593. Fu, Y.; Zhu, S.; Su, R.; Qiao, A.; Stoica, I.; and Zhang, H. 2024. Efficient LLM Scheduling by Learning to Rank. In Globersons, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J. M.; and Zhang, C., eds., Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. Gu, A.; Rozière, B.; Leather, H. J.; Solar-Lezama, A.; Syn- naeve, G.; and Wang, S. 2024. CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution. In Salakhut- dinov, R.; Kolter, Z.; Heller, K. A.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Forty-first Interna- tional Conference on Machine Learning, ICML 2024, Vi- enna, Austria, July 21-27, 2024, Proceedings of Machine Learning Research, 16568–16621. PMLR / OpenReview.net. Han, J.; Du, L.; Wu, Y.; Liang, G.; Zhou, X.; Zheng, W.; Han, D.; and Sun, Z. 2025. AdaV: Adaptive Text-visual Redirec- tion for Vision-Language Models. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the Association for Computational Linguistics, ACL 2025, Vi- enna, Austria, July 27 - August 1, 2025, volume ACL 2025 of Findings of ACL, 4985–4997. Association for Computa- tional Linguistics. Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021a. Measuring Massive Mul- titask Language Understanding. In 9th International Con- ference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021b. Measuring Mathematical Problem Solving With the MATH Dataset. In Vanschoren, J.; and Yeung, S., eds., Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual. Hu, C.; Huang, H.; Xu, L.; Chen, X.; Xu, J.; Chen, S.; Feng, H.; Wang, C.; Wang, S.; Bao, Y.; Sun, N.; and Shan, Y. 2024. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads. CoRR, abs/2401.11181. Huangyw, H.; Zhang, Y.; Cheng, N.; Li, Z.; Wang, S.; and Xiao, J. 2025. Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, volume ACL 2025 of Findings of ACL, 5174–5193. Association for Computational Linguistics. Jain, N.; Han, K.; Gu, A.; Li, W.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; and Stoica, I. 2025. Live- CodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In The Thirteenth Interna- tional Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. Jiang, G.; Feng, W.; Quan, G.; Hao, C.; Zhang, Y.; Liu, G.; and Wang, H. 2025. VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models. CoRR, abs/2509.19803. Jin, Y.; Wu, C.; Brooks, D.; and Wei, G. 2023. S 3 : Increas- ing GPU Utilization during Generative Inference for Higher Throughput. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Infor- mation Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. Kingma, D. P.; and Ba, J. 2015. Adam: A Method for Stochas- tic Optimization. In Bengio, Y.; and LeCun, Y., eds., 3rd In- ternational Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings. Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J.; Zhang, H.; and Stoica, I. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Flinn, J.; Seltzer, M. I.; Druschel, P.; Kaufmann, A.; and Mace, J., eds., Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, 611–626. ACM. Li, Z.; Zhong, J.; Zheng, Z.; Wen, X.; Xu, Z.; Cheng, Y.; Zhang, F.; and Xu, Q. 2025. Compressing Chain-of-Thought in LLMs via Step Entropy. CoRR, abs/2508.03346. Loshchilov, I.; and Hutter, F. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net. Nguyen, D.; Payani, A.; and Mirzasoleiman, B. 2025. Be- yond Semantic Entropy: Boosting LLM Uncertainty Quan- tification with Pairwise Semantic Similarity. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Findings of ACL, 4530–4540. Association for Computational Linguistics. Qiu, H.; Mao, W.; Patke, A.; Cui, S.; Jha, S.; Wang, C.; Franke, H.; Kalbarczyk, Z. T.; Basar, T.; and Iyer, R. K. 2024. Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction. CoRR, abs/2404.08509. Qiu, W.; and Srivastava, S. 2025. Batch Prompting Sup- presses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models. CoRR, abs/2511.04108. Ram, D.; Xia, W.; and Soatto, S. 2025. Learning to Focus: Fo- cal Attention for Selective and Scalable Transformers. CoRR, abs/2511.06818. Shaham, U.; Ivgi, M.; Efrat, A.; Berant, J.; and Levy, O. 2023. ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Un- derstanding. In Bouamor, H.; Pino, J.; and Bali, K., eds., Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Findings of ACL, 7977–7989. Association for Computational Lin- guistics. Shahout, R.; Malach, E.; Liu, C.; Jiang, W.; Yu, M.; and Mitzenmacher, M. 2025a. Don’t stop me Now: Embedding based Scheduling for LLMS. In The Thirteenth Interna- tional Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. Shahout, R.; Malach, E.; Liu, C.; Jiang, W.; Yu, M.; and Mitzenmacher, M. 2025b. Don’t stop me Now: Embedding based Scheduling for LLMS. In The Thirteenth Interna- tional Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. Sundararajan, M.; Taly, A.; and Yan, Q. 2017. Axiomatic Attribution for Deep Networks. In Precup, D.; and Teh, Y. W., eds., Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, 3319–3328. PMLR. Team, L. 2024. The Llama 3 Herd of Models. CoRR, abs/2407.21783. Wang, J.; Xu, W.; Yang, A.; Zhou, W.; Lu, L.; Li, H.; Wang, X.; and Zhu, J. 2025a. Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling. In Belgrave, D.; Zhang, C.; Montoya, L. N.; Lin, H.; Pascanu, R.; Koniusz, P.; Ghassemi, M.; Chen, N.; Ruíz, I. V. M.; and Loaiza-Bonilla, A., eds., Advances in Neural Informa- tion Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mex- ico, November 30 - December 5, 2025. Wang, S.; Zhang, G.; Zhang, L. L.; Shang, N.; Yang, F.; Chen, D.; and Yang, M. 2025b. LoongRL: Reinforcement Learn- ing for Advanced Reasoning over Long Contexts. CoRR, abs/2510.19363. Wu, T.; Zhang, J.; and Gao, Y. 2026. Furina: Fragmented Uncertainty-Driven Refusal Instability Attack. CoRR, abs/2605.26158. Xie, H.; Chen, Y.; Wang, L.; Hu, L.; and Wang, D. 2026. Predicting LLM Output Length via Entropy-Guided Repre- sentations. CoRR, abs/2602.11812. Xing, L.; Huang, Q.; Dong, X.; Lu, J.; Zhang, P.; Zang, Y.; Cao, Y.; He, C.; Wang, J.; Wu, F.; and Lin, D. 2024. PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction. CoRR, abs/2410.17247. Xu, B.; and Lu, Y. 2025. TECP: Token-Entropy Conformal Prediction for LLMs. CoRR, abs/2509.00461. Xu, Z.; Wang, Z.; Qian, Z.; Shi, D.; Tang, F.; Hu, M.; Su, S.; Zou, X.; Feng, W.; Mahapatra, D.; Peng, Y.; Lin, M.; and Ge, Z. 2026. Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding. CoRR, abs/2603.13366. Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Wei, H.; Lin, H.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Lin, J.; Dang, K.; Lu, K.; Bao, K.; Yang, K.; Yu, L.; Li, M.; Xue, M.; Zhang, P.; Zhu, Q.; Men, R.; Lin, R.; Li, T.; Xia, T.; Ren, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.; Wan, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; and Qiu, Z. 2024. Qwen2.5 Technical Report. CoRR, abs/2412.15115. Yang, N.; Lin, H.; Liu, Y.; Tian, B.; Liu, G.; and Zhang, H. 2025. Token-Importance Guided Direct Preference Opti- mization. CoRR, abs/2505.19653. Zhang, Q.; Cheng, A.; Lu, M.; Zhang, R.; Zhuo, Z.; Cao, J.; Guo, S.; She, Q.; and Zhang, S. 2025. Beyond Text- Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025, 20857–20867. IEEE. Zhang, Y.; Wang, Y.; Chen, J.; Qin, K.; Zhao, Y.; and Nguyen, C. 2026. LLM-based Embeddings: Attention Values En- code Sentence Semantics Better Than Hidden States. CoRR, abs/2602.01572. Zhou, J.; Lu, T.; Mishra, S.; Brahma, S.; Basu, S.; Luan, Y.; Zhou, D.; and Hou, L. 2023. Instruction-Following Evalua- tion for Large Language Models. CoRR, abs/2311.07911. Reproducibility Checklist Instructions for Authors: This document outlines key aspects for assessing repro- ducibility. Please provide your input by editing this .tex file directly. For each question (that applies), replace the “Type your response here” text with your answer. Example: If a question appears as Proofs of all novel claims are included (yes/partial/no) Type your response here you would change it to: Proofs of all novel claims are included (yes/partial/no) yes Please make sure to: • Replace ONLY the “Type your response here” text and nothing else. • Use one of the options listed for that question (e.g., yes, no, partial, or NA). • Not modify any other part of the command or any other lines in this document. You can this .tex file right before document of your main file or compile it as a stand-alone document. Check the instructions on your con- ference’s website to see if you will be asked to provide this checklist with your paper or separately. 1. General Paper Structure 1.1. Includes a conceptual outline and/or pseudocode de- scription of AI methods introduced (yes/partial/no/NA) yes 1.2. Clearly delineates statements that are opinions, hypoth- esis, and speculation from objective facts and results (yes/no) yes 1.3. Provides well-marked pedagogical references for less- familiar readers to gain background necessary to repli- cate the paper (yes/no) yes 2. Theoretical Contributions 2.1. Does this paper make theoretical contributions? (yes/no) yes If yes, please address the following points: 2.2. All assumptions and restrictions are stated clearly and formally (yes/partial/no) yes 2.3. All novel claims are stated formally (e.g., in theorem statements) (yes/partial/no) yes 2.4. Proofs of all novel claims are included (yes/par- tial/no) yes 2.5. Proof sketches or intuitions are given for complex and/or novel results (yes/partial/no) yes 2.6. Appropriate citations to theoretical tools used are given (yes/partial/no) yes 2.7. All theoretical claims are demonstrated empirically to hold (yes/partial/no/NA) yes 2.8. All experimental code used to eliminate or disprove claims is included (yes/no/NA) yes 3. Dataset Usage 3.1. Does this paper rely on one or more datasets? (yes/no) yes If yes, please address the following points: 3.2. A motivation is given for why the experiments are conducted on the selected datasets (yes/par- tial/no/NA) yes 3.3. All novel datasets introduced in this paper are in- cluded in a data appendix (yes/partial/no/NA) NA 3.4. All novel datasets introduced in this paper will be made publicly available upon publication of the pa- per with a license that allows free usage for research purposes (yes/partial/no/NA) NA 3.5. All datasets drawn from the existing literature (po- tentially including authors’ own previously pub- lished work) are accompanied by appropriate cita- tions (yes/no/NA) yes 3.6. All datasets drawn from the existing literature (po- tentially including authors’ own previously published work) are publicly available (yes/partial/no/NA) yes 3.7. All datasets that are not publicly available are de- scribed in detail, with explanation why publicly available alternatives are not scientifically satisficing (yes/partial/no/NA) NA 4. Computational Experiments 4.1. Does this paper include computational experiments? (yes/no) yes If yes, please address the following points: 4.2. This paper states the number and range of values tried per (hyper-) parameter during development of the paper, along with the criterion used for selecting the final parameter setting (yes/partial/no/NA) partial 4.3. Any code required for pre-processing data is included in the appendix (yes/partial/no) yes 4.4. All source code required for conducting and analyz- ing the experiments is included in a code appendix (yes/partial/no) yes 4.5. All source code required for conducting and analyz- ing the experiments will be made publicly available upon publication of the paper with a license that al- lows free usage for research purposes (yes/partial/no) yes 4.6. All source code implementing new methods have comments detailing the implementation, with ref- erences to the paper where each step comes from (yes/partial/no) yes 4.7. If an algorithm depends on randomness, then the method used for setting seeds is described in a way sufficient to allow replication of results (yes/par- tial/no/NA) yes 4.8. This paper specifies the computing infrastructure used for running experiments (hardware and soft- ware), including GPU/CPU models; amount of mem- ory; operating system; names and versions of relevant software libraries and frameworks (yes/partial/no) yes 4.9. This paper formally describes evaluation metrics used and explains the motivation for choosing these metrics (yes/partial/no) yes 4.10. This paper states the number of algorithm runs used to compute each reported result (yes/no) yes 4.11. Analysis of experiments goes beyond single- dimensional summaries of performance (e.g., aver- age; median) to include measures of variation, con- fidence, or other distributional information (yes/no) yes 4.12. The significance of any improvement or decrease in performance is judged using appropriate statisti- cal tests (e.g., Wilcoxon signed-rank) (yes/partial/no) partial 4.13. This paper lists all final (hyper-)parameters used for each model/algorithm in the paper’s experiments (yes/partial/no/NA) yes