Paper deep dive
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
Man Liu, Xingchen Liu, Xingjian Tian, Bing Lu, Shengkay Lyu, Shengquan Yin, Wenjing Huang, Zheng Wei, Hairui Zhao, Guangming Tan, Dingwen Tao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 8:00:03 AM
Summary
TACO (Tensor-parallel Adaptive COmmunication compression) is a robust FP8-based framework designed to mitigate communication overhead in large-scale Tensor-Parallel (TP) LLM training. It addresses the challenges of error accumulation and high-frequency communication by employing a data-driven reshaping strategy via an Adaptive Scale-Hadamard Transform and a Dual-Scale Quantization mechanism. TACO also features a highly fused compression operator to reduce memory traffic and kernel launch overhead. When integrated into a 3D-parallel training framework (combining DP, PP, and TP), TACO demonstrates significant throughput improvements (up to 1.87x for TP and 1.53x for 3D) on GPT and Qwen models while maintaining near-lossless accuracy.
Entities (9)
Relation Signals (5)
TACO → implements → FP8
confidence 100% · a robust FP8-based framework for compressing TP intermediate tensors.
TACO → integrateswith → SDP4Bit
confidence 100% · integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism (SDP4Bit)
TACO → uses → Adaptive Scale-Hadamard Transform
confidence 100% · First, we employ a data-driven reshaping strategy combined with an Adaptive Scale-Hadamard Transform
TACO → uses → Dual-Scale Quantization
confidence 100% · while its Dual-Scale Quantization mechanism ensures numerical stability throughout training.
TACO → improves → GPT Models
confidence 90% · Detailed experiments on GPT models and Qwen model demonstrate up to 1.87X end-to-end throughput improvement
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce significant computational overhead during compression. To this end, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework for compressing TP intermediate tensors. First, we employ a data-driven reshaping strategy combined with an Adaptive Scale-Hadamard Transform to enable high-fidelity FP8 quantization, while its Dual-Scale Quantization mechanism ensures numerical stability throughout training. Second, we design a highly fused compression operator to reduce memory traffic and kernel launch overhead, allowing efficient overlap with communication. Finally, we integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism to develop a compression-enabled 3D-parallel training framework. Detailed experiments on GPT models and Qwen model demonstrate up to 1.87X end-to-end throughput improvement while maintaining near-lossless accuracy, validating the effectiveness and efficiency of TACO in large-scale training.
Tags
Links
- Source: https://arxiv.org/abs/2604.24088v1
- Canonical: https://arxiv.org/abs/2604.24088v1
Trouble viewing inline? Open PDF directly →
Full Text
88,363 characters extracted from source content.
Expand or collapse full text
by TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training Man Liu Hangzhou Institute for Advanced Study, University of Chinese Academy of SciencesHangzhouChina liuman24@mails.ucas.ac.cn , Xingchen Liu Institute of Computing Technology, Chinese Academy of SciencesBeijingChina liuxingchen23s@ict.ac.cn , Xingjian Tian Hangzhou Institute for Advanced Study, University of Chinese Academy of SciencesHangzhouChina tianxingjian25@mails.ucas.ac.cn , Bing Lu Institute of Computing Technology, Chinese Academy of SciencesBeijingChina lubing@ict.ac.cn , Shengkai Lyu Institute of Computing Technology, Chinese Academy of SciencesBeijingChina lvshengkai23@mails.ucas.ac.cn , Shengquan Yin University of Science and Technology of ChinaHefeiChina yinshengquan@mail.ustc.edu.cn , Wenjing Huang Institute of Computing Technology, Chinese Academy of SciencesBeijingChina huangwenjing23@mails.ucas.ac.cn , Zheng Wei Institute of Computing Technology, Chinese Academy of SciencesBeijingChina weizheng@ncic.ac.cn , Hairui Zhao Institute of Computing Technology, Chinese Academy of SciencesBeijingChina zhaohairui@ict.ac.cn , Guangming Tan Institute of Computing Technology, Chinese Academy of SciencesBeijingChina tgm@ict.ac.cn and Dingwen Tao Institute of Computing Technology, Chinese Academy of SciencesBeijingChina taodingwen@ict.ac.cn (2026) Abstract. Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce significant computational overhead during compression. To this end, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework for compressing TP intermediate tensors. First, we employ a data-driven reshaping strategy combined with an Adaptive Scale–Hadamard Transform to enable high-fidelity FP8 quantization, while its Dual-Scale Quantization mechanism ensures numerical stability throughout training. Second, we design a highly fused compression operator to reduce memory traffic and kernel launch overhead, allowing efficient overlap with communication. Finally, we integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism to develop a compression-enabled 3D-parallel training framework. Detailed experiments on GPT models and Qwen model demonstrate up to 1.87× end-to-end throughput improvement while maintaining near-lossless accuracy, validating the effectiveness and efficiency of TACO in large-scale training. Tensor parallelism, quantization, distributed training, communication compression, large language model. †journalyear: 2026†copyright: c†conference: The 35th International Symposium on High-Performance Parallel and Distributed Computing; July 13–16, 2026; Cleveland, OH, USA†booktitle: The 35th International Symposium on High-Performance Parallel and Distributed Computing (HPDC ’26), July 13–16, 2026, Cleveland, OH, USA†doi: 10.1145/3806645.3807584†isbn: 979-8-4007-2640-8/2026/07†ccs: Software and its engineering Message passing†ccs: Theory of computation Data compression 1. Introduction The rapid scaling of large language models (LLMs) to tens of billions, hundreds of billions, and even trillion-parameter scales has driven the adoption of increasingly sophisticated distributed training strategies (Lamprecht et al., 2025; Chowdhery et al., 2023; Jiang et al., 2024; Smith et al., 2022). Among them, 3D parallelism — comprising data parallelism (DP), tensor parallelism (TP), and pipeline parallelism (P)—has emerged as the dominant paradigm for training ultra-large models (Zheng et al., 2024; Chen et al., 2024). However, as model scale grows, training performance becomes increasingly constrained by communication rather than computation. Recent system studies show that communication can account for over 50% of the total training time (Narayanan et al., 2021; Wang et al., 2024). Different parallelism strategies exhibit fundamentally different communication patterns and compression challenges. DP performs relatively low-frequency gradient synchronization, while P primarily relies on lightweight point-to-point communication without global synchronization, making both comparatively amenable to effective communication compression (Rajbhandari et al., 2020; Alistarh et al., 2017; Huang et al., 2019; Narayanan et al., 2019). In contrast, TP requires frequent, tightly synchronized communication to exchange intermediate tensors (sharded activations and gradients) during both forward and backward passes (Shoeybi et al., 2020; Brown et al., 2020; Touvron et al., 2023; Grattafiori and others, 2024; Workshop et al., 2023). As a result, TP communication lies on the critical execution path, accounts for over 50–60% of total communication time (see Figure 1(a)), and is notoriously difficult to overlap with computation (Narayanan et al., 2021; Shoeybi et al., 2020; Wang et al., 2024).This creates a fundamental challenge—compression overhead under frequent communication—as compression overhead must be minimized, which demands extreme operator-level optimization while carefully and strictly controlling error. To mitigate communication overhead, prior work has explored various compression techniques, including quantization (Zhang et al., 2023; Markov et al., 2023), sparsification (Chen et al., 2020; Li and Hoefler, 2022), and low-rank approximation (Vogels et al., 2019; Zhang et al., 2023). These approaches have been successfully applied in DP and P. Representative methods such as SDP4bit (Jia et al., 2024) and TahQuant (He et al., 2025) demonstrate that carefully designed compression strategies can significantly reduce communication cost while effectively preserving training stability in their respective domains (Dettmers et al., 2023). In parallel, many inference-oriented compression methods—such as SmoothQuant (Xiao et al., 2024), FlatQuant (Sun et al., 2025), QuaRot (Ashkboos et al., 2024b), and GPTQ (Frantar et al., 2023)—primarily leverage INT4/INT8 formats tailored for static weights and activations. However, communication compression for TP remains largely underexplored, particularly in training. Unlike the coarse-grained and relatively infrequent communication patterns in DP and P, TP necessitates multiple rounds of compression within every Transformer block communication, giving rise to a critical challenge: error accumulation under high-frequency communication. Even minor quantization errors in TP intermediate tensors can propagate through attention and residual connections, amplifying during backpropagation (Xu et al., 2024) and, as shown in Figure 1(b), potentially destabilizing training or causing divergence. Moreover, TP communication compression is much more challenging during training than inference. Unlike inference, pretraining involves thousands of forward and backward passes, repeatedly injecting and amplifying quantization noise, so directly applying inference-oriented compression to TP tensors often leads to catastrophic training failure. Through a systematic analysis of the distributions of TP intermediate tensors, we make two key observations that clearly and fundamentally highlight the fundamental difficulty of their compression. First, these tensors are strongly dominated by small-magnitude values with highly concentrated, zero-centered distributions. This characteristic constitutes the primary challenge for TP communication compression (distinct TP intermediate tensor characteristics), as conventional methods fail to adequately capture the numerical subtleties of these extremely small values, resulting in significant information loss and progressively severe error accumulation, especially under repeated synchronization along the TP computation path. Second, regarding quantization preprocessing, prior work commonly applies fixed transformations—such as rotations or Hadamard transforms—to balance value distributions (Chee et al., 2023; Tseng et al., 2024; Ashkboos et al., 2024a; Savkin, 2025; Egiazarian et al., 2026). While these transforms can spread variance across dimensions in other contexts, they are data-independent and lack the adaptability required to disperse the dense, zero-centered clusters inherent in TP intermediate tensors. Consequently, static transforms alone are insufficient to enable high-fidelity and reliable quantization for tightly synchronized TP communication. (a) Execution time breakdown across different models and scales (b) Validation loss comparison between the baseline and TP using TahQuant compression under 3D parallelism on GPT-350M Figure 1. Communication overhead and impact of quantization on training performance and convergence The left plot shows execution time breakdown across different models and scales. The right plot compares validation loss between baseline and tensor parallel training with TahQuant compression under 3D parallelism. Motivated by these challenges, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework explicitly designed to efficiently compress intermediate tensors within TP training. TACO is specifically designed to address the critical challenges of error accumulation during TP pretraining, the distinct characteristics of TP intermediate tensors, and the high compression overhead induced by tightly synchronized communication. Unlike prior works, TACO directly targets latency-critical intermediate tensor communication in TP training, where both numerical fidelity and efficiency are essential. It introduces a data-driven reshaping strategy that dynamically adapts to the statistical distributions of TP intermediate tensors, enabling high-fidelity FP8 quantization in practice. To ensure numerical stability, the framework employs a dual-scale quantization technique that preserves precision across all training stages. In addition, we develop a highly fused compression operator to minimize memory traffic and kernel launch overhead, significantly enhancing computational efficiency. To the best of our knowledge, TACO is the first 3D parallel communication compression system that achieves stable convergence in large-scale TP training scenarios. Our primary contributions are summarized as follows. • We systematically analyze TP intermediate tensors, whose dense, near-zero distributions render INT8 quantization inadequate while favoring FP8. Nevertheless, FP8’s representational capacity under low-bit remains limited, and standard Hadamard transform fails to sufficiently disperse these highly concentrated values. This analysis provides crucial guidance for designing effective TP intermediate tensor compression methods. • We propose TACO, a TP communication compression framework. TACO introduces the Adaptive Scale–Hadamard Transform to dynamically reshape the distribution of TP intermediate tensors for high-fidelity FP8 quantization, and implements Dual-Scale Quantization to keep all communicated values within the FP8 representable range across training, effectively preventing error amplification under frequent TP communication. • We develop a highly fused compression operator that combines all quantization steps in a single kernel, reuses warp-level reductions, and accesses metadata without extra copies. This design reduces memory traffic and kernel launches, accelerates computation, and enables efficient overlap with communication via tight integration with backends and fine-grained scheduling, thus mitigating high-frequency compression overhead. • We integrate TACO with state-of-the-art compression methods for DP (SDP4Bit) and P (TahQuant), enabling fully compression-enabled 3D-parallel training. In addition, evaluation results on GPT models and Qwen2.5-7B show that our method achieves up to 1.87× end-to-end throughput improvement for TP compression and up to 1.53× under full 3D-parallel compression, while maintaining near-lossless training accuracy. Figure 2. Illustration of TP in a Transformer block, including the AllReduce operations required during forward and backward passes. Diagram of tensor parallelism in a Transformer block showing how computation is split across devices and where AllReduce operations occur during forward and backward passes. 2. Background and Related Work 2.1. Tensor Parallelism and Its Communication Currently, distributed training of LLMs encounters increasingly significant communication bottlenecks stemming from different parallelism strategies, including DP, P, and TP. Among these, TP is the primary contributor to intra-node communication overhead, as it requires frequent synchronization of intermediate activations and gradients across GPUs. TP enables the scaling of Transformer-based LLMs across multiple GPUs by exploiting the inherent parallelism of large matrix multiplications (Shoeybi et al., 2020; Anthony et al., 2024). In TP, weight matrices are partitioned along the hidden dimension—either by rows or columns—allowing multiple devices to collaboratively compute a single Transformer block. As illustrated in Figure 2, mainstream implementations apply TP to both the MLP and Attention blocks. During the forward pass, column-parallel linear layers produce partial outputs on each GPU, which must be synchronized via collective communication before being consumed by subsequent layers. Similarly, during backpropagation, row-parallel layers require collective synchronization of gradients prior to parameter updates. Because TP collectives are invoked synchronously at every layer and iteration, TP constitutes the primary intra-node communication pathway in 3D-parallel training systems. Whether implemented as AllReduce or, when combined with sequence parallelism (SP), decomposed into AllGather and Reduce-Scatter operations along the sequence dimension (Li et al., 2023), TP communication is highly sensitive to both bandwidth and latency. As model widths scale to tens of thousands and depths grow to hundreds of layers, the communication volume incurred by TP increases rapidly, leading to severe bottlenecks even on advanced GPU interconnects (Shoeybi et al., 2020; Narayanan et al., 2021; Brown et al., 2020; Gale et al., 2022). Consequently, TP communication consistently accounts for a substantial fraction of the overall training communication cost. Figure 3. FP8 Format Specifications: E4M3 vs. E5M2. E4M3 provides a larger dynamic range with relatively lower precision, while E5M2 offers higher precision but a smaller dynamic range. Comparison of two FP8 formats. E4M3 uses fewer exponent bits and more mantissa bits, providing higher precision but a smaller dynamic range. E5M2 uses more exponent bits and fewer mantissa bits, resulting in a larger dynamic range but lower precision. 2.2. Communication Compression in Distributed Training Communication compression has been extensively studied in distributed training, particularly for data-parallel gradient synchronization. In DP, gradients exhibit relatively stable distributions and tolerate moderate approximation errors, enabling diverse compression techniques, including stochastic quantization (Alistarh et al., 2017), low-bit gradient methods (Lin et al., 2020), and low-precision optimizers (Zhang et al., 2023; Markov et al., 2023; Dettmers et al., 2023). Quantization-aware training methods, such as LLM-QAT (Liu et al., 2023), further enhance robustness under low-precision representations by incorporating quantization effects during training, mainly targeting weights and activations. More recent system-level designs further push DP gradient communication toward ultra-low precision. For example, SDP4Bit achieves near-4-bit communication in sharded DP training (Jia et al., 2024), EDGC uses entropy-driven dynamic gradient compression to reduce latency in GPT training (Yi et al., 2025), and TAGC introduces a Transformer-aware hierarchical compression scheme (Polyakov et al., 2025). Collectively, these methods demonstrate that aggressive communication compression is feasible in DP while preserving convergence. Beyond DP, communication compression has also been studied in P training, where reducing inter-stage communication can significantly improve throughput. Representative approaches include dynamic precision control (Chen et al., 2021), activation difference compression (Wang et al., 2022), and joint activation–gradient compression with error compensation (Rudakov et al., 2023). More recently, TahQuant introduces fine-grained activation quantization along the P communication path to improve accuracy preservation (He et al., 2025). While effective in P settings, these methods typically rely on relatively loose synchronization constraints. Recent system-level efforts further explore communication-efficient training via coordinated compression and execution optimization, e.g., Tango (Chen et al., 2023), which reduces computation and communication overhead through quantization-aware scheduling. In contrast, TP exhibits fundamentally different communication characteristics. TP intermediate tensors are smaller, exchanged at higher frequency, and reside on the critical computation path. Consequently, TP communication is highly sensitive to numerical perturbations, and directly applying DP- or P-oriented compression methods can severely degrade numerical stability and hinder convergence. Although recent efforts explore low-bit TP communication (Dong et al., 2024; Li et al., 2024), they primarily target inference workloads, which are more tolerant of approximation errors. Furthermore, in end-to-end 3D-parallel training systems, existing designs often adopt conservative strategies that leave TP communication uncompressed to avoid catastrophic training divergence (Xu et al., 2024). To date, there remains no communication compression scheme that enables stable convergence for TP training of large language models. 2.3. Numerical Challenges of Low-Bit Quantization in TP Communication Low-bit quantization is a common strategy for reducing communication and memory overhead in large-scale model training. In practice, INT8 quantization is widely adopted in inference due to its simplicity and favorable efficiency–accuracy trade-offs (Sun et al., 2025; Ashkboos et al., 2024b). However, its applicability to TP communication during pretraining is constrained by much stricter numerical stability requirements. INT8 maps floating-point values to a uniform integer grid using a single scaling factor, with quantization defined as QINT8(x)=round(x/Δ)Q_INT8(x)=round(x/ ) and Δ=|x|max/127 =|x|_ /127. While hardware-efficient, this uniform quantization provides limited effective resolution for values concentrated near zero, which are frequently observed in TP intermediate tensors (Kuzmin et al., 2022). In high-frequency, tightly synchronized TP communication, small mismatches in scaling factors across layers or shards can introduce rounding and saturation effects, leading to distorted gradient aggregation and unstable optimization dynamics (Micikevicius et al., 2018). In contrast, reduced-precision floating-point formats, such as FP8, preserve scale information through explicit exponent encoding, enabling more robust handling of heterogeneous and dynamically changing tensor values. An FP8 number is represented as x=(−1)s⋅2e−B⋅(1+f)x=(-1)^s· 2^e-B·(1+f), where s, e, and f denote the sign, exponent, and mantissa, respectively. Common FP8 variants include E4M3 and E5M2 (Shen et al., 2024). The bit-level structures of these variants are illustrated in Figure 3.This inherent adaptivity makes FP8 particularly well-suited for low-bit TP communication during pretraining, addressing key numerical limitations of INT8 in such settings (Micikevicius et al., 2022). Moreover, FP8 is currently being widely adopted in modern AI accelerators, including NVIDIA and AMD GPUs, and is expected to become a broadly supported type in the future—similar to how FP16 gained widespread hardware adoption—making it a timely and forward-looking choice for large-scale training. While FP8 mitigates several numerical limitations inherent to INT8 quantization, effective TP communication compression depends not only on the numerical format itself, but also critically on how well it aligns with the statistical properties and synchronization patterns of TP intermediate tensors. In the following section, we analyze why FP8 is particularly well suited for TP communication and how its representational characteristics can be leveraged to achieve both high compression efficiency and stable training. Figure 4. Histogram of TP communication data. Histogram showing the distribution of tensor parallel communication data values, illustrating how frequently different value ranges occur and revealing a skewed distribution pattern. 3. Why FP8 is Suitable for TP Communication Compression? 3.1. Distribution of TP Intermediate Tensors We analyze the intermediate tensors involved in TP AllReduce (see Figure 4). Our results indicate that these tensors are highly concentrated around zero while simultaneously exhibiting a long-tail distribution. Let the TP intermediate tensor be X=xii=1NX=\x_i\_i=1^N with probability density pX(x)p_X(x), such that the zero-centered region satisfies ∫−ϵpX(x)x≈1 _-ε^εp_X(x)\,dx≈ 1, where ϵ→0+ε→ 0^+. The dense and long-tail subsets can be defined as X0=xi∈X∣|xi|≤ϵX_0=\x_i∈ X |x_i|≤ε\ and XL=xi∈X∣|xi|>ϵX_L=\x_i∈ X |x_i|>ε\, capturing the extremely small and relatively large values, respectively. In other words, most values transmitted during TP communication are extremely small and densely clustered near zero, while only a few occupying the long tail. For such distributions, quantization must provide sufficient resolution around X0X_0; otherwise, small-magnitude values may collapse to the same quantization level, resulting in information loss. 3.2. INT8 Incompatibility with TP Communication Compression INT8 quantization typically adopts a fixed scale with zero-point, performing uniform quantization over the entire value range. However, this approach is unsuitable for TP intermediate tensors (see Figure 4). INT8 maps the range using a fixed step size Δ , while TP intermediate tensors are highly concentrated near zero, forming a sharp peak. Figure 5 shows the visual comparison of INT8 and FP8 representations. INT8 adopts uniform quantization with evenly spaced representable values, imposing the same quantization resolution on both high-density and low-density regions. As a result, the numerous small-magnitude values clustered around zero incur large relative quantization errors, leading to a pronounced degradation in training accuracy. Let the quantization error be (1) ei=xi−x^i,x^i=QINT8(xi),i=1,…,N.e_i=x_i- x_i, x_i=Q_INT8(x_i), i=1,…,N. Due to the uniform step, multiple dense zero-centered values may map to the same integer, leading to collisions: (2) ∃xi,xj∈X0,QINT8(xi)=QINT8(xj),i≠j.∃ x_i,x_j∈ X_0, Q_INT8(x_i)=Q_INT8(x_j), i≠ j. Consequently, the mean squared error in the dense region remains significant, severely degrading training accuracy. Our experiments on the tensors show that INT8 error is approximately uniformly distributed across the range (see Figure 6), consistent with its uniform step mechanism. Therefore, INT8 can’t provide fine-grained resolution in the dense zero region and is unsuitable for TP. Figure 5. Data distribution characteristics of INT8 and FP8. Distributions of values under INT8 and FP8 quantization. INT8 shows uniform quantization with evenly spaced bins, while FP8 exhibits non-uniform distribution with denser representation near zero and a wider dynamic range. 3.3. The Mathematical Suitability of FP8 FP8 addresses the limitations of INT8 by providing an exponentially scaled representation that achieves high precision near zero while maintaining a wide dynamic range. Figure 5 illustrates the one-dimensional density of FP8 representable values, characterized by a pronounced concentration near zero and a long tail, which highlights FP8’s suitability for compressing TP intermediate tensors. For values in the dense zero-centered region X0X_0, the quantization error is bounded by the unit in the last place (ULP): (3) |eiFP8|≤ULP(xi)=2⌊log2|xi|⌋−m,xi∈X0,|e_i^FP8| (x_i)=2 _2|x_i| -m, x_i∈ X_0, where m is the mantissa width. This implies that smaller values (lower exponent) are represented with finer precision and denser quantization points, closely matching the dense, small-magnitude peak observed in TP intermediate tensors. For large-magnitude values in the long-tail region XLX_L, FP8 step size grows exponentially, providing sufficient dynamic range. As a result, the overall tensor-level quantization error is significantly lower than INT8, while preserving high-fidelity representation near zero. This property can be approximated for E4M3 FP8 as (4) QFP8(x)≈sign(x)⋅2E−B⋅(1+M/2m),Q_FP8(x) (x)· 2^E-B· (1+M/2^m ), where the exponent E and mantissa M are encoded with 4 and 3 bits, respectively. The exponential step ensures minimal error for critical small values, while still covering the long-tail region, achieving local high precision and global dynamic range simultaneously. The quantization errors of int8 and FP8 to TP intermediate tensor are as shown in Figure 6, which demonstrates that the error of fp8 is obviously smaller than that of INT8. Overall, TP intermediate tensors exhibit distributions that are heavily skewed toward low-magnitude values, rendering uniform INT8 quantization susceptible to substantial precision loss. In contrast, FP8’s exponent-based representation provides finer granularity in the near-zero regime, making it more suitable for large-scale distributed training. Figure 6. Quantization errors of INT8 and FP8. Comparison of quantization errors between INT8 and FP8, showing how error magnitudes are distributed and highlighting differences in precision loss under the two formats. 4. Design of TACO Figure 7. Overview of TACO. TP intermediate tensors are compressed via adaptive rescaling, the ASH transform, and dual-scale quantization, with kernel fusion and communication–computation co-optimization. Decompression is performed using a fused kernel for efficient reconstruction. System overview of TACO. Tensor parallel intermediate tensors are first compressed using adaptive rescaling, the ASH transform, and dual-scale quantization. The system further applies kernel fusion and communication–computation co-optimization to improve efficiency. On the receiver side, a fused kernel is used for fast decompression and reconstruction of tensors. 4.1. High-Level Overview Motivated by the observations above, we propose TACO. TACO specifically targets the compression of intermediate tensors exchanged in TP communication. It enhances the adaptability of FP8 quantization by dynamically regulating the energy distribution of these tensors, effectively mitigating the quantization error arising from their highly concentrated nature near zero. Furthermore, it enables high-throughput data transmission on GPU clusters through a unified kernel design. The overview of TACO is shown in Figure 7. In the compression workflow, the limitations of standard Hadamard are first analyzed in Section 4.2.1. To address the identified “Zero-Collapse” issue, the Adaptive Scale-Hadamard (ASH) transform is detailed in Section 4.2. This process involves calculating the second-order raw moment and scale factors (①), followed by the ASH transform (②). In Section 4.3, the Dual-Scale (DS) quantization (③) is applied to map the transformed data to the FP8 format. Since executing these operations sequentially incurs significant memory overhead, kernel fusion (④) is required to combine variance computation, ASH transform, and DS quantization into a single operator. The specific implementation and kernel fusion methods are detailed in Section 4.4.1. Furthermore, to mitigate the overhead of compression latency, an overlap strategy (⑤) is employed to parallelize computation and communication. The system performance optimization and overlap details are discussed in Section 4.4.2. In the decompression workflow, the data is first processed by the dequantization step, followed by the inverse DS quantization and inverse ASH transformation. We optimize the decompression process by fusing these operations into a unified fused_ash _decompress_kernel. This fusion eliminates intermediate memory writes and reduces kernel launch overhead. 4.2. Adaptive Scale–Hadamard Transform 4.2.1. Motivation: Limitation of Standard Hadamard Despite the superior near-zero resolution of FP8 compared to INT8, its application to TP intermediate tensors is hindered by their extreme value concentration. We identify “Zero-Collapse” as the root cause of failure, where direct quantization or standard transformations cannot effectively map data into the representable range of FP8. Direct FP8 Quantization. Due to the lack of adaptive scaling, the vast majority of low-magnitude values fall into the subnormal range or underflow directly to zero, resulting in severe fidelity loss. Inefficacy of Standard Hadamard. While the standard Hadamard transform is effective for spreading outliers, it is fundamentally an isometric transformation that preserves the Euclidean norm (L2L_2 energy) of the input vector. Consequently, data blocks with inherently low energy remain confined to a narrow numerical range even after rotation, failing to occupy the effective bits of FP8. As illustrated in Figure 8, the distribution of the transformed data remains sharply peaked around zero, indicating that the standard Hadamard transform fails to sufficiently disperse these dense, low-magnitude clusters. Consequently, this results in persistent underutilization of FP8’s high-precision dynamic range. 4.2.2. ASH Transform To overcome the limitations of the standard Hadamard transform and prevent zero-collapse, we propose the Adaptive Scale–Hadamard Transform, which combines block-wise energy rescaling with orthogonal Hadamard rotation for high-fidelity quantization. This design amplifies low-magnitude blocks while appropriately scaling high-magnitude blocks, enabling the transformed data to fully exploit FP8’s representable range. Let ∈ℝNX ^N denote the flattened TP intermediate tensor prior to transfer. We partition X into M contiguous blocks of size B: (5) =[G1,G2,…,GM],Gk∈ℝB,X=[G_1,G_2,…,G_M], G_k ^B, where B is chosen to fit entirely in GPU shared memory for low-latency, block-local processing. The effect of ASH transform is clearly visualized in Figure 8. Compared to the standard Hadamard transform, ASH transform effectively disperses the dense near-zero clusters, expanding them into FP8’s high-precision quantization range and yielding a more balanced distribution. Block-wise Adaptive Rescaling. To explicitly control the numerical magnitude of each block, ASH begins with block-wise adaptive energy rescaling. We specifically adopt the second-order raw moment (energy) rather than variance to estimate the local scale. This is motivated by the observation that TP intermediate tensors are typically zero-centered; thus, the raw energy captures the effective signal magnitude required for quantization without incurring the computational overhead of mean centering. For block GkG_k, we calculate its root mean square (RMS) amplitude as follows: (6) σk=1B∑j=1BGk,j2+ϵ, _k= 1B _j=1^BG_k,j^2+ε, where ϵε is a small constant ensuring numerical stability for near-zero blocks. We then compute an adaptive scaling factor to map each block’s energy to a reference target level: (7) αk=τσk, _k= τ _k, where τ denotes a constant target energy aligned with FP8’s effective dynamic range. This scaling strategy amplifies low-magnitude blocks while attenuating higher-magnitude ones, effectively normalizing the dynamic range across all blocks. The rescaled block is computed element-wise as G~k=αk⋅Gk G_k= _k· G_k, where each block GkG_k has its own adaptive scale αk _k. This ensures that every block is normalized independently, providing distinguishable numerical ranges across blocks before the rotation stage. By doing so, we effectively maximize the utilization of FP8’s high-precision range and minimize quantization errors for near-zero values. Importantly, this computation involves only lightweight reductions and element-wise multiplications, making it highly efficient and naturally suited for massively parallel GPU execution, where thousands of blocks can be processed concurrently with minimal overhead. Figure 8. Distribution of tensors before and after Hadamard-based transformations. AS-Hadamard redistributes densely clustered near-zero values into FP8’s high-precision quantization range, unlike the standard Hadamard transform. Comparison of tensor distributions before and after Hadamard-based transformations. The AS-Hadamard transform redistributes values that are densely concentrated near zero into a more uniform distribution, better matching FP8 quantization precision, whereas the standard Hadamard transform does not achieve this effect as effectively. Orthogonal Hadamard Rotation. After rescaling, each block undergoes an orthogonal Walsh–Hadamard transform: (8) k=1BBG~k,Z_k= 1 BH_B G_k, where BH_B denotes the Hadamard matrix of order B. The factor 1/B1/ B is incorporated to explicitly enforce orthogonality, thereby preserving the energy of the transformation. Since the normalized Hadamard matrix is bothsymmetric and orthogonal (BT=B−1=1BBH_B^T=H_B^-1= 1 BH_B), the transformation is exactly invertible. Crucially, because the rotation is energy-preserving, it redistributes the dense low-magnitude clusters without disrupting the block-wise scaling established during the rescaling step. The transformed data exhibits an approximate zero-mean, Gaussian-like distribution, which naturally aligns with the non-uniform exponent density of FP8. As illustrated in Figure 8, ASH significantly improves the utilization of FP8’s effective dynamic range compared to the standard Hadamard transform. For efficiency, this operation is implemented using the Fast Walsh–Hadamard Transform (FWHT) in an in-place manner within shared memory, reducing computational complexity from O(B2)O(B^2) to O(BlogB)O(B B). 4.3. Dual-Scale FP8 Quantization Following the ASH transform, the rotated tensor kZ_k exhibits a more favorable Gaussian-like distribution. Nevertheless, careful scale management remains critical due to the rigid upper bound of the FP8 format. Without proper scaling, high-magnitude values can exceed the maximum representable range (QmaxQ_ ), causing overflow or severe saturation. Such numerical instability is a primary contributor to training divergence and abrupt loss spikes. Post-Rotation Quantization Scale. To ensure that all values strictly fit in the FP8 representable range, we compute a block-wise post-rotation scale based on the maximum absolute value in the block: (9) sk=max(|k|)Qmax,s_k= (|Z_k|)Q_ , where max(|k|) (|Z_k|) denotes the maximum value in block k, and QmaxQ_ is the largest representable FP8 value. The data are then quantized element-wise to FP8: (10) k=CvtFP8(ksk),q_k=Cvt_FP8\! ( Z_ks_k ), where CvtFP8Cvt_FP8 denotes the intrinsic conversion function (NVIDIA’s _nv_cvt_float_to_fp8). By strictly enforcing this mapping, numerical overflow is effectively prevented, and the available bit-width is maximally and efficiently utilized throughout computation. Algorithm 1 TACO: Tensor-Parallel Adaptive Communication Compression 1:Input intermediate tensor X; block size B; target energy τ; FP8 maximum value QmaxQ_ 2:Reconstructed intermediate tensor ′X 3:Sender-side: Fused Compression Kernel 4:Partition X into M contiguous blocks G1,G2,…,GM\G_1,G_2,…,G_M\ of size B 5:for each block GkG_k in parallel do 6: Load GkG_k into shared memory / registers 7: Block-wise Adaptive Rescaling 8: σk←1B∑j=1BGk,j2+ϵ _k← 1B _j=1^BG_k,j^2+ε 9: αk←τ/σk _k←τ/ _k 10: G~k←αk⋅Gk G_k← _k· G_k 11: Orthogonal Hadamard Rotation 12: k←1BFWHT(G~k)Z_k← 1 BFWHT( G_k) 13: Post-Rotation FP8 Quantization 14: sk←max(|k|)/Qmaxs_k← (|Z_k|)/Q_ 15: k←CvtFP8(k/sk)q_k _FP8(Z_k/s_k) 16:end for 17:Communicate (k,αk,sk)k=1M\(q_k, _k,s_k)\_k=1^M using TP collectives 18:Receiver-side: Decompression Kernel 19:for each received tuple (k,αk,sk)(q_k, _k,s_k) in parallel do 20: ^k←CvtFP32(k)⋅sk Z_k _FP32(q_k)· s_k 21: G^k←1BFWHT(^k) G_k← 1 BFWHT( Z_k) 22: Gk′←G^k/αkG _k← G_k/ _k 23: Append Gk′G _k to ′X 24:end for 25:return ′X Dual-Scale Reconstruction. To enable high-fidelity reconstruction, TACO utilizes two distinct scalars per block to separately control the data distribution and the quantization range. Both the adaptive rescaling factor αk _k (Eq. 7), which normalizes block energy, and the quantization scale sks_k (Eq. 9), which prevents numerical overflow, are transmitted alongside the compressed tensor. Reconstruction strictly follows the reverse order of the compression sequence: (11) ^k Z_k =CvtFP32(k)⋅sk, =Cvt_FP32(q_k)· s_k, (12) G^k G_k =1BB^k, = 1 BH_B Z_k, (13) Gk′ G _k =G^k/αk. = G_k/ _k. Specifically, Eq. 11 recovers the post-rotation magnitude, Eq. 12 applies the inverse orthogonal Hadamard transform to restore the original ordering of the data, and Eq. 13 reverses the adaptive rescaling. By decoupling these two factors, this mechanism mitigates the conflict between resolving small values and containing large values. It ensures that low-magnitude clusters are safeguarded against underflow, while high-magnitude blocks remain strictly bounded within the FP8 representable range throughout training. The communication overhead of transmitting two scalars per block is negligible compared to the substantial bandwidth savings achieved by FP8 compression. Despite this minimal overhead, the additional scaling metadata is critically important for maintaining numerical stability and ensuring stable training convergence. As a result, TACO enables highly efficient FP8-based TP communication, while having minimal impact on overall training accuracy. 4.4. Performance Optimization 4.4.1. System-Level Kernel Fusion Optimizations. TACO implements a fully fused compression kernel (Algorithm 1) that integrates adaptive rescale factor computation, the ASH transform, and DS quantization into a single GPU kernel (Figure 9). By eliminating redundant global memory accesses and kernel launches, this design significantly increases arithmetic intensity and improves overall efficiency. In contrast, naïve implementations typically launch separate reduction kernels to compute block-wise variance for adaptive energy normalization and the post-rotation maximum magnitude for FP8 scaling, incurring substantial overhead. Figure 9. Comparison of Naïve Multi-Kernel Execution vs. TACO Single Kernel Execution. The upper part illustrates the high memory traffic in traditional methods, while the lower part demonstrates the efficiency of fused kernel execution with parallel CTAs. Comparison between naive multi-kernel execution and TACO single-kernel execution. The naive approach requires multiple kernel launches and incurs high memory traffic due to repeated intermediate data movement. In contrast, TACO fuses operations into a single kernel with parallel CTAs, significantly reducing memory traffic and improving execution efficiency. TACO effectively eliminates this redundancy by coalescing both reductions within a single fused kernel using warp-level primitives and shared memory. Concretely, each thread locally computes both the squared input value for variance estimation and the absolute value of its rotated output for post-rotation scaling. Warp-level shuffle operations are then used to simultaneously aggregate partial sum and partial maximum with minimal synchronization overhead, after which shared memory efficiently finalizes the block-level reductions to produce σk _k and sks_k without launching auxiliary kernels. This design significantly reduces global memory traffic and kernel launch overhead, enabling fully block-local computation for both normalization and quantization scale derivation. 4.4.2. System performance optimization To better integrate TACO with communication protocols, we implement two optimizations: refining buffer distribution and integrating TACO with COCCL, a compressible collective communication library, followed by performance tuning, ensuring efficient TP communication across GPUs, which allows TACO to be seamlessly embedded into existing collective primitives with minimal protocol modification. First, each compressed block requires two scalar parameters: the adaptive pre-scaling factor αk _k and the post-rotation quantization scale sks_k. TACO stores these scalars contiguously with the corresponding FP8-compressed payload in global memory, enabling a zero-copy metadata layout. During decompression, the kernel retrieves αk _k and sks_k via simple pointer arithmetic, avoiding explicit metadata copies as well as additional communication launches for scale gathering. Moreover, this layout enables fully coalesced global memory accesses during reconstruction, as threads within a block read contiguous FP8 values followed immediately by their associated metadata,further minimizing memory latency. Second, integrating TACO into traditional ring- or tree-based communication algorithms results in frequent compression within a single communication, leading to significant performance degradation and accumulated compression errors. To reduce compression overhead, we integrate TACO into COCCL (Liu et al., 2026), a compressible collective communication protocol built on NCCL. We tune this library as shown in Figure 10 and find that, TP communication scenario—where all communication occurs within a node—using COCCL’s two-shot AllReduce as the communication algorithm achieves the best overall performance. The two-shot algorithm decomposes AllReduce into ReduceScatter and AllGather. The ReduceScatter phase consists of one compressed AlltoAll operation followed by a single local reduction. This design effectively reduces TACO execution to two operations per communication round, minimizing compression frequency and increasing compression block granularity. As a result, it significantly lowers execution overhead when integrating TACO with communication while preserving accuracy. In addition, COCCL incorporates a two-level overlap strategy. Through optimization of the overlap granularity, we determine that overlapping data blocks of size 64 MB yields optimal performance, further masking the computational latency of TACO. Figure 10. Throughput comparison. Left: Performance of different shot methods. Right: Impact of overlap chunk size on throughput. Throughput comparison across different configurations. The left plot evaluates the performance of different shot methods, while the right plot analyzes the impact of overlap chunk size on throughput, demonstrating the trade-off between communication-computation overlap and execution efficiency. 5. Evaluation 5.1. Experimental Setup Platforms. We evaluate TACO on a GPU-centric cluster comprising two nodes, each equipped with two Intel XEON(R) Platinum 8558 CPUs (192 cores) and eight NVIDIA H100 SXM5 GPUs with 80 GB memory. The nodes are interconnected via four 400Gbps InfiniBand links, providing an aggregate inter-node bandwidth of 1.6 Tbps. The software stack includes CUDA 12.6 and NVIDIA driver 550.90.07. Baselines. We compare TACO against the following communication strategies: Baseline (w/o Comp), which performs standard distributed training without communication compression; TahQuant (He et al., 2025), a P compression method serving as a representative baseline for communication quantization; SDP4bit (Jia et al., 2024), a 4-bit quantization method optimized for DP gradient compression. Models and Datasets. In Section 5.2, we evaluate compression algorithms on GPT-350M trained on the Pile dataset (Gao et al., 2020), using a learning rate schedule of (3×10−4→3×10−5)(3× 10^-4→ 3× 10^-5), with a global batch size of 256 and 10,000 training iterations. We further extend the evaluation to a larger model, GPT-6.7B, trained on the same Pile dataset. In addition, Qwen2.5-7B (Team, 2024) is trained on the Open-Web-Math dataset (Paster et al., 2023), following a learning rate schedule of (3×10−4→3×10−5)(3× 10^-4→ 3× 10^-5), with a global batch size of 64 for 10,000 iterations, to evaluate the generalization ability of the proposed method across different model families and data distributions. In Section 5.5, we conduct large-scale evaluations under a 3D parallel training configuration, where GPT-6.7B is trained from scratch on the Pile dataset using PyTorch (Paszke et al., 2019) v2.5.1 and Megatron-LM (Shoeybi et al., 2020), with parallelism configured as (TP = 4, P = 2, DP = 2). Metrics. We report End-to-end throughput (TFLOPS) to measure efficiency, and Model quality (Validation/Test Loss) to evaluate convergence. Degradation (Deg.) is reported as the relative percentage increase in loss relative to the BF16 baseline. Table 1. Accuracy comparison under TP=8 after 10,000 training iterations for the BF16 baseline, TahQuant, and TACO; degradation (Deg.) indicates the relative loss increase over the baseline. Method Val Loss ↓ Test Loss ↓ Val Deg. ↓ Test Deg. ↓ Baseline 2.389899 2.344701 – – TahQuant 2.458742 2.413642 +2.88 % +2.94 % TACO 2.395784 2.351210 +0.25 % +0.28 % 5.2. Evaluation of Accuracy with TP In this section, we systematically evaluate the impact of communication compression on model convergence under TP settings, where intermediate tensors are exchanged across GPUs during both forward and backward passes. First, we benchmark the overall End-to-End Performance in Section 5.2.1 under high-parallelism configurations (up to TP8) to comprehensively demonstrate robustness and convergence stability at scale. Next, we perform a Component-wise Analysis in Section 5.2.2 of TACO’s internal mechanisms—ASH and DS—to assess their contributions to numerical stability and precision recovery. We further extend this evaluation in Section 5.2.3 to large-scale models (GPT-6.7B and Qwen-2.5-7B), rigorously verifying that TACO maintains stable optimization dynamics and exhibits negligible accuracy degradation under aggressive TP compression. 5.2.1. End-to-End Convergence Comparison To evaluate the effectiveness of the proposed TACO framework for compressing TP intermediate tensors, we benchmark its end-to-end training accuracy against the state-of-the-art communication compression method TahQuant under a TP8 configuration. Table 1 reports the final validation and test losses for the uncompressed BF16 baseline, TahQuant, and TACO. After 10,000 training iterations, the BF16 baseline achieves a validation loss of 2.389899 and a test loss of 2.344701, serving as the reference for assessing compression-induced degradation. While TahQuant reduces communication volume, it incurs a substantial accuracy penalty, with validation and test losses increasing by +2.88%+2.88\% and +2.94%+2.94\%, respectively. This degradation suggests that quantization errors introduced into TP intermediate tensors accumulate across layers and iterations, ultimately impeding convergence toward the full-precision optimum. In contrast, TACO exhibits consistently near-lossless accuracy. As shown in Table 1, TACO attains a validation loss of 2.395784 and a test loss of 2.351210, corresponding to only +0.25%+0.25\% and +0.28%+0.28\% degradation relative to BF16. Compared to TahQuant, TACO improves fidelity by more than an order of magnitude. This robustness is particularly significant in TP settings, where intermediate tensors are exchanged at every layer and iteration, and even small numerical perturbations can rapidly propagate and destabilize training. Figure 11. Ablation of TACO components on convergence stability. The plot demonstrates the progressive stability gained by combining ASH and DS, with the final framework (purple) nearly overlapping the uncompressed baseline (blue). Ablation results on convergence stability. The study evaluates the effect of ASH and DS components individually and in combination. As components are added, training stability improves progressively, and the full TACO system closely matches the uncompressed baseline, demonstrating negligible impact on convergence. 5.2.2. Component-wise Analysis To elucidate how TACO preserves training stability, we systematically analyze its core architectural components: ASH and DS. Figure 11 presents the convergence of these configurations under TP4. The empirical evidence reveals that naively applying standard NVFP8 compression to TP intermediate tensors leads to immediate and catastrophic divergence, with the validation loss spiking to 5.605692. As shown by the NVFP8 curve in Figure 11, the loss quickly plateaus at an extremely high value (near 6.0) and completely fails to track the downward trend of the baseline. This indicates that standard 8-bit quantization without conditioning is clearly insufficient for the precision requirements of TP. This failure is rooted in the highly non-uniform nature of the numerical distributions within these tensors. TP intermediate tensors typically exhibit a dense clustering of values near zero; such a distribution fails to utilize the discrete representable points of the FP8 format effectively, resulting in significant quantization noise that destabilizes the gradient flow across parallel partitions. Integrating Dual-Scale (DS) quantization in isolation provides only partial stabilization of the training objective, reducing the validation loss to 3.300491. By partitioning TP intermediate tensors into finer sub-blocks and assigning independent scaling factors, DS mitigates precision loss caused by local dynamic range mismatches. However, the training trajectory (red diamonds in Figure 11) remains substantially above the uncompressed baseline, indicating that DS alone cannot fully recover full-precision performance. Without prior reshaping of the numerical distribution, the FP8 mantissa bits remain underutilized for the majority of densely clustered values, resulting in persistent information loss that limits convergence. Crucially, the full TACO configuration (NVFP8 + ASH + DS) achieves near-baseline performance, with a validation loss of 2.667557. As shown by the purple trajectory in Figure 11, this configuration closely tracks the baseline curve, particularly during the later stages of training, as highlighted in the magnified inset, and consistently outperforms TahQuant under the same setting. In this synergy, ASH acts as a preconditioner that disperses dense clusters and flattens the distribution, while DS explicitly aligns the transformed blocks to the FP8 representable range. Notably, ASH alone yields limited improvement, since transformed values can still exceed the FP8 range without DS’s adaptive scaling. These results demonstrate that the combination of ASH and DS is essential for enabling high-fidelity, low-bit communication in TP training. Figure 12. Validation loss comparison between the baseline (no compression) and TACO on GPT 6.7B and Qwen2.5-7B. Validation loss comparison between the BF16 baseline (no compression) and TACO on GPT 6.7B and Qwen2.5 7B. 5.2.3. Large-Scale Model Verification To evaluate the scalability and numerical robustness of the proposed framework, we further conduct experiments on two representative large-scale language models, GPT 6.7B and Qwen-2.5 7B. As shown in Figure 12, we compare the training loss trajectories between baseline (no compression) and TACO. For GPT 6.7B, TACO achieves a final validation loss of 2.570587, compared to 2.552718 for baseline, corresponding to a marginal degradation of +0.70%. Similarly, on Qwen-2.5 7B, TACO reaches a loss of 2.257332 versus 2.256751 for the baseline, resulting in an extremely small degradation of +0.03%. Across both model families, the loss curves remain stable throughout training, and the performance gap between TACO and full-precision training is negligible. These results demonstrate that TACO consistently preserves optimization stability under aggressive communication compression, and generalizes well across different architectures and scales without requiring additional tuning. 5.3. Ablation Study We systematically investigate the impact of Hadamard-based transform and the selection of low-bit formats on training convergence under TP4. In Section 5.3.1, we analyze the effectiveness of ASH by comparing it against standard Hadamard transform. In Section 5.3.2, we evaluate the sensitivity to different quantization formats. Figure 13. Validation and test loss for baseline, standard Hadamard, and ASH under TP4. Ablation study under TP4 comparing baseline, standard Hadamard, and ASH. Standard Hadamard degrades both validation and test performance, while ASH mitigates this issue and restores accuracy close to the uncompressed baseline, demonstrating its effectiveness in preserving training quality. 5.3.1. Component Ablation. Our analysis compares the uncompressed baseline against the standard Hadamard transform and the proposed ASH mechanism. Applying a standard Hadamard transform increases the validation loss from 2.663061 to 2.757688, corresponding to a +3.55%+3.55\% degradation. This accuracy drop occurs because the standard rotation does not sufficiently recondition the dense clusters of near-zero values inherent in TP intermediate tensors. As a result, a large fraction of values remains concentrated in a narrow range, severely underutilizing the FP8 mantissa. In contrast, ASH restores performance to near-baseline levels, achieving a validation loss of 2.667557 (a marginal +0.17%+0.17\% degradation) by spreading values more uniformly across the representable range, thereby minimizing quantization error for small-magnitude elements. As illustrated in Figure 13, the standard Hadamard transform exhibits a visible accuracy gap, whereas ASH effectively recovers the lost fidelity. These results confirm that adaptive distribution reshaping is a prerequisite for high-precision, low-bit communication. 5.3.2. Format Ablation. We further evaluate the effectiveness of different low-bit numerical formats when combined with ASH, with results shown in Figure 14. The choice of numerical format proves critical to training stability, as reflected by the markedly different loss trajectories observed in our ablation study. ASH+INT8 denotes INT8 quantization applied after ASH transform. As illustrated in Figure 14, this configuration leads to catastrophic divergence. Although the loss remains superficially stable during the initial ∼ 1,500 iterations, it subsequently exhibits a sharp exponential increase, ultimately reaching a validation loss of 68.10. This failure arises because INT8’s limited range cannot accommodate the broadened tensor distribution produced by ASH, resulting in saturation of high-magnitude values and collapse of small-magnitude ones. The floating-point formats offer significantly better robustness. The FP8 (E5M2) configuration partially mitigates the divergence seen in INT8, maintaining a stable downward trend. However, as shown in the magnified inset of Figure 14, the E5M2 curve (orange triangles) remains consistently higher than the baseline, settling at a validation loss of 3.305199. This +24.1%+24.1\% degradation indicates that while the E5M2 format provides sufficient dynamic range (5 bits for exponent), its limited 2-bit mantissa lacks the necessary resolution to represent the critical details of the reshaped tensors. In contrast, the FP8 (E4M3) format achieves the optimal balance between range and precision. As illustrated in Figure 14, the E4M3 curve (green squares) tracks the uncompressed baseline (blue circles) with remarkable fidelity throughout the entire training process, achieving a near-lossless validation loss of 2.667557, only a 0.19% degradation. These results highlight that both distribution conditioning via ASH and careful selection of the FP8 format are essential for maintaining stability in TP communication compression. Figure 14. Validation loss curves for ASH combined with different low-bit formats. INT8 causes complete divergence Ablation study of ASH with different low-bit formats. The results show that ASH is stable under FP8-based settings, while INT8 leads to complete divergence during training, demonstrating that aggressive quantization without sufficient representational range is incompatible with the proposed framework. Table 2. ASH block size ablation on accuracy and throughput. Throughput is measured in TFLOPS, with speedup relative to the baseline. Bold denotes the best result. Block Size Val Loss ↓ Test Loss ↓ Throughput ↑ Speedup ↑ Baseline 2.663061 2.650814 27.1 1.00× ASH (32) 2.670513 2.658282 29.5 1.09× ASH (64) 2.670663 2.658583 30.2 1.11× ASH (128) 2.668261 2.656122 37.9 1.40× ASH (256) 2.667557 2.655678 41.2 1.52× ASH (512) 2.670552 2.658692 38.1 1.41× 5.3.3. ASH Block Size Sensitivity The block size (B) of ASH is a critical hyperparameter that governs the trade-off between numerical stability and computational efficiency. Table 2 reports the validation loss, test loss, and training performance in a range of block sizes. These results provide several key insights into how the granularity of the transformation affects the statistical conditioning and performance of TP intermediate tensors. At small block sizes (B∈32,64B∈\32,64\), overly fine-grained partitioning introduces non-negligible overhead due to increased kernel invocations. Limited per-block computation leads to poor GPU utilization and suboptimal memory bandwidth efficiency, resulting in only modest throughput gains of 1.091.09–1.11×1.11× over the baseline. Moreover, the limited receptive field of the Hadamard transform at this granularity is insufficient to effectively reshape TP intermediate tensor distributions, leading to a slight convergence degradation. In contrast, moderate block sizes (B∈128,256B∈128,256) significantly improve arithmetic intensity and enable more efficient memory coalescing. As reported in Table 2, B=256B=256 achieves the highest speedup of 1.52×1.52× while preserving sufficient spatial scope to effectively align the reshaped TP intermediate tensor distribution with the FP8 dynamic range. This configuration strikes an optimal balance between computational efficiency and numerical fidelity, while maintaining near-baseline convergence behavior in practice. However, excessively large blocks (B=512B=512) degrade both throughput and numerical accuracy. Relative to the optimal block size of B=256B=256, the reduced speedup of 1.41×1.41× stems from thread-level workload imbalance and increased shared-memory bank conflicts, which limit effective parallelism and overall execution efficiency. Moreover, the expanded spatial scope of the transform diminishes the effectiveness of adaptive scaling by aggregating TP intermediate tensors with heterogeneous magnitudes into a single scaling region, resulting in a modest but systematic loss of numerical fidelity and slightly impaired convergence. In summary, a block size of B=256B=256 provides the best trade-off between computational efficiency and numerical robustness. Maximize GPU parallelism while ensuring that the reshaped TP intermediate tensor distribution remains well conditioned for FP8 quantization. This balance is crucial for TACO to provide high performance and almost lossless low-bit TP communication. Figure 15. End-to-end training throughput (TFLOPS) under different TP degrees on GPT-2.7B and GPT-6.7B. We compare standard Ring and Tree-based collectives with TahQuant and TACO. End-to-end training throughput (TFLOPS) under varying tensor parallel degrees for GPT-2.7B and GPT-6.7B. The study compares Ring and Tree-based collective communication with TahQuant and TACO, demonstrating improved scalability and higher throughput of TACO across different model sizes and parallel configurations. 5.4. Evaluation of Performance with TP We next analyze system-level performance under TP degrees of 2, 4, and 8 on GPT-2.7B and GPT-6.7B. Section 5.4.1 to quantify end-to-end throughput and communication scalability under different collective strategies, and Section 5.4.2 to uncover the underlying sources of TACO’s performance gains through a detailed decomposition of computation, communication, and compression overhead. 5.4.1. End-to-End Throughput Comparison We systematically evaluate end-to-end training throughput for GPT-2.7B and GPT-6.7B under TP degrees of 2, 4, and 8, comparing Ring AllReduce, tree-based collective communication, TahQuant, and TACO (see Figure 15. Throughput is measured in TFLOPS, and relative improvements are reported primarily as speedup over the Ring baseline. Ring and tree-based collectives are two widely used communication algorithms in distributed training: Ring AllReduce overlaps communication with computation but incurs latency that scales linearly with the TP degree, whereas tree-based collectives reduce the number of communication steps but often suffer from limited bandwidth utilization and increased synchronization overhead at scale. Across most configurations, TACO achieves the highest throughput and largest speedups over Ring AllReduce. On GPT-2.7B with TP=2, TACO improves throughput by approximately 1.23×1.23× over Ring, slightly surpassing TahQuant (1.21×1.21×). As the TP degree increases to 4, communication overhead becomes more pronounced: Ring throughput degrades significantly, whereas TACO maintains a 1.63×1.63× speedup, outperforming TahQuant’s 1.52×1.52×. At TP=8, TahQuant slightly outperforms TACO, achieving a 1.43×1.43× speedup versus 1.90×1.90× for TACO, reflecting reduced amortization efficiency of TACO’s aggressive kernel optimizations under extreme TP. In contrast, TACO consistently outperforms all baselines on GPT-6.7B across all TP degrees. At TP=2, it achieves a 1.29×1.29× speedup over Ring, surpassing TahQuant’s 1.25×1.25×. The advantage grows with higher TP degrees: at TP=4, TACO reaches a 1.70×1.70× speedup versus 1.54×1.54× for TahQuant, and at TP=8, it maintains a 1.87×1.87× improvement, compared to 1.40×1.40× for TahQuant. These results clearly highlight TACO’s increasing strong efficiency relative to alternatives as TP scales, thanks to its ability to compress TP intermediate tensors and effectively overlap communication with computation. Overall, throughput decreases as the TP degree increases due to the rapidly growing volume and frequency of TP intermediate tensor communication. TACO consistently mitigates this degradation by compressing TP intermediate tensors into FP8 and fusing compression, decompression, and communication into optimized kernels. By overlapping these operations asynchronously, TACO reduces the effective communication cost along the critical path, improving hardware utilization and delivering substantially better scalability across model sizes and TP configurations. Figure 16. Throughput comparison across different TP settings. The numbers above the bars indicate the speedup ratio relative to the previous baseline. Throughput comparison across different tensor parallel configurations. The results report both absolute throughput and relative speedup over the baseline, demonstrating consistent performance improvements of the proposed method under various parallel settings. 5.4.2. Performance Breakdown We evaluate TACO’s performance on GPT-6.7B under varying TP degrees, reporting end-to-end throughput in TFLOPS (Figure 16). The uncompressed baseline achieves 134.9, 63.7, and 31.3 TFLOPS for TP2, TP4, and TP8, respectively. Applying TACO without kernel fusion slightly reduces throughput due to the local computation overhead of compression, yielding 98.2, 54.8, and 22.0 TFLOPS—corresponding to relative speedups of 0.73×, 0.86×, and 0.70× compared to the baseline. Applying kernel fusion dramatically improves performance: TACO with fusion achieves 174.1, 108.3, and 58.5 TFLOPS for TP2, TP4, and TP8, corresponding to speedups of 1.77×, 1.98×, and 2.66× over TACO without fusion. Further co-optimization on top of kernel fusion provides additional speedups of 1.05×, 1.18×, and 1.33× for TP2, TP4, and TP8, respectively. These results demonstrate that TACO’s optimizations effectively exploit computation–communication co-design: while FP8 compression introduces minor local computation, kernel fusion and co-optimization maximize throughput, especially as TP communication dominates. Overall, Figure 16 highlights that TACO consistently delivers substantial performance gains over the baseline, with the benefits of combined optimizations growing as TP increases. 5.5. 3D Parallel Training Evaluation 5.5.1. Training Accuracy We evaluate TACO’s accuracy under full 3D parallelism on GPT-6.7B. To isolate the impact of TP compression, we consider three settings applied consistently to both models: (1) a baseline without compression, (2) 2D parallelism, where DP and P communications are quantized using SDP4bit and TahQuant respectively while TP remains uncompressed, and (3) 3D parallelism, where TACO additionally use in TP alongside SDP4bit and TahQuant. Figure 17 presents the validation loss curves of GPT-6.7B under these settings. The baseline model achieves a final loss of 2.663061, while the 2D-parallel configuration slightly degrades to 2.676003 due to quantization noise in DP and P communications. In contrast, TACO under full 3D parallelism closely tracks the baseline throughout training, reaching a final loss of 2.679906, corresponding to only a 0.14% loss increase over the 2D setting, demonstrating that adding TP compression does not introduce additional optimization instability. Across the entire training trajectory, TACO exhibits stable convergence behavior comparable to full-precision training, indicating that the proposed method effectively preserves gradient fidelity even under aggressive communication compression across all parallel dimensions. These results highlight that TACO enables end-to-end compression of DP, P, and TP communications in large-scale training while maintaining near-lossless optimization quality, demonstrating strong robustness and scalability in full 3D parallel training settings, and consistently delivering reliable performance across different model scales and training configurations. Figure 17. Validation loss of GPT-6.7B under full 3D parallelism. Validation loss curves of GPT-6.7B under full 3D parallel training. The results demonstrate stable convergence of the proposed approach, indicating strong numerical robustness and training consistency at large scale. 5.5.2. End-to-End Throughput Finally, we present an end-to-end training throughput evaluation of TACO on GPT models. As shown in Table 3, compared with the uncompressed baseline, TACO achieves consistent and significant performance improvements across different model sizes. For GPT models, the speedup reaches up to 1.53×. Notably, the improvements from 2D to full 3D parallelism are substantially amplified once TP communication compression is enabled, indicating that tensor parallel communication is a major bottleneck in large-scale distributed training. These results demonstrate that, under tightly synchronized 3D parallel training, reducing TP communication volume is critical for achieving high system efficiency. TACO’s optimized compression operators ensure that the additional quantization overhead is fully amortized by the reduction in communication cost, resulting in a net throughput gain. Moreover, the consistent speedups across GPT model scales suggest that TACO scales robustly with increasing model size and communication intensity. Training remains numerically stable throughout optimization. As shown in Figure 17, TACO preserves a validation loss trajectory that closely matches the full-precision baseline, confirming that the system-level gains are achieved without sacrificing convergence or accuracy. 6. Discussion Design objective. TACO is not designed to maximize raw communication speed, but to enable near-lossless compression of TP intermediate tensors while preserving convergence in large-scale training. This is motivated by the observation that existing methods often introduce optimization instability under high-frequency synchronization or require delicate tuning of error compensation and scaling strategies, limiting their robustness. To achieve this, TACO adopts FP8 as a practical precision point, striking a balance between compression efficiency and numerical fidelity, and enabling stable training under aggressive communication reduction. Generality across hardware. Although TACO is implemented with FP8 as the primary precision format, its design is not dependent on FP8-specific hardware support. Instead, FP8 serves as a target precision level that defines the compression semantics of TP communication. On platforms without native FP8 support, TACO degrades gracefully to an INT8-based implementation, where quantization is performed using our ASH together with DS. In this configuration, intermediate tensors are still compressed to low-bit integer representations during communication, while scaling and reconstruction are handled in a software-assisted manner. This design preserves the same communication semantics as FP8-based execution, while trading off additional lightweight arithmetic overhead for broader hardware compatibility. As a result, TACO maintains its communication reduction benefits across heterogeneous accelerators without requiring specialized FP8 support. Table 3. End-to-end throughput (TFLOPS) under 3D parallelism on GPT models (TP=4, P=2, DP=2). Models Size Baseline 2D (w/o TACO) 3D (w/ TACO) GPT 2.7B 39.9 40.3 (1.01×) 59.7 (1.50×) 6.7B 61.1 62.3 (1.02×) 93.3 (1.53×) 13B 73.8 75.2 (1.02×) 111.7 (1.51×) 7. Conclusion and Future Work In this paper, we presented TACO, an efficient framework for accelerating distributed training by optimizing the communication of TP intermediate tensors. By leveraging FP8-based compression and system-level kernel fusion, TACO effectively alleviates the communication bottlenecks inherent in high-degree TP. Our systematic evaluations across multiple model scales, including GPT and Qwen, demonstrate that TACO consistently achieves superior throughput, with speedups of up to 1.87×. Moreover, experiments under full 3D parallelism confirm that TACO delivers robust, architecture-agnostic performance gains while maintaining convergence nearly identical to uncompressed baselines. Future work will focus on extending TACO’s compression strategies to encompass gradient and optimizer state communication within 3D-parallel training stacks. Additionally, we aim to investigate adaptive quantization schemes that dynamically adjust the precision of TP intermediate tensors according to layer-wise sensitivity, further enhancing hardware utilization and efficiency in exascale distributed training systems. Acknowledgements.This work was supported by the National Key Research and Development Program of China (Grant No. 2025YFB3003702), the Innovation Funding of ICT, CAS (Grant No. E461050), and the National Natural Science Foundation of China (Grant Nos. 62032023 and T2125013). The AI-driven experiments, simulations, and model training were conducted on the robotic AI-Scientist platform at the Chinese Academy of Sciences. References D. Alistarh, D. Grubic, J. Z. Li, R. Tomioka, and M. Vojnovic (2017) QSGD: communication-efficient sgd via gradient quantization and encoding. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Red Hook, NY, USA, p. 1707–1718. External Links: ISBN 9781510860964 Cited by: §1, §2.2. Q. Anthony, B. Michalowicz, J. Hatef, L. Xu, M. Abduljabbar, A. Shafi, H. Subramoni, and D. Panda (2024) Demystifying the communication characteristics for distributed transformer models. External Links: 2408.10197, Link Cited by: §2.1. S. Ashkboos, I. Markov, E. Frantar, T. Zhong, X. Wang, J. Ren, T. Hoefler, and D. Alistarh (2024a) QUIK: towards end-to-end 4-bit inference on generative large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 3355–3371. External Links: Link, Document Cited by: §1. S. Ashkboos, A. Mohtashami, M. L. Croci, B. Li, P. Cameron, M. Jaggi, D. Alistarh, T. Hoefler, and J. Hensman (2024b) Quarot: outlier-free 4-bit inference in rotated llms. Advances in Neural Information Processing Systems 37, p. 100213–100240. Cited by: §1, §2.3. T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33, p. 1877–1901. Cited by: §1, §2.1. J. Chee, Y. Cai, V. Kuleshov, and C. M. De Sa (2023) Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems 36, p. 4396–4429. Cited by: §1. C. Chen, J. Ni, S. Lu, X. Cui, P. Chen, X. Sun, N. Wang, S. Venkataramani, V. V. Srinivasan, W. Zhang, et al. (2020) Scalecom: scalable sparsified gradient compression for communication-efficient distributed training. Advances in Neural Information Processing Systems 33, p. 13551–13563. Cited by: §1. J. Chen, L. Zheng, Z. Yao, D. Wang, I. Stoica, M. W. Mahoney, and J. E. Gonzalez (2021) ActNN: reducing training memory footprint via 2-bit activation compressed training. External Links: 2104.14129, Link Cited by: §2.2. S. Chen, D. Zheng, C. Ding, C. Huan, Y. Ji, and H. Liu (2023) TANGO: re-thinking quantization for graph neural network training on gpus. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’23, New York, NY, USA. External Links: ISBN 9798400701092, Link, Document Cited by: §2.2. Y. Chen, X. Pan, Y. Li, B. Ding, and J. Zhou (2024) E-llm: large-scale training and inference of early-exit large language models with 3d parallelism. In Proceedings of the 41st International Conference on Machine Learning, ICML’24, Vienna, Austria. Cited by: §1. A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. (2023) Palm: scaling language modeling with pathways. Journal of Machine Learning Research 24 (240), p. 1–113. Cited by: §1. T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer (2023) Qlora: efficient finetuning of quantized llms. Advances in neural information processing systems 36, p. 10088–10115. Cited by: §1, §2.2. H. Dong, T. Johnson, M. Cho, and E. Soroush (2024) Towards low-bit communication for tensor parallel llm inference. External Links: 2411.07942, Link Cited by: §2.2. V. Egiazarian, R. L. Castro, D. Kuznedelev, A. Panferov, E. Kurtic, S. Pandit, A. Marques, M. Kurtz, S. Ashkboos, T. Hoefler, and D. Alistarh (2026) Bridging the gap between promise and performance for microscaling fp4 quantization. External Links: 2509.23202, Link Cited by: §1. E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh (2023) GPTQ: accurate post-training quantization for generative pre-trained transformers. External Links: 2210.17323, Link Cited by: §1. T. Gale, D. Narayanan, C. Young, and M. Zaharia (2022) MegaBlocks: efficient sparse training with mixture-of-experts. External Links: 2211.15841, Link Cited by: §2.1. L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy (2020) The pile: an 800gb dataset of diverse text for language modeling. External Links: 2101.00027, Link Cited by: §5.1. A. Grattafiori et al. (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §1. G. He, Y. Cao, Y. He, T. Bai, K. Yuan, and B. Yuan (2025) TAH-quant: effective activation quantization in pipeline parallelism over slow network. External Links: 2506.01352, Link Cited by: §1, §2.2, §5.1. Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen (2019) GPipe: efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA. Cited by: §1. J. Jia, C. Xie, H. Lu, D. Wang, H. Feng, C. Zhang, B. Sun, H. Lin, Z. Zhang, X. Liu, et al. (2024) Sdp4bit: toward 4-bit communication quantization in sharded data parallelism for llm training. Advances in Neural Information Processing Systems 37, p. 8734–8759. Cited by: §1, §2.2, §5.1. Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong, Y. Jia, S. He, H. Chen, Z. Bai, Q. Hou, S. Yan, D. Zhou, Y. Sheng, Z. Jiang, H. Xu, H. Wei, Z. Zhang, P. Nie, L. Zou, S. Zhao, L. Xiang, Z. Liu, Z. Li, X. Jia, J. Ye, X. Jin, and X. Liu (2024) MegaScale: scaling large language model training to more than 10,000 gpus. External Links: 2402.15627, Link Cited by: §1. A. Kuzmin, M. Van Baalen, Y. Ren, M. Nagel, J. Peters, and T. Blankevoort (2022) Fp8 quantization: the power of the exponent. Advances in Neural Information Processing Systems 35, p. 14651–14662. Cited by: §2.3. I. Lamprecht, A. Karnieli, Y. Hanani, N. Giladi, and D. Soudry (2025) Tensor-parallelism with partially synchronized activations. Note: NeurIPS 2025 PosterAccepted as NeurIPS 2025 Poster External Links: Link Cited by: §1. Q. Li, B. Zhang, L. Ye, Y. Zhang, W. Wu, Y. Sun, L. Ma, and Y. Xie (2024) Flash communication: reducing tensor parallelization bottleneck for fast large language model inference. External Links: 2412.04964, Link Cited by: §2.2. S. Li, F. Xue, C. Baranwal, Y. Li, and Y. You (2023) Sequence parallelism: long sequence training from system perspective. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, Canada, p. 2391–2404. Cited by: §2.1. S. Li and T. Hoefler (2022) Near-optimal sparse allreduce for distributed deep learning. In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, PPoPP ’22, New York, NY, USA, p. 135–149. External Links: ISBN 9781450392044, Link, Document Cited by: §1. Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally (2020) Deep gradient compression: reducing the communication bandwidth for distributed training. External Links: 1712.01887, Link Cited by: §2.2. X. Liu, H. Kong, H. Zhao, S. Lyu, Z. Wei, M. Liu, X. Tian, L. Zhao, Z. Chen, F. Wang, Z. Chen, Z. Wang, G. Tan, and D. Tao (2026) COCCL: a collective communication library supporting easy integration and configuration of customized compression for scalable llm training. In Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP ’26, New York, NY, USA, p. 384–397. External Links: ISBN 9798400723100, Link, Document Cited by: §4.4.2. Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y. Mehdad, Y. Shi, R. Krishnamoorthi, and V. Chandra (2023) LLM-qat: data-free quantization aware training for large language models. External Links: 2305.17888, Link Cited by: §2.2. I. Markov, A. Vladu, Q. Guo, and D. Alistarh (2023) Quantized distributed training of large models with convergence guarantees. In Proceedings of the 40th International Conference on Machine Learning, ICML’23, Honolulu, Hawaii, USA. Cited by: §1, §2.2. P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu (2018) Mixed precision training. External Links: 1710.03740, Link Cited by: §2.3. P. Micikevicius, D. Stosic, N. Burgess, M. Cornea, P. Dubey, R. Grisenthwaite, S. Ha, A. Heinecke, P. Judd, J. Kamalu, N. Mellempudi, S. Oberman, M. Shoeybi, M. Siu, and H. Wu (2022) FP8 formats for deep learning. External Links: 2209.05433, Link Cited by: §2.3. D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia (2019) PipeDream: generalized pipeline parallelism for dnn training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles, SOSP ’19, New York, NY, USA, p. 1–15. External Links: ISBN 9781450368735, Link, Document Cited by: §1. D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia (2021) Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’21, New York, NY, USA. External Links: ISBN 9781450384421, Link, Document Cited by: §1, §1, §2.1. K. Paster, M. D. Santos, Z. Azerbayev, and J. Ba (2023) OpenWebMath: an open dataset of high-quality mathematical web text. External Links: 2310.06786, Link Cited by: §5.1. A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala (2019) PyTorch: an imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA. Cited by: §5.1. I. Polyakov, A. Dukhanov, and E. Spirin (2025) TAGC: optimizing gradient communication in distributed transformer training. In Proceedings of the 5th Workshop on Machine Learning and Systems, EuroMLSys ’25, New York, NY, USA, p. 254–260. External Links: ISBN 9798400715389, Link, Document Cited by: §2.2. S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He (2020) ZeRO: memory optimizations toward training trillion parameter models. External Links: 1910.02054, Link Cited by: §1. M. I. Rudakov, A. N. Beznosikov, Ya. A. Kholodov, and A. V. Gasnikov (2023) Activations and gradients compression for model-parallel training. Doklady Mathematics 108 (S2), p. S272–S281. External Links: ISSN 1531-8362, Link, Document Cited by: §2.2. S. Savkin (2025) Quantization methods for matrix multiplication and efficient transformers. Ph.D. Thesis, MASSACHUSETTS INSTITUTE OF TECHNOLOGY. Cited by: §1. H. Shen, N. Mellempudi, X. He, Q. Gao, C. Wang, and M. Wang (2024) Efficient post-training quantization with fp8 formats. Proceedings of Machine Learning and Systems 6, p. 483–498. Cited by: §2.3. M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro (2020) Megatron-lm: training multi-billion parameter language models using model parallelism. External Links: 1909.08053, Link Cited by: §1, §2.1, §2.1, §5.1. S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, E. Zhang, R. Child, R. Y. Aminabadi, J. Bernauer, X. Song, M. Shoeybi, Y. He, M. Houston, S. Tiwary, and B. Catanzaro (2022) Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. External Links: 2201.11990, Link Cited by: §1. Y. Sun, R. Liu, H. Bai, H. Bao, K. Zhao, Y. Li, J. Hu, X. Yu, L. Hou, C. Yuan, X. Jiang, W. Liu, and J. Yao (2025) FlatQuant: flatness matters for llm quantization. External Links: 2410.09426, Link Cited by: §1, §2.3. Q. Team (2024) Qwen2.5: a party of foundation models. External Links: Link Cited by: §5.1. H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample (2023) LLaMA: open and efficient foundation language models. External Links: 2302.13971, Link Cited by: §1. A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. De Sa (2024) QuIP#: even better llm quantization with hadamard incoherence and lattice codebooks. In Proceedings of the 41st International Conference on Machine Learning, ICML’24, Vienna, Austria. Cited by: §1. T. Vogels, S. P. Karimireddy, and M. Jaggi (2019) PowerSGD: practical low-rank gradient compression for distributed optimization. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA. Cited by: §1. H. Wang, C. Ruan, J. He, J. Ruan, C. Tang, X. Ma, and C. Li (2024) Hiding communication cost in distributed llm training via micro-batch co-execution. External Links: 2411.15871, Link Cited by: §1, §1. J. Wang, B. Yuan, L. Rimanic, Y. He, T. Dao, B. Chen, C. Ré, and C. Zhang (2022) Fine-tuning language models over slow networks using activation quantization with guarantees. Advances in Neural Information Processing Systems 35, p. 19215–19230. Cited by: §2.2. B. Workshop, T. L. Scao, A. Fan, C. Akiki, et al. (2023) BLOOM: a 176b-parameter open-access multilingual language model. External Links: 2211.05100, Link Cited by: §1. G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han (2024) SmoothQuant: accurate and efficient post-training quantization for large language models. External Links: 2211.10438, Link Cited by: §1. L. Xu, Q. Anthony, Q. Zhou, N. Alnaasan, R. Gulhane, A. Shafi, H. Subramoni, and D. K. D. Panda (2024) Accelerating large language model training with hybrid gpu-based compression. In 2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid), Philadelphia, PA, USA, p. 196–205. Cited by: §1, §2.2. Q. Yi, J. Duan, H. Hu, Q. Hua, H. Zhao, S. Qian, D. Yang, J. Cao, J. Tang, Y. Yu, C. Liao, K. Wang, and L. Zhang (2025) EDGC: entropy-driven dynamic gradient compression for efficient llm training. External Links: 2511.10333, Link Cited by: §2.2. L. Zhang, L. Zhang, S. Shi, X. Chu, and B. Li (2023) Evaluation and optimization of gradient compression for distributed deep learning. In 2023 IEEE 43rd International Conference on Distributed Computing Systems (ICDCS), Hong Kong, China, p. 361–371. Cited by: §1, §2.2. H. Zheng, P. Liang, Y. Tang, Y. Shi, L. Qiao, and D. Li (2024) 3D parallelism for transformers via integer programming. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, p. 6440–6444. Cited by: §1.