Paper deep dive
Neural Network Learning of One-Bit Protocols for Qubit Measurement Simulation
Josep Escrig, Mani Zartab, Giulio Gasbarri, Estel Ferrer, Ramon Muñoz-Tapia, Gael Sentís
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/1/2026, 10:36:17 AM
Summary
This paper investigates the classical communication complexity required to simulate qubit prepare-and-measure (PM) quantum statistics. While two bits are necessary and sufficient for arbitrary qubit measurements, the authors use neural networks to demonstrate that a single bit can achieve high accuracy for specific, symmetric families of measurements. They identify that POVMs with uniformly weighted elements, particularly those forming regular polyhedra on the Bloch sphere, are amenable to 1-bit simulation. By analyzing the neural network's learned patterns, they derive an analytical protocol that is extremely accurate for finite informationally complete symmetric configurations and becomes exact in the limit of continuous isotropic measurements.
Entities (6)
Relation Signals (5)
Neural Network → usedtodiscover → One-Bit Protocol
confidence 95% · We use a neural network procedure to demonstrate that a single bit can achieve high average accuracy for specific measurement families.
One-Bit Protocol → approximates → Quantum Statistics
confidence 93% · a single bit can achieve high average accuracy for specific measurement families.
Regular Polyhedra → exhibits → High Simulation Accuracy
confidence 92% · symmetric measurements with uniformly weighted elements, such as those forming regular polyhedra, are particularly amenable to this restricted communication.
Toner-Bacon Protocol → uses → Single Bit Communication
confidence 90% · Toner and Bacon [23] subsequently showed that shared classical randomness together with a single bit of communication suffices to reproduce the correlations of local projective measurements on a maximally entangled pair.
Renner et al. → proved → Two-Bit Sufficiency
confidence 88% · Renner et al. [18] extended the simulation of qubit statistics from projective to generalized measurements, proving that any qubit PM scenario can be simulated exactly with a worst-case communication cost of two bits.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Communication complexity provides a natural framework for quantifying the classical resources required to reproduce quantum statistics. In the qubit prepare-and-measure scenario, two classical bits have been shown to be necessary and sufficient to simulate arbitrary qubit states and arbi- trary quantum measurements exactly. However, this result does not exclude the possibility that restricted families of measurements may admit accurate 1-bit classical approximations. We use a neural network procedure to demonstrate that a single bit can achieve high average accuracy for specific measurement families. A performance analysis of our neural network reveals that symmet- ric measurements with uniformly weighted elements, such as those forming regular polyhedra, are particularly amenable to this restricted communication. By analyzing the patterns learned by the neural network, we derive an analytical protocol that is extremely accurate for finite information- ally complete symmetric configurations and becomes exact in the limit of a continuous isotropic measurement.
Tags
Links
- Source: https://arxiv.org/abs/2607.23645v1
- Canonical: https://arxiv.org/abs/2607.23645v1
Trouble viewing inline? Open PDF directly →
Full Text
73,270 characters extracted from source content.
Expand or collapse full text
Neural Network Learning of One-Bit Protocols for Qubit Measurement Simulation Josep Escrig Fundació i2CAT, internet i innovació digital a Catalunya, 08034 Barcelona, Spain Universitat Politècnica de Catalunya, 08005 Barcelona, Spain Mani Zartab Física Teòrica: Informació i Fenòmens Quàntics, Departament de Física, Universitat Autònoma de Barcelona, 08193 Bellaterra (Barcelona), Spain Giulio Gasbarri Naturwissenschaftlich-Technische Fakultät, Universität Siegen, Siegen 57068, Germany Estel Ferrer Fundació i2CAT, internet i innovació digital a Catalunya, 08034 Barcelona, Spain Universitat Politècnica de Catalunya, 08005 Barcelona, Spain Ramon Muñoz-Tapia Física Teòrica: Informació i Fenòmens Quàntics, Departament de Física, Universitat Autònoma de Barcelona, 08193 Bellaterra (Barcelona), Spain Gael Sentís Física Teòrica: Informació i Fenòmens Quàntics, Departament de Física, Universitat Autònoma de Barcelona, 08193 Bellaterra (Barcelona), Spain Abstract Communication complexity provides a natural framework for quantifying the classical resources required to reproduce quantum statistics. In the qubit prepare-and-measure scenario, two classical bits have been shown to be necessary and sufficient to simulate arbitrary qubit states and arbitrary quantum measurements exactly. However, this result does not exclude the possibility that restricted families of measurements may admit accurate 1-bit classical approximations. We use a neural network procedure to demonstrate that a single bit can achieve high average accuracy for specific measurement families. A performance analysis of our neural network reveals that symmetric measurements with uniformly weighted elements, such as those forming regular polyhedra, are particularly amenable to this restricted communication. By analyzing the patterns learned by the neural network, we derive an analytical protocol that is extremely accurate for finite informationally complete symmetric configurations and becomes exact in the limit of a continuous isotropic measurement. I Introduction Understanding how much classical communication is required to reproduce quantum statistics provides an operational way to compare quantum and classical information processing [2, 3]. This perspective is especially relevant because quantum states serve as information carriers in communication protocols, cryptography, and dimension-limited information-processing tasks. In this setting, simulation costs provide a concrete benchmark for the expressive power of quantum communication: they indicate when a low-dimensional quantum system can encode correlations that would require substantially larger classical messages. The prepare-and-measure(PM) framework has been used to propose dimension witnesses [11], the certification of quantum devices [8], the analysis of communication advantages [16], and the design and verification of quantum networks [1], among other applications. Several advances in this direction have shown that certain quantum statistics admit surprisingly efficient classical simulations. In Ref. [5], the authors introduced a coding scheme for the classical teleportation of a qubit, requiring on average only 2.19 bits of classical information to reproduce the statistics of projective measurements. Toner and Bacon [23] subsequently showed that shared classical randomness together with a single bit of communication suffices to reproduce the correlations of local projective measurements on a maximally entangled pair. In the PM scenario, they further showed that two classical bits are sufficient to simulate the statistics of arbitrary projective measurements on a qubit. Degorre et al. [9] later observed that the communication step in the above protocol can be replaced by a source of biased shared randomness depending on the qubit state, thereby identifying a key ingredient underlying classical simulations of quantum statistics. Building on these ideas, Renner et al. [18] extended the simulation of qubit statistics from projective to generalized measurements, proving that any qubit PM scenario can be simulated exactly with a worst-case communication cost of two bits. Extensions to higher-dimensional systems have also been investigated [14, 19, 24, 20], including approximate simulation protocols and analyses of average communication cost. Ultimately, the existence of an exact classical simulation protocol for PM scenarios in dimension d≥3d≥ 3 using a finite amount of classical communication remains open. The 2-bit worst-case simulation of arbitrary qubit PM statistics, however, does not preclude more economical descriptions in less demanding regimes. One natural relaxation is to move from worst-case to average communication cost. Indeed, recent work has shown that, by encoding the messages sent by Alice, qubit PM statistics can be simulated with an average communication cost of 1.89 bits [20]. This suggests that the apparent 2-bit barrier may be a consequence of demanding a uniform worst-case guarantee over all state preparations and measurements, and that more efficient simulations may be possible once the relevant operational figure of merit is adapted. A second route toward lower communication cost is to restrict the family of quantum states and/or measurements to be simulated. This is particularly natural in PM scenarios, where practical communication tasks often involve structured ensembles of preparations and measurements rather than the full set of qubit statistics. From this perspective, the central question becomes which physically or operationally meaningful subsets of qubit statistics already admit an even more efficient classical simulability. Identifying such subsets would help clarify which features of quantum PM experiments are responsible for their classical communication cost. Further motivation comes from recent numerical evidence in the Bell scenario. In Ref. [21], neural network methods were used to search for 1-bit classical protocols simulating the statistics of local projective measurements on bipartite entangled two-qubit states. Strikingly, although one bit of communication is known to be insufficient in full generality, the authors found that 1-bit protocols approximate the target quantum correlations extremely well across the tested instances, and no explicit counterexample was identified by their search. This suggests that one bit of communication can already capture a surprisingly large portion of quantum statistics, extending the intuition provided by the Toner–Bacon protocol beyond the maximally entangled singlet case [23]. These observations motivate the search for 1-bit simulation protocols in PM scenarios. In this work, we focus on restricted families of measurements and ask whether their qubit statistics can be reproduced, or accurately approximated, using only a single bit of classical communication. To this end, we adopt a neural network(N)- procedure, in the spirit of Ref. [21], using the network as a tool to discover candidate protocols rather than imposing a fixed analytical ansatz from the outset. When successful, such searches provide evidence for hidden structure in the corresponding family of measurements and point toward compact classical descriptions of their quantum statistics. Thus, our goal is to map regimes of 1-bit-simulability in qubit PM scenarios and to use protocols, discovered by N-procedure, as evidence for new analytically tractable classical simulations. Following this approach, we first train a neural network on randomly generated PM instances and use it to identify structured families of measurements that are strong candidates for 1-bit simulability. The families singled out by the optimization are highly-symmetric POVMs whose elements are distributed close to uniformly on the Bloch sphere. By inspecting the behavior learned by the network, we then extract an explicit analytical protocol, thereby moving from a black-box numerical search to a concrete classical simulation strategy. Our analysis shows that the extracted 1-bit protocol, while it is not exact for arbitrary finite POVMs, yields highly accurate approximations to the quantum probabilities for informationally-complete symmetric POVMs, and the approximation improves as the number of outcomes increases. In the continuous limit, corresponding to the covariant qubit POVM, the protocol becomes exact, placing this measurement within the class of exactly 1-bit-simulable PM resources. These results provide evidence that symmetry can substantially reduce the communication needed to simulate quantum statistics, and they illustrate how neural network searches can help uncover analytical structure in PM simulation problems. This paper is structured as follows. In Section I we introduce the theoretical background and definitions used throughout this paper. In Section I we describe the N-procedure used to learn quantum statistics under the restrictions of classical shared randomness and one bit of classical communication, and in Section IV we present the numerical results. In Section V we infer the classical 1-bit protocol and study its properties. We end with the conclusions and include the technical details in the Appendices. I Background A general m-outcome quantum measurement is described by a POVM, i.e., a set of positive semidefinite operators ℳ:=Mii=1mM:=\M_i\_i=1^m satisfying the completeness condition ∑i=1mMi= _i=1^mM_i= . The probability of obtaining outcome i when applying ℳM on a generic state ρ is given by the Born rule PQ(i|ρ,ℳ)=Tr(Miρ)P_Q(i|ρ,M)=Tr(M_iρ). For qubits, each POVM element can be written in the Bloch representation as Mi=pi(+→⋅σ→)M_i=p_i ( + y_i· σ ), where pi≥0p_i≥ 0, ∑ipi=1 _ip_i=1, and |y→i|≤1| y_i|≤ 1. Similarly, we can express a qubit state as ρ=(+→⋅σ→)/ρ=( + r· σ)/2. The probability PQ(i|ρ,ℳ)P_Q(i|ρ,M) then takes the form PQ(i|ρ,ℳ)=pi(1+y→i⋅r→). P_Q(i|ρ,M)=p_i (1+ y_i· r ). (1) Let us now introduce the quantum PM scenario. In this setup, Alice prepares a qubit state ρr→ _ r, represented by its Bloch vector r→ r, and sends it to Bob through a quantum communication channel. Bob then performs a POVM ℳYM_Y specified by Y:=(pi,y→i)i=1mY:=(p_i, y_i)_i=1^m, and reports the outcome i with probability given by (1). The scenario is illustrated in Fig. 1. Figure 1: Quantum PM scenario. Alice prepares the qubit state ρr→ _ r and sends it to Bob, who performs the POVM ℳYM_Y and reports the outcome i. A classical simulation replaces the transmitted qubit by a classical message, assisted by shared randomness. Alice receives the state description r→ r, Bob receives the measurement description Y, and they share a random variable independent of both. Alice sends a classical messace c to Bob. Bob, uses c, Y, and the share randomness to generate an outcome i. The protocol exactly simulates the quantum PM scehnario if, for all states and measurements, the resulting classical probability PC(i|r→,Y)P_C(i| r,Y) coincides with the quantum one i.e, PC(i|r→,Y)=PQ(i|ρr→,ℳY)∀i,r→,Y. P_C(i| r,Y)=P_Q(i| _ r,M_Y) ∀ i, r,Y. (2) If the equality holds only up to a chosen error metric, evaluated on a spceific class of states and measurements, we will refer the protocol as an approximate simulation. Renner et al. [18] showed that the qubit PM scenario admits an exact classical simulation with two bits of communication, and that two bits are necessary in the worst case. For comparison with the results of the following sections, we briefly recall their protocol. Without loss of generality, it is enough to consider pure states, |r→|=1| r|=1, since mixed states can be decomposed as a convex combination of pure states, and the corresponding classical randomness can be incorporated into the shared randomness of the protocol. Similarly, for measurements, it is enough to consider POVM elements proportional to rank-one projectors, (i.e., |y→i|=1,∀i | y_i |=1,\,∀ i), since general qubit POVMs can be obtained from rank-one refinements by classical post-processing. The protocol is illustrated in Fig. 2 and proceed as follows: 1. Alice is given the Bloch vector r→ r, and Bob is given the POVM specified by Y. In addition, they share two random unit vectors λ→1,λ→2∈2 λ_1, λ_2 ^2. 2. Alice computes the classical bits c1=H(λ→1⋅r→),c2=H(λ→2⋅r→), c_1=H( λ_1· r), c_2=H( λ_2· r), (3) and sends them to Bob, where H(x)H(x) is the Heaviside step function. 3. Bob defines the effective hidden variables λ→i′=(2ci−1)λ→i, λ_i =(2c_i-1) λ_i, (4) that is, he flips λ→i λ_i whenever ci=0c_i=0. 4. Bob selects y→i y_i with probability pip_i. He then sets λ→′=λ→1′ λ = λ_1 if |λ→1′⋅y→i|≥|λ→2′⋅y→i|| λ_1 · y_i|≥| λ_2 · y_i|, and λ→′=λ→2 λ = λ_2 otherwise. 5. Bob reports the outcome i with probability PB(i|λ→′,Y)=pi(y→i⋅λ→′)H(y→i⋅λ→′)∑j=1mpj(y→j⋅λ→′)H(y→j⋅λ→′). P_B(i| λ ,Y)= p_i( y_i· λ )\,H( y_i· λ ) _j=1^mp_j( y_j· λ )\,H( y_j· λ ). (5) Figure 2: 2-bit classical simulation protocol of Ref. [18] for the PM scenario. Alice and Bob share λ→1,λ→2∈2 λ_1, λ_2 ^2, Alice sends the bits c1,c2c_1,c_2, and Bob outputs i according to the protocol. As shown in Ref. [18], the statistics generated by this protocol can be written as PC(i|r→,Y)=∫2λ→ρ(λ→|r→,Y)PB(i|λ→,Y), P_C(i| r,Y)= _S^2d λ\;ρ( λ| r,Y)\,P_B(i| λ,Y), (6) where dλ→=dΩ/(4π)d λ=d /(4π) denotes the normalized uniform measure on the sphere. The effective distribution ρ(λ→|r→,Y)ρ( λ| r,Y), which encodes the statistical effect of the flipping and comparison steps, is given by: ρ(λ→|r→,Y)=8H(r→⋅λ→)∑j=1mpj(y→j⋅λ→)H(y→j⋅λ→).ρ( λ| r,Y)=8H( r· λ) _j=1^mp_j( y_j· λ)\,H( y_j· λ). (7) I Neural Network learning procedure Another important result of Ref. [18] is that no classical protocol using fewer than two classical bits is able to reproduce the statistics of all possible quantum input setups in this setting. However, this does not exclude the possibility that restricted families of measurements may admit simpler classical simulations. A priori, it remains unclear which structural properties of a POVM may enable such a reduction, and which protocols could possibly achieve it. Our aim is to use a N-procedure, as a data-driven probe for identifying such regimes. For a given POVM ℳYM_Y, Alice and Bob share a random unit vector λ→∈2 λ ^2, sampled uniformly from the Bloch sphere. Alice receives the input state ρr→ _ r and sends to Bob the single bit c=H(λ→⋅r→)c=H( λ· r). The neural network takes as input the pair (c,λ→)(c, λ) and outputs a probability distribution over the outcomes of the POVM, PNN(i|c,λ→,Y)P_ N(i|c, λ,Y). This output plays the role of Bob’s response function. The dependence on Y is implicit in the trained network: for each POVM ℳYM_Y considered, a separate network is trained. The statistics generated by the N-procedure are obtained by averaging the response function over the shared randomness. For a fixed input state ρr→ _ r, the averaged prediction is PNN(i|r→,Y)=∫2λ→PNN(i|H(λ→⋅r→),λ→,Y). P_ N(i| r,Y)= _S^2d λ\>P_ N (i|H( λ· r), λ,Y ). (8) In practice, this integral is estimated by Monte Carlo sampling. The training set consists of M Haar-random pure input states r→(j)j=1M\ r^(j)\_j=1^M. For each state r→(j) r^(j), we sample N independent shared random vectors λ→(j,k)k=1N\ λ^(j,k)\_k=1^N and define the corresponding bits c(j,k)=H(λ→(j,k)⋅r→(j))c^(j,k)=H ( λ^(j,k)· r^(j) ). For each state j and outcome i, the Monte Carlo average of the neural network response is p¯i(j)=1N∑k=1NPNN(i|c(j,k),λ→(j,k),Y). p_i^(j)= 1N _k=1^NP_ N (i|c^(j,k), λ^(j,k),Y ). (9) The target probability is the Born probability PQ(i|ρr→(j),ℳY)=pi(1+y→i⋅r→(j)). P_Q(i| _ r^(j),M_Y)=p_i (1+ y_i· r^(j) ). (10) The network is trained by minimizing the mean squared error(MSE) between the averaged neural network probabilities and the corresponding Born probabilities, ℒMSE=1M∑j=1M∑i=1m[p¯i(j)−PQ(i|ρr→(j),ℳY)]2. _ MSE= 1M _j=1^M _i=1^m [ p_i^(j)-P_Q(i| _ r^(j),M_Y) ]^2. (11) We use MSE as the training loss because it is smooth and penalizes large deviations more strongly, which is convenient for numerical optimization. To evaluate the accuracy of the trained network, we use the mean absolute error(MAE), MAE=1M∑j=1M∑i=1m|p¯i(j)−PQ(i|ρr→(j),ℳY)|.MAE= 1M _j=1^M _i=1^m | p^(j)_i-P_Q(i| _ r^(j),M_Y) |. (12) Unlike the MSE, the MAE directly quantifies the average absolute discrepancy between the predicted and target probabilities, providing a more standard figure of merit for comparing the accuracy of different protocols. Two important remarks are in order. First, for each POVM whose statistics we wish to simulate, the neural network must be retrained. Nevertheless, as we show below, the analysis of the resulting network responses allows us to infer a POVM-independent 1-bit protocol. Second, because the MAE is evaluated over randomly sampled states and shared random vectors, it quantifies an average simulation accuracy. The training methodology and network architecture are described in more detail in Appendix A. IV Numerical performance We now analyze the numerical performance of the neural network, introduced in the previous section, with a focus on how it varies depending on the characteristics of the POVMs. Throughout the numerical experiments, in this section, unless otherwise stated, we fix the parameters to N=4000N=4000 shared randomness samples and M=2000M=2000 input states during the training stage and N=4000N=4000, M=1000M=1000 during the test stage. We found these values to provide a good compromise between numerical accuracy and computational cost. Additional implementation details are given in Appendix A. We first study the N-procedure on randomly generated POVMs with different numbers of outcomes. Random POVMs can be sampled in several inequivalent ways, depending on the chosen measure over the space of effects [12]. In the present work, we use a constructive procedure based on random Bloch-sphere directions followed by a numerical search for nonnegative weights satisfying the POVM completeness condition. Fig. 3 summarizes the MAE obtained in 100100 experiments over 100100 different randomly generated POVMs with numbers of outcomes ranging from m=3m=3 to m=8m=8. The figure reports the median, mean, and interquartile range of the MAE values obtained with the N-procedure. For comparison, we also show the MAE obtained by direct finite-sample estimation of the Born probabilities. This Born rule curve should be interpreted as a sampling baseline: it quantifies the error due to finite sampling when the target probabilities are known. Figure 3: MAE of the probabilities calculated with the N-procedure for POVMs with different numbers of elements m, compared with the mean MAE sampled directly from the Born rule. The most immediate observation is that the MAE of the N-procedure is systematically larger than the MAE obtained by direct Born rule sampling. Nevertheless, the results also show that some POVM instances are reproduced much more accurately than others. In particular, for certain values of m the best cases approach the Born rule baseline rather closely. This suggests that the performance of the 1-bit protocol may depend on structural properties of the measurement. By direct inspecting the well-performing POVMs, we observe that the uniformity in the weight pii=1m\p_i\_i=1^m distribution across the POVM elements seems to play an important role. To investigate this hypothesis, we plot the MAE as a function of the standard deviation of the POVM weights. A smaller standard deviation corresponds to a more uniform distribution of the weights, with Std=0Std=0 corresponding to the equal-weight case pi=1/mp_i=1/m for all i. Fig. 4 shows that the accuracy of the N-procedure improves significantly as this standard deviation decreases, for any number of outcomes m. The most accurate cases are concentrated near the standard deviation Std=0Std=0. Figure 4: MAE of the probabilities obtained from the N-procedure for different POVMs as a function of StdStd of the weights pii=1m\p_i\_i=1^m of the elements of each POVM. POVMs with different numbers of elements m are represented with different colors. Motivated by this numerical trend, we focus from this point onwards on POVMs with equal weights, i.e., pi=1/m,∀ip_i=1/m\,,∀ i. In the rank-one case, these measurements can be written as 2m|ψi⟩⟨ψi|i=1m, \ 2m | _i \! _i | \_i=1^m, where each |ψi⟩ | _i is a normalized vector. The POVM normalization condition is therefore ∑i=1m|ψi⟩⟨ψi|=(m/2) _i=1^m | _i \! _i |=(m/2) , which implies that the vectors |ψi⟩i=1m\ | _i \_i=1^m form a unit-norm tight frame. We refer to this class as equal-trace rank-one POVMs or e-POVMs. Accordingly, they have also been termed unit-norm tight-frame POVMs in the literature [10, 4]. Geometrically, under the Bloch-sphere representation, the pure states associated with an e-POVM correspond to m points on the sphere whose centroid lies at the origin. Thus, the POVM normalization condition admits a simple interpretation as a balance condition on the corresponding Bloch-sphere configuration ∑i=1my→i=0→ _i=1^m y_i= 0, with |yi|=1|y_i|=1, where the POVM elements take the form Mi=1m(+→⋅σ→)M_i= 1m ( + y_i· σ ). When needed, we refer to the informationally complete(IC) members of this class as eIC-POVMs. Within the class of e-POVMs, it is natural to examine highly regular configurations [22] of Bloch vectors. We therefore compare the performance of the N-procedure on regular polygons and regular polyhedra inscribed in the Bloch sphere. Both classes have equal weights and balanced Bloch vectors, but they differ in an important respect: regular polygons lie in a plane and are therefore not informationally complete, whereas regular polyhedral configurations span ℝ3R^3 and define eIC-POVMs. The set of regular polygons and polyhedra inscribed in the Bloch sphere111These structures also belong to the class of ”highly-symmetric POVMs” defined in Ref. [22]. Fig. 5 compares the MAE of the N-procedure with the MAE obtained by direct sampling from the Born rule for several such configurations. The comparison reveals a clear qualitative distinction. For polygonal configurations such as the triangle, square, pentagon, and hexagon, the MAE of the N-procedure remains noticeably above the Born rule baseline. By contrast, for regular polyhedra configurations, the gap between the N-procedure and the Born rule baseline is significantly smaller. This suggests that the symmetric three-dimensional configuration of measurement vectors, representing an informationally complete POVM plays an important role. Figure 5: MAE of the N-procedure (dark gray) and of the Born rule (light gray) for various e-POVMs with m outcomes, taking 4000 samples (value of N) and averaging over 1000 different states (value of M). In the left plot, we have polygon configurations, while on the right plot we have regular polyhedra. Overall, the numerical results support two main observations. First, 1-bit N-procedure can reproduce the statistics of certain POVMs with high average accuracy, even though one bit cannot simulate the full qubit PM scenario exactly. Second, this performance is not uniform across all measurements. It strongly depends on the distribution of the POVM weights. In particular, the protocol performs best for eIC-POVMs in highly symmetric configurations. V Inferred classical 1-bit protocol The numerical results of the previous section indicate that the N-procedure performs particularly well for eIC-POVMs whose Bloch vectors form highly isotropic three-dimensional configurations. This motivates the search for a transparent analytical description of the behavior learned by the neural network. We first inspect the neural-network response and extract a simple analytical rule. We then define the corresponding 1-bit protocol and study its numerical performance for finite eIC-POVMs. Finally, we show that the protocol, although not exact for finite eIC-POVMs in general, becomes exact for regular polyhedral configurations when the input state aligns with one of the measurement outcomes, and also, in the continuous isotropic limit. V.1 Extraction and definition of the 1-bit protocol To identify the structure learned by the N-procedure, we inspect the neural-network output for a representative eIC-POVM. We use the six-direction Cartesian configuration, namely the POVM whose Bloch vectors are ±x^± x, ±y^± y, and ±z^± z. This is the octahedral configuration on the Bloch sphere and corresponds to the union of the three mutually unbiased qubit bases. Fig. 6 shows the conditional response of the neural network for the outcome associated with +z^+ z, and for the transmitted bit c=1c=1, PNN(i=5|c=1,λ→,Y)P_N(i=5|c=1, λ,Y), as a function of the shared vector λ→ λ, parametrized by the polar angles (θ,ϕ)(θ,φ). The red dot indicates the polar direction of the POVM element y→5 y_5. The left panel shows the response of a single trained neural network. Although the response is smooth and shows a clear geometric dependence on λ→ λ, the fluctuations associated with a particular training run prevent sharp boundaries between the decision regions, making difficult to infer a simple analytical rule directly. To reduce these network-dependent fluctuations, we average the response probabilities over 100100 independently trained networks. The resulting ensemble-averaged response is shown in the central panel. The corresponding response maps for all outcomes and for both values of the communicated bit are provided in Appendix E. Figure 6: Conditional response probabilities for outcome i=5i=5 of the six-direction Cartesian, or octahedral, eIC-POVM, conditioned on a transmitted bit c=1c=1. The left panel shows the output of a single trained neural network (k=1k=1), the center panel shows the output averaged over an ensemble of k=100k=100 independently trained networks, and the right panel shows the response of the analytical 1-bit protocol in (17). The probabilities are represented as a function of the shared random vector λ→ λ, parametrized by the polar angles (ϕ,θ)(φ,θ). The red marker indicates the Bloch vector y→5 y_5 associated with the selected outcome. The ensample-averaged response reveals a simple pattern. For c=1c=1, the probability associated with the POVM element y→i y_i is approximately zero on the hemisphere opposite to y→i y_i, and increases approximately linearly with the scalar product y→i⋅λ→ y_i· λ on the hemisphere centered on y→i y_i. This suggests the proportionality rule PNN(i|c=1,λ→,Y)∝(y→i⋅λ→)H(y→i⋅λ→),P_N(i|c=1, λ,Y) ( y_i· λ)\,H( y_i· λ), (13) For c=0c=0, the same pattern is observed after the flip of the hidden variable λ, i.e., λ→↦−λ→ λ - λ, leading to PNN(i|c=0,λ→,Y)∝(−y→i⋅λ→)H(−y→i⋅λ→).P_N(i|c=0, λ,Y) (- y_i· λ)\,H(- y_i· λ). (14) Equations (13) and (14) suggest that the dependence on the communicated bit can be described by a sign flip of the shared vector λ. This motivates the following 1-bit protocol: 1. Alice receives the description of a qubit state ρr→ _ r, represented by its Bloch vector r→ r, and Bob receives the description of an e-POVM ℳYM_Y specified by Y (Remember, pi=1/mp_i=1/m ∀i∀ i). In addition, they share a random unit vector λ→∈2 λ ^2, uniformly distributed over the Bloch sphere. 2. Alice computes the bit c=H(λ→⋅r→),c=H( λ· r), (15) and sends it to Bob. 3. Bob defines the effective hidden variable λ→′=(2c−1)λ→. λ^\, =(2c-1) λ. (16) Thus, Bob flips the shared vector whenever c=0c=0. 4. Bob reports the outcome i with probability P1−bit(i|λ→′,Y)=(y→i⋅λ→′)H(y→i⋅λ→′)∑j=1m(y→j⋅λ→′)H(y→j⋅λ→′).P_1-bit(i| λ^\, ,Y)= ( y_i· λ^\, )\,H( y_i· λ^\, ) _j=1^m( y_j· λ^\, )\,H( y_j· λ^\, ). (17) The right panel of Fig. 6 shows the response of (17) for the octahedral outcome associated to i=5i=5. Its geometric structure closely reproduces the ensamble-averaged neural network response: the probability of outcome i is maximal at the Bloch vector y→i y_i and concentrated in the hemisphere defined by it and increase with the positive overlap y→i⋅λ′→ y_i· λ . V.2 Numerical performance We tested the analytical protocol in (17) on six e-POVMs: the tetrahedron (m=4m=4), the octahedron (m=6m=6), the hexahedron (m=8m=8), the icosahedron (m=12m=12), the dodecahedron (m=20m=20), and the icosidodecahedron222A quasi-regular polyhedron (m=30m=30). The probabilities generated by the protocol are compared with the target Born probabilities and with the probabilities obtained by direct finite sampling from the Born rule. In this section, we quantify the discrepancy using the Kullback-Leibler divergence(KLD), averaged over randomly chosen input states. In contrast with the MAE used in Sec. I to define the training loss and measures the average absolute deviation between probabilities, the KLD is more sensitive to the full shape of the probability distribution, and in particular to relative errors in outcomes with small probabilities. For this reason, the numerical values obtained from the KLD can differ significantly from those obtained with the MAE. While both measures lead to neural networks that performed similarly both in training and testing, the extra sensitivity of the KLD is useful to us now to closely compare the complete output distributions generated by different protocols. The results are shown in Fig. 7. For small sample sizes, the KLD is dominated by statistical fluctuations, and both curves decrease as the number of samples increases. The Born-rule sampling baseline continues to approach zero, whereas the KLD of the 1-bit protocol eventually reaches a nonzero plateau. This plateau represents the intrinsic bias of the protocol with respect to the Born probabilities and shows that the finite-POVM simulation is not exact. The agreement is nevertheless very accurate for all regular polyhedral configuration studied. Moreover the intrinsic discrepancy becomes small for the POVMs with larger and more isotropically distributed sets of Bloch vectors, such as the dodecahedral and the icosidodecahedral configurations. Figure 7: KLD of the analytical 1-bit protocol and of direct Born-rule sampling for six regular eIC-POVMs: tetrahedron, octahedron, hexahedron, icosahedron, dodecahedron, and icosidodecahedron (m=4,6,8,12,20,30m=4,6,8,12,20,30, respectively). Each curve is averaged over M=50M=50 randomly sampled input states. The blue curves represent the Born-rule sampling baseline, while the yellow curves represent the analytical 1-bit protocol. We also compare the analytical 1-bit protocol directly with the neural network output. Fig. 8 shows the mean KLD of the analytical protocol of a single trained neural network, an ensemble average over 1010 independently trained neural networks, and direct sampling from the Born rule. The analytical protocol gives a smaller KLD than both the single-network response and the ensemble-average neural network response, showing that the explicit rule in (17) better reproduces the Born statistics. These results indicate that, while the neural network learns the correct qualitative structure, the explicit analytical rule provides a more accurate realization of this structure at the level of the full output distribution. Note also that the averaged neural network response has a similar mean KLD than that of a single network. Despite not improving statistical accuracy, averaging over independently trained networks was key to distill the geometric response pattern condensed in Eqs. (13) and (14). Figure 8: Mean KLD for the analytical protocol, the N-procedure, an ensemble of ten N-procedures, and direct sampling from the Born rule. For each configuration, the mean is computed over M=20M=20 random input states r→ r, using N=105N=10^5 samples per state. Fig. 8 also shows that the KLD of the 1-bit protocol approaches the KLD obtained by direct sampling from the Born rule as the number of POVM elements increases. This trend is consistent with Fig. 7: at 10510^5 samples, the 1-bit protocol is already extremely close to the Born rule sampling benchmark, especially for the dodecahedral and icosidodecahedral POVMs. This supports the interpretation that regular polyhedral eIC-POVMs with more elements provide better finite approximations to the isotropic continuum construction discussed below. In conclusion, the 1-bit protocol improves on the numerical performance of the trained N-procedures while retaining the same qualitative feature: 1-bit communication reproduces the target statistics with high accuracy for eIC-POVMs. V.3 The continuous POVM limit The analytical expression for the probability distribution over outcomes of the 1-bit protocol takes the following form. The probability density of the shared vector λ→ λ after step 3 of the protocol [cf. (16)] is ρ(λ→|r→)=2H(r→⋅λ→),ρ( λ| r)=2H( r· λ)\,, (18) and the probability of obtaining outcome i given a state with Bloch vector r→ r reads [cf. (6)] P1−bit(i|r→,Y)=2∫λ→H(r→⋅λ→)(y→i⋅λ→)H(y→i⋅λ→)∑j=1m(y→j⋅λ→)H(y→j⋅λ→),P_1-bit(i| r,\;Y)=2 d λ H( r· λ)\,( y_i· λ)\,H( y_i· λ) _j=1^m( y_j· λ)\,H( y_j· λ), (19) where recall that dλ→=dΩ/4πd λ=d /4π is the Haar measure in 2S^2. Although (19) provides a remarkably accurate approximation to the eIC-POVM probabilities, it is not exact in general. This can be seen from a simple example. For instance, in the octahedral configuration, consider the probability of obtaining the outcome +z^+ z when Alice’s state points along the diagonal in the x-z plane. In this case, one can show analytically that the quantum probability and the probability generated by the 1-bit protocol have a discrepancy of (10−3)O(10^-3). At the other extreme, if Alice’s state happens to align with one of Bob’s measurement outcomes, i.e., r→=y→i r= y_i for some i, one can show that the 1-bit protocol is exact for all outcomes i (See Appendix C.). This notwithstanding, in the limit of a continuous uniform POVM ℳ=2|m→⟩⟨m→|dm→m→∈2M=\2 | m \! m |d m\_ m ^2, dm→=dΩ/4πd m=d /4π, (19) yields the exact quantum probability density, which reads PQ(m→|ρr→,ℳ)=1+r→⋅m→.P_Q( m| _ r,M)=1+ r· m\,. (20) Indeed, by rotational symmetry the denominator in (19) becomes the number ∫m→(m→⋅λ→)H(m→⋅λ→)=14. d m\,( m· λ)\,H( m· λ)= 14. (21) Using xH(x)=(x+|x|)/2x\,H(x)=(x+|x|)/2, we get P1−bit(m→|r→,Y)=4∫λ→H(r→⋅λ→)[(m→⋅λ→)+|m→⋅λ→|].P_1-bit( m| r,Y)=4 d λH( r· λ) [\,( m· λ)+| m· λ| ]\,. (22) The first term is proportional to r→⋅m→ r· m, while the second is again a number, i.e., P1−bit(m→|r→,Y)=γ1+γ2r→⋅m→P_1-bit( m| r,Y)= _1+ _2 r· m. Taking m→=−r→ m=- r, the integral vanishes and one gets γ1=γ2 _1= _2. For m→=r→ m= r we get γ1=1 _1=1, and we recover (20). We have thus so far established that the 1-bit protocol provides accurate but not exact eIC-POVM statistics in general, and that it simulates exactly the statistics of the continuous POVM. We bridge these facts by giving evidence of the decreasing behavior of the L1L_1 distance L1(P1−bit,PQ)≔∑j=1m|P1−bit(j|r→,Y)−PQ(j|ρr→,ℳY)|L_1(P_1-bit,P_Q) _j=1^m|P_1-bit(j| r,Y)-P_Q(j| _ r,M_Y)| of the 1-bit protocol for eIC-POVMs with increasing number of outcomes. Note that MAE is the average L1L_1 distance over batch of inputs, hence the choice. We do this in two ways. For spherical t-designs (a subset of the class of eIC-POVMs, where m≥t2/4m≥ t^2/4), we prove analytically that the L1L_1 distance scales as (1/t)O(1/ t) (see Appendix D.). In addition, we perform a numerical worst-case search for the regular polyhedra considered in Fig. 7. That is, for each POVM configuration, we search for the state in the Bloch sphere that maximizes the L1L_1 distance when the PM scenario is simulated using the 1-bit protocol. Table 1 provides the results. Worst-case L1(P1−bit,PQ)L_1(P_1-bit,P_Q) Configuration m L1L_1 Projective 2 0.2138 Tetrahedron 4 0.0093 Octahedron 6 0.0159 Hexahedron 8 0.0085 Icosahedron 12 0.0034 Dodecahedron 20 0.0029 Icosidodecahedron 30 0.0028 Table 1: Worst-case L1L_1 distance for the measurement configurations considered. The order of magnitude of error, decreases by increasing the number of outcomes, starting from (10−1)O(10^-1) for m=2m=2 up to (10−3)O(10^-3) for m=30m=30. VI Conclusions We have shown that one bit of classical communication can approximate with high accuracy qubit prepare-and-measure statistics for restricted, highly structured families of measurements. Using a search by N-procedure, we identified equal-trace rank-one POVMs—and, in particular, informationally complete symmetric configurations—as the measurement families most amenable to 1-bit classical simulation. The numerical results indicate that uniform weights and three-dimensional symmetry are the key features behind this enhanced simulability. Guided by the structure learned by the network, we derived an explicit analytical 1-bit protocol and analyzed its performance on finite symmetric POVMs. Although the protocol is not exact for finite configurations, its error decreases as the measurement becomes more isotropic. In the continuous limit, corresponding to the covariant qubit POVM, the protocol becomes exact. This proves that the continuous isotropic qubit measurement is exactly 1-bit simulable in the prepare-and-measure scenario. Our results demonstrate that the 2-bit communication cost required for arbitrary qubit POVMs is not representative of all measurement families. Symmetry and isotropy can substantially reduce the classical resources needed to reproduce quantum statistics. This result loosely aligns with other recent observations pointing out that, perhaps contrary to prior common understanding, SIC-POVMs do not comprise the ”most quantum” representatives among all measurements: they are not the hardest to simulate by convex combinations of projective measurements [6], and they possess the least intrinsic randomness [7]. More broadly, this work contributes to the growing effort to identify low-communication classical descriptions of nonclassical correlations. Related questions have been studied in Bell-simulation settings, including 1-bit protocols, communication-assisted models, and neural network based approaches to nonclassicality problems [15, 17, 13]. Our results show that neural networks can also help uncover analytically tractable 1-bit protocols in prepare-and-measure scenarios, and point toward a systematic classification of measurement families with reduced classical communication cost. VII Acknowledgements RMT, GS, and MZ acknowledge financial support from MCIN grant PID2022-141283NB-I00 funded by MCIN/AEI/10.13039/501100011033. MZ acknowledges Some Sankar Bhattacharya and support from the Ministerio de Ciencia e Innovación of the Spanish Government, contract PRE2020-093634. G acknowledges DAAD, the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation, project numbers 447948357 and 440958198), the Sino-German Center for Research Promotion (Project M-0294), and the German Ministry of Education and Research (Project QuKuK, BMBF Grant No. 16KIS1618K). GS acknowledges funding from the Programa Talent UAB - Banco de Santander. EF acknowledges support by the Generalitat de Catalunya and the European Social Fund grant Joan Oró 2025 FI-1 00848. JE acknowledges the funding received by i2CAT from Department de Recerca i Universitats of the Generalitat de Catalunya. References [1] J. Bowles, N. Brunner, and M. Pawłowski (2015-08) Testing dimension and nonclassicality in communication networks. Phys. Rev. A 92, p. 022351. External Links: Document, Link Cited by: §I. [2] G. Brassard (2003) Quantum communication complexity. Foundations of Physics 33, p. 1593–1616. Cited by: §I. [3] H. Buhrman, R. Cleve, S. Massar, and R. De Wolf (2010) Nonlocality and communication complexity. Reviews of Modern Physics 82, p. 665–698. Cited by: §I. [4] P. G. Casazza and G. Kutyniok (2013) Finite frames: theory and applications. Springer Science & Business Media. Cited by: §IV. [5] N. J. Cerf, N. Gisin, and S. Massar (2000) Classical teleportation of a quantum bit. Physical Review Letters 84 (11), p. 2521. Cited by: §I. [6] G. Cobucci, R. Brinster, S. Khandelwal, H. Kampermann, D. Bruß, N. Wyderka, and A. Tavakoli (2026-02) Maximally nonprojective measurements are not always symmetric informationally complete. Phys. Rev. Lett. 136, p. 060201. External Links: Document, Link Cited by: §VI. [7] F. Curran (2026) Quantum randomness beyond projective measurements. arXiv preprint quant-ph/2605.18291. Cited by: §VI. [8] C. de Gois, G. Moreno, R. Nery, S. Brito, R. Chaves, and R. Rabelo (2021) General method for classicality certification in the prepare and measure scenario. PRX Quantum 2, p. 030311. Cited by: §I. [9] J. Degorre, S. Laplante, and J. Roland (2005) Simulating quantum correlations as a distributed sampling problem. Physical Review A 72, p. 062314. Cited by: §I. [10] Y.C. Eldar and G.D. Forney (2002) Optimal tight frames and quantum measurement. IEEE Transactions on Information Theory 48 (3), p. 599–610. External Links: Document Cited by: §IV. [11] R. Gallego, N. Brunner, C. Hadley, and A. Acín (2010) Device-independent tests of classical and quantum dimensions. Physical Review Letters 105, p. 230501. Cited by: §I. [12] T. Heinosaari, M. A. Jivulescu, and I. Nechita (2020) Random positive operator valued measures. Journal of Mathematical Physics 61, p. 042202. External Links: Document Cited by: §IV. [13] T. Kriváchy, Y. Cai, D. Cavalcanti, A. Tavakoli, N. Gisin, and N. Brunner (2020) A neural network oracle for quantum nonlocality problems in networks. npj Quantum Information 6, p. 70. Cited by: §VI. [14] A. Montina (2011) Approximate simulation of entanglement with a linear cost of communication. Physical Review A—Atomic, Molecular, and Optical Physics 84 (4), p. 042307. Cited by: §I. [15] A. Montina (2012) Epistemic view of quantum states and communication complexity of quantum channels. Physical Review Letters 109, p. 110501. Cited by: §VI. [16] M. Pawłowski and N. Brunner (2011-07) Semi-device-independent security of one-way quantum key distribution. Phys. Rev. A 84, p. 010302(R). External Links: Document, Link Cited by: §I. [17] M. J. Renner and M. T. Quintino (2023) The minimal communication cost for simulating entangled qubits. Quantum 7, p. 1149. Cited by: §VI. [18] M. J. Renner, A. Tavakoli, and M. T. Quintino (2023) Classical cost of transmitting a qubit. Physical Review Letters 130, p. 120801. Cited by: Figure 10, §A.3, §I, Figure 2, §I, §I, §I. [19] T. Rudolph (2006) Ontological models for quantum mechanics and the kochen-specker theorem. arXiv preprint quant-ph/0608120. Cited by: §I. [20] S. Schlösser and M. Kleinmann (2026) Bounding the classical cost of simulating quantum behaviors in the prepare-and-measure scenario. arXiv preprint arXiv:2603.01255. Cited by: §I, §I. [21] P. Sidajaya, A. D. Lim, B. Yu, and V. Scarani (2023) Neural network approach to the simulation of entangled states with one bit of communication. Quantum 7, p. 1150. Cited by: §I, §I. [22] W. Slomczynski and A. Szymusiak (2014) Highly symmetric povms and their informational power. arXiv preprint quant-ph/1402.0375. Cited by: §IV, footnote 1. [23] B. F. Toner and D. Bacon (2003) Communication cost of simulating bell correlations. Physical Review Letters 91, p. 187904. Cited by: §I, §I. [24] M. Zartab, G. Gasbarri, G. Sentís, and R. Muñoz-Tapia (2026) Prepare-and-measure and entanglement simulation beyond qubits. Scientific Reports 16, p. 21297. Cited by: §I. [25] C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal (1997) Algorithm 778: l-bfgs-b: fortran subroutines for large-scale bound-constrained optimization. ACM Transactions on mathematical software (TOMS) 23 (4), p. 550–560. Cited by: Appendix B. Appendix A Neural network training methodology and hyperparameter selection A.1 Training data and batch-averaged loss For a given POVM ℳYM_Y, the training dataset is constructed as follows. First, we generate M random pure qubit states sampled uniformly on the Bloch sphere. We denote their Bloch vectors by r→(j) r^(j), with j=1,…,Mj=1,…,M. For each state r→(j) r^(j), the target quantum probability distribution is computed using the Born rule, PQ(i|ρr→(j),ℳY)=pi(1+y→i⋅r→(j)),i=1,…,m.P_Q(i| _ r^(j),M_Y)=p_i (1+ y_i· r^(j) ), i=1,…,m. (23) These probabilities constitute the target distributions of the training procedure. Next, for each state r→(j) r^(j), we generate N realizations of the shared randomness. We denote the k-th random vector associated with the state r→(j) r^(j) by λ→(j,k) λ^(j,k), with k=1,…,Nk=1,…,N. Each vector is sampled uniformly on the unit sphere. The corresponding communicated bit is then computed as c(j,k)=H(λ→(j,k)⋅r→(j)).c^(j,k)=H ( λ^(j,k)· r^(j) ). (24) Therefore, each state r→(j) r^(j) gives rise to N neural network inputs, each of them formed by the three components of λ→(j,k) λ^(j,k) and the bit c(j,k)c^(j,k). The full training dataset contains M×NM× N input instances. However, these instances are naturally grouped into M batches, each corresponding to one quantum state. This grouping is essential because the neural network is not required to reproduce the quantum probability distribution for each individual realization of the shared randomness. Instead, the relevant quantity is the output distribution obtained after averaging over the N realizations associated with the same state r→(j) r^(j). Let y^i(j,k) y_i^(j,k) denote the neural network prediction for outcome i for the k-th shared randomness realization associated with the state r→(j) r^(j). The batch-averaged prediction for that state is y¯i(j)=1N∑k=1Ny^i(j,k)=p¯i(j), y_i^(j)= 1N _k=1^N y_i^(j,k)= p_i^(j), (25) where p¯i(j) p_i^(j) is the same as (9). This quantity is compared with the target quantum probability yi(j):=PQ(i|ρr→(j),ℳY).y_i^(j):=P_Q(i| _ r^(j),M_Y). (26) The loss associated with the state r→(j) r^(j) is defined by Lj=∑i=1m[p¯i(j)−PQ(i|ρr→(j),ℳY)]2.L_j= _i=1^m [ p_i^(j)-P_Q(i| _ r^(j),M_Y) ]^2. (27) The full loss function, according to (11) is then obtained by averaging over the M sampled states: ℒMSE=1M∑j=1MLj=1M∑j=1M∑i=1m[p¯i(j)−PQ(i|ρr→(j),ℳY)]2.L_ MSE= 1M _j=1^ML_j= 1M _j=1^M _i=1^m [ p_i^(j)-P_Q(i| _ r^(j),M_Y) ]^2. (28) In this way, M is the number of quantum states used during training, whereas N is the number of shared randomness samples used for each state. Increasing M improves the coverage of the Bloch sphere during training, while increasing N improves the numerical estimate of the averaged output distribution. A.2 Network architecture and optimization For each POVM, we train an independent feed-forward neural network. The input layer has four neurons, corresponding to the three Cartesian components of λ→ λ and the communicated bit c. The network has two hidden layers with 10 and 16 neurons, respectively. The output layer has m neurons, where m is the number of outcomes of the POVM. ReLU activation functions are used in the hidden layers. The output layer uses a softmax activation function, ensuring that the neural network output is a normalized probability distribution over the m possible outcomes. The training-set size is determined by M and N, corresponding to the number of randomly produced states, and number of independent Monte Carlo samples of shared randomness used per each state, respectively. For the number of input states, we set Mval=M4M_val= M4 for validation, and Mtest=1000M_test=1000 for testing. For N, in both validation and testing phases, we have Ntest=Nval=N_test=N_val=N. The network is trained using the Adam optimizer with a learning rate of 0.0010.001. In all simulations, the number of training epochs is fixed to 100. An example of the training convergence is shown in Fig. 9. The loss defined in (28) converges to approximately 1.5×10−31.5× 10^-3 after 100 epochs for both the training and validation datasets. Figure 9: Loss as a function of the number of epochs for the training and validation datasets. A.3 Performance metric and selection of M and N The performance of the N-procedure is evaluated using MAE between the batch-averaged neural network prediction and the target quantum probability distribution: MAE=1M∑j=1M∑i=1m|1N∑k=1Ny^i(j,k)−yi(j)|=1M∑j=1M∑i=1m|p¯i(j)−PQ(i|ρr→(j),ℳY)|.MAE= 1M _j=1^M _i=1^m | 1N _k=1^N y_i^(j,k)-y_i^(j) |= 1M _j=1^M _i=1^m | p_i^(j)-P_Q(i| _ r^(j),M_Y) |. (29) As in the loss function, the average over the N shared randomness realizations is necessary because the relevant quantity is the probability distribution predicted for each state, rather than the output of the network for a single realization of λ→ λ. To select suitable values of M and N, we analyze the performance of the N-procedure for a POVM with m=6m=6 elements. Fig. 10 shows the MAE for different values of M and N. The error decreases on average as both parameters increase, although the improvement is more pronounced when increasing N. This is expected because a larger value of N gives a more accurate estimate of the batch-averaged distribution that the network is trained to reproduce, whereas a larger value of M improves the coverage of the state space during training. Figure 10: MAE as a function of the batch size N for different sample sizes M, compared with the MAE of the protocol of Renner et al. [18]. For comparison, Fig. 10 also includes the MAE obtained from the exact 2-bit protocol of Ref. [18]. Although that protocol reproduces the quantum probabilities exactly in the infinite-sampling limit, its numerical implementation also depends on the number of samples used to estimate the corresponding probabilities. Therefore, finite-N effects are also present in that case. The comparison shows that the 1-bit N-procedure remains less accurate than the exact 2-bit simulation, but nevertheless provides a nontrivial approximation to the target quantum statistics. For all remaining experiments, we fix N=4000N=4000, and M=2000M=2000, which provides a good compromise between numerical accuracy and computational cost. Appendix B Generation of random qubit POVMs Random rank-one qubit POVMs are generated by sets of random points, y→ y in the unit sphere 2S^2. These can be simply obtained from sampling two independent random variables u,v∼(0,1)u,v (0,1) and defining ϕ=2πuφ=2π u and θ=arccos(1−2v)θ= (1-2v). The random unit vector in spherical coordinates y→=(sinθcosϕ,sinθsinϕ,cosθ) y=( θ φ,\, θ φ,\, θ) is then assured to be uniformly sampled. In the implementation, we generated a pool of 10610^6 points. For a POVM with m outcomes, we then selected m directions y→jj=1m\ y_j\_j=1^m at random from this pool. Once the directions y→jj=1m\ y_j\_j=1^m were fixed, the weights pjj=1m\p_j\_j=1^m were obtained numerically by enforcing the POVM conditions: ∑j=1mpj=1,∑j=1mpjy→j=0→. _j=1^mp_j=1, _j=1^mp_j y_j= 0. (30) We used the numerical routine L-BFGS-B [25] that allowed us to fix the initial point to pj=1/mp_j=1/m and to enforce that no weight became too small. If a sampled set of directions did not admit an admissible solution, a new set of directions was generated and the procedure was repeated. Appendix C Exactness of the protocol for the regular polyhedra, where the state aligns with one the measurement outcomes. In this appendix, we prove that for the regular polyhedra, the protocol gives exact solution whenever r→=y→i r= y_i for some i. We start from the analytical description of the protocol from (19) of the main text: P1−bit(k|r→=y→i,Y)=2∫2λ→H(y→i⋅λ→)y→k⋅λ→H(y→k⋅λ→)∑j=1my→j⋅λ→H(y→j⋅λ→). P_1-bit(k| r= y_i,Y)=2 _S^2d λ H( y_i· λ) y_k· λ\>H( y_k· λ) _j=1^m y_j· λ\>H( y_j· λ). (31) Now, by ∑j=1my→j=0 _j=1^m y_j=0, using the identity xH(x)=|x|+x2xH(x)= |x|+x2, and applying H(y→i⋅λ→)H( y_i· λ), over the integral region, we get: P1−bit(k|r→=y→i,Y)=2∫y→i⋅λ→≥0λ→|y→k⋅λ→|∑j=1m|y→j⋅λ→|+2∫y→i⋅λ→≥0λ→y→k⋅λ→∑j=1m|y→j⋅λ→|≔I1+I2. P_1-bit(k| r= y_i,Y)=2 _ y_i· λ≥ 0d λ\> | y_k· λ| _j=1^m| y_j· λ|+2 _ y_i· λ≥ 0d λ\> y_k· λ _j=1^m| y_j· λ| I_1+I_2. (32) Because the Integrand in I1I_1 is even under λ→↦−λ→ λ - λ, we have: I1=∫2λ→|y→k⋅λ→|∑j=1m|y→j⋅λ→|=1m I_1= _S^2d λ | y_k· λ| _j=1^m| y_j· λ|= 1m (33) where we have used the symmetry of the regular polyhedra. For I2I_2 we have: I2=2y→k⋅∫y→i⋅λ→≥0λ→λ→∑j=1m|y→j⋅λ→|≔2y→k⋅w→, I_2=2 y_k·\> _ y_i· λ≥ 0d λ λ _j=1^m| y_j· λ| 2 y_k· w, (34) where w→ w is an effective vector. Now, let’s see how it behaves under the symmetry rotation Riw→R_i w, where RiR_i is the nontrivial rotation around the axis y→i=Riy→i y_i=R_i y_i, that maps the regular polyhedral shape to itself, up to relabeling the indices: Riw→=∫y→i⋅λ→≥0λ→Riλ→∑j=1m|y→j⋅λ→|=∫(y→iTRiTRiλ→)≥0|det(Ri)|λ→Riλ→∑j=1m|y→jTRiTRiλ→|=w→. R_i w= _ y_i· λ≥ 0d λ R_i λ _j=1^m| y_j· λ|= _( y_i^TR_i^TR_i λ)≥ 0|det(R_i)|d λ R_i λ _j=1^m| y_j^TR_i^TR_i λ|= w. (35) The last equality results from changing the variable of integration by λ→′≔Riλ→ λ R_i λ and using the fact |det(Ri)|=1|det(R_i)|=1. Therefore, we see that w→=αy→i w=α y_i. To derive α, we apply the dot product with y→i y_i: y→i⋅w→=α=∫y→i⋅λ→≥0λ→y→i⋅λ→∑j=1m|y→j⋅λ→|=12m, y_i· w=α= _ y_i· λ≥ 0d λ y_i· λ _j=1^m| y_j· λ|= 12m, (36) where we have used the fact that in the region of integration, y→i⋅λ→=|y→i⋅λ→| y_i· λ=| y_i· λ|, and (33). Now, by putting it into (34), we have: I2=1my→k⋅y→i, I_2= 1m y_k· y_i, (37) which alongside with I1I_1 can be combined and substituted into (32), which leads to the expression: P1−bit(k|r→=y→i,Y)=1m(1+y→i⋅y→k)=PQ(k|ρr→=y→i,ℳY).□ P_1-bit(k| r= y_i,Y)= 1m(1+ y_i· y_k)=P_Q(k| _ r= y_i,M_Y). (38) Appendix D Scaling of L_1 distance for spherical t-designs In this appendix, we study the asymptotic behavior of L1L_1 distance of the analytical 1-bit protocol from quantum predictions, and provide an upper bound for spherical t−t-design measurements that scales as (1/t)O(1/ t). As any spherical t−t-design is a particular example of an e-POVM, we start by writing: Mi=1m(I+y→i⋅σ→),y→i∈2, M_i= 1m (I+ y_i· σ ), y_i ^2, (39) with ∑i=1my→i=0→ _i=1^m y_i= 0. Without loss of generality, we also restrict to pure input states, so that r→∈2 r ^2. The Born rule probability is then PQ(i|ρr→,ℳY)=1m(1+y→i⋅r→). P_Q(i| _ r,M_Y)= 1m (1+ y_i· r ). (40) It is useful to write the quantum probabilities in integral form as 1m(1+y→i⋅r→)=8m∫λ′→H(r→⋅λ→′)y→i⋅λ→′H(y→i⋅λ→′)=8m∫r→⋅λ→′≥0λ′→y→i⋅λ→′H(y→i⋅λ→′). 1m (1+ y_i· r )= 8m d λ \,H( r· λ ) y_i· λ \,H( y_i· λ )= 8m _ r· λ ≥ 0d λ \, y_i· λ \,H( y_i· λ )\,. (41) Recall that the 1-bit protocol probability (19) reads P1−bit(i|r→,Y)=2∫λ→′H(r→⋅λ→′)(y→i⋅λ→′)H(y→i⋅λ→′)∑j=1m(y→j⋅λ→′)H(y→j⋅λ→′),P_1-bit(i| r,\;Y)=2 d λ H( r· λ )\,( y_i· λ )\,H( y_i· λ ) _j=1^m( y_j· λ )\,H( y_j· λ ), (42) where the denominator can be written as ∑j=1m(y→i⋅λ→′)H(y→i⋅λ→′)=12∑j=1m|y→j⋅λ→′|≔SY(λ→′) _j=1^m( y_i· λ )H( y_i· λ )= 12 _j=1^m| y_j· λ | S_Y( λ ) (43) In the above we have used the identity xH(x)=(x+|x|)/2xH(x)=(x+|x|)/2 and the t-design property ∑iyi→=0 _i y_i=0. It is also convenient to define the quantities Q(i|λ→′,Y)=4m(y→i⋅λ→′)H(y→i⋅λ→′), Q(i| λ ,Y)= 4m( y_i· λ )H( y_i· λ ), (44) and ZY(λ→′):=∑i=1mQ(i|λ→′,Y)=2m∑j=1m|y→j⋅λ→′|=4mSY(λ→′), Z_Y( λ ):= _i=1^mQ(i| λ ,Y)= 2m _j=1^m| y_j· λ |= 4mS_Y( λ ), (45) to have |P1−bit(i|r→,Y)−PQ(i|ρr→,ℳY)|=2|∫r→⋅λ→′≥0Q(i|λ→′,Y)[1ZY(λ→′)−1]dλ→′.|≤2∫r→⋅λ→′≥0Q(i|λ→′,Y)|1ZY(λ→′)−1|dλ→′. |P_1-bit(i| r,Y)-P_Q(i| _ r,M_Y)|=2 | _ r· λ ≥ 0Q(i| λ ,Y) [ 1Z_Y( λ )-1 ]d λ . |≤ 2 _ r· λ ≥ 0Q(i| λ ,Y) | 1Z_Y( λ )-1 |d λ . (46) We now use the expansion of the absolute function |x||x| in terms of a series of Legendre polynomials |x|=∑k=0∞c2kP2k(x), |x|= _k=0^∞c_2kP_2k(x), (47) where the coefficients read c2k=(−1)k−1(4k+1)(2k−3)!!(2k+2)!!, c_2k=(-1)^k-1(4k+1) (2k-3)!!(2k+2)!!, (48) (note that by analytical extension c0=1/2c_0=1/2) to write the terms |y→i⋅λ→′|| y_i· λ | in (43) in the Legendre polynomials basis: SY(λ→′)=12∑j=1m[∑k=0∞c2kP2k(y→j⋅λ→′)]. S_Y( λ )= 12 _j=1^m[ _k=0^∞c_2kP_2k( y_j· λ )]. (49) Notice that for spherical t−t-designs one has: SY(λ→′)=12∑j=1m(∑k=0⌊t/2⌋c2kP2k(y→j⋅λ→′)+∑k=⌊t/2⌋+1∞c2kP2k(y→j⋅λ→′)) S_Y( λ )= 12 _j=1^m ( _k=0 t/2 c_2kP_2k( y_j· λ )+ _k= t/2 +1^∞c_2kP_2k( y_j· λ ) ) =m2∫2y→(∑k=0⌊t/2⌋c2kP2k(y→⋅λ→′))+12∑j=1m∑k=⌊t/2⌋+1∞c2kP2k(y→j⋅λ→′) = m2 _S^2d y\> ( _k=0 t/2 c_2kP_2k( y· λ ) )+ 12 _j=1^m _k= t/2 +1^∞c_2kP_2k( y_j· λ ) =m2y→[(∑k=0⌊t/2⌋c2kP2k(y→⋅λ→′))]+12∑j=1m∑k=⌊t/2⌋+1∞c2kP2k(y→j⋅λ→′). = m2\;E_ y [ ( _k=0 t/2 c_2kP_2k( y· λ ) ) ]+ 12 _j=1^m _k= t/2 +1^∞c_2kP_2k( y_j· λ ). (50) We note that all Legendre polynomials up to ⌊t/2⌋ t/2 satisfy [P2k(y→⋅λ→′)]=0E[P_2k( y· λ )]=0, except for k=0k=0, which gives [P0(y→⋅λ→′)]=1E[P_0( y· λ )]=1, resulting in the first term reproducing a constant value of m/4m/4, and therefore, one has: SY(λ→′)=m4+Rt(λ→′), S_Y( λ )= m4+R_t( λ ), (51) where Rt(λ→′):=12∑j=1m∑k=⌊t/2⌋+1∞c2kP2k(y→j⋅λ→′). R_t( λ ):= 12 _j=1^m _k= t/2 +1^∞c_2kP_2k( y_j· λ ). (52) Additionally, as: 4m|Rt(λ→′)|=2m|∑j=1m|y→j⋅λ→′|−m2|<1, 4m|R_t( λ )|= 2m | _j=1^m| y_j· λ |- m2 |<1, (53) substituting it into (46) leads to: |P1−bit(i|r→,Y)−PQ(i|ρr→,ℳY)| |P_1-bit(i| r,Y)-P_Q(i| _ r,M_Y) | ≤2∫r→⋅λ→′≥0Q(i|λ→′,Y)|1(1+4mRt(λ→′))−1|λ→′ ≤ 2 _ r· λ ≥ 0Q(i| λ ,Y) | 1(1+ 4mR_t( λ ))-1 |d λ ≤supλ→′|4Rt(λ→′)m+4Rt(λ→′)|PQ(i|ρr→,ℳY). _ λ | 4R_t( λ )m+4R_t( λ ) |P_Q(i| _ r,M_Y). (54) Now, let’s analyze F(Rt(λ→′)):=|4Rt(λ→′)m+4Rt(λ→′)|F(R_t( λ )):=| 4R_t( λ )m+4R_t( λ )|, as a function of RtR_t. To find an upper bound, from (53) it follows that: supλ→′∈2[F(Rt(λ→′)]≤4supλ→′∈2[|Rt(λ→′)|]m−4supλ→′∈2[|Rt(λ→′)|] sup_ λ ^2[F(R_t( λ )]≤ 4\> sup_ λ ^2[|R_t( λ )|]m-4\> sup_ λ ^2[|R_t( λ )|] (55) On the other hand: |Rt(λ→′)| |R_t( λ )| ≤12∑j=1m∑k=⌊t/2⌋+1∞|c2k|⋅|P2k(y→j⋅λ→′)|≤m2∑k=⌊t/2⌋+1∞|c2k| ≤ 12 _j=1^m _k= t/2 +1^∞|c_2k|·|P_2k( y_j· λ )|≤ m2 _k= t/2 +1^∞|c_2k| (56) as |P2k(⋅)|≤1|P_2k(·)|≤ 1 for all k’s. Substituting this inequality into (55), for large enough t, where ∑k=⌊t/2⌋+1∞|c2k|<1/2 _k= t/2 +1^∞|c_2k|<1/2, results in: supλ→′∈2[F(Rt(λ→′)]≤2supλ→′∈2|Rt(λ→′)|1−2supλ→′∈2|Rt(λ→′)|≤2∑k=⌊t/2⌋+1∞|c2k|1−2∑k=⌊t/2⌋+1∞|c2k|. sup_ λ ^2[F(R_t( λ )]≤ 2\> sup_ λ ^2|R_t( λ )|1-2\> sup_ λ ^2|R_t( λ )|≤ 2 _k= t/2 +1^∞|c_2k|1-2 _k= t/2 +1^∞|c_2k|. (57) We finally upper bound the coefficients |c2k||c_2k|. For this, we rewrite (48) as: |c2k|=4k+14k2+2k−2⋅(2k−1)!!(2k)!! |c_2k|= 4k+14k^2+2k-2· (2k-1)!!(2k)!! (58) and treat the first and second term separately. For the first term, we have: 4k+14k2+2k−2=54for k=1, and4k+14k2+2k−2≤1kfor k>1. 4k+14k^2+2k-2= 54 for $k=1$, \ and\ \ \ 4k+14k^2+2k-2≤ 1k for $k>1$. (59) For the second term, we use the Wallis identity 2π∫0π/2xcos2kx=(2k−1)!!(2k)!!, 2π _0^π/2dx ^2kx= (2k-1)!!(2k)!!\,, (60) together with the inequality log(cosx)≤−x2/2 ( x)≤-x^2/2, to obtain (2k−1)!!(2k)!!≤2π∫0π/2xe−kx2<∫0∞xe−kx2=1πk (2k-1)!!(2k)!!≤ 2π _0^π/2dx\,e^-kx^2< _0^∞dx\,e^-kx^2= 1 π k (61) Consequently, each term of (59) is upper bounded by: |c2|≤54π;|c2k>2|≤1k3π. |c_2|≤ 54 π; |c_2k>2|≤ 1 k^3π. (62) Further, the summation ∑k=⌊t/2⌋+1∞|c2k| _k= t/2 +1^∞|c_2k| is further upper bounded by using the fact that 1k3π 1 k^3π is a monotonic decreasing function of k and the relationship between the sum and the corresponding integral ∑k=⌊t/2⌋+1∞|c2k|≤1π∫⌊t/2⌋∞k−3/2k=2π⌊t/2⌋. _k= t/2 +1^∞|c_2k|≤ 1 π _ t/2 ^∞k^-3/2dk= 2 π t/2 . (63) Therefore, by putting it into (57), one gets: supλ→′∈2[F(Rt(λ→′)]≤4π⌊t/2⌋⋅π⌊t/2⌋π⌊t/2⌋−4=4π⌊t/2⌋−4. sup_ λ ^2[F(R_t( λ )]≤ 4 π t/2 · π t/2 π t/2 -4= 4 π t/2 -4. (64) and further, by substituting it into (D), one has: |P1−bit(i|r→,Y)−PQ(i|ρr→,ℳY)| |P_1-bit(i| r,Y)-P_Q(i| _ r,M_Y)| ≤4π⌊t/2⌋−4PQ(i|ρr→,ℳY). ≤ 4 π t/2 -4P_Q(i| _ r,M_Y). (65) Finally, by summing over all outcomes, the upper bound for L1(P1−bit,PQ)L_1(P_1-bit,P_Q) reads: L1(P1−bit,PQ)=∑i=1m|P1−bit(i|r→,Y)−PQ(i|ρr→,ℳY)| L_1(P_1-bit,P_Q)= _i=1^m |P_1-bit(i| r,Y)-P_Q(i| _ r,M_Y) | ≤4π⌊t/2⌋−4=(1/t)as t→∞.□ ≤ 4 π t/2 -4=O(1/ t) as t→∞. (66) Appendix E Additional Plots The main text focuses on a representative conditional response map, namely the outcome y→5 y_5 for the transmitted bit c=1c=1, in order to illustrate how the ensemble-averaged neural network output motivates the analytical 1-bit protocol. Here, we provide the corresponding complete set of response maps for the six-direction Cartesian (octahedral) eIC-POVM. In particular, we display all six outcomes and both possible values of the transmitted bit. Each figure is organized in six rows and four columns. Row i corresponds to the POVM element y→i y_i, while the first three columns show the neural network conditional response probabilities obtained from ensembles of size k=1k=1, k=5k=5, and k=100k=100, respectively. The last column displays the corresponding response of the analytical 1-bit protocol defined in (17). In each panel, the probability is shown as a function of the shared random vector λ→ λ, parametrized by the polar angles (ϕ,θ)(φ,θ). The red marker indicates the direction of the Bloch vector associated with the outcome shown in that row. The common color scale, ranging from 0 to 11, allows a direct comparison of the different ensemble sizes, outcomes, and analytical predictions. Figure 11: Conditional response maps for the six outcomes of the six-direction Cartesian (octahedral) eIC-POVM, conditioned on the transmitted bit c=1c=1. Each row corresponds to one POVM element y→i y_i, with i=1,…,6i=1,…,6. The first three columns show the neural network response probabilities PNN(i|c=1,λ→,Y)P_N(i|c=1, λ,Y) for ensemble sizes k=1k=1, k=5k=5, and k=100k=100, respectively. The fourth column shows the analytical response P1−bit(i|λ→,′,Y)P_1-bit(i| λ^, ,Y) from (17), where c=1c=1. The red marker in each panel identifies the Bloch vector y→i y_i associated with the displayed outcome. Figure 12: Conditional response maps for the six outcomes of the six-direction Cartesian (octahedral) eIC-POVM, conditioned on the transmitted bit c=0c=0. The arrangement is identical to that in Fig. 11: rows correspond to the outcomes y→1,…,y→6 y_1,…, y_6, and columns show the neural network responses for ensemble sizes k=1k=1, k=5k=5, and k=100k=100, followed by the analytical 1-bit response. In this case, the analytical protocol uses the flipped hidden variable λ→′=−λ→ λ =- λ, as prescribed by (16). The maps show that the same response structure is recovered for c=0c=0, with the support of each outcome shifted according to this sign flip. Red markers indicate the directions of the corresponding POVM Bloch vectors.