Paper deep dive
Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds
Keito Inoshita
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/22/2026, 2:38:43 AM
Summary
The paper introduces Bias-corrected Affective Ceiling Estimation (BACE), a framework to estimate the irreducible error (Bayes error) in emotion recognition tasks by accounting for human label variation and annotation noise. Using an anchored Dirichlet-mixture empirical Bayes estimator and noise deconvolution, the authors demonstrate that while unconstrained point estimates vary widely, at least 33% of the error on the GoEmotions dataset is irreducible due to intrinsic ambiguity.
Entities (9)
Relation Signals (7)
BACE → appliedto → GoEmotions
confidence 98% · at least about 33% of a representative classifier's error on GoEmotions is irreducible
BACE → decomposes → error
confidence 95% · separates irreducible from reducible error... separates classifier error into aleatoric and epistemic components
SamLowe/roberta-base-go_emotions → evaluatedon → GoEmotions
confidence 95% · The evaluated classifier is the representative public SamLowe/roberta-base-go_emotions... on GoEmotions
Fano's inequality → provides → lower bound
confidence 95% · Fano’s inequality links conditional entropy to a lower bound on any predictor’s error
BACE → uses → Dirichlet-mixture empirical Bayes
confidence 95% · An anchored Dirichlet-mixture empirical Bayes estimator... recovers the human-consensus distribution
BACE → appliedto → EPIC
confidence 90% · the same pattern recurring on... irony... EPIC
BACE → appliedto → MD-Agreement
confidence 90% · the same pattern recurring on offensiveness... MD-Agreement
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline. Our aim is not to pin this ceiling to a single number, but to quantify how far it depends on finite annotation, estimator choice, annotation noise, and the evaluation protocol, and thereby to discipline how confidently saturation can be claimed. We propose Bias-corrected Affective Ceiling Estimation (BACE), an analysis framework that estimates a bias-corrected ceiling, separates irreducible from reducible error, and disciplines the resulting claims. An anchored Dirichlet-mixture empirical Bayes estimator, bracketed between plug-in and NSB, recovers the human-consensus distribution; an annotator split, a noise deconvolution, and a fixed claim gate then attribute error without circularity. Methodologically, unconstrained point estimates place reachability anywhere from 0.38 to 1.03, so saturation cannot be decided by any single estimator. Substantively, the only assertion passing the claim gate is that at least about 33% of a representative classifier's error on GoEmotions is irreducible, with the same pattern recurring on offensiveness and irony.
Tags
Links
- Source: https://arxiv.org/abs/2608.15619v1
- Canonical: https://arxiv.org/abs/2608.15619v1
Trouble viewing inline? Open PDF directly →
Full Text
68,875 characters extracted from source content.
Expand or collapse full text
Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds Keito Inoshita Affiliation: Faculty of Business and Commerce, Kansai University Email: inosita.2865@gmail.com Abstract Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline. Our aim is not to pin this ceiling to a single number, but to quantify how far it depends on finite annotation, estimator choice, annotation noise, and the evaluation protocol, and thereby to discipline how confidently saturation can be claimed. We propose Bias-corrected Affective Ceiling Estimation (BACE), an analysis framework that estimates a bias-corrected ceiling, separates irreducible from reducible error, and disciplines the resulting claims. An anchored Dirichlet-mixture empirical Bayes estimator, bracketed between plug-in and NSB, recovers the human-consensus distribution; an annotator split, a noise deconvolution, and a fixed claim gate then attribute error without circularity. Methodologically, unconstrained point estimates place reachability anywhere from 0.380.38 to 1.031.03, so saturation cannot be decided by any single estimator. Substantively, the only assertion passing the claim gate is that at least about 3333% of a representative classifier’s error on GoEmotions is irreducible, with the same pattern recurring on offensiveness and irony. Keywords: emotion recognition; human label variation; uncertainty decomposition; Bayes-error ceiling; empirical Bayes. 1 Introduction Driven by advances in deep learning, emotion recognition from text has continued to improve on benchmarks, and the field has entered a stage in which powerful classifiers compete over marginal differences in performance. However, whether a ceiling exists on the accuracy that can be achieved, and how close current classifiers are to that ceiling, has rarely been asked. Unless it is determined whether residual errors stem from limited model capacity or from ambiguity intrinsic to human judgment, one cannot tell whether a reported improvement is genuine progress or mere fine-tuning toward saturation. Emotion is subjective, and multiple annotators routinely assign different labels; this human label variation is a natural candidate for the irreducible error. While perspectivism treats such variation as signal (26; 28; 7), ChaosNLI (22) and CIFAR-10H (25) use annotation entropy as an informal ceiling, and soft-label frameworks (3) address it, inter-annotator agreement is not an upper bound on model performance (27). Yet emotion annotation typically provides only about three annotators per instance, and the resulting instability of ceiling estimation under such sparse annotation has not been addressed. Whether saturation can be claimed safely is obstructed by four challenges. First, finite-sample entropy estimation is downward-biased by concavity (17; 23), yet informal ceilings do not correct for it. Second, the ceiling shifts substantially with the estimator and can yield the opposite conclusions of “room remaining” and “saturation” for the same classifier, yet this variation has not been quantified. Third, building the ceiling and the evaluation from the same annotations introduces circularity, so the ceiling overfits to noise and can suggest superhuman performance. Fourth, observed disagreement conflates intrinsic ambiguity with annotation noise, and treating all of it as irreducible overestimates the ceiling and falsely suggests saturation. In this study, Bias-corrected Affective Ceiling Estimation (BACE) is proposed. Its aim is not to estimate the ceiling as a single point, but to quantify how much ceiling estimation depends on finite annotation, estimator choice, noise, and the evaluation protocol, and thereby to discipline how far saturation can be claimed. BACE estimates the human consensus distribution and an information-theoretic ceiling (Estimate), separates classifier error into aleatoric and epistemic components without circularity (Separate), and admits only robust claims under a fixed discipline (Discipline). Through bias-corrected entropy estimation with systematic bracketing, annotator splitting, deconvolution, a claim gate, and predictability maps over emotion, offensiveness, and irony, it is shown that reachability depends heavily on the estimator, so that saturation cannot be asserted on its own, while a fixed fraction of the error nonetheless remains irreducible as a conservative lower bound. This shifts benchmark interpretation from a pursuit of absolute accuracy toward separating, in information-theoretic terms, the reducible room from the irreducible limit. 2 Related Work 2.1 Human Label Variation and Learning from Disagreement Treating annotator disagreement in subjective tasks as signal rather than noise is now established. Plank (26) showed that human label variation is pervasive in NLP and questioned single-gold evaluation, and Uma et al. (28) surveyed learning from disagreement across hard-label aggregation, soft labels, and annotator modeling. Pavlick and Kwiatkowski (24) found disagreement in natural language inference to be intrinsic, and Aroyo and Welty (2) criticized the single-truth myth, noting that for subjective tasks such as emotion, aggregated hard labels discard essential information. Perspectivist work further models annotator individuality: Davani et al. (7) predicted individual labels before aggregation, the LeWiDi shared task (15) standardized soft-label evaluation of disagreement, and reliability estimators such as MACE (13) separated annotator noise from genuine ambiguity. These works preserve disagreement but neither quantify the ceiling it imposes on attainable accuracy nor address the finite-sample bias of estimating entropy from few annotations, which is where we depart by analyzing the estimator dependence of converting disagreement into a predictability ceiling. 2.2 Bayes Error Ceilings and Entropy Estimation The accuracy ceiling has long been studied in information theory. Fano’s inequality (9; 6) links conditional entropy to a lower bound on any predictor’s error, and for a known conditional distribution the tight 00–11 lower bound is the complement of the maximum posterior probability. Recent direct Bayes-error estimation bypasses entropy: Ishida et al. (14) estimated the binary Bayes error from class-uncertainty labels, and Ushio et al. (29) extended this to soft labels and showed that calibration alone is insufficient. Richie et al. (27) showed by simulation that inter-annotator agreement is not an upper bound on model performance, but offered no principled alternative. Estimating entropy from finite samples is itself hard. The plug-in estimator is downward-biased by concavity (17), as characterized by Paninski (23); the NSB estimator (21) is low-bias in undersampled regimes, with further options in coverage adjustment (5), shrinkage (12), and a general Bayesian framework (1), while Wolpert and Wolf (30) and Minka (18) provide closed-form and fixed-point Dirichlet–multinomial estimators. None of these has been applied systematically to the emotion-predictability ceiling, where informal ceilings often use uncorrected plug-in entropy. We instead integrate them into a systematic bracket via a hierarchical empirical Bayes over the cross-instance structure of emotion and, unlike direct Bayes-error estimation for the binary case, take an entropy and Fano path with explicit finite-sample correction for multi-class emotion; the direct estimator coincides with the plug-in lower end and so underestimates the bound (Section 4.3). 2.3 Aleatoric and Epistemic Uncertainty in Affective Computing Decomposing uncertainty into an irreducible aleatoric and a reducible epistemic component is widely adopted, with ChaosNLI (22) in natural language inference and CIFAR-10H (25) in image classification both treating human disagreement as a lower bound on uncertainty. In affective computing, Baan et al. (3) evaluated model calibration against human uncertainty, while datasets provide multiply annotated resources: GoEmotions (8) with annotator identifiers, SemEval emotion tasks (19) and BRIGHTER (20) for multilingual emotion, and MD-Agreement (16), EPIC (10), and MultiPICo (4) for offensiveness and irony. Yet none organizes both systematic and sample uncertainty across emotion benchmarks while blocking circular reasoning and overclaiming, and some use only entropy estimation (21) or disagreement deconvolution (11) in isolation. In contrast, we jointly provide instance-level entropy estimation, bias correction, rigorous Fano and Bayes ceilings, aleatoric/epistemic decomposition, and claim gating. 3 Methodology 3.1 Problem Definition Let xii=1M\x_i\_i=1^M denote the set of instances, Y the label space, and K=||K=|Y| its cardinality. Each instance xix_i is independently labeled by nin_i annotators, and the number of annotators who assign label y is denoted by ciyc_iy, where ∑yciy=ni _yc_iy=n_i. The human-consensus distribution pi(y):=p(y∣xi)p_i(y):=p(y x_i) is operationally defined as the probability that a single annotator drawn at random from population A assigns label y under guideline G and context C. Consequently, pip_i is relative to the four-tuple (A,G,C,)(A,G,C,Y), and this relativity is clarified in the Limitations section. The conditional entropy of instance i and its dataset-level average are given respectively by Hi:=−∑y∈pi(y)log2pi(y)H_i:=- _y p_i(y) _2p_i(y) and H(Y∣X):=1M∑i=1MHiH(Y X):= 1M _i=1^MH_i, where HiH_i is the intrinsic ambiguity of instance xix_i. Because a ceiling becomes meaningful only once the evaluation protocol is fixed, we define three protocols separately. Under the random-annotator 00–11 protocol P1, the ground truth Y is a single sample from pip_i, and the error rate is Pe:=1M∑i=1MPr(y^i≠Yi)P_e:= 1M _i=1^M ( y_i≠ Y_i). Under the majority-vote protocol P2, the ground truth is the majority vote over the realized annotators, itself a random variable under finite samples. Under the soft-label protocol P3, the model distribution qiq_i is scored by its cross-entropy to pip_i. Because conflating protocols produces category inconsistencies, every results table carries a protocol column. When pip_i is known, an exact lower bound on the 00–11 error under P1 is given directly without invoking Fano’s inequality, and ei∗:=1−maxy∈pi(y),C:=1M∑i=1Mei∗e_i^*:=1- _y p_i(y), C:= 1M _i=1^Me_i^* (1) gives the predictability ceiling as the exact Bayes error, where ei∗e_i^* denotes the unreachable error lower bound at xix_i and C its instance average. The classifier error ErrSOTA(P)Err_SOTA(P) is decomposed into the irreducible component C(P)C(P) and the reducible component G(P):=ErrSOTA(P)−C(P)G(P):=Err_SOTA(P)-C(P), and the reachability is R(P):=C(P)ErrSOTA(P)∈(0,1],R(P):= C(P)Err_SOTA(P)∈(0,1], (2) where C(P)C(P) is aleatoric uncertainty, G(P)G(P) is epistemic uncertainty, and R(P)R(P) is the irreducible fraction of the error. The central question is how stably R(P)R(P) can be determined from finite annotations. 3.2 Overview of BACE When pip_i is estimated from few annotations, the plug-in estimator biases the entropy downward owing to concavity, reaching the order of 11 bit for ni=3n_i=3 with an effective label count of 55 to 1010, so an uncorrected point estimate cannot serve as the primary estimator. BACE is organized as three core layers, Estimate, Separate, and Discipline, followed by a cross-task application (Figure 1), and it handles this bias systematically through four technical components detailed below: anchored Dirichlet-mixture empirical Bayes; instance-level Fano and exact Bayes ceilings; aleatoric/epistemic decomposition based on annotator splits; and deconvolution of annotation noise together with a claim gate. Together these prevent circular reasoning and overclaiming while producing a cross-task predictability map. By the relativity of the four-tuple in Section 3.1, these ceilings are predictability with respect to a specific annotator population, observation channel, and taxonomy, not claims about the unknowability of emotion itself. Figure 1: Overview of BACE: three core layers (Estimate, Separate, Discipline) followed by a cross-task application. 3.3 Anchored Dirichlet-Mixture Empirical Bayes The goal of this estimator is not entropy estimation per se, but the stable estimation of the downstream predictability ceiling and reachability. In what follows, the bias of entropy estimation is therefore controlled as a systematic error source that can distort the ceiling and reachability. The plug-in entropy carries a systematic bias [H^plug]−H≈−(K∗−1)/(2nln2)E[ H^plug]-H≈-(K^*-1)/(2n 2) (17; 23), where K∗K^* denotes the effective number of labels. Underestimating entropy inflates the epistemic gap, whereas overestimating it overclaims that the classifier has reached the ceiling, so these two dangerous directions point opposite to each other. Accordingly, every principal quantity is bracketed between the conservative plug-in lower side and the NSB (21) upper side rather than reported as a single point estimate. The main estimator is an empirical Bayes that shares the inter-instance structure of the emotion label distribution as a hierarchical prior, transferring the base rate and confusion-pair structure learned from all instances to low-n instances for which entropy is unrecoverable at ni=3n_i=3. A single asymmetric Dirichlet prior cannot express a mixture of near-unanimous and split instances and, on real data, exceeds the NSB upper end at low granularity, so a finite Dirichlet-mixture prior pi p_i ∼∑s=1SwsDir(τsms), _s=1^Sw_s\,Dir( _sm_s), (3) ci c_i ∼Multinomial(ni,pi) (n_i,p_i) is adopted, where wsw_s is the component weight, ms∈ΔK−1m_s∈ ^K-1 is the base measure of component s, and τs>0 _s>0 is its concentration. Assignment uses the responsibility ris∝wsDM(ci∣τsms)r_is w_s\,DM(c_i _sm_s), where DMDM denotes the Dirichlet–multinomial marginal likelihood, so leakage from double use of the observed counts does not arise structurally. Closed forms for the posterior entropy and maxypiy _yp_iy, the Minka-type fitting of (τs,ms)( _s,m_s) (30; 18), category-wise anchoring, and the holdout selection of S are given in Appendix D. The ordering plug-in≤EB≤NSBplug -in is a diagnostic expectation rather than a theorem; when a reversal persists even after correction, the empirical Bayes point is demoted to a reference value and claims are made only at both ends of the systematic bracket. 3.4 Fano and Exact Bayes Ceilings For a deterministic predictor y^=g(x) y=g(x) forming the Markov chain Y–X–Y Y, Fano’s inequality (9; 6) gives, per instance, Hi≤hb(ei)+eilog2(K−1),ei≥fK(Hi),H_i≤ h_b(e_i)+e_i _2(K-1), e_i≥ f_K(H_i), (4) where hbh_b is the binary entropy, ei:=Pr(y^i≠Yi∣xi)e_i:= ( y_i≠ Y_i x_i) is the instance error, and fK:=φK−1f_K:= _K^-1 is the convex increasing inverse of φK(t):=hb(t)+tlog2(K−1) _K(t):=h_b(t)+t _2(K-1). The Fano lower bound is always looser than the exact Bayes error, so the main ceiling for P1 is taken to be the exact Bayes error of Eq. (1), and the Fano-type lower bound is reported alongside it in three complementary roles. These include unifying 00–11 error and soft-label evaluation through [CE]≥H(Y∣X)E[CE]≥ H(Y X) (Appendix E). For multi-label emotion, the joint distribution has 2K2^K configurations and cannot be estimated from ni≤5n_i≤ 5, so only the marginal quantities required by the evaluation metrics are estimated. Under the binary decomposition (Yk=1[k∈S]Y_k=1[k∈ S], pik:=Pr(Yk=1∣xi)p_ik:= (Y_k=1 x_i)), log2(K−1)=0 _2(K-1)=0, so Fano degenerates to the entropy condition hb(e)≥Hh_b(e)≥ H, the lower inverse becomes exactly eik≥min(pik,1−pik)e_ik≥ (p_ik,1-p_ik), and the entropy-based lower bound coincides with the exact Bayes error, so the per-emotion ceiling is exact. Its instance-wise lower bound is obtained in closed form through a per-label Beta–Binomial empirical Bayes prior and the regularized incomplete beta function (Appendix E). F1 is not decomposable and its Bayes-optimal value depends on the joint distribution, so what is reported is the oracle plug-in F1 obtained by optimal thresholding of the marginal probabilities (31), which is a marginal reference value rather than a proof of the upper bound. 3.5 Aleatoric/Epistemic Decomposition via Annotator Splitting Constructing both the ceiling and the evaluation from the same annotations introduces the circular reasoning that the ceiling overfits to annotation noise. To prevent this, the annotators of each instance are split into two disjoint groups. That is, the analytic ceiling and the oracle predictor are constructed from p^iA p_i^A on the estimation split A, and the classifier under evaluation and the oracle are scored on the evaluation split B, where |Ai|=⌈ni/2⌉|A_i|= n_i/2 , |Bi|=⌊ni/2⌋|B_i|= n_i/2 , and an odd extra annotator is assigned to A in order to prioritize the estimation bottleneck. This split is made deterministic by a fixed-seed permutation over the lexicographically sorted record table (Appendix F). The reachability R(P)R(P) of Eq. (2) is taken as the principal reported quantity, and its irreducible component C(P)C(P) is a lower bound that applies to any model predicting labels of the same input channel and annotator population. For protocol consistency, the ceiling is computed under the same protocol as the classifier. As a diagnostic, the oracle constructed from A and scored on B agrees with the analytic ceiling, supporting the estimation. When the classifier falls below the ceiling beyond the confidence interval (G^<0 G<0), this is treated, by a fixed rule, as an alarm for underestimation of the ceiling, leakage of the test annotations, or protocol mismatch, mechanically ruling out the erroneous conclusion of superhuman performance. 3.6 Annotation-Noise Deconvolution and the Claim Gate Because the observed disagreement mixes true ambiguity and annotation error, the entropy of the raw p p overestimates ambiguity. Following the perspectivist standpoint (26; 28), we do not remove it completely but fit a lightweight model in which annotator a reports the signal with probability 1−εa1- _a and a granularity-matched noise ν with probability εa _a, Pr(yt=y)=(1−εat)p~it(y)+εatν(y), (y_t=y)=(1- _a_t)\, p_i_t(y)+ _a_t\,ν(y), (5) by EM (11). Here ν is restricted to uniform and base-rate types, with no confusion matrix that would absorb true ambiguity, and εa _a (shared across annotators) and p~i p_i (shared across instances) are identified from the crossed structure. Monotonicity is not assumed: denoising lowers the ceiling for plug-in but, for mixture empirical Bayes and NSB, can raise it above the raw data through the interaction with finite-sample correction or hierarchical smoothing, so both conditions are reported and the ordering depends on the estimator. Only the post-deconvolution plug-in serves as the conservative lower bound, where denoising and the plug-in downward bias compound to the smallest reachability, while the post-deconvolution mixture empirical Bayes and NSB are a sensitivity analysis. The EM updates, the reuse of ε^a _a across granularities, and the recovery check are in Appendix H. The claim gate disciplines claims a priori. A reachability claim that the classifier attains X% of the ceiling is asserted only if it holds at the conservative ends of both the systematic bracket and the 9595% confidence interval, at both granularities, and in particular at the post-deconvolution plug-in end, where the ceiling is smallest and reachability hardest to attain; otherwise it is reported as an interval without a point claim. The gate thereby secures, at the methodological level, the central claim that an informal human ceiling moves substantially under the choice of estimator. 4 Experiments 4.1 Datasets and Experiment Design Our primary foundation is GoEmotions (8), which fully releases annotator-identified raw annotations over 2727 emotions plus neutral, letting us relate the granularity of the emotion space to the ceiling. For robustness we also use the binary offensiveness task MD-Agreement (16), the binary irony task EPIC (10), and the English portion of BRIGHTER (20) (statistics in Appendix A). GoEmotions has only 3.413.41 annotators per instance and base rates spanning two orders of magnitude, from 0.2620.262 for neutral to 0.0030.003 for grief, so the per-instance entropy is noisy and a hierarchical empirical Bayes bias correction is indispensable. The evaluated classifier is the representative public SamLowe/roberta-base-go_emotions, which we do not claim to be the single best model; bhadresh-savani/bert-base-go-emotion and monologg/bert-base-cased-goemotions-original are also evaluated in Section 4.5 against the same ceiling and instances. The ceiling estimators plug-in, Miller–Madow (17), NSB (21), and the proposed mixture empirical Bayes share a common implementation, and the direct and soft-label Bayes-error estimators (14; 29) and bias-corrected entropy estimators (5; 12) serve as external baselines on synthetic distributions with known entropy (Appendix B, I). We report the exact Bayes error for P1, lower bounds for Hamming and macro/micro-F1, the cross-entropy lower bound for P3, and the reachability R, with bootstrap intervals separated from the systematic estimator bracket. To prevent leakage, the decomposition uses only the official GoEmotions test instances and the prior is fitted on the out-of-test A-side annotations. 4.2 The Need for a Bias-Corrected Bracket Table 1 shows the GoEmotions coarse-granularity ceilings under plug-in, mixture empirical Bayes, and NSB, together with the deconvolved plug-in end and the A/B-track decomposition. At the L2 Ekman 77-class level, the systematic bracket spans [0.258,0.391][0.258,0.391], confirming on real data that a single point estimate swings the result by roughly 1313 points in error rate. Mixture empirical Bayes recovers the ordering plug-in≤EB≤NSBplug -in at L2, but a reversal remains at L1, so the empirical Bayes point is demoted to a reference value (Section 3.3). The deconvolved plug-in end is 0.1970.197 at L2 and 0.1790.179 at L1, providing the conservative end for reachability claims. As the claim gate requires, the all-annotation track and the A/B tracks agree within the systematic range. The differences among these estimators are not mere numerical fluctuation: since the model error is fixed, where the ceiling is placed directly determines whether the same classifier appears far from the ceiling or already at it. Table 1: Predictability ceilings and aleatoric/epistemic decomposition for GoEmotions (P1 exact Bayes error). Gran. plug-in NSB dec.×plug ErrSOTAErr_SOTA ceiling gap R cons. [CI] L2 (K=7K=7) 0.258 0.391 0.197 0.388 0.150 0.238 0.386 [.368,.403] L1 (K=4K=4) 0.227 0.328 0.179 0.350 0.134 0.216 0.383 [.366,.402] 4.3 Exact Lower Bounds from Binary Decomposition and Per-Label Predictability Under the binary decomposition, the sample deficiency vanishes for n≥2n≥ 2, the lower bound becomes an exact Bayes error with zero slack, and the Hamming, macro, and micro lower bounds coincide at 0.03690.0369. The existing direct and soft-label estimators (14; 29) reduce to the lower plug-in end and underestimate this bound (Appendix B). The per-label irreducible-error lower bound ek∗e^*_k ranges from 0.00330.0033 for grief to 0.2180.218 for the frequent and ambiguous neutral, giving an independent map of which emotions are in principle hard to predict. The bounds for all 2828 labels and the oracle plug-in F1 are given in Appendix J. 4.4 Estimator Dependence of Reachability and a Robust Lower Bound Having established that the ceiling itself depends on the estimator, we verify whether this dependence affects the substantive conclusion of saturation. Table 1 decomposes the error of the evaluated classifier into irreducible and reducible components. At the default smoothing β=0.7β=0.7, the conservative reachability is about 0.380.38 at both granularities (L2 0.3860.386, L1 0.3830.383), and its minimum across β∈0.5,0.7,0.9β∈\0.5,0.7,0.9\ and split seeds falls to about 0.340.34 (0.33750.3375 at L1 with β=0.5β=0.5; Appendix F). The only statement that passes the claim gate is therefore that at least about 3333% of the P1 error on GoEmotions is an irreducible aleatoric component, and this is a conservative lower bound rather than a point estimate. Because the number of observed annotators is small, the denoising effect and the finite-sample bias correction cannot be fully separated, so this lower bound is interpreted on the conservative side (Appendix H). Figure 2 visualizes how the reachability R depends on the choice of estimator and deconvolution. For the same evaluation error, the reachability swings from 0.380.38 at the conservative deconvolved plug-in end, through about 0.520.52 at the raw plug-in end, to between 0.900.90 and 1.031.03 at the mixture empirical Bayes and NSB ends. The effect of deconvolution is not one-directional: it lowers the reachability for plug-in but raises it for mixture empirical Bayes and NSB, so the direction of change differs by estimator. It is thus the core of this work that the verdict of saturation holds only at the empirical Bayes and NSB ends and fails at the plug-in end. The binary F1 and the soft-label P3 likewise span a wide range (Appendix J). The diagnostic quantity G G is positive at the conservative ends and turns only slightly negative at the L2 upper-end NSB (raw counts) and the deconvolved mixture empirical Bayes and NSB, which is an alarm rather than superhuman performance. Figure 2: The reachability R depends on the combination of estimator, noise handling, and finite-sample correction, and the direction of its change is not constant. 4.5 Dependence on the Evaluated Model (Multi-Model) Table 2 reports the reachability of four fine-tuned classifiers (one of which is ModernBERT-large) and one GPT-4o-mini few-shot classifier, decomposed against the same L2 ceiling, and shows that the verdict of saturation also changes with the classifier placed in the denominator. The raw empirical Bayes reachability spans [0.732,1.009][0.732,1.009] and crosses 11 through the denominator alone, so a larger backbone does not break the ceiling and no single value of R decides the question of saturation. The details of the multi-model decomposition are given in Appendix C. Table 2: Multi-model decomposition against the same GoEmotions L2 ceiling (P1). Classifier ErrErr RconsR_cons [CI] REBR_EB [CI] bhadresh-BERT 0.382 0.392 [.375,.411] 1.009 [.977,1.043] ModernBERT-large 0.386 0.388 [.371,.405] 0.997 [.967,1.028] SamLowe-RoBERTa 0.388 0.386 [.368,.403] 0.992 [.964,1.023] monologg-BERT 0.419 0.357 [.342,.372] 0.918 [.891,.946] GPT-4o-mini (fs) 0.526 0.285 [.273,.297] 0.732 [.714,.750] 4.6 Cross-Task Predictability Map Figure 3 shows the cross-task predictability map across emotion, offensiveness, and irony, controlling for base rate as cross-task comparison requires. At a base rate of about 0.30.3, the irreducible ambiguity is largest for irony at [0.206,0.261][0.206,0.261], exceeding that of offensiveness and emotion presence. The lower bound also rises monotonically with the number of classes. The deconvolved conservative ends are 0.1390.139 for irony and 0.1370.137 for offensiveness, and the mean noise rate is largest for irony at 0.2570.257 and 0.1510.151 for offensiveness. The details of deconvolution and BRIGHTER are given in Appendix J. Table 3 decomposes an off-the-shelf public classifier on each corpus against its conservative ceiling and shows that two core findings recur across tasks. Namely, the conservative irreducible component remains positive for offensiveness, irony, and emotion, and the reachability again depends strongly on the estimator. Hence the positive C and this estimator dependence are task-general rather than an artifact of GoEmotions. However, because these classifiers are not fine-tuned on the target corpora, these reachabilities are lower bounds. The details are given in Appendix K. Figure 3: Cross-task predictability map controlling for base rate. Table 3: Cross-task SOTA decomposition against the conservative ceiling. Task ErrSOTAErr_SOTA CconsC_cons RconsR_cons [CI] REBR_EB RNSBR_NSB Offensiveness (MD-Ag.) 0.332 0.137 0.413 [.404,.423] 0.640 0.618 Irony (EPIC) 0.388 0.139 0.359 [.342,.375] 0.672 0.603 Emotion (BRIGHTER) 0.275 0.116 0.422 [.416,.429] 0.482 0.543 5 Discussion 5.1 Reachability Is Estimator-Dependent The central finding is that the informal ceiling depends very strongly on the estimator, so a plateau judgment cannot be entrusted to a single point estimate. For the same evaluation error, reachability swings from 0.380.38 to 1.031.03, and even with the ceiling and estimator fixed it flips across 11 depending on which strong public classifier is the denominator (Section 4.5). Hence the only assertion passing the claim gate is that at least about 3333% of the GoEmotions evaluation error is irreducible, a lower bound the systematic bracket states explicitly as an interval, and this estimator dependence recurs on offensiveness, irony, and emotion (Appendix K). 5.2 What Is Irreducible in Emotion The per-label lower bounds differ greatly, ranging from below 0.010.01 for rare, high-agreement emotions to 0.2180.218 for neutral (Section 4.3). This is internal structure that average accuracy conceals. Once base rates are controlled for, irony carries the largest irreducible ambiguity; however, the lower bounds for irony and offensiveness nearly converge after deconvolution, so much of the raw disagreement is annotation noise, and treating disagreement directly as irreducible ambiguity overestimates the ceiling. Predictability is thus structured by label, task, and granularity, and because the lower bound rises with the number of classes, the design of the taxonomy itself governs it. 5.3 Implications for Benchmarking and Annotation Budgets Because reachability depends on the estimator, a leaderboard’s residual error is not uniformly a reducible epistemic gap, and since at least about 3333% of the GoEmotions error is irreducible, the room for improvement is often overestimated. Benchmark reports should therefore position achieved accuracy relative to an irreducible lower bound with a systematic bracket. Since the bracket width depends strongly on the number of annotators, about 0.050.05 for BRIGHTER and 0.100.10 for the sparse L1 granularity of GoEmotions, designers can set annotators per instance from a target precision, and released identifiers let deconvolution separate genuine ambiguity from annotation noise. 6 Conclusion We addressed with BACE the accuracy race in emotion recognition that proceeds while the predictability ceiling remains unknown. Bracketing an anchored Dirichlet-mixture empirical Bayes between the plug-in and NSB, and decomposing error through an annotator split, a deconvolution, and a fixed claim gate, BACE estimates a bias-corrected ceiling from human label variation and blocks circular reasoning and over-claiming. Reachability swings from 0.380.38 to 1.031.03 with the estimator and the evaluated classifier, so a plateau verdict is impossible without estimator discipline. The only statement passing the claim gate is that at least about 3333% of the evaluation error on GoEmotions is irreducible, a pattern that recurs on offensiveness and irony. In future work, we will extend the framework to more languages and taxonomies. Limitations Our ceilings are relative to a specific annotator population, observation channel, and taxonomy: a model with extra-textual information is not a counterexample to Fano, and the ceiling concerns this pool’s label distribution, not the unknowability of emotion. In the L1 sentiment setting the estimator ordering reversal persists after correction, so the mixture empirical Bayes point is used only as a reference and claims are made at the two ends of the systematic bracket. At the observed annotator counts, deconvolution denoising cannot be separated from finite-sample bias, so the deconvolved plug-in end serves only as a conservative lower bound and the reported 3333% is specific to GoEmotions; relatedly, G^<0 G<0 occurs only at the upper-end L2 ceilings and signals overestimation or leakage, not superhuman performance. Annotator-level overlap may remain because classifiers can see the same annotators’ labels in training, and BRIGHTER carries no annotator identifiers, so only its raw-data ceiling is reported. Finally, a single LLM (GPT-4o-mini, one few-shot configuration) was evaluated, so the capability band is not exhaustive. All numbers are reported as they are. Ethical Considerations Because this study addresses emotion labels that are subjective and can be contested, several ethical considerations are made explicit. First, our framework treats disagreement among annotators as signal rather than noise, and avoids collapsing perspectival differences into a single majority-vote label. Second, the estimated ceiling is a quantity relative to a specific annotator population, annotation guideline, observation channel, and taxonomy, and is not a claim about any individual’s “true emotion.” Using our ceiling to determine the emotional state of a particular person is therefore neither intended nor appropriate. Regarding data, only publicly released datasets were used, in accordance with their respective licenses. GoEmotions is released under Apache-2.0, MD-Agreement and EPIC under non-commercial licenses, and BRIGHTER is restricted to research use. Our code is available at https://github.com/keito-git/affectceiling-bace; the datasets are not bundled with it, and retrieval scripts are distributed instead. No new human-subject data collection was conducted, no personally identifying information was added, and no annotator was re-identified. A foreseeable misuse is to interpret a predictability ceiling as a fundamental limit on the ability to read an individual’s emotions. To guard against this, the ceiling is presented at the level of datasets and populations and as a conservative lower bound. This conservative framing is intended to curb over-claiming in affective computing and to encourage an evaluation culture that honestly reports where achieved accuracy stands relative to an irreducible floor. References Archer et al. (2014) E. Archer, I. M. Park, and J. W. Pillow Bayesian entropy estimation for countable discrete distributions. The Journal of Machine Learning Research 15 (1), p. 2833–2868. Cited by: §2.2. Aroyo and Welty (2015) L. Aroyo and C. Welty Truth is a lie: crowd truth and the seven myths of human annotation. AI Magazine 36 (1), p. 15–24. External Links: Document Cited by: §2.1. Baan et al. (2022) J. Baan, W. Aziz, B. Plank, and R. Fernández Stop measuring calibration when humans disagree. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, p. 1892–1915. External Links: Document Cited by: §1, §2.3. Casola et al. (2024) S. Casola, S. Frenda, S. M. Lo, E. Sezerer, A. Uva, V. Basile, C. Bosco, A. Pedrani, C. Rubagotti, V. Patti, and D. Bernardi MultiPICo: multilingual perspectivist irony corpus. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, p. 16008–16021. External Links: Document Cited by: §2.3. Chao and Shen (2003) A. Chao and T. Shen Nonparametric estimation of shannon’s index of diversity when there are unseen species in sample. Environmental and Ecological Statistics 10, p. 429–443. External Links: Document Cited by: Table 5, Appendix B, §2.2, §4.1. Cover and Thomas (2006) T. M. Cover and J. A. Thomas Elements of information theory. Wiley-Interscience. External Links: Document Cited by: §2.2, §3.4. Davani et al. (2022) A. M. Davani, M. Díaz, and V. Prabhakaran Dealing with disagreements: looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics 10, p. 92–110. External Links: Document Cited by: §1, §2.1. Demszky et al. (2020) D. Demszky, D. Movshovitz-Attias, J. Ko, A. Cowen, G. Nemade, and S. Ravi GoEmotions: a dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 4040–4054. External Links: Document Cited by: §2.3, §4.1. Fano (1961) R. M. Fano Transmission of information: a statistical theory of communications. The MIT Press, Cambridge, MA. Cited by: §2.2, §3.4. Frenda et al. (2023) S. Frenda, A. Pedrani, V. Basile, S. M. Lo, A. T. Cignarella, R. Panizzon, C. Marco, B. Scarlini, V. Patti, C. Bosco, and D. Bernardi EPIC: multi-perspective annotation of a corpus of irony. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, p. 13844–13857. External Links: Document Cited by: §2.3, §4.1. Gordon et al. (2021) M. L. Gordon, K. Zhou, K. Patel, T. Hashimoto, and M. S. Bernstein The disagreement deconvolution: bringing machine learning performance metrics in line with reality. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, p. 1–14. External Links: Document Cited by: Appendix H, §2.3, §3.6. Hausser and Strimmer (2009) J. Hausser and K. Strimmer Entropy inference and the James-Stein estimator, with application to nonlinear gene association networks. Journal of Machine Learning Research 10, p. 1469–1484. Cited by: Table 5, Appendix B, §2.2, §4.1. Hovy et al. (2013) D. Hovy, T. Berg-Kirkpatrick, A. Vaswani, and E. Hovy Learning whom to trust with MACE. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 1120–1130. Cited by: §2.1. Ishida et al. (2023) T. Ishida, I. Yamane, N. Charoenphakdee, G. Niu, and M. Sugiyama Is the performance of my deep network too good to be true? A direct approach to estimating the Bayes error in binary classification. In The Eleventh International Conference on Learning Representations, External Links: Document Cited by: Table 5, Appendix B, §2.2, §4.1, §4.3. Leonardelli et al. (2023) E. Leonardelli, G. Abercrombie, D. Almanea, V. Basile, T. Fornaciari, B. Plank, V. Rieser, A. Uma, and M. Poesio SemEval-2023 task 11: learning with disagreements (LeWiDi). In Proceedings of the 17th International Workshop on Semantic Evaluation, p. 2304–2318. External Links: Document Cited by: §2.1. Leonardelli et al. (2021) E. Leonardelli, S. Menini, A. P. Aprosio, M. Guerini, and S. Tonelli Agreeing to disagree: annotating offensive language datasets with annotators’ disagreement. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 10528–10539. External Links: Document Cited by: §2.3, §4.1. Miller (1955) G. A. Miller Note on the bias of information estimates. In Information Theory in Psychology Problems and Methods, Vol. 2, p. 95–100. Cited by: §1, §2.2, §3.3, §4.1. Minka (2000) T. P. Minka Estimating a dirichlet distribution. Technical report Microsoft Research. Cited by: Appendix D, §2.2, §3.3. Mohammad et al. (2018) S. Mohammad, F. Bravo-Marquez, M. Salameh, and S. Kiritchenko SemEval-2018 task 1: affect in tweets. In Proceedings of the 12th International Workshop on Semantic Evaluation, p. 1–17. External Links: Document Cited by: §2.3. Muhammad et al. (2025) S. H. Muhammad, N. Ousidhoum, I. Abdulmumin, J. P. Wahle, T. Ruas, M. Beloucif, C. de Kock, N. Surange, D. Teodorescu, I. S. Ahmad, D. I. Adelani, A. F. Aji, F. D. M. A. Ali, I. Alimova, V. Araujo, N. Babakov, N. Baes, A. Bucur, A. Bukula, G. Cao, R. T. Cardenas, R. Chevi, C. I. Chukwuneke, A. Ciobotaru, D. Dementieva, M. S. Gadanya, R. Geislinger, B. Gipp, O. Hourrane, O. Ignat, F. I. Lawan, R. Mabuya, R. Mahendra, V. Marivate, A. Panchenko, A. Piper, C. H. P. Ferreira, V. Protasov, S. Rutunda, M. Shrivastava, A. C. Udrea, L. D. A. Wanzare, S. Wu, F. V. Wunderlich, H. M. Zhafran, T. Zhang, Y. Zhou, and S. M. Mohammad BRIGHTER: BRIdging the gap in human-annotated textual emotion recognition datasets for 28 languages. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, p. 8895–8916. External Links: Document Cited by: §2.3, §4.1. Nemenman et al. (2002) I. Nemenman, F. Shafee, and W. Bialek Entropy and inference, revisited. In Proceedings of the 15th Neural Information Processing Systems, p. 471–478. External Links: Document Cited by: §2.2, §2.3, §3.3, §4.1. Nie et al. (2020) Y. Nie, X. Zhou, and M. Bansal What can we learn from collective human opinions on natural language inference data?. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, p. 9131–9143. External Links: Document Cited by: §1, §2.3. Paninski (2003) L. Paninski Estimation of entropy and mutual information. Neural Computation 15 (6), p. 1191–1253. External Links: Document Cited by: §1, §2.2, §3.3. Pavlick and Kwiatkowski (2019) E. Pavlick and T. Kwiatkowski Inherent disagreements in human textual inferences. Transactions of the Association for Computational Linguistics 7, p. 677–694. External Links: Document Cited by: §2.1. Peterson et al. (2019) J. C. Peterson, R. M. Battleday, T. L. Griffiths, and O. Russakovsky Human uncertainty makes classification more robust. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 9616–9625. External Links: Document Cited by: §1, §2.3. Plank (2022) B. Plank The “problem” of human label variation: on ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, p. 10671–10682. External Links: Document Cited by: §1, §2.1, §3.6. Richie et al. (2022) R. Richie, S. Grover, and F. (. Tsui Inter-annotator agreement is not the ceiling of machine learning performance: evidence from a comprehensive set of simulations. In Proceedings of the 21st Workshop on Biomedical Language Processing, p. 275–284. External Links: Document Cited by: §1, §2.2. Uma et al. (2021) A. N. Uma, T. Fornaciari, D. Hovy, S. Paun, B. Plank, and M. Poesio Learning from disagreement: a survey. Journal of Artificial Intelligence Research 72. External Links: Document Cited by: §1, §2.1, §3.6. Ushio et al. (2026) R. Ushio, T. Ishida, and M. Sugiyama Practical estimation of the optimal classification error with soft labels and calibration. In The Fourteenth International Conference on Learning Representations, External Links: Document Cited by: Table 5, Appendix B, §2.2, §4.1, §4.3. Wolpert and Wolf (1995) D. H. Wolpert and D. R. Wolf Estimating functions of probability distributions from a finite set of samples. Physical Review E 52, p. 6841. External Links: Document Cited by: Appendix D, §2.2, §3.3. Ye et al. (2012) N. Ye, K. M. A. Chai, W. S. Lee, and H. L. Chieu Optimizing F-measure: a tale of two approaches. In Proceedings of the 29th International Conference on Machine Learning, p. 1555–1562. Cited by: §3.4. Appendix A Dataset Statistics Table 4 summarizes the statistics of the datasets used in Section 4.1. Table 4: Dataset statistics. Dataset Inst. Ann. nin_i (mean/med.) Labels GoEmotions 57,828 82 3.41 / 3 multi-label 27+neutral MD-Agreement 10,753 819 5 / 5 binary offensiveness EPIC 3,000 74 4.72 / 5 binary irony BRIGHTER eng 5,655 no ID 8.52 / 8 6 emotions, ordinal Appendix B Comparison with Existing Bayes-Error Estimators Table 5 relates the proposed bracket to the existing Bayes-error and entropy estimators of Section 4.3. For the binary GoEmotions floor, which is the macro floor and equals the Hamming and cell micro floors, the direct estimator of Ishida et al. (14) and the simplified soft-label isotonic estimator of Ushio et al. (29), whose calibration is nearly the identity map on empirical soft labels, both agree with the plug-in lower end of 0.02740.0274. In contrast, the bias-corrected Beta–Binomial point rises to 0.03690.0369 and the NSB upper end rises to 0.1120.112, so the direct estimator inherits the finite-sample downward bias and underestimates the floor. On the categorical side, the mixture empirical Bayes and NSB ceilings lie inside the range spanned by the coverage-adjusted estimator of Chao and Shen (5) and the shrinkage estimator of Häusser and Strimmer (12), where these informal ceilings are read off through the Fano inverse and the shrinkage estimator lies closest to NSB. Table 5: Comparison with existing Bayes-error and entropy estimators. Block Estimator Value Binary floor Ishida direct (14) 0.0274 Binary floor Ushio soft-label (29) 0.0274 Binary floor BACE Beta-Binomial EB (point) 0.0369 Binary floor BACE NSB K=2K=2 (upper) 0.1120 L2 H / ceil. plug-in (informal) 0.713 / 0.259 L2 H / ceil. Chao–Shen (5) 1.062 / 0.205 L2 H / ceil. Häusser shrink. (12) 1.547 / 0.385 L2 H / ceil. BACE mixture EB 1.327 / 0.379 L2 H / ceil. BACE NSB (upper) 1.393 / 0.390 Appendix C Multi-Model Reachability The multi-model decomposition of Section 4.5 (Table 2) decomposes four fine-tuned public classifiers spanning three base-size transformers and one ModernBERT-large, together with a GPT-4o-mini few-shot classifier belonging to a different capability class, against the same GoEmotions ceiling and the same 5,3305,330 evaluation instances, and Figure 4 visualizes their reachability. At the conservative end all five classifiers fall below 11, yet at the raw empirical Bayes end the saturation verdict flips on the classifier alone. Specifically, the interval of bhadresh-BERT straddles 11 whereas the interval of monologg-BERT lies entirely below 11, and this holds even with the ceiling and the estimator held fixed. The larger ModernBERT-large backbone does not escape this cluster either. It attains essentially the same error as the base sizes (0.3860.386), a raw empirical Bayes reachability of 0.9970.997 and an NSB reachability of 1.0091.009, placing it inside the same R≈1R≈ 1 saturation band as the base-size transformers. Enlarging the backbone therefore does not lower the error on the collapsed categories, and the base-size fine-tuned classifiers already sit at the reachability ceiling. Whether a stronger model breaks the ceiling is thus answered negatively within this range. The GPT-4o-mini classifier (openai/gpt-4o-mini, via OpenRouter, temperature 00, an 88-shot few-shot prompt with JSON output, no parsing failures across all 5,3305,330 instances) is the weakest on the 2828-class task and hence has the largest error, so under the same ceiling it attains the smallest reachability, pushing the raw empirical Bayes band down to 0.7320.732 and the conservative band down to 0.2850.285. Reachability therefore spans a broad capability range from fine-tuned transformers to LLM few-shot, and no single value of R decides saturation. Figure 4: Multi-model reachability at L2. Appendix D Mixture Empirical Bayes: Closed Form and Fitting The instance-level posterior quantities of the anchored Dirichlet mixture empirical Bayes of Section 3.3 are obtained as a responsibility-weighted sum of per-component closed forms (30). In particular, the posterior entropy is given by [Hi∣ci]=∑s=1SrisDir(ci+τsms)[H],E[H_i c_i]= _s=1^Sr_is\,E_Dir(c_i+ _sm_s)[H], (6) where each term on the right-hand side denotes the expected entropy under the posterior Dirichlet of component s, evaluated by the Wolpert–Wolf closed form. The expectation of maxypiy _yp_iy entering the exact Bayes error is evaluated by allocating Monte Carlo samples in proportion to the responsibilities ris∝wsDM(ci∣τsms)r_is w_s\,DM(c_i _sm_s). The hyperparameters (ws,τs,ms)(w_s, _s,m_s) are fitted by marginal-likelihood maximization, and the concentration and base measure of each component are obtained by Minka-style fixed-point iteration (18). To avoid divergence on nearly binary labels for which the shared Minka update drives τs→∞ _s→∞, a locally robust fit is used with an upper bound τs≤104 _s≤ 10^4 and a non-finite guard. Model selection over the number of components starts from K components anchored to each category and one global splitting component, and S is adopted only when it improves the mean holdout log marginal likelihood over S=1S=1. On a synthetic population with peaked and diffuse modes, this holdout criterion improves the log marginal likelihood by +0.111+0.111 over S=1S=1, and on GoEmotions the selected value is S=8S=8 at L2 and S=5S=5 at L1, both equal to K+1K+1. Figure 5 visualizes the ceiling on the irreducible error of GoEmotions under the three estimators and shows that it is the estimator, not the data, that sets the target. Figure 5: Ceiling on the irreducible error of GoEmotions under the plug-in, mixture empirical Bayes, and NSB estimators. Appendix E Fano Bound: Three Roles and the Per-Instance Limit At the dataset level, the per-instance Fano inequality of Eq. (4) gives Pe≥(1/M)∑ifK(Hi)≥fK(H(Y∣X))P_e≥(1/M) _if_K(H_i)≥ f_K(H(Y X)), where the second inequality follows from Jensen’s inequality applied to the convex inverse fKf_K, so that applying the lower bound per instance before averaging is tighter than a single application to the aggregated entropy. The exact Bayes error of Eq. (1) is the primary ceiling for P1, but the looser Fano-type lower bound is reported alongside it in three complementary roles. First, it makes explicit a model-agnostic guarantee, namely a theoretical skeleton that lower-bounds any future model. Second, against the upward plug-in bias of maxyp _yp, it provides an independent estimation path fK(H^i)f_K( H_i) via entropy, used to diagnose the discrepancy between the two paths. Third, through the exact lower bound on log loss [CE]≥H(Y∣X)E[CE]≥ H(Y X), it unifies with P3, since the same conditional entropy ties 00–11 error and soft-label evaluation into a single quantity. In the binary decomposition of Section 3.4, Fano degenerates to eik≥min(pik,1−pik)e_ik≥ (p_ik,1-p_ik) and the entropy-based lower bound equals the exact Bayes error. The instance-level floor for binary label k is then given in closed form by [min(p,1−p)∣c]=μI1/2(a′+1,b′)+(1−μ)(1−I1/2(a′,b′+1)),E[ (p,1-p) c]=\\ μ\,I_1/2(a +1,b )+(1-μ) (1-I_1/2(a ,b +1) ), (7) where I is the regularized incomplete beta function, a′=cik+aka =c_ik+a_k, b′=ni−cik+bkb =n_i-c_ik+b_k, and μ=a′/(a′+b′)μ=a /(a +b ), and where (ak,bk)(a_k,b_k) denotes the per-label Beta–Binomial empirical Bayes prior. Appendix F Annotator-Split and Bootstrap Details The A/B split of Section 3.5 is made deterministic as follows. The annotation records are sorted lexicographically by instance identifier and annotator identifier, the annotators of each instance are permuted with a single fixed-seed random number generator, and the resulting split table is stored so that all subsequent experiments refer only to this table. The frozen split uses seed 2026072320260723. Its structural A ratio is 0.6130.613, which arises from assigning the odd extra annotator to A, and the per-annotator A ratio has mean 0.5930.593 and standard deviation 0.0780.078. Confidence intervals are obtained by instance-resampling bootstrap with the prior fit and deconvolution held fixed, because fully refitting the mixture empirical Bayes and deconvolution across a thousand resamples is impractical. The uncertainty arising from the choice of estimator is therefore carried separately by the systematic bracket rather than by the bootstrap, and the two sources are reported side by side, so that sampling uncertainty and estimator-selection uncertainty are never conflated. The calibration of these intervals is verified by a coverage simulation on synthetic data with a known ground truth. As shown in Appendix G, the plug-in and NSB ends respectively under-cover and over-cover by design, whereas the bias-corrected mixture empirical Bayes gives the best coverage and the systematic bracket captures the true irreducible error on every dataset. The robustness of the conservative reachability to these choices is quantified directly. When an outer bootstrap of 200200 resamples re-estimates the noise rate through a cluster bootstrap on the A-side instances in addition to resampling the evaluation instances, the full uncertainty interval of the conservative reachability at the default decay coefficient β=0.7β=0.7 is [0.367,0.401][0.367,0.401] at L2 and [0.366,0.401][0.366,0.401] at L1, essentially the same width as the instance-only interval. This is because the conservative end depends only on the noise rate and not on the prior fit, so prior-fitting uncertainty does not propagate. Across the nine combinations of the decay coefficient β∈0.5,0.7,0.9β∈\0.5,0.7,0.9\ and three split seeds, the conservative numerator has mean 0.1470.147 and standard deviation 0.0100.010 at L2 and mean 0.1310.131 and standard deviation 0.0090.009 at L1. A smaller β removes more annotation noise and hence lowers the ceiling, so the conservative reachability itself spans [0.345,0.411][0.345,0.411] at L2 and [0.338,0.409][0.338,0.409] at L1, with a minimum of 0.33750.3375 at L1 with β=0.5β=0.5, which is the value underpinning the headline lower bound of at least about 3333% across the smoothing-and-seed grid. Finally, the NSB upper end is not a grid artifact. Across integration grids of 120120, 500500, and 20002000 nodes the ceiling drifts by at most 0.00160.0016 at L2 and 0.00140.0014 at L1, below the Monte Carlo noise, so the upper-end reachability near 1.031.03 is genuine. Finally, we delineate the statistical framing honestly. The comparisons in this work are estimation rather than between-group significance testing, and the reported intervals are per-quantity instance-bootstrap 9595% confidence intervals. The cross-label and cross-task comparisons are descriptive and not significance claims from multiple testing, so they are not subject to multiple-comparison correction. The effect size is the reported absolute magnitude itself, namely the bracket width, the spread of the reachability R from 0.380.38 to 1.031.03, and the ceiling gap between the plug-in and corrected estimators, together with their attendant confidence intervals, and not a standardized test statistic. For reproducibility, classifier inference is run on an RTX PRO 6000 Blackwell 96GB (transformers 4.44.2, torch 2.2) and ceiling estimation is run on CPU. The fixed dependencies (requirements.txt) and the full code are released, so all numbers are reproducible from the frozen split (seed 2026072320260723). Appendix G Confidence-Interval Coverage Simulation Whether the bootstrap confidence intervals of Appendix F attain nominal coverage is verified directly on synthetic data with a known ground truth. Label distributions pip_i are drawn from a known generative distribution G, and its dataset-level irreducible error Ctrue=G[1−maxypiy]C_true=E_G[1- _yp_iy] is fixed by large-scale Monte Carlo. G is a K=7K=7 anchored peak-and-splitting mixture with a neutral-dominant base measure that mimics the skew of the GoEmotions L2 marginal, and it belongs to the same distributional family assumed by the mixture empirical Bayes. With 2,000,0002,000,000 prior draws we obtain Ctrue=0.261C_true=0.261 with a Monte Carlo standard error of 1.7×10−41.7× 10^-4. Each synthetic dataset consists of 2,0002,000 instances, and each instance is given n∈3,5n∈\3,5\ finite annotations by a multinomial, consistent with the mean annotator count of 3.43.4 in GoEmotions. For each estimator we compute the same instance-level estimator as in Appendix D and the same B=1000B=1000 instance-bootstrap 9595% confidence interval as in Appendix F, and across Msim=100M_sim=100 independent synthetic datasets we measure the empirical coverage as the fraction of confidence intervals that contain CtrueC_true. Msim=100M_sim=100 was chosen from the synchronous CPU-only run time, and the standard error of the coverage near nominal is about 0.0220.022. The frozen seed is 2026072920260729. Table 6 shows the results. The plug-in coverage is 00 at both n=3n=3 and n=5n=5, because the plug-in is a negatively biased lower bound owing to the finite-sample upward bias of 1−maxyp^1- _y p, so its confidence interval lies entirely below CtrueC_true on every dataset, with the bias shrinking with the sample size from −0.072-0.072 at n=3n=3 to −0.044-0.044 at n=5n=5. The NSB coverage is also 00, but because NSB is designed as an upper bound with positive bias +0.064+0.064 and +0.040+0.040, its confidence interval lies entirely above CtrueC_true on every dataset. The systematic bracket [plug-in,NSB][plug-in,\ NSB] spanned by the two ends therefore contains CtrueC_true on all 100100 datasets, consistent with the framing of this work in which it is the bracket rather than a point estimator that captures the truth. The bias-corrected mixture empirical Bayes is nearly unbiased, with bias of only +0.001+0.001 and +0.003+0.003, so it attains the best coverage of the three, reaching 0.750.75 at n=3n=3 and 0.890.89 at n=5n=5 and approaching the nominal 0.950.95 with the annotation count. The residual under-coverage below nominal nonetheless stems from finite-sample residual bias and slight prior misspecification, and it supports the design choice of reporting the systematic bracket rather than relying on the confidence interval of a single point estimator alone. Table 6: Empirical coverage of the 9595% bootstrap confidence intervals. Ground truth Ctrue=0.261C_true=0.261, Msim=100M_sim=100 synthetic datasets, B=1000B=1000. Bold marks the best coverage at each n. Estimator n Coverage CI width Bias plug-in 3 0.00 0.028 −0.072-0.072 mixture EB 3 0.75 0.020 +0.001+0.001 NSB 3 0.00 0.023 +0.064+0.064 plug-in 5 0.00 0.027 −0.044-0.044 mixture EB 5 0.89 0.024 +0.003+0.003 NSB 5 0.00 0.027 +0.040+0.040 Appendix H Annotation-Noise Deconvolution: EM and Recovery The mixture model of Eq. (5) is fitted by EM (11). The E step is implemented as a leave-one-out prediction step, because in-sample posterior-mean imputation degenerates at ni=3n_i=3, as each observation explains itself and ε collapses to zero. A slightly smoothed leave-one-out prediction with decay coefficient β=0.7β=0.7 recovers the identifiability of ε . The update of p~ p is a regularized posterior-mean imputation rather than the maximum a posteriori value, and it therefore has no strict monotonicity guarantee, so monotonicity is monitored empirically and damping is introduced upon a violation. The noise rate εa _a is estimated once at L2 granularity, and the same ε^a _a together with the corresponding ν is reused for the other granularities and the binary decomposition, which avoids the inconsistency of ε taking different values across granularities. The validity of this approximation is supported by a recovery experiment on synthetic data. At n=3n=3 the noise rate is recovered with mean absolute error 0.0390.039 and correlation 0.930.93, and when the true noise is ε=0 =0 no false positives arise, with an estimated mean of 0.0150.015. Because the removal effect is confounded with the finite-sample downward bias of the plug-in at n=3n=3, the isolation of the removal effect is verified at n=15n=15, where the absolute entropy bias decreases from 0.2700.270 for the raw plug-in to 0.1120.112 after deconvolution. Furthermore, a recovery experiment that draws the signal from a prior fitted to the actual GoEmotions L2 counts and the annotator counts from the actual distribution (mean 3.43.4 across all tracks and 2.32.3 on the A side) shows that at these observed counts the noise rate is overestimated with mean absolute error of about 0.160.16 and a mean of 0.3750.375 against a true 0.2120.212, yet the rank correlation stays high between 0.880.88 and 0.960.96, and the removal effect cannot be separated from the finite-sample bias until n=15n=15, where the error drops to 0.0180.018. In all real-sample settings the post-deconvolution plug-in floor lies below the true Bayes error, for example 0.2160.216 against 0.3730.373. This confirms that the post-deconvolution plug-in end is a conservative lower bound and that the reported irreducible fraction is a lower bound rather than a point estimate. Figure 2 visualizes how the reachability R resulting from Section 4.4 depends on the choice of estimator and deconvolution, and shows that no single reachability value exists. Appendix I Validation of the Estimators on Synthetic Data The estimators are validated on synthetic distributions with known entropy. On the uniform distribution with K=7K=7 and n=3n=3, the entropy bias [H^]−HtrueE[ H]-H_true in bits is −1.498-1.498 for the plug-in, −1.118-1.118 for Miller–Madow, −0.789-0.789 for NSB, and −0.781-0.781 for the symmetric Dirichlet with concentration 0.50.5, and on the uniform distribution with K=28K=28 and n=3n=3 the plug-in bias reaches −3.298-3.298, confirming a bias exceeding 11 bit in the severely undersampled regime. The Fano inverse at K=2K=2 agrees with min(p,1−p) (p,1-p) to an error of 1.4×10−141.4× 10^-14, and the closed-form binary Bayes-error floor of Eq. (7) agrees with quadrature evaluation to a maximum error of 3.75×10−123.75× 10^-12. Appendix J Irreducible-Error Floors for All Labels Table 7 lists the instance-level irreducible-error floor ek∗e^*_k for all 2828 GoEmotions labels of the binary decomposition, of which the main text in Section 4.3 reports only representative values. The floor is smallest for rare, high-agreement emotions and largest for frequent, ambiguous emotions. Figure 6 shows the per-label headroom of Section 4.5, where reaching the oracle F1 is unrelated to the height of the floor. For binary F1, the oracle marginal-threshold floor is 0.3570.357 at macro and 0.4070.407 at micro, whereas the evaluated classifier at threshold 0.50.5 attains 0.3330.333 at macro and 0.4380.438 at micro, with the micro reachability exceeding 11 because the floor is a marginal reference value. The soft-label P3 reachability against the cross-entropy floor spans 0.280.28 to 0.850.85 at L2 and 0.300.30 to 0.860.86 at L1. As a diagnostic, G G takes negative values from −0.001-0.001 to −0.013-0.013 at the upper L2 ceiling, namely the raw NSB counts and the post-deconvolution mixture empirical Bayes and NSB, and this is reported as an alarm rather than as superhuman performance, whereas at the conservative end all gaps are positive. Figure 6: Per-label headroom. Table 7: Irreducible-error floors ek∗e^*_k for all 2828 labels (ascending in the floor). Label ek∗e^*_k Label ek∗e^*_k Label ek∗e^*_k Label ek∗e^*_k grief 0.0033 desire 0.0175 sadness 0.0296 curiosity 0.0407 pride 0.0063 gratitude 0.0220 amusement 0.0326 realization 0.0415 relief 0.0063 surprise 0.0242 confusion 0.0336 disapproval 0.0523 nervousness 0.0087 disgust 0.0243 anger 0.0358 admiration 0.0620 remorse 0.0106 love 0.0258 joy 0.0359 annoyance 0.0622 embarrassment 0.0114 excitement 0.0264 optimism 0.0393 approval 0.0827 fear 0.0133 caring 0.0271 disappoint. 0.0393 neutral 0.2178 Figure 3 visualizes the cross-task predictability map of Section 4.6 together with the base rates, and shows that without controlling for the base rate the ceilings are not comparable across tasks. For the binary tasks with a base rate of about 0.30.3, the irreducible ambiguity is largest for irony at [0.206,0.261][0.206,0.261], followed by offensiveness at [0.173,0.212][0.173,0.212] and the presence of emotion at [0.116,0.149][0.116,0.149], which is interpreted as arising because irony is the most context-dependent and the most dispersed across annotators. Deconvolution lowers the floor on every binary task, from 0.2060.206 to 0.1390.139 for irony and from 0.1730.173 to 0.1370.137 for offensiveness, so that irony and offensiveness nearly converge. The mean noise rate ε¯ of irony is also the largest at 0.2570.257, suggesting that a substantial part of the observed disagreement stems from annotation noise. When emotion intensity is treated as an ordinal K=4K=4, the BRIGHTER floor rises to [0.173,0.221][0.173,0.221], confirming that the floor rises monotonically with the number of classes. Because BRIGHTER lacks annotator identifiers, deconvolution and the A/B split cannot be applied, and its multiply annotated data has a narrow bracket width of about 0.050.05, confirming that the impact of undersampling is small. Appendix K Cross-Task SOTA Decomposition Table 3 decomposes off-the-shelf public classifiers on each cross-task corpus against their conservative post-deconvolution plug-in ceiling, extending the reachability analysis of Section 4.4 beyond GoEmotions. The offensiveness classifier is cardiffnlp/twitter-roberta-base-offensive and the irony classifier is cardiffnlp/twitter-roberta-base-irony, each thresholded to the binary positive class. The emotion row is SamLowe/roberta-base-go_emotions (macro across six emotions), mapped and averaged onto the six BRIGHTER emotions through a noisy-OR over the GoEmotions labels. Two observations transfer from GoEmotions. First, the conservative irreducible component remains strictly positive on all three corpora, at 0.1370.137 for offensiveness, 0.1390.139 for irony, and 0.1160.116 for emotion, so a positive irreducible error C>0C>0 is not a GoEmotions artifact. Second, the reachability again depends strongly on the estimator. The conservative post-deconvolution plug-in end is the smallest on every task, and the raw empirical Bayes and NSB ends raise it by roughly 1.31.3 to 1.91.9 times, from Rcons=0.413R_cons=0.413 to REB=0.640R_EB=0.640 for offensiveness, from 0.3590.359 to 0.6720.672 for irony, and from 0.4220.422 to RNSB=0.543R_NSB=0.543 for emotion, so the estimator dependence of the saturation verdict generalizes across tasks. These reachabilities are lower bounds and are honestly scoped in two respects. The classifiers are public models that are not fine-tuned on the target corpora, so their error is inflated relative to an in-domain model, and the reported R therefore underestimates the reachability that a fine-tuned classifier would attain. Accordingly, the cross-task RconsR_cons of 0.360.36 to 0.420.42 is comparable to the conservative band of 0.290.29 to 0.390.39 for GoEmotions, whereas the raw ends of 0.480.48 to 0.670.67 fall below the nearly unit in-domain reachability of GoEmotions in Section 4.5. The emotion row further involves a domain and label mismatch. The six BRIGHTER emotions are matched through a synthetic mapping in which categories such as fear–anxiety and social-warmth are composites of several GoEmotions labels, which lowers their reachability further. The per-emotion reachabilities are not cherry-picked and span from 0.2970.297 for the composite fear–anxiety category to 0.5400.540 for anger, with both ends reported. The cross-task decomposition therefore supports the generality of the two pillars as a qualitative phenomenon rather than as a transfer of the GoEmotions numbers.