Paper deep dive
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
Seitaro Ono, Senna Ross, Jun Saiki
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/10/2026, 3:33:33 AM
Summary
This paper proposes Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step to calibrate the Word Embedding Association Test (WEAT). The authors argue that anisotropy in embedding spaces violates the isotropy assumption required for cosine similarity to accurately measure semantic associations, leading to unreliable bias measurements. By transforming the covariance matrix of the embedding space to the identity matrix, ZCA whitening restores isotropy. Evaluations across seven models and ten test suites show that this calibration reduces anisotropy, improves semantic similarity benchmarks, and significantly alters WEAT results, with over 30% of results changing significance status. The findings suggest that previous bias measurements in anisotropic spaces may be over- or under-estimated.
Entities (10)
Relation Signals (10)
ZCA Whitening → calibrates → WEAT
confidence 95% · We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT).
ZCA Whitening → reduces → Anisotropy
confidence 95% · The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models.
Cosine Similarity → assumes → Isotropy
confidence 90% · It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic.
GPT-2 → evaluatedwith → ZCA Whitening
confidence 90% · We evaluate our approach on ten standard WEAT test suites and seven models... GPT-2
GloVe → evaluatedwith → ZCA Whitening
confidence 90% · We evaluate our approach on ten standard WEAT test suites and seven models... GloVe
BERT → evaluatedwith → ZCA Whitening
confidence 90% · We evaluate our approach on ten standard WEAT test suites and seven models... BERT-base-cased
RoBERTa → evaluatedwith → ZCA Whitening
confidence 90% · We evaluate our approach on ten standard WEAT test suites and seven models... RoBERTa-base
ZCA Whitening → restores → Isotropy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. It relies on cosine similarity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. However, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening transforms the covariance of the embedding space into the identity matrix while minimizing perturbation to the original vectors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten standard WEAT test suites and seven models spanning three architectural families, yielding 70 model-task combinations. The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models. Particularly for highly anisotropic models, we further observe improvements on standard semantic similarity benchmarks, indicating that the calibrated space better captures semantic associations. After calibration, over 30% of WEAT results change significance status, and effect sizes shift in both directions depending on bias category. These shifts suggest that uncalibrated measurements may both overestimate and underestimate the associations encoded in the embedding space. These findings indicate that previously reported bias measurements in anisotropic embedding spaces should be interpreted with caution and may benefit from re-evaluation with calibrated methods. Our approach contributes to restoring the measurement foundation of WEAT across both computational social science and AI fairness research.
Tags
Links
- Source: https://arxiv.org/abs/2608.06908v1
- Canonical: https://arxiv.org/abs/2608.06908v1
Trouble viewing inline? Open PDF directly →
Full Text
77,287 characters extracted from source content.
Expand or collapse full text
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests Seitaro Ono 1 , Senna Ross 2 , Jun Saiki 1 1 Graduate School of Human and Environmental Studies, Kyoto University, Kyoto, Japan 2 Brock University, St. Catharines, Ontario, Canada ono@cv.jinkan.kyoto-u.ac.jp, sz23hr@brocku.ca, saiki.jun.8e@kyoto-u.ac.jp Abstract We propose Zero-phase Component Analysis (ZCA) whiten- ing as a geometric pre-processing step for the Word Em- bedding Association Test (WEAT). WEAT is a bias mea- surement method widely used in both computational social science and AI fairness research. It relies on cosine simi- larity as a measure of semantic association, which assumes that the embedding space is approximately isotropic. How- ever, prior work has reported that many widely used language models do not satisfy this assumption, raising concerns about the reliability of bias measurements. ZCA whitening trans- forms the covariance of the embedding space into the iden- tity matrix while minimizing perturbation to the original vec- tors. This transformation restores the isotropy condition on which WEAT relies. We evaluate our approach on ten stan- dard WEAT test suites and seven models spanning three ar- chitectural families, yielding 70 model–task combinations. The results show that ZCA whitening substantially reduces the anisotropy of the embedding spaces across all models. Particularly for highly anisotropic models, we further observe improvements on standard semantic similarity benchmarks, indicating that the calibrated space better captures semantic associations. After calibration, over 30% of WEAT results change significance status, and effect sizes shift in both di- rections depending on bias category. These shifts suggest that uncalibrated measurements may both overestimate and un- derestimate the associations encoded in the embedding space. These findings indicate that previously reported bias mea- surements in anisotropic embedding spaces should be inter- preted with caution and may benefit from re-evaluation with calibrated methods. Our approach contributes to restoring the measurement foundation of WEAT across both compu- tational social science and AI fairness research. Code — https://github.com/seigit/zca-weat Accepted at the 9th AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026). This is an extended version in- cluding appendices; the final published version will appear in the AAAI Digital Library. 1 Introduction Word embeddings (Mikolov et al. 2013; Pennington, Socher, and Manning 2014) encode statistical regularities of lan- guage, including social biases that can propagate to down- stream applications (Bolukbasi et al. 2016; Caliskan, Bryson, and Narayanan 2017; Garg et al. 2018). Measur- ing these biases accurately is necessary for developing equi- table NLP systems. The Word Embedding Association Test (WEAT; Caliskan, Bryson, and Narayanan 2017) transposes the logic of the Implicit Association Test (IAT; Greenwald, McGhee, and Schwartz 1998) from human reaction-time studies to the geometry of word embedding spaces, measur- ing differential cosine-similarity associations between target concepts (e.g., European-American vs. African-American names) and attribute concepts (e.g., pleasant vs. unpleasant words). The WEAT has attracted significant attention in both computational social science and AI fairness research, and has been applied in a large number of studies to measure racial, gender, age, and other biases across diverse models and modalities. However, the validity of any measurement metric depends on the assumptions under which it operates. WEAT assumes that cosine similarity faithfully captures semantic associa- tions between words. This assumption holds when the em- bedding space is approximately isotropic (Ethayarajh 2019; Timkey and Van Schijndel 2021; Rudman et al. 2022). In an isotropic space, vectors are distributed with roughly uniform variance in all directions. Under this condition, the angle be- tween two vectors reflects their semantic relationship rather than a geometric artifact. Recent work has shown that this isotropy assumption fails in many language models: static embeddings exhibit domi- nant mean directions (Mu and Viswanath 2018), contextu- alized representations occupy narrow cones with near-unity pairwise similarities (Ethayarajh 2019), and a small number of “rogue dimensions” dominate cosine similarity in Trans- formers (Timkey and Van Schijndel 2021). Despite this accumulating evidence, the consequences of anisotropy for bias measurement have received surprisingly little direct attention. Wolfe, Hiniker, and Howe (2024) re- cently introduced ML-EAT, which provides a comprehen- sive multilevel evaluation framework and correctly identi- fies anisotropy as a threat to the validity of cosine-based bias tests. However, the purpose of ML-EAT was to improve the interpretability of EAT measurements, and proposing a cor- rection for anisotropy was outside its scope. A calibration method that restores the geometric conditions under which cosine-based bias measurement becomes reliable remains a significant challenge. arXiv:2608.06908v1 [cs.CL] 7 Aug 2026 GPT-2 GPT-2 Before Whitening Mean: 0.9967, SD: 0.0032 GPT-2 After ZCA Whitening Mean: 0.0013, SD: 0.0417 GloVe GloVe Before Whitening Mean: 0.1918, SD: 0.2114 GloVe After ZCA Whitening Mean: 0.0402, SD: 0.1128 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 Figure 1: Pairwise cosine similarity matrices for 2,000 randomly sampled embedding vectors from WikiText-103 before (left column) and after (right column) ZCA whitening. Top: GPT-2, whose extreme anisotropy (mean cosine similarity = 0.997) renders the matrix a uniform red field. Bottom: GloVe (mean = 0.192), which exhibits moderate anisotropy before whitening. After calibration, both models converge to near-zero mean similarity with meaningful pairwise variation restored. In this paper, we introduce a calibration step that removes anisotropy from language model embedding spaces, restor- ing the isotropy assumption under which cosine-based bias measurement is valid. Specifically, we propose ZCA (Zero- phase Component Analysis) whitening as a pre-processing calibration for embedding association tests. ZCA whitening transforms the embedding space to have identity covariance while minimizing the perturbation to the original vectors, making it well suited for effectively reducing anisotropy. Moreover, whitening has been shown to improve the quality of embeddings on standard semantic similarity benchmarks (Huang et al. 2021; Su et al. 2021), suggesting that calibrated embedding spaces may better capture the semantic associa- tions on which bias measurement depends. We evaluate our approach across seven embedding mod- els spanning three architectural families (static, contextu- alized, and contrastive unsupervised models) using the ten standard WEAT test suites from Caliskan, Bryson, and Narayanan (2017). We make the following contributions: 1. We propose ZCA whitening as a novel pre-processing calibration method for embedding association tests and show that it substantially reduces anisotropy across all seven models in the embedding space. 2. We show that cosine similarity in the calibrated space more faithfully captures semantic associations on stan- dard similarity benchmarks, supporting the validity of bias measurements in the calibrated space. 3. We systematically show that calibration changes WEAT outcomes in both directions across 70 model–task com- binations, reducing measured associations in some con- figurations and increasing them in others. 2 Related Work 2.1 The Embedding Association Test Caliskan, Bryson, and Narayanan (2017) introduced the Word Embedding Association Test (WEAT), a measurement of intrinsic bias in word embeddings drawing on the de- sign of the IAT (Greenwald, McGhee, and Schwartz 1998). The WEAT quantifies the relative association of two target groups (such as European-American and African-American names) with two attribute groups (such as pleasant and un- pleasant words). Given target sets X and Y and attribute sets A and B, the individual association for each target word w is: s(w,A,B) = 1 |A| X a∈A cos(⃗w,⃗a)− 1 |B| X b∈B cos(⃗w, ⃗ b) (1) The test statistic aggregates individual associations across the two target sets: S(X,Y,A,B) = X x∈X s(x,A,B)− X y∈Y s(y,A,B)(2) The effect size d normalizes this difference: d = mean x∈X s(x,A,B)− mean y∈Y s(y,A,B) std w∈X∪Y s(w,A,B) (3) This effect size is analogous to Cohen’s d (Cohen 1992), where values of 0.2, 0.5, and 0.8 are conventionally inter- preted as small, medium, and large effects. Statistical sig- nificance is assessed via a one-sided permutation test that computes the p-value as: p = Pr i S(X i ,Y i ,A,B) > S(X,Y,A,B) (4) where (X i ,Y i ) denotes the set of all equally sized parti- tions of X ∪ Y . Since its introduction, the WEAT has been extended across modalities and architectures, in order to measure social biases in AI systems. Text-based variants include SC-WEAT for single-target associations (Caliskan, Bryson, and Narayanan 2017), SEAT for sentence-level embeddings (May et al. 2019), a pretraining-objective-based adapta- tion (Kurita et al. 2019), and CEAT, which treats contex- tualization as a random effect at the word level (Guo and Caliskan 2021). The framework has also been extended be- yond text to image encoders (iEAT; Steed and Caliskan 2021), grounded vision-and-language models (Grounded- WEAT; Ross, Katz, and Barbu 2021), CLIP (Wolfe and Caliskan 2022; Wolfe, Banaji, and Caliskan 2022a; Rad- ford et al. 2021), speech models (SpEAT; Slaughter et al. 2023), and text-to-video generators (VEAT; Sun et al. 2026). Recently, Wolfe, Hiniker, and Howe (2024) proposed ML- EAT, which provides multilevel interpretable bias measure- ments and includes an anisotropy diagnostic. All of these variants rely on cosine similarity and are therefore subject to the same geometric artifacts that this paper addresses. Their widespread adoption across NLP, computer vision, and speech processing underscores the importance of ensur- ing that their measurements are geometrically sound. 2.2 Applications of EATs in Social Science Embedding association tests have been widely adopted in the social sciences as a tool for studying societal phenom- ena at scale. A growing body of work uses EATs to probe contemporary cultural patterns in large corpora, including gender defaults and stereotypes in internet English (Caliskan et al. 2022; Bailey, Williams, and Cimpian 2022), the rela- tionship between gender bias and societal structure (Napp 2023), and temporal shifts in Wikipedia (Schmahl et al. 2020). Cross-cultural and cross-linguistic studies have ex- tended this reach to dozens of languages (Lewis and Lupyan 2020; Charlesworth et al. 2021; Mukherjee et al. 2023). Historical corpora have enabled the study of how these as- sociations evolve over time, including a century of gen- der and ethnic stereotypes (Garg et al. 2018), 200 years of social-group stereotypes (Charlesworth, Caliskan, and Ba- naji 2022), intersectional stereotypes (Charlesworth et al. 2024a; Borenstein et al. 2023), xenophobia in 19th–20th century travel literature (Sunsay 2023), and the expansion of the societal “moral circle” (Leach, Kitchin, and Sut- ton 2023). Collectively, these applications demonstrate that EATs can serve not only as a measure of contemporary as- sociations but also as a tool for tracing long-term societal change. In applied domains, EATs have been used to mea- sure biases in biomedical research (Rios, Joshi, and Shin 2020), clinical notes (Cobert et al. 2024), legal opinions (Matthews, Hudzina, and Sepehr 2022), and court proceed- ings (Dutta et al. 2023). Gray and Wu (2025) further pro- posed SD-WEAT, a variant handling multi-level attribute groups for healthcare bias benchmarking. Recent work has also linked embedding-based measurements to human-level implicit associations, demonstrating meaningful correspon- dence between embedding biases and societal stereotypes (Charlesworth et al. 2024b; Morehouse et al. 2023; Manzini et al. 2019). These diverse applications underscore that the validity of EAT measurements is not a narrow technical concern but a matter of scientific integrity across multiple disciplines. However, several studies have identified potential problems with methods that rely on cosine similarity between embed- ding vectors, suggesting that the geometric properties of the embedding space may affect the reliability of such measure- ments. 2.3 Embedding Anisotropy and Its Implications for WEAT Anisotropy in word embeddings describes a geometric prop- erty in which vectors cluster in a narrow region of the em- bedding space (Cai et al. 2021), causing cosine similarities between arbitrary word pairs to be systematically high. Etha- yarajh (2019) demonstrated that contextualized representa- tions from BERT, ELMo, and GPT-2 occupy a narrow cone with average cosine similarities approaching 1.0 in upper layers, while Mu and Viswanath (2018) identified dominant principal components in word2vec and GloVe, and Timkey and Van Schijndel (2021) showed that a few “rogue dimen- sions” dominate cosine computation in Transformers. Fur- ther work has documented additional effects (Zhou et al. 2022; Godey, de la Clergerie, and Sagot 2024; Machina and Mercer 2024; Rajaee and Pilehvar 2022). For WEAT, such geometric distortion can compress or in- flate the differential cosine associations that define its ef- fect size, producing measurements that do not accurately reflect the underlying bias structure. Wolfe, Hiniker, and Howe (2024) incorporated an anisotropy diagnostic into ML-EAT that flags unreliable models, but ML-EAT aims at interpretability rather than correction. Mu and Viswanath (2018) proposed removing top principal components, but this requires explicitly discarding information and selecting a model-dependent hyperparameter, limiting its practical ap- plicability. Thus, a calibration method that restores the reli- ability of cosine-based measurement without requiring such choices is desirable. 3 Approach 3.1 How Anisotropy Distorts Cosine-Based Association Tests Cosine similarity between two vectors ⃗u,⃗v ∈R d is defined as: cos(⃗u,⃗v) = ⃗u·⃗v ∥⃗u∥⃗v∥ (5) For cosine similarity to serve as a reliable measure of se- mantic association, the embedding space should be approx- imately isotropic (Ethayarajh 2019; Wolfe, Hiniker, and Howe 2024). Specifically, the covariance matrix Σ of the embedding vectors should be approximately proportional to the identity matrix (Rudman et al. 2022): Σ≈ σ 2 I(6) When Σ departs from proportionality to I , cosine simi- larity may become systematically distorted. Because cosine similarity is dominated by high-variance directions, vectors can appear more similar than they semantically are, while meaningful differences along low-variance directions are masked (Mu and Viswanath 2018; Timkey and Van Schi- jndel 2021). Since the WEAT computes differential cosine-similarity associations between target and attribute word sets (Equa- tions 1–3), the effect size d may be affected by anisotropy if the distortion is not uniform across the word sets involved in the test. This suggests that anisotropy could both inflate and compress genuine bias signals, depending on the rela- tionship between the word sets and the covariance structure of the embedding space. To address this potential distortion, a calibration step that transforms the covariance matrix of the embedding space to- ward the identity matrix is needed, thereby restoring the geo- metric conditions under which cosine similarity can function as a reliable measure of semantic association. 3.2 ZCA Whitening We adopt Zero-phase Component Analysis (ZCA) whiten- ing as the pre-processing calibration step. Given a set of embedding vectorsx 1 ,x 2 ,...,x n ⊂R d , we first compute the sample mean μ and the sample covari- ance matrix Σ: μ = 1 n n X i=1 x i ,Σ = 1 n− 1 n X i=1 (x i − μ)(x i − μ) ⊤ (7) We then center each vector by subtracting the mean and ap- ply the ZCA whitening matrix W ZCA to obtain the trans- formed vectors: z i = W ZCA (x i − μ), W ZCA = Σ −1/2 (8) where Σ −1/2 denotes the symmetric matrix square root of Σ −1 . To compute this matrix, we perform the eigendecom- position of the covariance matrix Σ = U ΛU ⊤ , where U is an orthogonal matrix whose columns are the eigenvectors and Λ is a diagonal matrix of the corresponding eigenvalues. The whitening matrix then takes the explicit form: W ZCA = U Λ −1/2 U ⊤ (9) This transformation rescales each eigenvector direction by the inverse square root of its eigenvalue: dimensions with disproportionately large variance are compressed, while di- mensions with small variance are expanded, resulting in uni- form variance across all directions. The transformed vectorsz i have identity covariance by construction: Cov(z) = W ZCA ΣW ⊤ ZCA = Σ −1/2 Σ Σ −1/2 = I,(10) where the final equality follows from the symmetry of W ZCA = Σ −1/2 (see Appendix A for a full derivation). Since Cov(z) = I , the calibrated embedding space satisfies the isotropy condition (Equation 6 with σ 2 = 1) under which cosine similarity can function as a reliable measure of se- mantic association. Among the family of whitening transformations that pro- duce identity covariance, ZCA uniquely minimizes the ex- pected squared distance between original and transformed vectors (Kessy, Lewin, and Strimmer 2018), making it par- ticularly suitable for bias measurement calibration where preserving the original semantic structure is essential. More- over, whitening has been shown to substantially improve scores on standard semantic textual similarity benchmarks (Huang et al. 2021; Su et al. 2021), suggesting that the cal- ibrated embedding space more faithfully captures seman- tic associations between words and may therefore provide a more reliable basis for extracting social associations that were obscured by anisotropy in the original space. 3.3 Calibration Pipeline Our full measurement pipeline operates in three stages. In the first stage, we estimate the whitening matrix from a large reference sample of embedding vectors. We adopt WikiText-103 (Merity et al. 2017) as the reference corpus for collecting these samples. From this corpus, we draw 100,000 samples and feed them into the model under evaluation to obtain the corresponding embedding vectors. Using the re- sulting embeddings, we compute the sample mean μ and the covariance matrix Σ (Equation 7), from which we estimate the ZCA whitening matrix W ZCA (Equation 9). This estima- tion is performed once per model and the resulting whitening statistics are reused across all test suites. In the second stage, we apply the estimated whitening transformation to the specific word vectors involved in the WEAT test suites. Each word vector is centered by subtract- ing the reference mean μ and then transformed by W ZCA to obtain the calibrated vector (Equation 8). In the third stage, we compute WEAT scores (effect size d and permutation p-value) using the calibrated vectors. This pipeline is modular: it can be applied to any em- bedding model and any cosine-similarity-based bias metric without modifying the metric itself. 4 Experimental Setup 4.1 Models We evaluate seven embedding models spanning three archi- tectural families. Static embeddings. GloVe (Pennington, Socher, and Manning 2014) trained on Wikipedia and Gigaword (6B to- kens, 100 dimensions), and word2vec (Mikolov et al. 2013) trained on Google News (approximately 100B tokens, 300 dimensions). We use the 6B-token, 100-dimensional GloVe model rather than the 840B-token, 300-dimensional variant used in Caliskan, Bryson, and Narayanan (2017). Our pur- pose is not to replicate their specific results but to evaluate the impact of anisotropy calibration, and the 6B model ex- hibits sufficient anisotropy (Table 2) to serve this purpose. Table 1: The ten standard WEAT test suites (Caliskan, Bryson, and Narayanan 2017). Each suite specifies two target sets (X , Y ) and two attribute sets (A, B).|·| denotes the number of stimulus words in each set, following the lists provided in Caliskan, Bryson, and Narayanan (2017). EA and A denote European-American and African-American, respectively. ID Bias Type Target XTarget YAttribute A Attribute B |X| |Y| |A| |B| W1ValenceFlowersInsectsPleasantUnpleasant25252525 W2ValenceInstrumentsWeaponsPleasantUnpleasant25252525 W3RaceEA namesAA namesPleasantUnpleasant32322525 W4RaceEA namesAA namesPleasantUnpleasant16162525 W5RaceEA namesAA namesPleasantUnpleasant161688 W6GenderMale namesFemale namesCareerFamily8888 W7GenderMathArtsMale termsFemale terms8888 W8GenderScienceArtsMale termsFemale terms8888 W9HealthMental diseasePhysical diseaseTemporaryPermanent6677 W10AgeYoung namesOld namesPleasantUnpleasant8888 Contextualized embeddings. BERT-base-cased (Devlin et al. 2019), RoBERTa-base (Liu et al. 2019), and GPT- 2 (Radford et al. 2019), each with 12 Transformer layers and 768-dimensional hidden representations. For all Trans- former models, we use representations from the final layer, consistent with prior work applying EATs to contextual- ized models (May et al. 2019; Guo and Caliskan 2021; Wolfe, Hiniker, and Howe 2024). Word-level embeddings are obtained following the Aggregated procedure of Bom- masani, Davis, and Cardie (2020): for each target or attribute word, we sample up to 20 sentences containing that word from a preprocessed WikiText-103 corpus. Preprocessing re- tains sentences of 7–75 tokens and excludes section head- ers, yielding up to 200,000 indexed sentences with a fixed random seed for reproducibility. Within each sentence, sub- word tokens of the target word are mean-pooled into a sin- gle contextualized vector, and the resulting vectors across the 20 sentences are then averaged to produce one word- level representation. Bommasani, Davis, and Cardie (2020) reported that bias evaluation results were fairly stable across n i ∈20, 50, 100 contexts per word. We therefore adopted n i = 20 to maintain consistency with their protocol and to minimize computational overhead. We applied the canon- ical WEAT to these word-level representations rather than sentence-level variants such as SEAT (May et al. 2019) or distributional variants such as CEAT (Guo and Caliskan 2021). This choice allows us to use a unified measurement procedure across all seven models, including static embed- dings for which SEAT and CEAT are not applicable. Unsupervised contrastive embeddings. We adopt two unsupervised contrastive models from the SimCSE frame- work (Gao, Yao, and Chen 2021): unsup-SimCSE- BERT-base-uncased (unsup-BERT) and unsup-SimCSE- RoBERTa-base (unsup-RoBERTa). Both models share the same Transformer architecture as their supervised counter- parts (12 layers, 768-dimensional hidden representations), and word-level embeddings are obtained using the same aggregation procedure of Bommasani, Davis, and Cardie (2020) described above. These models differ from the stan- dard contextualized models in two key aspects: training sig- nal (unsupervised vs. supervised) and objective function (contrastive vs. language modeling). We include them in our analysis to examine how these factors influence the degree of anisotropy and the patterns of bias detected by WEAT. Whitening statistics. The whitening matrix W ZCA is es- timated from WikiText-103 (Merity et al. 2017). For static models, we sample up to 100,000 vocabulary items. For con- textualized and unsupervised contrastive models, we sam- ple up to 100,000 sentences. These contextualized mod- els are pretrained on sentence-level inputs. Feeding isolated words may therefore produce out-of-distribution representa- tions (Bommasani, Davis, and Cardie 2020). To avoid this issue, we use sentence-level inputs to estimate the covari- ance matrix. To ensure numerical stability when computing the whitening matrix (Equation 9), we regularize the eigen- values by replacing Λ with Λ+εI before inversion, yielding: W ZCA = U (Λ + εI) −1/2 U ⊤ (11) where ε = 10 −3 . This prevents near-zero eigenvalues from producing excessively large entries in the whitening matrix. 4.2 WEAT Test Suites We use the ten standard WEAT test suites from Caliskan, Bryson, and Narayanan (2017) (Table 1). These test suites correspond to well-established IAT findings in social psy- chology and have been used in numerous studies to bench- mark bias in NLP models (Kurita et al. 2019; May et al. 2019; Guo and Caliskan 2021; Wolfe, Hiniker, and Howe 2024). 4.3 Evaluation Protocol Experiment 1: Anisotropy measurement. We quantify anisotropy by computing the mean and standard deviation of pairwise cosine similarities among 2,000 randomly sam- pled embedding vectors from WikiText-103, both before and after ZCA whitening. An isotropic space yields a mean pair- wise cosine similarity near zero. In addition, we qualitatively compare the embedding spaces before and after whitening by examining the explained variance ratio of the covari- ance matrices, which illustrates how variance is distributed across dimensions. We also visualize pairwise cosine simi- larity matrices over the sampled vectors. 1020304050 Principal Component Index 10 3 10 2 10 1 10 0 Explained Variance Ratio Before Whitening 1020304050 Principal Component Index After ZCA Whitening GloVe word2vec BERT RoBERTa GPT-2 unsup-BERT unsup-RoBERTa Figure 2: Explained variance ratio of the top 50 principal components (log scale) before (left) and after (right) ZCA whitening for all seven models. Left: Contextualized models (GPT-2, RoBERTa, BERT) show extreme variance concentration in the first few components, while static models (GloVe, word2vec) and contrastive models (unsup-BERT, unsup-RoBERTa) exhibit flatter distributions. Right: After whitening, all models converge to a near-uniform eigenvalue distribution. Experiment 2: Impact of anisotropy on semantic struc- ture. We evaluate how anisotropy affects the quality of cosine-similarity-based semantic measurements by com- puting Spearman rank correlations (ρ) on WordSim-353 (Finkelstein et al. 2002), SimLex-999 (Hill, Reichart, and Korhonen 2015), and STS-B (Cer et al. 2017) before and after whitening. Each benchmark provides word or sen- tence pairs with human-assigned similarity scores. A higher Spearman correlation indicates that cosine similarity better captures the semantic associations in these benchmarks. Sta- tistical significance is assessed via paired bootstrap resam- pling (10,000 iterations). If calibration improves these cor- relations, it suggests that cosine similarity in the calibrated space more faithfully captures semantic associations, sup- porting the validity of bias measurements conducted in that space. Experiment 3: Bias measurement with WEAT. For each model–task combination, we compute the WEAT effect size d and permutation p-value (100,000 permutations) on both raw and whitened embeddings. Results are classified based on the change in statistical significance at p < 0.05: Sta- ble, significance unchanged; Disappearing, significant be- fore whitening but non-significant after; Emerging, non- significant before but significant after whitening. We then analyze, for each bias category, how these classifications dis- tribute across models and identify the systematic patterns of distortion induced by anisotropy. 5 Results 5.1 ZCA Whitening Reduces Anisotropy Table 2 presents the mean pairwise cosine similarity be- fore and after ZCA whitening for all seven models. Static embeddings exhibit mild anisotropy (GloVe: 0.192; word2vec: 0.059), while contextualized models display se- Table 2: Anisotropy measured by mean pairwise cosine similarity (± standard deviation, SD) among 2,000 ran- domly sampled vectors before (Raw) and after (White) ZCA whitening. ModelRaw Mean Raw SD Wh. Mean Wh. SD GloVe0.1920.2110.0400.113 word2vec0.0590.0740.0130.067 BERT0.8210.0680.0020.043 RoBERTa0.9540.0240.0010.048 GPT-20.9970.0030.0010.042 unsup-BERT0.3860.0940.0010.042 unsup-RoBERTa0.4750.0900.0010.041 vere anisotropy, with GPT-2 reaching an extreme value of 0.997 (SD = 0.003) in which virtually all vectors point in nearly the same direction. This pattern is consistent with prior reports of cosine similarities concentrating near 1.0 between arbitrary word pairs in GPT-2 Base (Timkey and Van Schijndel 2021; Wolfe, Hiniker, and Howe 2024). The unsupervised contrastive models show intermediate anisotropy (unsup-BERT: 0.386; unsup-RoBERTa: 0.475), consistent with the uniformity-promoting nature of con- trastive learning (Gao, Yao, and Chen 2021) but still sub- stantial enough to distort cosine-based measurements. After ZCA whitening, all models converge to near-zero mean pair- wise cosine similarity (0.001–0.040), confirming that cali- bration reduces anisotropy regardless of the initial degree of distortion. This reduction is also visible in the pairwise co- sine similarity matrices (Figure 1). GPT-2’s matrix appears as a nearly uniform red field before whitening, while GloVe exhibits moderate anisotropy. After calibration, both center near zero with restored pairwise variation. The same qualita- tive pattern holds for the other five models (see Appendix B). Figure 2 provides further confirmation through the ex- plained variance ratio of each model’s embedding covari- ance matrix. Before whitening, the contextualized models show extreme variance concentration in the first few prin- cipal components, whereas the static and contrastive mod- els exhibit flatter but still non-uniform distributions. After whitening, the explained variance ratio flattens substantially across all models, indicating that dominant directional com- ponents have been removed and variance is distributed uni- formly across dimensions. 5.2 Calibrated Embeddings Better Capture Semantic Associations Table 3 reports Spearman rank correlations on three seman- tic similarity benchmarks before and after ZCA whitening. For models with high anisotropy (GPT-2, RoBERTa, and BERT), whitening produces substantial improvements on word-level benchmarks (WordSim-353 and SimLex-999). The most dramatic gains appear for GPT-2, where Spear- man ρ on WordSim-353 increases from 0.263 to 0.620 (∆ = +0.358,p < 0.001) and on SimLex-999 from 0.097 to 0.411 (∆ = +0.314, p < 0.001), suggesting that the original embedding space was so anisotropic that cosine similarity was largely non-functional and whitening restored much of its discriminative capacity. This pattern is consistent with Timkey and Van Schijndel (2021), who showed that cor- recting for rogue dimensions substantially improves cosine similarity as a measure of semantic association. RoBERTa and BERT show similar word-level gains (∆ = +0.054 to +0.191), while their STS-B scores are either unchanged (RoBERTa) or modestly decreased (BERT: ∆ =−0.055). For word2vec, which has the lowest initial anisotropy (mean cosine 0.059), the results are mixed and small in magnitude (|∆| ≤ 0.027), suggesting that whitening nei- ther substantially improves nor degrades an already near- isotropic space. The unsupervised contrastive models show essentially unchanged word-level scores but small, signifi- cant decreases on STS-B (∆ = −0.055 for unsup-BERT; ∆ = −0.034 for unsup-RoBERTa), possibly reflecting a domain mismatch between the reference corpus (WikiText- 103) and the data on which these models were originally optimized. Taken together, these results indicate that ZCA whiten- ing preserves the semantic structure of the embedding space. In many cases, it also enables cosine similarity to more faithfully capture semantic associations. This improvement is particularly pronounced for highly anisotropic models, where calibration is most needed. The modest STS-B de- creases warrant acknowledgment but are small relative to the word-level gains. Because WEAT measures bias through co- sine similarity, these findings support the interpretation that WEAT measurements in the calibrated space are more re- liable indicators of genuine semantic associations, particu- larly for models with high initial anisotropy. 5.3 Anisotropy Distorts WEAT in Both Directions Figure 3 visualizes WEAT effect sizes before and after ZCA whitening across all 70 model–task combinations, and Fig- ure 4 plots whitened versus raw effect sizes grouped by bias Table 3: Spearman ρ on semantic similarity benchmarks be- fore (Raw) and after (White) ZCA whitening. ∆ denotes the change. Significance: ∗ p < 0.001, ∗ p < 0.01, ∗ p < 0.05, n.s. = not significant (paired bootstrap, 10k iterations). ModelBench. Raw White∆ Sig. GloVe WS-353.477.586+.109 ∗ SL-999.297.380+.083 ∗ STS-B.560.599+.039 ∗ word2vec WS-353.686.661 −.025 ∗ SL-999.441.469+.027 ∗ STS-B.707.684 −.023 ∗ BERT WS-353.560.658+.098 ∗ SL-999.412.465+.054 ∗ STS-B.633.578 −.055 ∗ RoBERTa WS-353.446.559+.113 ∗ SL-999.266.457+.191 ∗ STS-B.650.648 −.002n.s. GPT-2 WS-353.263.620+.358 ∗ SL-999.097.411+.314 ∗ STS-B.428.596+.168 ∗ unsup- BERT WS-353.718.734+.016n.s. SL-999.536.545+.009n.s. STS-B.840.786 −.055 ∗ unsup- RoBERTa WS-353.536.532 −.004n.s. SL-999.436.407 −.029n.s. STS-B.845.811 −.034 ∗ type. Full numerical results, including effect sizes, p-values, and significance change classifications for all combinations, are provided in Appendix C. Across the 70 combinations, we observe 48 Stable cases (68.6%), 12 Disappearing cases (17.1%), and 10 Emerging cases (14.3%). Effect size shifts also occur within Stable cases. Among the Stable cases that remain significant in both spaces, 7 exhibit |∆d| > 0.5, including GPT-2 on W3 (d = 1.24 → 0.57) and unsup- RoBERTa on W6 (d = 1.61 → 0.99). In total, 29 of 70 combinations (41.4%) show either a significance change or a substantial effect-size shift (|∆d| > 0.5) after calibration, suggesting that anisotropy is a practical source of measure- ment error. Valence (W1–W2). Most valence measurements cluster near or above the diagonal in Figure 4(a), indicating that ef- fect sizes are preserved or increase after whitening. Static models retain significance throughout, though word2vec shows compression of inflated effects (W2: d = 1.63 → 1.03). The contextualized models show a markedly differ- ent pattern: four of six contextualized model–task combi- nations exhibit Emerging bias (BERT W2, RoBERTa W2, GPT-2 W1, GPT-2 W2). The most striking case is GPT-2 on W2 (d = −0.87 → d = 1.05). BERT and RoBERTa on W2, and GPT-2 on W1, follow the same pattern (see Ap- pendix C). Because valence associations such as flowers– pleasant and weapons–unpleasant are among the most ro- bust findings in the IAT literature (Greenwald, McGhee, and Schwartz 1998), their absence in uncalibrated contextual- WEAT 1WEAT 2WEAT 3WEAT 4WEAT 5WEAT 6WEAT 7WEAT 8WEAT 9WEAT 10 WEAT Test Suite GloVe word2vec BERT RoBERTa GPT-2 unsup- BERT unsup- RoBERTa ValenceRaceGenderHealth Age Static Contextualized Contrastive R W Upper-left: Raw (R) / Lower-right: Whitened (W) Significant (p < . 05): bold + * + border Not significant: no border 1.17* 1.18* 1.30* 1.16* 0.95* 0.24 1.32* 0.81* 1.32* 0.81* 1.72* 1.63* 1.28* 1.27* 1.09* 0.94* 1.30* 1.08* 0.75 0.39 1.54* 1.32* 1.63* 1.03* 0.67* 0.22 1.31* 0.78* 1.31* 0.78* 1.89* 1.81* 0.97* 0.09 1.14* 0.59 1.38* 1.14* 0.72 0.33 0.69* 1.10* 0.34 1.11* 0.55* -0.30 1.19* 0.31 1.18* 0.41 0.99* 1.21* 1.54* 1.51* 0.48 1.15* 0.92 0.83 1.45* 1.00* 0.87* 1.26* 0.11 0.81* 0.48* 0.60* 0.32 0.51 0.24 0.54 1.44* 1.17* 0.03 0.68 -0.28 0.50 0.16 -0.47 1.17* 1.00* 0.25 1.13* -0.87 1.05* 1.24* 0.57* 0.94* -0.33 0.91* -0.27 0.30 0.59 -0.58 0.69 -0.54 0.81 -0.01 1.55* 1.25* 1.15* 1.18* 1.26* 1.25* 1.34* 0.50* 0.32 0.21 0.43 0.22 0.46 1.55* 1.12* 0.47 1.18* 0.59 0.98* 0.96 1.38* -0.07 0.09 1.23* 1.23* 1.47* 0.15 0.73* 0.93* 0.86* 1.07* 0.65* 1.04* 1.61* 0.99* 0.10 0.52 -0.00 0.70 1.12* -0.78 0.49 0.92* 1.5 1.0 0.5 0.0 0.5 1.0 1.5 Effect size d (clipped to [-1.5, 1.5]) Figure 3: WEAT effect sizes (d) before (upper-left triangle) and after (lower-right triangle) ZCA whitening across seven models and ten test suites. Color intensity encodes effect size magnitude (clipped to [−1.5, 1.5]). Bold values with asterisks and bor- dered cells indicate statistical significance (p < 0.05); regular-weight values without borders indicate non-significant results. Models are grouped by architectural family: static (top), contextualized (middle), and unsupervised contrastive (bottom). Test suites are grouped by bias type: Valence (W1–W2), Race (W3–W5), Gender (W6–W8), Health (W9), and Age (W10). ized embeddings likely reflects geometric distortion rather than a genuine lack of association. The contrastive mod- els are mostly Stable, with one Disappearing case (unsup- RoBERTa W2: d = 1.47→ 0.15). Race (W3–W5). Race measurements are predominantly located below the diagonal in Figure 4(b), indicating sys- tematic decreases in effect size after whitening. This cate- gory concentrates the bulk of overestimation: 8 of 12 Disap- pearing cases (66.7%) occur in W3–W5. Both static models show Disappearing bias on W3 (e.g., GloVe: d = 0.95 → 0.24), with reduced but still significant effects on W4 and W5. Among the contextualized models, BERT exhibits Dis- appearing bias across all three race tests, and GPT-2 on W4 and W5 (full transitions in Appendix C); RoBERTa is the exception, maintaining stable significance throughout. The contrastive models show only one Disappearing case within the race tests (unsup-BERT W3), and unsup-RoBERTa even shows slightly increased race effects after whitening (W4: d = 0.86→ 1.07). Gender (W6–W8). Gender measurements show consid- erable scatter around the diagonal in Figure 4(c), reflect- ing coexistence of overestimated and underestimated bias. Static models retain strong significance on W6 (word2vec: d = 1.89→ 1.81; GloVe: d = 1.72→ 1.63), but word2vec shows Disappearing cases on W7 and W8. The contextu- alized models reveal a more nuanced pattern: only one of nine combinations shows Emerging bias (BERT on W8: d = 0.48 → 1.15). GPT-2 shows directional reversals on W7 (d = −0.58 → 0.69) and W8 (d = −0.54 → 0.81), where raw associations flip from negative to positive after whitening. This suggests that severe anisotropy can invert the apparent sign of association. The contrastive models pro- vide further evidence: unsup-BERT shows Emerging bias on both W7 (d = 0.47 → 1.18) and W8 (d = 0.59 → 0.98). These emerging effect sizes are comparable to or larger than those in the standard contextualized models, suggesting that while contrastive training reduces geometric distortion, it does not eliminate the bias encoded in the pre-training data. Health and Age (W9–W10). For health (W9), the most notable result is GPT-2’s Emerging case (d =−0.01→ d = 1.55), where whitening uncovers a strong association that was geometrically masked in the raw space. Unsup-BERT also shows an Emerging case and unsup-RoBERTa a Disap- pearing one on W9, while static models show mixed patterns (see Appendix C). For age (W10), six of seven models show Stable results, with one Emerging case (unsup-RoBERTa: d = 0.49→ d = 0.92); the contextualized models maintain large significant effects (d > 1.0 after whitening). 5.4 The Role of Anisotropy Severity The frequency of significance changes appears broadly con- sistent with the degree of initial anisotropy: GPT-2 (mean cosine 0.997) shows changes in 5 of 10 tests, BERT (0.821) 10123 Effect size d (Raw) 1.5 1.0 0.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Effect size d (Whitened) Overestimated Underestimated (a) Valence 10123 Effect size d (Raw) 1.5 1.0 0.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Overestimated Underestimated (b) Race 10123 Effect size d (Raw) 1.5 1.0 0.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Overestimated Underestimated (c) Gender 10123 Effect size d (Raw) 1.5 1.0 0.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Overestimated Underestimated (d) Other (Health & Age) Disappearing Stable Emerging Static Contextualized Contrastive Figure 4: Scatter plots of WEAT effect sizes before (x-axis) and after (y-axis) ZCA whitening, separated by bias type: (a) Va- lence, (b) Race, (c) Gender, and (d) Other (Health & Age). The diagonal line represents no change; points above the line indicate increased effect sizes after whitening (underestimation in the raw space), and points below indicate decreased effect sizes (over- estimation). Marker color denotes significance change: Disappearing (red), Stable (pink), and Emerging (blue). Marker shape denotes architectural family: circle (static), triangle (contextualized), and square (contrastive). in 5, unsup-BERT (0.386) in 4, while word2vec (0.059) shows changes in 3 of 10 tests. However, RoBERTa (0.954) shows only 1 change despite its high anisotropy, and GloVe (0.192) likewise shows only 1 change, indicating that the re- lationship between anisotropy severity and the frequency of significance changes is not strictly monotonic and may de- pend on additional factors such as the specific geometry of the embedding space relative to the WEAT stimulus words. 5.5 Patterns of Distortion Across Bias Categories and Embedding Families The category-level analysis reveals two systematic asymme- tries, summarized visually in Figure 4. First, distortion di- rection depends on bias category: Disappearing cases con- centrate in the race tests, which account for 8 of 12 (66.7%), while Emerging cases are most frequent in the valence tests (4 of 10), with the remainder split evenly between gender (3) and health/age (3). This asymmetry suggests that anisotropy may interact differently with different stimulus configura- tions. However, the specific mechanism behind this pattern is not fully clear from our experimental results alone. One possibility is that the geometric relationship between target sets and the dominant variance directions of the embedding space differs across bias categories. As a result, anisotropy may amplify cosine differences in some cases and compress them in others. Identifying the precise factors that determine the direction of distortion is an important direction for future work. Second, distortion direction depends on architectural fam- ily. The contextualized models, which exhibit the highest anisotropy (mean cosine 0.821–0.997), account for 6 of 10 Emerging cases (60.0%), consistent with the interpreta- tion that severe anisotropy compresses the cosine similar- ity range until genuine differential associations fall below the significance threshold. The static models, despite their relatively low anisotropy, contribute 4 of 12 Disappearing cases (33.3%), indicating that even moderate distortion can inflate specific measurements. The contrastive models ex- hibit both directions of distortion across categories, indicat- ing that contrastive training reduces geometric distortion but neither prevents it uniformly nor eliminates the bias encoded in the pre-training data. 6 Discussion 6.1 Implications for Computational Social Science Our findings have direct implications for intrinsic bias measurement. The observation that approximately 30% of WEAT measurements change significance status after geo- metric calibration suggests that a substantial fraction of as- sessments based on uncalibrated embeddings may be un- reliable. Practitioners conducting bias audits should there- fore perform an isotropy check before applying cosine-based measurements and calibrate when anisotropy is detected. The WEAT has been widely used in computational so- cial science to study societal biases that are difficult or costly to measure experimentally, including historical shifts in stereotypes over 200 years (Charlesworth, Caliskan, and Banaji 2022), 100 years of gender and ethnic stereotypes (Garg et al. 2018), gender stereotypes across 25 languages (Lewis and Lupyan 2020), and intersectional stereotypes (Charlesworth et al. 2024a). Our results suggest that con- clusions drawn from highly anisotropic models may need to be revisited with calibrated measurements to deter- mine whether reported biases were overestimated, underesti- mated, or accurately captured. This does not invalidate prior findings; rather, geometric calibration offers a tool to assess their robustness. The bidirectional nature of the distortion makes this con- cern particularly pressing. If anisotropy only inflated bias measurements, uncalibrated results could be interpreted as upper bounds. However, because anisotropy can also mask genuine biases, uncalibrated results cannot be treated as con- servative estimates either. Our work parallels methodological debates in psychology about how implicit bias should be measured. Just as Green- wald et al. (2022) provided recommended best practices for research using the IAT, we argue that computational bias measures require analogous methodological discipline, in- cluding correction for known geometric distortions. More- house, Swaroop, and Pan (2025) likewise called for import- ing social science best practices into LLM bias probing. Our ZCA calibration addresses one specific aspect of measure- ment invariance: ensuring that WEAT scores are comparable across models with differing degrees of anisotropy. The importance of calibrating intrinsic bias measurements is reinforced by growing evidence that such biases propagate to downstream behavior. EAT-measured biases in vision- language models correlate with downstream task perfor- mance and propagate to zero-shot retrieval (Ghate et al. 2025), image classification (Wolfe, Banaji, and Caliskan 2022b), visual question answering, captioning, and gener- ation (Wolfe and Caliskan 2022; Friedrich et al. 2025), and sentiment classification (Mei, Fereidooni, and Caliskan 2023). Biased AI outputs have also been shown to shape humans’ own implicit associations (Sim et al. 2025), estab- lishing a pathway from embedding-level bias to changes in human implicit associations. If intrinsic measurements are distorted by anisotropy, the resulting inaccuracies may lead to erroneous assessments of downstream risk, making geo- metric calibration a prerequisite for reliable bias auditing. As AI systems come under regulatory oversight, bias au- diting tools must meet high standards of reliability. The EU AI Act (European Parliament and Council of the European Union 2024) mandates bias examination as part of its data governance requirements for datasets used in high-risk AI systems, and the NIST AI Risk Management Framework (National Institute of Standards and Technology 2023) em- phasizes valid and reliable measurement in AI risk assess- ment. ZCA calibration could be integrated into such frame- works as a standard preprocessing step for embedding-based bias measurement, ensuring comparable audit results across models. 6.2 Compatibility with EAT Variants ZCA whitening operates as a pre-processing calibration step that is agnostic to the specific form of the association test applied afterward. Because it transforms only the embed- ding space geometry without modifying the test procedure itself, it is compatible with any cosine-similarity-based bias metric. It can thus be applied to sentence embeddings in SEAT (May et al. 2019), contextualized vectors in CEAT (Guo and Caliskan 2021), image embeddings in iEAT (Steed and Caliskan 2021), and speech representations in SpEAT (Slaughter et al. 2023). For ML-EAT (Wolfe, Hiniker, and Howe 2024), our approach provides a corrective step that complements its anisotropy diagnostic: when ML-EAT iden- tifies a model as highly anisotropic, ZCA whitening can restore the conditions under which cosine-based evalua- tion becomes meaningful. More broadly, any future variant of embedding association tests relying on cosine similarity would benefit from the same calibration procedure. 6.3 Limitations Several limitations should be acknowledged. First, the whitening matrix is estimated from WikiText-103, and dif- ferent reference corpora may yield different transforma- tions; the sensitivity of WEAT outcomes to the choice and size of the reference corpus warrants systematic investiga- tion. Second, our evaluation is limited to seven models, and whether the calibration behavior we observe generalizes to a broader range of models requires further investigation. Third, our implementation uses a fixed regularization param- eter ε for numerical stability, and we do not provide a sen- sitivity analysis of how this choice affects calibrated WEAT scores. Fourth, the relationship between calibrated WEAT measurements and downstream task bias remains an open question. Prior work has reported mixed findings: Goldfarb- Tarrant et al. (2021) found frequent divergence, while Ghate et al. (2025) demonstrated correlation in vision-language models. Whether calibration improves this relationship re- quires future investigation. Addressing these open questions constitutes an important direction for future research. Beyond these limitations, we emphasize the scope of our contribution. This work evaluates the geometric reliability of WEAT measurements. We examine whether cosine sim- ilarity is a valid metric within the embedding space, not how the associations it measures shape the outputs of a de- ployed system. We do not evaluate the alignment between WEAT measurements and human implicit association mea- sures. The latter is a separate question that applies equally to raw and calibrated measurements and has been examined in prior work (Caliskan, Bryson, and Narayanan 2017; More- house et al. 2023; Charlesworth et al. 2024b). Extending hu- man alignment analysis to calibrated measurements is a nat- ural next step. 7 Conclusion This research proposed ZCA whitening as a pre-processing calibration step for the Word Embedding Association Test (WEAT). The method transforms the covariance of the em- bedding space into the identity matrix while minimizing per- turbation to the original vectors, restoring the isotropy con- dition on which WEAT relies. We evaluated our approach on ten WEAT test suites and seven models (70 model–task combinations). ZCA whitening reduced anisotropy across all models, and for highly anisotropic models it also im- proved standard semantic similarity benchmark scores. Af- ter calibration, over 30% of WEAT results changed signif- icance status, and effect sizes shifted in both directions de- pending on bias category. These results suggest that bias measurements obtained in strongly anisotropic spaces may partly reflect geometric properties of the space rather than social associations alone. We do not claim that previously reported findings are in- valid. Rather, studies conducted under such conditions may benefit from re-examination with a calibrated metric. ZCA whitening operates as a model-agnostic pre- processing step. It can be applied not only to WEAT but also to its variants, such as CEAT, iEAT, and SpEAT, and in prin- ciple to any bias metric that relies on cosine similarity. The reliability of such measurements is inseparable from the ge- ometry of the space in which they are computed, and our approach offers a practical step toward geometrically valid bias measurement in computational social science and AI fairness research. Ethical Statement This work aims to improve the reliability of bias measure- ment in word embeddings. However, several ethical consid- erations deserve attention. Risk of misinterpretation. Our finding that some previ- ously significant bias measurements may not survive geo- metric calibration could be misinterpreted as evidence that AI systems are less biased than previously reported. We cau- tion against this interpretation. Calibration reveals that while some specific measurements were inflated, others were de- flated, and the overall picture is one of measurement unre- liability rather than bias absence. The appropriate response is to re-measure with calibrated tools, not to conclude that bias is less prevalent. This concern is particularly acute given the bidirectional nature of the distortion we document: prac- titioners, journalists, or model developers who selectively cite Disappearing cases as evidence of model improvement would be misrepresenting our findings. As Blodgett et al. (2020) argued, NLP bias research must be grounded in clear normative reasoning about who is harmed and how; our cal- ibration method addresses one specific source of measure- ment error but does not resolve deeper questions about what constitutes bias or how bias measurements should inform practice. Scope of intrinsic bias metrics. Even with proper geo- metric calibration, intrinsic bias measurements capture only one dimension of the complex ways in which AI sys- tems can perpetuate social inequities. Embedding-level bias metrics should be used alongside downstream task evalua- tions (Goldfarb-Tarrant et al. 2021; Cabello, Jørgensen, and Søgaard 2023), qualitative audits, and participatory assess- ments involving affected communities (Selbst et al. 2019). The communities most likely to be harmed by biased AI systems include racial, gender, and other minoritized groups represented in WEAT target sets. Their perspectives should inform what counts as a meaningful bias measurement in the first place. Responsibility in auditing contexts. As discussed in Sec- tion 6.1, embedding-based bias measurements are increas- ingly used in regulatory and auditing contexts. This dual-use character introduces specific responsibilities: a calibration method that changes which biases are flagged as significant could, if applied uncritically, either expose previously hid- den harms or obscure documented ones. We therefore rec- ommend that calibrated measurements be reported alongside uncalibrated ones rather than replacing them, so that audit trails preserve both perspectives and allow stakeholders to assess the geometric reliability of each measurement. Acknowledgments We thank the reviewers of AIES 2026 for their thoughtful and constructive feedback. We also thank our colleagues for helpful discussions throughout this work. References Bailey, A. H.; Williams, A.; and Cimpian, A. 2022. Based on Billions of Words on the Internet, People = Men. Science Advances, 8(13): eabm2463. Blodgett, S. L.; Barocas, S.; Daum ́ e I, H.; and Wallach, H. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5454–5476. Online: Association for Computational Linguistics. Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016.Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Ad- vances in Neural Information Processing Systems, 29. Bommasani, R.; Davis, K.; and Cardie, C. 2020. Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4758–4781. Online: Association for Computational Linguis- tics. Borenstein, N.; Sta ́ nczak, K.; Rolskov, T.; K ̈ afer, N. K.; da Silva Perez, N.; and Augenstein, I. 2023. Measuring in- tersectional biases in historical documents. In Findings of the Association for Computational Linguistics: ACL 2023, 2711–2730. Cabello, L.; Jørgensen, A. K.; and Søgaard, A. 2023. On the independence of association bias and empirical fairness in language models. In Proceedings of the 2023 ACM con- ference on fairness, accountability, and transparency, 370– 378. Cai, X.; Huang, J.; Bian, Y.; and Church, K. 2021. Isotropy in the Contextual Embedding Space: Clusters and Mani- folds. In International Conference on Learning Represen- tations. Caliskan, A.; Ajay, P. P.; Charlesworth, T.; Wolfe, R.; and Banaji, M. R. 2022. Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Seman- tics. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, 156–170. Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017. Se- mantics derived automatically from language corpora con- tain human-like biases. Science, 356(6334): 183–186. Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L. 2017. SemEval-2017 Task 1: Semantic Textual Similar- ity Multilingual and Crosslingual Focused Evaluation. In Proceedings of the 11th International Workshop on Seman- tic Evaluation (SemEval-2017), 1–14. Vancouver, Canada: Association for Computational Linguistics. Charlesworth, T. E.; Caliskan, A.; and Banaji, M. R. 2022. Historical representations of social groups across 200 years of word embeddings from Google Books. Proceedings of the National Academy of Sciences, 119(28): e2121798119. Charlesworth, T. E.; Ghate, K.; Caliskan, A.; and Banaji, M. R. 2024a. Extracting intersectional stereotypes from em- beddings: Developing and validating the Flexible Intersec- tional Stereotype Extraction procedure. PNAS Nexus, 3(3): pgae089. Charlesworth, T. E.; Yang, V.; Mann, T. C.; Kurdi, B.; and Banaji, M. R. 2021. Gender stereotypes in natural language: Word embeddings show robust consistency across child and adult language corpora of more than 65 million words. Psy- chological Science, 32(2): 218–240. Charlesworth, T. E. S.; Morehouse, K.; Rouduri, V.; and Cunningham, W. A. 2024b. Echoes of Culture: Relation- ships of Implicit and Explicit Attitudes with Contemporary English, Historical English, and 53 Non-English Languages. Social Psychological and Personality Science, 15(7): 812– 823. Cobert, J.; Mills, H.; Lee, A.; Gologorskaya, O.; Espejo, E.; Jeon, S. Y.; Boscardin, W. J.; Heintz, T. A.; Kennedy, C. J.; Ashana, D. C.; Chapman, A. C.; and Lee, S. J. 2024. Mea- suring Implicit Bias in ICU Notes Using Word-Embedding Neural Network Models. Chest, 165(6): 1481–1490. Cohen, J. 1992. A Power Primer. Psychological Bulletin, 112(1): 155–159. Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. Minneapolis, Min- nesota: Association for Computational Linguistics. Dutta, S.; Srivastava, P.; Solunke, V.; Nath, S.; and Khud- aBukhsh, A. R. 2023. Disentangling Societal Inequality from Model Biases: Gender Inequality in Divorce Court Proceedings. In Proceedings of the Thirty-Second Interna- tional Joint Conference on Artificial Intelligence (IJCAI-23), 5959–5967. Ethayarajh, K. 2019. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 conference on empirical methods in natural language pro- cessing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), 55–65. European Parliament and Council of the European Union. 2024. Regulation (EU) 2024/1689 of the European Par- liament and of the Council of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (AI Act). Of- ficial Journal of the European Union, L 1689. https://eur- lex.europa.eu/eli/reg/2024/1689/oj. Accessed: 2026-08-04. Finkelstein, L.; Gabrilovich, E.; Matias, Y.; Rivlin, E.; Solan, Z.; Wolfman, G.; and Ruppin, E. 2002.Placing Search in Context: The Concept Revisited. ACM Transac- tions on Information Systems, 20(1): 116–131. Friedrich, F.; Brack, M.; Struppek, L.; Hintersdorf, D.; Schramowski, P.; Luccioni, S.; and Kersting, K. 2025. Au- diting and Instructing Text-to-Image Generation Models on Fairness. AI and Ethics, 5(3): 2103–2123. Gao, T.; Yao, X.; and Chen, D. 2021. SimCSE: Simple Con- trastive Learning of Sentence Embeddings. In Moens, M.-F.; Huang, X.; Specia, L.; and Yih, S. W.-t., eds., Proceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, 6894–6910. Online and Punta Cana, Do- minican Republic: Association for Computational Linguis- tics. Garg, N.; Schiebinger, L.; Jurafsky, D.; and Zou, J. 2018. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sci- ences, 115(16): E3635–E3644. Ghate, K.; Charlesworth, T.; Diab, M. T.; and Caliskan, A. 2025. Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the As- sociation for Computational Linguistics: ACL 2025, 18562– 18580. Vienna, Austria: Association for Computational Lin- guistics. ISBN 979-8-89176-256-5. Godey, N.; de la Clergerie, ́ E.; and Sagot, B. 2024. Anisotropy Is Inherent to Self-Attention in Transformers. In Graham, Y.; and Purver, M., eds., Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 35–48. St. Julian’s, Malta: Association for Computational Linguis- tics. Goldfarb-Tarrant, S.; Marchant, R.; Mu ̃ noz S ́ anchez, R.; Pandya, M.; and Lopez, A. 2021. Intrinsic Bias Metrics Do Not Correlate with Application Bias. In Proceedings of the 59th Annual Meeting of the Association for Compu- tational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Pa- pers), 1926–1940. Association for Computational Linguis- tics. Gray, M.; and Wu, L. 2025. Benchmarking Bias in Embed- dings of Healthcare AI Models: Using SD-WEAT for Detec- tion and Measurement Across Sensitive Populations. BMC Medical Informatics and Decision Making, 25(1): 258. Greenwald, A. G.; Brendl, M.; Cai, H.; Cvencek, D.; Do- vidio, J. F.; Friese, M.; Hahn, A.; Hehman, E.; Hofmann, W.; Hughes, S.; Hussey, I.; Jordan, C.; Kirby, T. A.; Lai, C. K.; Lang, J. W. B.; Lindgren, K. P.; Maison, D.; Ostafin, B. D.; Rae, J. R.; Ratliff, K. A.; Spruyt, A.; and Wiers, R. W. 2022. Best Research Practices for Using the Implicit Association Test. Behavior Research Methods, 54(3): 1161–1180. Greenwald, A. G.; McGhee, D. E.; and Schwartz, J. L. 1998. Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6): 1464. Guo, W.; and Caliskan, A. 2021. Detecting emergent in- tersectional biases: Contextualized word embeddings con- tain a distribution of human-like biases. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 122–133. Hill, F.; Reichart, R.; and Korhonen, A. 2015. SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Es- timation. Computational Linguistics, 41(4): 665–695. Huang, J.; Tang, D.; Zhong, W.; Lu, S.; Shou, L.; Gong, M.; Jiang, D.; and Duan, N. 2021. WhiteningBERT: An easy unsupervised sentence embedding approach. In Findings of the association for computational linguistics: EMNLP 2021, 238–244. Kessy, A.; Lewin, A.; and Strimmer, K. 2018. Optimal whitening and decorrelation.The American Statistician, 72(4): 309–314. Kurita, K.; Vyas, N.; Pareek, A.; Black, A. W.; and Tsvetkov, Y. 2019. Measuring bias in contextualized word representations. In Proceedings of the first workshop on gen- der bias in natural language processing, 166–172. Leach, S.; Kitchin, A. P.; and Sutton, R. M. 2023. Word em- beddings reveal growing moral concern for people, animals and the environment. British Journal of Social Psychology, 62(4): 1925–1938. Lewis, M.; and Lupyan, G. 2020. Gender stereotypes are re- flected in the distributional structure of 25 languages. Nature Human Behaviour, 4(10): 1021–1028. Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692. Machina, A.; and Mercer, R. 2024. Anisotropy is Not Inher- ent to Transformers. In Duh, K.; Gomez, H.; and Bethard, S., eds., Proceedings of the 2024 Conference of the North Amer- ican Chapter of the Association for Computational Linguis- tics: Human Language Technologies (Volume 1: Long Pa- pers), 4892–4907. Mexico City, Mexico: Association for Computational Linguistics. Manzini, T.; Yao Chong, L.; Black, A. W.; and Tsvetkov, Y. 2019. Black is to Criminal as Caucasian is to Police: De- tecting and Removing Multiclass Bias in Word Embeddings. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceed- ings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 615–621. Minneapolis, Minnesota: Association for Compu- tational Linguistics. Matthews, S.; Hudzina, J.; and Sepehr, D. 2022. Gender and racial stereotype detection in legal opinion word embed- dings. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 12026–12033. May, C.; Wang, A.; Bordia, S.; Bowman, S.; and Rudinger, R. 2019. On measuring social biases in sentence encoders. In Proceedings of the 2019 Conference of the North Ameri- can Chapter of the Association for Computational Linguis- tics: Human Language Technologies, Volume 1 (Long and Short Papers), 622–628. Mei, K.; Fereidooni, S.; and Caliskan, A. 2023. Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks. In Proceed- ings of the 2023 ACM Conference on Fairness, Accountabil- ity, and Transparency (FAccT ’23), 1699–1710. Chicago, IL, USA: ACM. Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017. Pointer Sentinel Mixture Models. In Proceedings of the 5th International Conference on Learning Representations (ICLR 2017). Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G.; and Dean, J. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Informa- tion Processing Systems, volume 26. Morehouse, K.; Swaroop, S.; and Pan, W. 2025. Position: Rethinking LLM Bias Probing Using Lessons from the So- cial Sciences. In Proceedings of the 42nd International Con- ference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 81841–81860. PMLR. Morehouse, K. N.; Charlesworth, T. E. S.; Rouduri, V.; and Cunningham, W. A. 2023.Traces of Human Attitudes in Contemporary and Historical Word Embeddings (1800– 2000). Research Square. Preprint. Mu, J.; and Viswanath, P. 2018. All-but-the-Top: Simple and Effective Post-processing for Word Representations. In Proceedings of the 6th International Conference on Learn- ing Representations (ICLR 2018). Mukherjee, A.; Raj, C.; Zhu, Z.; and Anastasopoulos, A. 2023.Global voices, local biases: Socio-cultural preju- dices across languages. In Proceedings of the 2023 Confer- ence on Empirical Methods in Natural Language Process- ing, 15828–15845. Napp, C. 2023. Gender stereotypes embedded in natural lan- guage are stronger in more economically developed and in- dividualistic countries. PNAS Nexus, 2(11): pgad355. National Institute of Standards and Technology. 2023. Ar- tificial Intelligence Risk Management Framework (AI RMF 1.0). Technical Report NIST AI 100-1, National Institute of Standards and Technology. https://doi.org/10.6028/NIST. AI.100-1. Pennington, J.; Socher, R.; and Manning, C. D. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, 1532–1543. Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from nat- ural language supervision. In International Conference on Machine Learning, 8748–8763. PMLR. Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019.Language Models are Unsupervised Multitask Learners. OpenAI blog, 1(8): 9. Rajaee, S.; and Pilehvar, M. T. 2022. An Isotropy Analy- sis in the Multilingual BERT Embedding Space. In Mure- san, S.; Nakov, P.; and Villavicencio, A., eds., Findings of the Association for Computational Linguistics: ACL 2022, 1309–1316. Dublin, Ireland: Association for Computational Linguistics. Rios, A.; Joshi, R.; and Shin, H. 2020. Quantifying 60 Years of Gender Bias in Biomedical Research with Word Embed- dings. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing (BioNLP 2020), 1–13. As- sociation for Computational Linguistics. Ross, C.; Katz, B.; and Barbu, A. 2021. Measuring social biases in grounded vision and language embeddings. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 998–1008. Rudman, W.; Gillman, N.; Rayne, T.; and Eickhoff, C. 2022. IsoScore: Measuring the Uniformity of Embedding Space Utilization. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Findings of the Association for Computational Linguis- tics: ACL 2022, 3325–3339. Dublin, Ireland: Association for Computational Linguistics. Schmahl, K. G.; Viering, T. J.; Makrodimitris, S.; Jahfari, A. N.; Tax, D.; and Loog, M. 2020. Is Wikipedia succeed- ing in reducing gender bias? Assessing changes in gender bias in Wikipedia using word embeddings. In Proceedings of the fourth workshop on natural language processing and computational social science, 94–103. Selbst, A. D.; Boyd, D.; Friedler, S. A.; Venkatasubrama- nian, S.; and Vertesi, J. 2019. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency, 59–68. Sim, M.; Brigham, N. G.; Kohno, T.; Charlesworth, T. E. S.; and Caliskan, A. 2025. Biased AI Outputs Can Impact Hu- mans’ Implicit Bias: A Case Study of the Impact of Gender- Biased Text-to-Image Generators. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), volume 8, 2375–2386. Slaughter, I.; Greenberg, C.; Schwartz, R.; and Caliskan, A. 2023. Pre-trained speech processing models contain human- like biases that propagate to speech emotion recognition. In Findings of the Association for Computational Linguistics: EMNLP 2023, 8967–8989. Steed, R.; and Caliskan, A. 2021. Image representations learned with unsupervised pre-training contain human-like biases. In Proceedings of the 2021 ACM conference on fair- ness, accountability, and transparency, 701–713. Su, J.; Cao, J.; Liu, W.; and Ou, Y. 2021. Whitening Sen- tence Representations for Better Semantics and Faster Re- trieval. arXiv:2103.15316. Sun, Y.; Saxon, M.; Yang, I.; Gueorguieva, A.-M.; and Caliskan, A. 2026. VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation. arXiv:2601.00996. Sunsay, C. 2023.A historical evaluation of the disease avoidance theory of xenophobia.PLOS ONE, 18(12): e0294816. Timkey, W.; and Van Schijndel, M. 2021. All bark and no bite: Rogue dimensions in transformer language mod- els obscure representational quality. In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, 4527–4546. Wolfe, R.; Banaji, M. R.; and Caliskan, A. 2022a. Evidence for hypodescent in visual semantic AI. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1293–1304. Wolfe, R.; Banaji, M. R.; and Caliskan, A. 2022b. Marked- ness in Visual Semantic AI. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1269–1279. Wolfe, R.; and Caliskan, A. 2022. American = White in Multimodal Language-and-Image AI. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, 800–812. Wolfe, R.; Hiniker, A.; and Howe, B. 2024.ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science.In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 1608–1620. Zhou, K.; Ethayarajh, K.; Card, D.; and Jurafsky, D. 2022. Problems with Cosine as a Measure of Embedding Similar- ity for High Frequency Words. In Proceedings of the 60th Annual Meeting of the Association for Computational Lin- guistics (Volume 2: Short Papers), 401–423. Dublin, Ireland: Association for Computational Linguistics. Supplementary Material for: Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests A Derivation of the Identity Covariance Property of ZCA-Whitened Vectors This appendix provides the complete derivation showing that the ZCA-whitened vectors z i = W ZCA (x i − μ) have identity covariance, as summarized in Equation 10 of the main paper. Starting from the definition of the sample covariance of z i : Cov(z) = 1 n− 1 n X i=1 z i z ⊤ i = 1 n− 1 n X i=1 W ZCA (x i − μ) W ZCA (x i − μ) ⊤ = 1 n− 1 n X i=1 W ZCA (x i − μ)(x i − μ) ⊤ W ⊤ ZCA = W ZCA 1 n− 1 n X i=1 (x i − μ)(x i − μ) ⊤ ! W ⊤ ZCA = W ZCA ΣW ⊤ ZCA = Σ −1/2 Σ Σ −1/2 = I.(A1) The third equality uses the transpose identity (AB) ⊤ = B ⊤ A ⊤ . The sixth equality substitutes W ZCA = Σ −1/2 , using the fact that this matrix is symmetric (so W ⊤ ZCA = W ZCA ). The final equality follows from Σ −1/2 Σ Σ −1/2 = Σ −1/2 Σ 1/2 Σ 1/2 Σ −1/2 = I. B Pairwise Cosine Similarity Heatmaps for All Models Figures A1–A3 show pairwise cosine similarity matrices for 2,000 randomly sampled embedding vectors, grouped by ar- chitectural family. Across all seven models, ZCA whitening substantially reduces pairwise similarity regardless of the initial anisotropy severity. GloVe Mean: 0.1918, SD: 0.2114Mean: 0.0402, SD: 0.1128 word2vec Mean: 0.0589, SD: 0.0739Mean: 0.0130, SD: 0.0669 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 Before WhiteningAfter ZCA Whitening Figure A1: Pairwise cosine similarity matrices for static em- bedding models (GloVe, word2vec) before (left) and after (right) ZCA whitening. BERT Mean: 0.8212, SD: 0.0677Mean: 0.0015, SD: 0.0428 RoBERTa Mean: 0.9539, SD: 0.0237Mean: 0.0014, SD: 0.0477 GPT-2 Mean: 0.9967, SD: 0.0032Mean: 0.0013, SD: 0.0417 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 Before WhiteningAfter ZCA Whitening Figure A2: Pairwise cosine similarity matrices for contextu- alized models (BERT, RoBERTa, GPT-2) before (left) and after (right) ZCA whitening. unsup-BERT Mean: 0.3861, SD: 0.0942Mean: 0.0013, SD: 0.0415 unsup-RoBERTa Mean: 0.4746, SD: 0.0896Mean: 0.0012, SD: 0.0410 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 Before WhiteningAfter ZCA Whitening Figure A3: Pairwise cosine similarity matrices for con- trastive models (unsup-BERT, unsup-RoBERTa) before (left) and after (right) ZCA whitening. C Full WEAT Results Table A1: WEAT effect sizes (d) and permutation p-values (100,000 permutations, one-sided) before (Raw) and after (White) ZCA whitening across seven models and ten test suites (see Table 1 for full definitions). Significance at p < 0.05. S = Stable, D = Disappearing (significant → non-significant), E = Emerging (non-significant → significant). Bold d values indicate p < 0.05. p-values smaller than 10 −4 are reported as their order of magnitude. Static Embeddings GloVeword2vec ID Biasd R p R d W p W Chgd R p R d W p W Chg W1Valence 1.17 10 −5 1.18 10 −5 S 1.54 10 −5 1.32 10 −5 S W2Valence 1.30 10 −5 1.16 10 −5 S 1.63 10 −5 1.03 10 −5 S W3Race0.95 10 −5 0.24 .1660D 0.67 .0031 0.22 .1989D W4Race1.32 10 −5 0.81 .0098S 1.31 10 −5 0.78 .0125S W5Race1.32 10 −5 0.81 .0095S 1.31 10 −5 0.78 .0131S W6Gender 1.72 10 −5 1.63 10 −5 S 1.89 10 −5 1.81 10 −5 S W7Gender 1.28 .0020 1.27 .0023S 0.97 .0233 0.09 .4353D W8Gender 1.09 .0091 0.94 .0280S 1.14 .0100 0.59 .1263D W9Health 1.30 .0069 1.08 .0254S 1.38 .0025 1.14 .0148S W10 Age0.75 .0691 0.39 .2294S0.72 .0769 0.33 .2636S Contextualized Embeddings GPT-2BERTRoBERTa ID Biasd R p R d W p W Chgd R p R d W p W Chgd R p R d W p W Chg W1Valence0.25 .1869 1.13 10 −5 E 0.69 .0063 1.10 10 −5 S0.87 .0008 1.26 10 −5 S W2Valence −0.87 .9993 1.05 10 −5 E0.34 .1211 1.11 10 −5 E0.11 .3486 0.81 .0018E W3Race1.24 10 −5 0.57 .0109S 0.55 .0142 −0.30 .8829D0.48 .0272 0.60 .0082S W4Race0.94 .0022 −0.33 .8209D 1.19 10 −5 0.31 .1886D0.32 .18670.51 .0772S W5Race0.91 .0032 −0.27 .7706D 1.18 .00010.41 .1256D0.24 .25300.54 .0644S W6Gender0.30 .29010.59 .1259S 0.99 .0217 1.21 .0051S1.44 .0008 1.17 .0077S W7Gender −0.58 .86410.69 .0867S 1.54 .0003 1.51 .0007S0.03 .48030.68 .0880S W8Gender −0.54 .85310.81 .0525S0.48 .1742 1.15 .0084E −0.28 .71130.50 .1658S W9Health −0.01 .5044 1.55 .0011E0.92 .05730.83 .0800S0.16 .3974 −0.47 .7795S W10 Age1.25 .0057 1.15 .0058S 1.45 .0004 1.00 .0208S1.17 .0060 1.00 .0225S Unsupervised Contrastive Embeddings unsup-BERTunsup-RoBERTa ID Biasd R p R d W p W Chgd R p R d W p W Chg W1Valence 1.18 .0008 1.26 10 −5 S1.23 10 −5 1.23 10 −5 S W2Valence 1.25 10 −5 1.34 10 −5 S1.47 10 −5 0.15 .2990D W3Race0.50 .0220 0.32 .1002D0.73 .0014 0.93 10 −5 S W4Race0.21 .2792 0.43 .1132S0.86 .0059 1.07 .0007S W5Race0.22 .2725 0.46 .1006S0.65 .0315 1.04 .0009S W6Gender 1.55 .0005 1.12 .0099S1.61 10 −4 0.99 .0218S W7Gender0.47 .1887 1.18 .0055E0.10 .42360.52 .1553S W8Gender0.59 .1277 0.98 .0185E −0.00 .50380.70 .0864S W9Health0.96 .0517 1.38 .0020E1.12 .0228 −0.78 .9073D W10 Age −0.07 .5591 0.09 .4343S0.49 .18230.92 .0314E