Paper deep dive
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
Qingjie Zhang, Ziqi Tang, Jie Zhang, Gelei Deng, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Tianwei Zhang, Han Qiu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/16/2026, 3:06:48 AM
Summary
This paper introduces SAMPLED-BPE, a lightweight, token-level auditing pipeline designed to detect and quantify web pollution in large-scale Chinese corpora. By sampling a small subset of data and training a BPE tokenizer, the method efficiently surfaces polluted tokens without full corpus scanning. The authors apply this pipeline to 11 open Chinese corpora and 6 Common Crawl snapshots (2021-2026), revealing widespread but uneven pollution, with OSCAR and mC4 showing high contamination rates. The study also releases a hierarchical dataset of 630k+ token records to support transparency and tracing of internet pollution.
Entities (25)
Relation Signals (18)
OSCAR → hashighpollution → Adult Content
confidence 95% · OSCAR is dominated by Adult Content, which accounts for 76.10%.
mC4 → hashighpollution → Online Gambling
confidence 95% · mC4 has a different profile: Online Gambling alone reaches 23.17%.
SAMPLED-BPE → releasesdataset → Chinese Web Token Dataset
confidence 95% · We further release a hierarchical Chinese web token dataset with 660k+ token records
SAMPLED-BPE → uses → BPE
confidence 95% · We propose SAMPLED-BPE... train a byte-pair encoding (BPE) tokenizer
SAMPLED-BPE → audits → WanJuan
confidence 90% · WanJuan... remain below 1% total pollution
SAMPLED-BPE → audits → MAPCC
confidence 90% · MAPCC... remain below 1% total pollution
SAMPLED-BPE → audits → SkyPile
confidence 90% · SkyPile... remain below 1% total pollution
SAMPLED-BPE → audits → CCI3
confidence 90% · CCI3... remain below 1% total pollution
SAMPLED-BPE → audits → WuDao
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3) Chinese web pollution is implicit and rapidly changing. We propose Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens. Experiments show that Sampled-BPE preserves usable estimates while substantially reducing runtime and memory: a 148.4 $\times$ speedup and a 35.8 $\times$ memory reduction induce only 4.25% relative error for pollution categories. We apply the pipeline to 11 open Chinese corpora and 6 Chinese Common Crawl snapshots from 2021 to 2026. The audit reveals widespread but uneven pollution across open corpora, as well as highly polluted and temporally shifting Chinese web content. We further release a hierarchical Chinese web token dataset with 660k+ token records, each with web context, category, and explanation fields, organized as trees to support review and tracing of pollution.
Tags
Links
- Source: https://arxiv.org/abs/2608.10678v1
- Canonical: https://arxiv.org/abs/2608.10678v1
Trouble viewing inline? Open PDF directly →
Full Text
78,039 characters extracted from source content.
Expand or collapse full text
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics Caution: this paper may include offensive and upsetting content. Qingjie Zhang 1 , Ziqi Tang 1 , Jie Zhang 2 , Gelei Deng 3 , Jinfeng Li 4 , Yuefeng Chen 4 , Yitong Yang 4 , Hui Xue 4 , Tianwei Zhang 3 , and Han Qiu 1* 1 Tsinghua University 2 SiliconProspect AI 3 Nanyang Technological University 4 Alibaba Group Emails: qj-zhang24@mails., qiuhan@tsinghua.edu.cn * Corresponding author Abstract Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3) Chinese web pollution is implicit and rapidly changing.We propose SAMPLED-BPE, a lightweight token-level auditing pipeline that samples a small subset and trains BPE tok- enizer to surface polluted tokens. Experiments show that SAMPLED-BPE preserves usable es- timates while substantially reducing runtime and memory: a 148.4×speedup and a 35.8× memory reduction induce only 4.25% relative error for pollution categories. We apply the pipeline to 11 open Chinese corpora and 6 Chi- nese Common Crawl snapshots from 2021 to 2026. The audit reveals widespread but un- even pollution across open corpora, as well as highly polluted and temporally shifting Chinese web content. We further release a hierarchical Chinese web token dataset 1 with 630k+ token records, each with web context, category, and explanation fields, organized as 92k+ trees to support review and tracing of pollution. 1 Introduction Chinese web pollution has surfaced in LLMs, in- cluding spam- and pornography-related Chinese tokens in ChatGPT’s vocabulary (Yang, 2024) and Chinese gambling content in Codex outputs (Ope- nAI Developer Community, 2026). Together with recent work linking polluted Chinese tokens to cor- pus pollution (Zhang et al., 2025), these observa- tions motivate auditing upstream Chinese corpora to understand where such pollution originates, how prevalent it is, and how it may propagate into down- stream datasets. However, auditing upstream Chinese corpora remains difficult. First, mainstream Chinese cor- 1 https://github.com/qingjiesjtu/SampledBPE Chinese Web-scale Corpus SampleAuditing TokensCorpus Profile Figure 1: Intuition from marine pollution monitoring: sample a subset, train a tokenizer, map token-level pol- lution signals, and estimate a corpus profile. pora are often at web scale, consisting of hun- dreds of gigabytes to terabytes of text (Yuan et al., 2021; Chen et al., 2023; He et al., 2023; Oepen et al., 2025). Fully scanning such corpora is costly (Penedo et al., 2023, 2024; Abbas et al., 2023), e.g., 99 days for indexing 83TB data on a 128- core CPU node (Xu et al., 2025). Second, existing corpus analyses often remain coarse-grained such as language distribution, source domains, or qual- ity scores at document-level (Dodge et al., 2021; Kreutzer et al., 2022; Soldaini et al., 2024; Hagar and Bandy, 2025). These views describe broad corpus composition, but cannot reveal fine-grained pollution at token-level. Third, Chinese web pollu- tion is often implicit and rapidly changing (MacA- vaney et al., 2019; Jiang et al., 2021; Zhang et al., 2025). A predefined keyword list can miss newly emerging polluted tokens and quickly become stale. Together, these challenges call for a lightweight, token-level, periodic auditing pipeline. Inspired by marine pollution monitoring (illus- trated in Figure 1), which estimates overall pollu- tion from sampled observations (Karydis and Kit- siou, 2013), we propose SAMPLED-BPE, a sim- ple yet effective pipeline for auditing web-scale Chinese corpora when full inspection is too costly. arXiv:2608.10678v1 [cs.CL] 11 Aug 2026 Specifically, we sample a small subset, train a byte- pair encoding (BPE) tokenizer (Sennrich et al., 2016) to surface and count high-frequency tokens, map these tokens to content categories using Inter- net search evidence, and aggregate token statistics into corpus-level profiles. This design avoids re- lying on a fixed keyword list, keeps the audit at token level, and makes repeated auditing feasible as Chinese web pollution evolves. Experiments show that SAMPLED-BPE pre- serves usable token-level and category-level esti- mates while substantially reducing runtime and memory. Using this pipeline, we find that pollu- tion is widespread but uneven across open Chinese corpora. We further show that Chinese Common Crawl (Common Crawl, 2011) is highly polluted and shifts over time, indicating the need for peri- odic upstream auditing. Finally, we release a hier- archical Chinese web token dataset to make these token-level findings reviewable. This work makes three contributions: • Method. We propose and validate SAMPLED- BPE, a lightweight pipeline for token-level au- diting of web-scale Chinese corpora. It reduces token-level auditing cost from months to hours while preserving usable estimates. • Results. We systematically audit 11 open Chi- nese corpora and 6 upstream Chinese Common Crawl snapshots from 2021 to 2026, revealing widespread and shifting pollution. For exam- ple, 68.72% of 2026 Chinese Common Crawl snapshot is Adult Content. • Dataset. We release a hierarchical Chinese web token dataset containing 630,684 token records, each with web context, category, and explana- tion fields. The dataset organizes tokens into 92,972 trees, supporting review and tracing of Internet pollution. 2 Background In this section, we first summarize the corpora au- dited in this paper, then discuss related work on corpus auditing and Chinese web pollution. 2.1 Open Chinese Corpora and Landscape LLM corpora span from MB-scale documents, to GB-scale tuning data, TB-scale pretraining corpora, and PB-scale upstream web archives (Wang et al., 2018; Longpre et al., 2023; Brown et al., 2020; Common Crawl, 2011). In this work, we use web- scale to refer to corpora in the TB range (Li et al., ~PB C Web Source 1.4TB OSCAR Extracted C 71GB mC4 Cleaned C 200GB WuDao Rule-cleaned web 518GB CCI3 Two-stage-filtered web 620GB SkyPile Model-filtered web 259GB ROOTS Curated source mix 641GB CulturaX Cleaned mC4+OSCAR 5.4TB HPLT Filtered Archive+C 1.4TB CWT Quality-scored web 1TB WanJuan Refined multi-source web 1.4TB MAPCC Heuristic-filtered C Fine-tuning Scale Task/domain tuning data Data Frontier Limits of raw corpora Pretraining Scale Scale of GPT-3 Document Scale One long-form text Web - Scale ~MB ~GB ~TB ~PB Figure 2: Scale landscape of Chinese web corpora au- dited in this work. 2025; Roziewski and Stokowiec, 2016). Figure 2 places the Chinese corpora audited in this work within this scale landscape. Mainstream open Chinese corpora include Chi- nese portions of broad multilingual corpora, such as OSCAR (Abadji et al., 2022), mC4 (Xue et al., 2021), HPLT (Oepen et al., 2025), Cul- turaX (Nguyen et al., 2024), and ROOTS (Lau- rençon et al., 2022); and Chinese-specific releases, such as CWT (Chen et al., 2023), WanJuan (He et al., 2023), MAPCC (Du et al., 2024b), SkyP- ile (Wei et al., 2023), CCI3 (Wang et al., 2024), and WuDao (Yuan et al., 2021). Detailed descrip- tions are provided in Appendix A. 2.2 Related Work Corpus Curation and Auditing.Web-scale cor- pora are usually curated from web text through source selection, language identification, quality filtering, safety or blocklist filtering, and dedupli- cation (Penedo et al., 2024; Soldaini et al., 2024; Weber et al., 2024; Penedo et al., 2023). Prior doc- umentation and auditing work further improves transparency and reveals residual risks such as unexpected sources, low-quality text, duplication, contamination, and filtering side effects (Lee et al., 2022; Kreutzer et al., 2022; Dodge et al., 2021; Gebru et al., 2021). These studies mainly exam- ine source-level, document-level, quality-level, or deduplication-level properties. In contrast, our work audits corpora at token-level. Category-level Token-level Figure 3: Conceptual overview of the SAMPLED-BPE auditing pipeline. It samples web-scale corpora, trains and counts BPE tokens, maps tokens to content categories with Internet evidence, and aggregates token statistics into corpus pollution profiles. Chinese Web Pollution and Polluted Tokens. The research community has reported multiple cases where Chinese web pollution surfaces in LLMs. For example, it appears in GPT vocabu- laries: after GPT-4o was released, researchers re- ported that many of its longest Chinese tokens were not ordinary words, but long phrases associated with spam, pornography, gambling, or scams (Yang, 2024). Moreover, it can even appear in outputs for otherwise normal tasks: users reported Codex outputs that unexpectedly inserted Chinese gam- bling or lottery-like strings during coding contexts, including aroundapply_patchoutputs and mal- formed tool-call text (OpenAI Developer Commu- nity, 2026). Recent work studies this issue through polluted tokens in LLM vocabularies, connecting their presence to possible training corpora pollu- tion (Zhang et al., 2025). However, these observa- tions inspect downstream models rather than up- stream corpora. 3 Sampled-BPE To enable low-cost and accurate pollution audit- ing of web-scale Chinese corpora, SAMPLED-BPE does not aim to reconstruct the full vocabulary. In- stead, it approximates the token-level and category- level statistics of corpora. 3.1 Auditing Pipeline Figure 3 gives a conceptual overview of the four- stage auditing pipeline. Step 1: streaming sampling.Given a target cor- pus, we sample documents in a single sequential pass. Compared with random sampling which re- quires enumerating or shuffling document identi- fiers, streaming sampling makes the sampling cost nearly linear in the input scan. This system choice reduces I/O and avoids full materialization of large corpus, while yielding audit accuracy comparable to random sampling (see detailed comparison in Appendix C). Step 2: BPE training & counting. We train a BPE tokenizer on the selected corpus and use the same tokenizer to count token frequencies (Sen- nrich et al., 2016). We use BPE for three reasons. First, because BPE is a compression method that re- peatedly merges frequent adjacent units, its learned vocabulary surfaces recurrent lexical patterns that characterize the corpus. Second, high-frequency polluted tokens can be exposed without relying on a fixed keyword list, allowing new pollution terms to be discovered (MacAvaney et al., 2019; Jiang et al., 2021; Zhang et al., 2025). Third, BPE is efficient for web-scale auditing: training is driven by local adjacent-pair frequency statistics under a fixed merge budget, and counting only requires a linear pass with the learned tokenizer rather than expensive document-level semantic analysis. Step 3: category mapping. Polluted tokens are often subtle, abbreviated, or context-dependent: the token alone may not reveal its category. Therefore, inspired by prior work on Chinese polluted-token detection (Zhang et al., 2025), we map tokens, with Internet search results as contextual evidence, into one of six content categories: Normal Content, Adult Content, Online Gambling, Online Gaming, Online Video, and Anomalous. To automate this step, we use GLM-4-32B, an open-source Chinese LLM with strong Chinese comprehension ability, as the base category classifier. The model is fine- tuned to predict a category from each token and its corresponding Internet search results, using a train/test split constructed from expert annotations of GPT Chinese vocabularies in Zhang et al. (2025). It achieves 97.32% classification accuracy on the test split after fine-tuning. Sampling Rate (%)Sampling Rate (%) S c o r e ( % ) W e i g h t e d R e l a t i v e E r r o r ( % ) Figure 4: Sampling preserves auditing accuracy. Left: token-level coverage and token-ratio correlations, with the 0.25% sampling rate annotated. Middle: category-level weighted relative error, with the 0.25% sampling rate annotated. Right: category composition of 0.25% sampling rate compared with Full. Step 4: corpus profiling.We combine token fre- quencies with category assignments to build a cor- pus profile. At the token level, the profile records each token’s ratio, predicted category, Internet ev- idence, and classification rationale. This enables later analyses of shared polluted tokens, corpus- specific polluted terms, and high-frequency lexical evolution. At the category level, we aggregate to- ken ratios into category prevalence and total pollu- tion ratios, which are the quantities reported in later corpus-level comparisons and temporal analyses. We refer to the unsampled corpus asFull. In the following experiments, each sampled corpus is au- dited with the same pipeline and compared against Fullto evaluate whether sampling preserves the token-level and category-level statistics needed for pollution analysis. 3.2 Sampling Preserves Auditing Accuracy We first ask whether SAMPLED-BPE preserves the token-level and category-level statistics needed to audit corpus pollution, that is, whether sampling gives an acceptable approximation to Full. For each sampling rate, we compare the sam- pled corpus withFull. LetV Full denote the token set extracted fromFull, and letV sample denote the token set extracted from a sampled corpus. At the token level, we measure token coverage as |V Full ∩ V sample |/|V Full |, and compute Spearman and Pearson correlations between token ratios over shared tokensV Full ∩V sample . At the category level, after mapping tokens into content categories, we compute weighted relative error both over shared tokens and over all tokens. Figure 4 (Left) shows that sampled corpora still provide an acceptable token-level approximation toFullat low sampling rates. For example, at the 0.25% sampling rate, token coverage remains: the sampled corpus recovers 76.83% ofFullto- 0.05 0.1 0.25 0.5 125 102550 Sampling rate (%) 2x 5x 10x 20x 50x 100x 200x Improvement 273.8 209.4 148.4 98.9 62.5 39.1 18.9 9.6 4.1 1.9 46.9 38.0 35.8 30.1 28.6 22.3 12.1 6.6 3.0 1.8 Runtime speedup Peak RSS reduction Figure 5: Runtime and memory improvement across sampling rates for SAMPLED-BPE. kens. More importantly, the ratios of the covered tokens closely matchFull, with Spearman reach- ing 86.88% and Pearson reaching 99.93%. This suggests that sampling preserves an acceptable approximation of token-level statistics. Figure 4 (Middle) shows that the same trend appears after aggregating tokens into content cat- egories. Even when token coverage is imperfect, the weighted relative error remains around 5% at the lowest sampling rate, for both shared tokens and all tokens. Figure 4 (Right) further illustrates this at the 0.25% sampling rate, where the category composition remains close to theFulldistribution. This shows that sampling also preserves category- level statistics for pollution auditing. 3.3 Sampling Reduces Auditing Cost Having shown that sampling preserves the statistics needed for pollution auditing, we next examine its runtime speedup and memory reduction. We measure end-to-end auditing cost as BPE training time plus token-counting time, and we report peak RSS as the memory footprint. All runs are conducted on the same machine with 96 CPU cores and 768 GB RAM. Figure 5 shows that smaller samples sharply re- OSCARmC4HPLTCulturaXCWTROOTSWanJuanMAPCCSkyPileCCI3WuDao Pollution83.3825.414.953.362.352.020.750.690.610.550.50 Adult76.100.311.950.150.050.230.040.100.060.160.10 Gambling0.5623.171.612.450.220.000.150.080.040.000.01 Gaming0.160.460.250.160.290.000.170.090.050.010.06 Video5.020.320.720.180.060.000.050.040.040.000.01 Anomalous1.531.150.430.421.731.790.330.380.420.370.31 Table 1: Pollution ratios (%) for the Chinese portions of 11 representative open corpora, sorted by total pollution. duce both runtime and memory. At the 0.25% sam- pling rate, the auditing pipeline achieves a 148.4× runtime speedup. The memory reduction is also substantial, with peak RSS reduced by 35.8×. Ex- tremely small samples provide even larger savings, but Figure 4 shows that their token coverage and category estimates are less accurate. Larger sam- ples improve accuracy, but their computational ad- vantage decreases. Overall, sampling makes web-scale corpus au- diting substantially cheaper while retaining ac- ceptable accuracy for pollution analysis. The time cost of auditing 1TB corpus can be reduced from months to hours. 4 Pollution in Open Chinese Corpora This section audits pollution in 11 open Chinese corpora or the Chinese portions of multilingual corpora as stated in Section 2.1. Following (Zhang et al., 2025), percentages are computed over tokens containing at least three Chinese characters over six categories. 4.1 Pollution Is Widespread but Uneven Table 1 shows that each corpus contains polluted tokens, however, the pollution ratio varies sharply, ranging from 0.50% in the cleanest corpora to 83.38%. OSCAR is the highest-pollution corpus, and mC4 is also heavily polluted at 25.41%. By contrast, WanJuan, MAPCC, SkyPile, CCI3, and WuDao all remain below 1% total pollution. This shows that Chinese pollution is widespread across open corpora, but is highly uneven across corpus families and construction pipelines. The lower pollution ratios in several corpora sug- gest that cleaning and curation substantially reduce pollution. Broad multilingual web pipelines such as OSCAR, mC4, HPLT, and CulturaX have the highest pollution ratios in Table 1. In contrast, cor- pora such as WanJuan, MAPCC, SkyPile, CCI3, and WuDao, which emphasize Chinese-specific collection, trusted sources, or quality filtering, are much cleaner. At the same time, even the lowest- pollution corpora retain measurable polluted to- kens, indicating that cleaning helps, but does not eliminate pollution. To support this claim, Ap- pendix D gives two more focused comparisons be- tween cleaned and uncleaned corpus variants. The category rows in Table 1 show that pollution is not monolithic. OSCAR is dominated by Adult Content, which accounts for 76.10%. mC4 has a different profile: Online Gambling alone reaches 23.17%. HPLT contains both Adult Content and Online Gambling at non-trivial levels. Cleaner cor- pora exhibit a different residual profile: their re- maining pollution is usually not dominated by a single category; instead, the largest residual compo- nent is often Anomalous, reflecting rare, peculiar, or contextually irrelevant phrases that survive qual- ity filtering. These contrasts show that the com- position of pollution varies across corpora. We provide word-cloud views to show these corpus differences in Appendix G. 4.2 Pollution Is Shared, Yet Mostly Unique Since open Chinese corpora differ sharply in both total pollution and category composition, we fur- ther investigate whether the same polluted tokens recur across corpora or whether each corpus has its own lexical artifacts. To compare polluted tokens across the 11 repre- sentative corpus in Table 1, we define a directional coverage metric: Coverage(A→ B) = |P A ∩ P B | |P A | . HereP A andP B denote the polluted tokens in cor- pora A and B. Figure 6 shows that pairwise overlap depends on the types of corpora being compared. Cleaner or more curated Chinese corpora tend to overlap sub- stantially with one another: MAPCC and SkyPile cover more than half of each other’s polluted tokens (similar for WanJuan, MAPCC, SkyPile, CCI3, and WuDao). This suggests that cleaning and curation OSCAR mC4 HPLT CulturaX CWT ROOTS WanJuan MAPCC SkyPile CCI3 WuDao OSCAR mC4 HPLT CulturaX CWT ROOTS WanJuan MAPCC SkyPile CCI3 WuDao 0.411.160.380.130.030.110.120.070.040.06 2.776.3013.82.300.552.952.201.760.861.22 27.321.925.712.34.6415.414.011.28.339.77 12.969.237.018.27.6523.620.719.512.113.5 4.7112.619.319.97.0725.317.918.911.412.7 0.792.295.576.395.415.806.498.398.295.60 9.1935.653.457.055.916.850.152.330.633.1 10.930.555.957.545.621.657.755.836.941.9 6.4723.843.552.646.727.158.554.342.540.9 4.7414.841.341.736.234.343.845.854.346.6 9.0326.159.957.549.528.658.564.264.557.5 0 10 20 30 40 50 60 Coverage (%) Figure 6: Pairwise coverage of polluted tokens across corpora. Each cell (omitting the diagonal) reports the percentage of polluted tokens in the row corpus that also appear in the column corpus. Shared Across k Corpora P r o p o r t i o n o f P o l l u t e d T o k e n s ( % ) At least k corpora Exactly k corpora Figure 7: Proportion of polluted tokens shared across corpora. Blue line show the proportion appearing in exactlykcorpora, and the red line shows the cumulative proportion appearing in at least k corpora. strategies still miss a shared set of residual pol- luted tokens, revealing a common limitation of existing filtering pipelines. By contrast, broader web corpora show more asymmetric containment. For example, 69.2% of CulturaX polluted tokens appear in mC4, but only 13.8% of mC4 polluted tokens appear in CulturaX. This suggests that some broad web corpora subsume the polluted tokens of other corpora while also introducing a much larger additional tail. To measure sharing beyond pairwise views, Fig- ure 7 aggregates over all polluted tokens by count- ing how many corpora each token appears in. Across the 11 corpora, we observe 106,671 pol- luted tokens. Among them, 102,816 (96.386%) appear in only one corpus. Only 1,542 (1.446%) ap- pear in at least three. This shows that most polluted CategoryExampleExplanation 老司机 (old driver) Someone who knows or shares adult resources. 威尼斯人 (Venetian) One of the world’s largest casino resorts. Adult Gambling 王者荣耀 (Honor of Kings) A popular multiplayer online game. Gaming 在线观看 (watch online) The act of watching video content on the web. Video 自己的小 (one’s own little) Semantically incomplete, lacking a clear referent. Anomalous Figure 8: Representative examples from the shared core of polluted tokens. The shared core contains tokens that appear in at least 8 of the 11 corpora in Table 1. tokens are corpus-specific, with a small shared core recurring across multiple datasets. Figure 8 lists representative examples of the small shared core. 252 polluted tokens meet this criterion, yet they span all five polluted categories. These tokens are shared because they reflect long- lived lexical anchors. For example, “老司机(old driver)” is a persistent euphemism in adult-resource pages, referring someone who knows or shares adult content resources; “威尼斯人(Venetian)” is repeatedly reused as one of the world’s largest casino resorts. In contrast, Appendix E lists corpus- specific polluted tokens. Together, these results show that polluted tokens form a small shared core with much larger corpus- specific tails. 5 Evolution of Chinese Web Content The previous section shows that polluted tokens appear, to varying degrees, in mainstream open Chinese corpora. Since many of these corpora are directly or indirectly derived from Common Crawl (Common Crawl, 2011; Xue et al., 2021; Abadji et al., 2022; Nguyen et al., 2024; Penedo et al., 2023), we next ask what the upstream source looks like: How polluted is Chinese Common Crawl it- self, and how does this pollution evolve over time? Common Crawl is updated monthly, with each release containing web text at terabyte scale. For this temporal case study, we audit six Chinese Com- mon Crawl snapshots from 2021 to 2026. For each snapshot, we run the same SAMPLED-BPE audit- ing pipeline with 1% sampling rate, which yields lower than 3% relative error according to Figure 4. 5.1 Pollution Is High and Shifting Table 2 shows that Chinese Common Crawl con- tains strikingly high pollution, with Adult Content as the dominant source in several years. In 2026, to- 202120222023202420252026 Pollution35.3561.8762.9034.8338.0179.52 Adult24.7354.5355.8628.0130.5368.72 Gambling5.311.050.541.480.740.21 Gaming0.400.160.230.330.220.07 Video3.934.754.893.865.058.39 Anomalous0.981.381.371.161.472.12 Table 2: Temporal pollution ratios (%) in Chinese Com- mon Crawl snapshots. Chinese web content is highly polluted, and its pollution profile shifts over time tal pollution even reaches 79.52%, with Adult Con- tent alone accounting for 68.72%. This indicates that pollution in the upstream Chinese web source is not a marginal tail phenomenon, but becomes comparable to or even exceeds normal content. Besides, different pollution categories evolve in different directions. Adult Content is both high and volatile, dropping from 55.86% in 2023 to 28.01% in 2024 before rising to 68.72% in 2026. Online Gambling follows the opposite direction, decreasing from 5.31% in 2021 to 0.21% in 2026, a trend that is consistent with intensified enforce- ment against cross-border gambling activities tar- geting Chinese users 2 . Online Video gradually in- creases from 3.93% to 8.39%, while Anomalous rises modestly and Online Gaming remains con- sistently small. These trends show that Chinese web content is highly polluted, and its pollution profile shifts over time. Appendix word clouds in Appendix G further visualize the changing surface vocabulary across Common Crawl snapshots. 5.2 Some Tokens Turn Over, Others Persist We further examine the evolution of tokens. For each category, we compute Jaccard overlap of to- kens for adjacent years and for the five-year span, which captures how tokens turn over and persist. Figure 9 shows that tokens in Normal Content, Online Gaming, and Online Video are replaced less frequently over time. Their adjacent-year over- lap is relatively high, and their five-year overlap also remains non-trivial. This means these cate- gories have less year-to-year token turnover and more long-lived tokens, suggesting that ordinary web content, online gaming, and online video are expressed through more stable forms on the web. By contrast, Adult Content, Online Gambling, and Anomalous show lower adjacent-year overlap and much weaker five-year persistence. The low adjacent-year overlap indicates frequent short-term 2 https://english.scio.gov.cn/pressroom/2025-02/07/ content_117699840.html 1-year gap 5-year gap NormalGamingVideoAdultAnomalousGambling J a c c a r d O v e r l a p ( % ) Figure 9: Jaccard overlap of tokens across years. Filled bars show average adjacent-year overlap, and hollow bars show the five-year (2021–2026) overlap. Category Short-lived tokensPersistent tokens Normal Adult Gambling Gaming Video Anomalous 世界杯 World Cup 国际完美世界下载 International Perfect World download 盗墓笔记第二季 The Lost Tomb Season 2 水蜜桃无码国产精品 peach uncensored domestic premium 乐橙体育 Lecheng Sports 古风君子以泽 ancient-style Junzi Yize 联系我们 contact us 王者荣耀 Honor of Kings 在线观看 watch online 亚洲精品 Asian premium content 娱乐城 casino 大香蕉 big banana Figure 10: Representative tokens illustrating turnover and persistence in Chinese Common Crawl snapshots. token turnover, while the low five-year overlap in- dicates that few tokens persist over longer hori- zons. This may suggest that Adult Content and Online Gambling sites continually reappear on the Chinese web and change their surface forms even after being repeatedly banned or blocked, while Anomalous reflects the fast-changing vocabulary of subcultural or irregular web expressions (see examples in Appendix F). Figure 10 provides representative tokens to inter- pret the contrast between stable and fast-changing categories. In the more stable categories such as Normal Content, Online Gaming, and Online Video, persistent tokens are reusable templates or stable entry points, such as “联系我们(contact us)” and “王者荣耀(Honor of Kings)”. Their short-lived tokens tend to be concrete events, ti- tles, or platform-specific phrases rather than broad changes in the category’s surface form. In the faster-changing categories, short-lived to- kens more often include adult-site phrases, gam- bling platforms, or subcultural expressions, such as “水蜜桃无码国产精品(peach uncensored do- mestic premium)”, “乐橙体育(Lecheng Sports)”, and “古风君子以泽(ancient-style Junzi Yize)”. Their persistent tokens mostly capture recurring templates, such as “亚洲精品(Asian premium content)” and “娱乐城 (casino)”. Online gambling 菲律宾申博 An online gambling brand Normal content 菲律宾 Philippines Normal content 申博 Apply for a PhD Adult content 浪小辉 A gay porn actor Adult content 猛男浪小辉 A muscular hunk Adult content 中国男同浪小辉 A Chinese gay porn actor A polluted online gambling token family rooted at "菲律宾申博" Representativetoken: Linked to the offshore gambling platform Sunbet and its account- registration sites, a typical gambling promotion term. ② Polluted tokens can emerge from composing normal tokens. Representative token: "Lang Xiaohui" is a performer in Chinese gay adult videos, a token directly linked to sexually explicit content. ① Roots capture minimal recurring polluted subtokens. A polluted adult content token family rooted at "浪小辉" Adult content 浪小辉 GAY 国产 Domestic gay porn star Online gamblingNormal + Normal Figure 11: Examples of token trees in the hierarchical Chinese web token dataset. Left: roots identify minimal recurring subtokens of longer surface variants. Right: two normal subtokens can emerge pollution from combination. Thus, Common Crawl’s Chinese pollution ex- hibits both category-level temporal drift and token- level lexical renewal. Static keyword lists or one- time filtering strategies can therefore quickly be- come stale, while SAMPLED-BPE supports peri- odic auditing of upstream web corpora. 6 Hierarchical Chinese Web Tokens Building on our token-level audits of Chinese open corpora and Common Crawl, we release a hier- archical Chinese web token dataset with 630,684 tokens, so that these token-level findings can be reviewed, reused, and extended. 6.1 Tokens Are Organized as Trees Chinese web tokens contain many substring rela- tions. A short token may appear as a subtoken inside many longer tokens, and these longer tokens often form related surface variants. Instead of re- leasing a flat token list, we organize tokens as a forest (shown in Figure 11), consisting of 92,972 trees. Each tree is rooted at the shortest subto- ken in the family, and its descendants are longer tokens that contain the root. Each token node contains Internet search evi- dence and a category label. Each tree further in- cludes representative tokens and explanations for their categories. The dataset is therefore not only a collection of token strings, but also an auditable resource that records the contextual evidence and classification rationale. 6.2 Hierarchy Traces Pollution We give two ways to trace pollution leveraging this hierarchical structure. First, when all tokens in a tree are from the same pollution family, the root identifies a minimal re- curring subtoken. This lets us summarize a set of surface variants with a shorter representative to- ken, rather than treating every longer variant as an unrelated keyword. Figure 11 (Left) shows an ex- ample of “浪小辉(A gay porn actor)” which forms a large pollution family containing long variants such as “浪小辉GAY国产”. Second, trees with tokens from different cate- gories reveal pollution that emerges from token composition. A subtoken may be normal in iso- lation, but become polluted when combined with another normal token. Figure 11 (Right) shows an example that “菲律宾(Philippines)” and “申 博 (Apply for a PhD)” can appear as normal to- kens separately, while “菲律宾申博(An online gambling brand)” is a gambling-related polluted token. Such cases are difficult to capture because the risk emerges from the composition rather than from either subtoken alone. Overall, the hierarchy makes the dataset more informative than a flat list of polluted tokens. It traces polluted token families back to minimal recurring subtoken and exposes hidden compo- sitional cases. This provides evidence for future review, cleaning, and periodic auditing. 7 Conclusion Motivated by Chinese web pollution surfacing in ChatGPT’s vocabularies and Codex’s outputs, this work presents SAMPLED-BPE, a lightweight token-level auditing pipeline that trains BPE tok- enizers on sampled corpora to estimate Chinese corpora pollution profiles. Experiments show that SAMPLED-BPE substantially reduces runtime and memory while preserving usable estimates. We audit 11 open Chinese corpora and 6 Chi- nese Common Crawl snapshots. Results show that pollution is widespread but uneven across open Chinese corpora, and upstream Chinese web con- tent is highly polluted and temporally shifting. A hierarchical Chinese web token dataset with 630k+ token records, each with web context, category, and explanation fields, is released for reviewing and tracing Chinese web pollution. Limitations Pollution in other languages. This work fo- cuses on polluted tokens in Chinese web-scale cor- pora. We do not investigate polluted tokens in other languages, nor do we claim that our auditing pipeline transfers directly beyond Chinese. Ex- tending SAMPLED-BPE to other languages would require multilingual annotators, language-specific search evidence, and possibly different pollution categories. We hope this work encourages re- searchers to audit web-scale upstream corpora in more languages. Readability to non-native Chinese readers. This paper necessarily includes many Chinese to- kens, some of which are offensive, euphemistic, abbreviated, or context-dependent. We provide En- glish glosses for representative examples when pos- sible, but many tokens cannot be translated literally without losing their web-specific usage or pollu- tion cues. Our goal is not to study Chinese linguis- tic expression itself, but to use Chinese web-scale corpora as a concrete case for studying upstream corpora pollution. We therefore treat translations as readability aids rather than complete semantic equivalents, and future work with broader multilin- gual expertise can improve these explanations. Ethics Statement ACL Ethics Policy is respected in this work. This work studies pollution in open Chinese corpora and upstream Chinese Common Crawl snapshots. We use publicly available corpora and web evi- dence for research purposes, and we respect the terms, conditions, and copyright requirements of the corresponding data sources. Because the paper analyzes web pollution, it necessarily includes ex- amples related to adult content, gambling, spam, anomalous expressions, and other potentially of- fensive or upsetting material. These examples are presented only for documenting and auditing cor- pus pollution, and should be used with caution in future research. We adhere to the Association for Computational Linguistics (ACL) guidelines on responsible NLP research 3 , with particular attention to transparency, research-use framing, and responsible handling of harmful content. 3 https://aclrollingreview.org/responsibleNLPresearch/ Use of AI Assistants The authors used AI assistants for language pol- ishing, LaTeX editing, and phrasing suggestions during paper preparation. All substantive claims, experimental results, analyses, citations, and final text were reviewed and verified by the authors. References Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, and Benoît Sagot. 2022. Towards a cleaner document- oriented multilingual crawled corpus. In Proceedings of the Thirteenth Language Resources and Evalua- tion Conference, pages 4344–4355. Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos. 2023. Semdedup: Data- efficient learning at web-scale through semantic dedu- plication. arXiv preprint arXiv:2303.09540. Allen Institute for AI. 2024. C4: Colossal clean crawled corpus. https://huggingface.co/datasets/allenai/c4. Hugging Face dataset card. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901. Jianghao Chen, Pu Jian, Tengxiao Xi, Dongyi Yi, Qian- long Du, Chenglin Ding, Guibo Zhu, Chengqing Zong, Jinqiao Wang, and Jiajun Zhang. 2023. Chi- nesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model. arXiv preprint arXiv:2311.01149. Common Crawl. 2011.Common crawl.https:// commoncrawl.org/. Jesse Dodge, Maarten Sap, Ana Marasovi ́ c, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 conference on empirical methods in natural language processing, pages 1286–1305. Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xing- wei Qu, Jiaheng Liu, Tianyu Zheng, Xinchen Luo, Guorui Zhou, Wenhu Chen, and Ge Zhang. 2024a. Chinese tiny llm: Pretraining a chinese-centric large language model. Preprint, arXiv:2404.04167. Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xingwei Qu, Jiaheng Liu, Tianyu Zheng, et al. 2024b. Chinese tiny llm: Pretraining a chinese-centric large language model. arXiv preprint arXiv:2404.04167. Timnit Gebru, Jamie Morgenstern, Briana Vec- chione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM, 64(12):86– 92. Nick Hagar and Jack Bandy. 2025. Practical datasets for analyzing llm corpora derived from common crawl. In Proceedings of the International AAAI Conference on Web and Social Media, volume 19, pages 2454– 2464. Conghui He, Zhenjiang Jin, Chao Xu, Jiantao Qiu, Bin Wang, Wei Li, Hang Yan, Jiaqi Wang, and Dahua Lin. 2023. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large mod- els. arXiv preprint arXiv:2308.10755. High Performance Language Technologies Project. 2025. HPLT monolingual datasets 3.0. https://hplt- project.org/datasets/v3.0. Dataset catalogue. Menghan Jiang, Xiang Ying Shen, Kathleen Ahrens, and Chu-Ren Huang. 2021. Neologisms are epi- demic: Modeling the life cycle of neologisms in china 2008-2016. PloS one, 16(2):e0245984. Michael Karydis and Dimitra Kitsiou. 2013. Marine wa- ter quality monitoring: A review. Marine pollution bulletin, 77(1-2):23–36. Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan Van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50– 72. Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Ed- uardo González Ponferrada, Huu Nguyen, et al. 2022. The bigscience roots corpus: A 1.6 tb composite mul- tilingual dataset. Advances in Neural Information Processing Systems, 35:31809–31826. Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445. Jeffrey Li, Mohammadreza Armandpour, Seyed Iman Mirzadeh, Sachin Mehta, Vaishaal Shankar, Raviteja Vemulapalli, Samy Bengio, Oncel Tuzel, Mehrdad Farajtabar, Hadi Pouransari, et al. 2025. Tic-lm: A web-scale benchmark for time-continual llm pretrain- ing. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 32231–32273. Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. In International conference on machine learning, pages 22631–22648. PMLR. Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019. Hate speech detection: Challenges and solutions. PloS one, 14(8):e0221152. Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2024. Cul- turaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 4226– 4237, Torino, Italia. ELRA and ICCL. Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charp- entier, Pinzhen Chen, Mariya Fedorova, Ona de Gib- ert, et al. 2025. Hplt 3.0: Very large-scale multilin- gual resources for llm and mt. mono-and bi-lingual data, multilingual evaluation, and pre-trained models. arXiv preprint arXiv:2511.01066. OpenAI Developer Community. 2026. Chinese gam- bling characters in Codex CLI message and code output?OpenAI Developer Community. Forum thread, first posted January 27, 2026. OSCARProject.2023.OSCAR-2301-HPC. https://huggingface.co/datasets/oscar-corpus/ oscar-2301-hpc. Hugging Face dataset card. PaddleNLP Contributors. 2025. WuDaoCorpus2.0 base corpus. https://paddlenlp.readthedocs.io/en/latest/ llm/tools/preprocess/docs/WuDaoCorpusBase.html. Documentation page. Guilherme Penedo, Hynek Kydlí ˇ cek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, Thomas Wolf, et al. 2024. The fineweb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems, 37:30811–30849. Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only.arXiv preprint arXiv:2306.01116. Aleksandra Piktus, Christopher Akiki, Paulo Villegas, Hugo Laurençon, Gérard Dupont, Sasha Luccioni, Yacine Jernite, and Anna Rogers. 2023. The roots search tool: Data transparency for llms. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 304–314. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the lim- its of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67. Szymon Roziewski and Wojciech Stokowiec. 2016. Languagecrawl: A generic tool for building lan- guage models upon common-crawl. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 2789– 2793. Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th annual meeting of the association for computational linguis- tics (volume 1: long papers), pages 1715–1725. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bo- gin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. 2024. Dolma: An open corpus of three trillion tokens for language model pretraining research. In Proceedings of the 62nd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 15725–15788. Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP, pages 353– 355. Liangdong Wang, Bo-Wen Zhang, Chengwei Wu, Hanyu Zhao, Xiaofeng Shi, Shuhao Gu, Jijie Li, Quanyue Ma, TengFei Pan, and Guang Liu. 2024. Cci3. 0-hq: a large-scale chinese dataset of high qual- ity designed for pre-training large language models. arXiv preprint arXiv:2410.18505. Maurice Weber, Daniel Y Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xi- aozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, et al. 2024. Redpajama: an open dataset for training large language models. Advances in neural information processing systems, 37:116462–116492. Tianwen Wei, Liang Zhao, Lichang Zhang, Bo Zhu, Lijie Wang, Haihua Yang, Biye Li, Cheng Cheng, Weiwei Lü, Rui Hu, et al. 2023. Skywork: A more open bilingual foundation model. arXiv preprint arXiv:2310.19341. BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili ́ c, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luc- cioni, François Yvon, et al. 2022. Bloom: A 176b- parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100. Hao Xu, Jiacheng Liu, Yejin Choi, Noah A Smith, and Hannaneh Hajishirzi. 2025. Infini-gram mini: Exact n-gram search at the internet scale with fm-index. In Proceedings of the 2025 Conference on Empiri- cal Methods in Natural Language Processing, pages 24955–24980. Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 conference of the North American chap- ter of the association for computational linguistics: Human language technologies, pages 483–498. Zeyi Yang. 2024. GPT-4o’s Chinese token-training data is polluted by spam and porn websites. MIT Tech- nology Review. Published May 17, 2024. Sha Yuan, Hanyu Zhao, Zhengxiao Du, Ming Ding, Xiao Liu, Yukuo Cen, Xu Zou, Zhilin Yang, and Jie Tang. 2021. Wudaocorpora: A super large-scale chinese corpora for pre-training language models. Ai Open, 2:65–68. Qingjie Zhang, Di Wang, Haoting Qian, Liu Yan, Tian- wei Zhang, Ke Xu, Qi Li, Minlie Huang, Hewu Li, and Han Qiu. 2025. Speculating llms’ chinese train- ing data pollution from their tokens. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26124–26144. A Details of Open Chinese Corpora Following Section 2.1, this section summarizes Common Crawl and the 11 representative open Chinese corpora audited in our experiments. For each corpus, we describe its source, public scale, audited Chinese portion when available, and rela- tion to LLM training when the release is explicitly tied to a model. These details provide context for the corpus landscape in Figure 2 and the pollution ratios in Table 1. C (Common Crawl Foundation, founded in 2007; data collected since 2008): Common Crawl is an open web-crawl archive maintained by a nonprofit organization. It has collected public web data since 2008 and, according to the official overview, now contains more than 10 PiB of archived data, with crawls published approximately once a month and each crawl typically containing more than two billion web pages.Common Crawl provides raw HTTP responses in WARC files, extracted metadata in WAT files, and extracted plaintext in WET files, making it the upstream source for many web-derived training corpora. In our WET-file accounting, each monthly crawl is approximately 6 TB compressed and 20 TB decompressed. Because Common Crawl does not publish official Chinese-only sizes, we estimate the language distribution by reading the WET headerWARC-Identified-Content-Language and counting the first tag as the record’s primary language; English accounts for 45.14% (about 9 TB decompressed) and Chinese for 5.19% (about 1 TB decompressed), as shown in Figure 12. We audit Chinese WET-derived snapshots from 2021–2026 by selecting the February crawl in each year and sampling 1% of the Chinese content, about 10 GB per crawl. Filtered or processed Common Crawl data has been used directly or indirectly in major models and corpora, including T5 through C4, GPT-3 through filtered Common Crawl, BLOOM through OSCAR/ROOTS, and Falcon through RefinedWeb (Common Crawl, 2011; Raffel et al., 2020; Workshop et al., 2022; Penedo et al., 2023). OSCAR (OSCAR project, Inria/University of Mannheim/DFKI, 2023): OSCAR-2301-HPC is a document-oriented multilingual web corpus de- rived from Common Crawl WET files with the Un- goliant pipeline. The public OSCAR-2301-HPC table reports a Chinese portion of 138.48M docu- ments, 44.38B space-separated words, and 1.4TB Figure 12: Language distribution in our Common Crawl WET accounting. of content. OSCAR is intended as raw multilin- gual pretraining data; earlier OSCAR releases were also a major source in ROOTS, the corpus used to train BLOOM (Abadji et al., 2022; OSCAR Project, 2023; Laurençon et al., 2022; Workshop et al., 2022). mC4 (Google Research, 2021): Multilingual C4 is the multilingual variant of C4, built by cleaning Common Crawl text and extending the C4 filtering recipe to many languages. The public C4 release lists mC4 as a 9.7TB multilingual collection, and the Chinese files audited in this work total 71.42GB. mC4 was introduced as the pretraining corpus for mT5, a multilingual text-to-text Transformer cover- ing 101 languages (Xue et al., 2021; Allen Institute for AI, 2024). HPLT (European HPLT initiative, 2025): HPLT 3.0 is a very large multilingual web corpus built from Internet Archive and Common Crawl data from 2012–2024. HPLT 3.0 applies HTML extrac- tion, language identification, deduplication, quality scoring, metadata annotation, and filtering; its Sim- plified Chinese Mandarin portion,cmn_Hans, is re- ported as 5.40TB, 2.21B documents, 2.97T tokens, and 4.14T characters. The release is designed as a broad LLM and MT resource and is accompanied by HPLT-trained reference models and multilingual evaluations (Oepen et al., 2025; High Performance Language Technologies Project, 2025). CulturaX (UONLP and collaborators, 2024): A cleaned multilingual dataset constructed by com- bining mC4 and OSCAR releases, followed by lan- guage identification, URL filtering, metric-based cleaning, document refinement, and MinHash dedu- plication. The full corpus contains 6.3T tokens in 167 languages; the Chinese partition contains 218.62M documents and 227.06B tokens, and the public Chinese files used for our audit total 641.02GB. CulturaX is released as reusable multi- lingual LLM training data rather than as a dataset paired with one flagship model (Nguyen et al., 2024). CWT (CASIA-LM, 2023): ChineseWebText is a Chinese web corpus extracted with the EvalWeb toolchain from noisy Common Crawl data. Eval- Web combines deduplication, language identifica- tion, rule-based filtering, and a BERT-based quality evaluation model, and assigns a quality score to each text. The official release contains 1.42TB of scored Chinese web text and a cleaner 600GB sub- set with quality scores above 90%, intended for Chinese LLM pretraining and threshold-based data selection (Chen et al., 2023). ROOTS (BigScience, 2022): The Responsible Open-science Open-collaboration Text Sources cor- pus is a 1.6TB multilingual collection spanning 59 languages, assembled from OSCAR and man- ually documented sources. The ROOTS Search Tool reports 259.01GB for the Chinese language group used as our official Chinese-size denomi- nator. ROOTS was the training corpus for the 176B-parameter BLOOM multilingual language model (Laurençon et al., 2022; Piktus et al., 2023; Workshop et al., 2022). WanJuan (Shanghai AI Laboratory/OpenDataLab, 2023): The text portion of Intern-WanJuan 1.0 is a cleaned pretraining corpus with more than 500M documents and more than 1TB of data. It inte- grates web pages, encyclopedias, books, patents, textbooks, and exam questions, converting hetero- geneous HTML, text, PDF, and EPUB sources into a unified JSONL format after fine-grained cleaning, deduplication, and safety-oriented filtering. Wan- Juan was used in the training of InternLM (He et al., 2023). MAPCC (M-A-P/CT-LLM team, 2024): MAP- C is a Chinese-centric pretraining corpus released with the Chinese Tiny LLM project. MAP-C con- tains about 800B Chinese tokens and is organized into components such aszh-ccfrom Chinese Com- mon Crawl, encyclopedic data, papers, books, and other sources; thezh-ccportion audited here has a public size of 1.408TB. MAP-C serves as the main Chinese pretraining data for CT-LLM, a 2B Chinese-centric language model trained together with English and code data (Du et al., 2024b,a). SkyPile (Skywork, 2023): SkyPile-150B is a large- scale Chinese web dataset for pretraining lan- guage models. The public portion contains about 233M unique web pages, around 150B tokens, and 620GB of plain text. Its construction applies fil- tering, deduplication, sensitive-data filtering, and low-quality filtering with tools such as fastText and BERT. The release is associated with the Sky- work bilingual foundation-model effort (Wei et al., 2023). CCI3 (BAAI, 2024): CCI3-HQ is a high-quality subset of Chinese Corpora Internet 3.0 designed for LLM pretraining. The release is described as an approximately 500GB Chinese corpus, and the public files used for our audit total 518GB. It is produced from trusted Chinese Internet data using a two-stage hybrid filtering pipeline. The techni- cal report evaluates the dataset by training a 0.5B model from scratch on 100B tokens and comparing against CCI3.0, SkyPile, and WanJuanV1 (Wang et al., 2024). WuDao (Beijing Academy of Artificial Intelli- gence, 2021): WuDaoCorpus2.0 Base is the public base subset of WuDaoCorpora, a large-scale Chi- nese corpus collected for pretraining. The full Wu- DaoCorpora is described as a 3TB Chinese corpus, while the publicly released WuDaoCorpus2.0 Base contains 200GB of plain text. It is derived from large-scale Chinese web pages and other sources, with text extracted from high-density pages and cleaned through a rule-based pipeline. It is com- monly used as a Chinese pretraining baseline and is associated with the WuDao large-model line (Yuan et al., 2021; PaddleNLP Contributors, 2025). B Sampling Rates of Chinese Corpora Following Section 4, this section reports the ac- tual sampling rates used for auditing the 11 open Chinese corpora and the corresponding full-token weighted relative errors estimated from Figure 4. Sampling rates use official public corpus-size de- nominators where available. Weighted errors are obtained by log-linear interpolation over the full- token category weighted-error curve. For ROOTS, the denominator is the com- plete Chinese ROOTS size, not only the audited roots_zh_uncorpusslice. For WanJuan, the of- ficial denominator is reported only as larger than 1TB, so the sampling rate is an upper bound and the error is interpolated at 0.53%. Table 3 shows that we use small sampling rates for most corpora, ranging from 0.19% for HPLT CorpusOSCARmC4HPLTCulturaXCWTROOTSWanJuanMAPCCSkyPileCCI3WuDao Sampling rate1.53%14.00%0.19%0.94%0.76%0.46%0.53%0.36%1.94%2.03%1.34% Weighted relative error2.34%0.24%3.63%2.75%2.93%2.90%3.02%2.68%2.06%2.00%2.46% Table 3: Sampling rates and estimated full-token weighted relative errors for auditing the 11 open Chinese corpora. Time CostToken CoverageCategory WRE Random1731.7s70.19%2.65% Streaming77.4s ▼22.4× 69.79% ▼0.57% 2.57% ▼2.69% Table 4: Streaming and random sampling comparison on C at 0.10% sampling rate. to 2.03% for CCI3, with mC4 as an exception be- cause the audited Chinese slice is much smaller. Despite these low sampling rates, the estimated full-token weighted relative errors remain small: all are below 4%, and most fall between 2% and 3%. This supports the result that SAMPLED-BPE can provide usable corpus pollution estimates with- out full-corpus auditing. C Streaming vs. Random Sampling Following Section 3.1, this section compares the prefix-style streaming sampler used in SAMPLED- BPE with an index-based random sampler on the same C reference setting. The random sampler first materializes a shuffled index set and then reads the selected rows, whereas the streaming sampler reads a contiguous prefix and can stop once the requested sample size has been reached. We use the 0.10% sample size relative to the 30% C refer- ence as a representative setting. In Table 4, arrows in the Streaming row show relative changes from Random. Table 4 shows that streaming sampling substan- tially reduces sampling time while preserving audit quality. Compared with random sampling, stream- ing sampling is 22.4×faster, while token coverage changes only from 70.19% to 69.79%. The cat- egory weighted relative error is also comparable, decreasing slightly from 2.65% to 2.57%. These results support our use of streaming sampling as a low-cost system choice for web-scale corpus audit- ing. D Cleaning Reduces but Does Not Eliminate Pollution Following Section 4, this section provides focused comparisons supporting the claim that cleaning and curation reduce, but do not eliminate, Chinese ChineseWebTextROOTS FullCleanerUncorpusPublic Pollution2.350.502.021.37 Adult0.050.050.230.19 Gambling0.220.020.000.01 Gaming0.290.050.000.01 Video0.060.020.000.00 Anomalous1.730.371.791.16 Table 5: Pollution ratios (%) for cleaned and less- cleaned Chinese corpus variants. web pollution. We compare ChineseWebText full and cleaner variants, and two ROOTS Chinese vari- ants with different source and curation composition. Values follow the same normalization as Table 1. Table 5 shows that cleaning sharply lowers pollu- tion in ChineseWebText: total pollution drops from 2.35% in CWT-Full to 0.50% in CWT-Cleaner. The reduction is especially clear for Anomalous con- tent, which decreases from 1.73% to 0.37%, and Online Gambling, which decreases from 0.22% to 0.02%. ROOTS shows a similar but weaker pat- tern: ROOTS-Public has lower total pollution than ROOTS-Uncorpus (1.37% vs. 2.02%), but still re- tains measurable residual pollution, mainly from Anomalous and Adult Content. These comparisons support the main-text claim that cleaning helps, but does not remove polluted tokens entirely. E Corpus-Specific Polluted Tokens Following Section 4.2, this section complements the shared-core analysis by showing representative polluted tokens that are unique to one audited cor- pus slice. These examples are grouped by predicted category. Figure 13 shows that corpus-specific polluted tokens are not limited to a single pollution type. Different corpus slices contain distinct Adult Con- tent, Online Gambling, Online Gaming, Online Video, and Anomalous tokens, reflecting source- specific and pipeline-specific lexical artifacts. This provides qualitative support that most polluted to- kens form corpus-specific tails rather than a shared vocabulary across all corpora. AdultGambling OSCAR 久精品国产 Jiujiu premium domestic 性欧美 European porn 禁网站 banned site 免费一级 free first-level 年买特马最准网站 most accurate site for buying special number picks 澳门特马今晚开奖 tonight’s Macau special number draw Gaming 完美世界前传下载 Perfect World prequel download 国际完美世界下载 International Perfect World download 完美世界国际版下载 Perfect World International version download Video 盗墓笔记第二季 The Lost Tomb Season 2 完美世界txt下载 Perfect World TXT download Anomalous 小说零 zero novels 亚洲人成色 Asian skin tone 人禽杂交 human-animal hybridization 最近中文字幕 recent Chinese subtitles mC4 免费黄片视频在线观看 free online porn videos 一级A做爰片免费视频 free X-rated adult film videos 凯发直播 Kaifa live streaming 搏天堂 Bo heaven 凯发官网 Kaifa official website 真人游戏下载手机版 mobile download for live- action games 天堂官方网站 Heaven official website 网络真人游戏 online live-action games 真人游戏在线 live-action games online 在线快 online fast 鎏焯檬只 liu chao meng zhi 环亚官网 Huanya official website 计划大师i苗军 Plan Master i Miaojun 娱乐官网免费下载 free download from entertainment official site 凯发官网平台 Kaifa official platform HPLT 韩国女主播 Korean female streamer 人 person 撸一撸 masturbate 重彩 chongcai pc蛋幸运 PC egg lucky 新开传奇私服 newly opened Legend private server 迷失传奇 Lost Legend 迷失传奇私服 Lost Legend private server 影音先 audio-video first 在播放 currently playing 圣王中文 Shengwang chinese 张布施 Zhang Bushi 第一王风 first Wang Feng 在看 watching 分彩 minute-by-minute lottery CulturaX 直播 live streaming 性增长 sexual growth 污小说 dirty novels 澳门百家 Macau hundreds 万达娱乐平台 Wanda entertainment platform 游戏游戏 game game 电竞下载 esports download 完美世界电竞app Perfect World Esports app 电竞比分直播 live esports scores 好织梦 good Zhimeng 企和 qi he 建和 jian he 弘扬民族优秀文化传统的互 联网视听节目 internet audiovisual programs promoting outstanding ethnic cultural traditions 百家 hundreds CWT n大学生 n college students n生活中 n in daily life n前列腺 n prostate 体育官方app下载 official sports app download n爱游戏 n love gaming n王者荣耀 n Honor of Kings n全民奇迹 n Quanmin Qiji n阴阳师 n Onmyoji 全场录像 full-game recording 全场录像回放 full-game replay n上一篇 n previous post n上一条 n previous item n下一条 n next item n短视频 n short videos 开云体育app Kaiyun Sports app ROOTS 性虐待和 sexual abuse and 性和透明度 sex and transparency 性旅游 sex tourism 捕鱼活动 fishing-game activities 可赔区 payout area 放置武器 placing weapons 格查尔 Quetzal 对消除 match-elimination 联合国网播 UN Web TV 和录像 and recordings 肯尼亚和 Kenya and 议员和 members of parliament and 国际贸易和 international trade and 下载今天的 download today's WanJuan 性站点 adult sites 性好的 sex is okay 轮过后 after the rounds 好彩网帝 good lottery web emperor 澳门游戏网站 Macau gaming website 策略手游 strategy mobile games AG超玩会 AG Super Play 仙侠手游 fantasy cultivation mobile games 视频加载中 video loading 视频中心 video center 向天岚 Xiang Tianlan 空包网 empty-package site 牧玄瑜 Mu Xuanyu 免费足球 free football 抽奖机会 lottery-draw chance MAPCC 用道具 use props 窥视卡 peeping card 原色地带 primary-color zone 六彩 six-color lottery 通吃岛证券网 take-all island securities site 暗夜精灵族 Night Elf race 网游之 Online Game: 泡堂 Crazy Arcade 下载点通 download Diandiantong 影视速递 film and TV express 只看该作者大 view only this author big 此主题相关图片如下 related images for this topic are as follows 点击观看 click to watch 六合采 Mark Six SkyPile 性消费 sexual consumption 性和准确性 sex and accuracy A级以上 above grade A 雷火网址 Leihuo website 扑克牌游戏 poker-card games 创意小游戏 creative mini-games 外挂软件 cheat software 小游戏的 mini-game’s 电子竞技游戏开播量 esports-game livestream starts 视频传输 video transmission 网通社从 netcom news agency from 损失以 loss by Bob直播间 Bob livestream room 排列三第 arrangement three number CCI3 性特征 sexual characteristics 性和创造性 sexuality and creativity 和学生的 with the student’s — 团大战 regiment battle 云十六 cloud sixteen 看动画片 watch cartoons 了学生的 le student’s 了学生 le student 股易舒 Guyishu WuDao 阳痿早泄 erectile dysfunction and premature ejaculation 男忄生问题 male problems 男忄生 male 多万字 many ten-thousand characters 游戏大小 game size 游戏加载 game loading 游戏加载完毕点击 click after game loading completes 视频接口 video interface 全村有 the whole village has 重梦境神话 heavy dream mythology 施绮 Shi Qi 日起任 ri qi ren 浏览过此新闻的网友还阅读 了以下新闻人 readers of this news also read the following news person 依法开展互联网视听节目服务 legally provide internet audiovisual program services Figure 13: Representative dataset-specific polluted tokens, with English glosses. Each cell lists high-ratio tokens that appear in only one of the 11 representative corpus slices, grouped by predicted category. 0.00.10.20.30.40.50.6 Top-100 token Jaccard Online Video Online Gaming Normal Content Adult Content Anomalous Online Gambling Longer year gaps expose token replacement 1-year gap 5-year gap 01020304050607080 Average yearly entrants in top-100 Online Video Online Gaming Normal Content Adult Content Anomalous Online Gambling 25.6 44.6 44.8 42.8 61.2 71.2 Fast-changing categories replace more tokens each year Figure 14: Additional top-100 token evolution sig- nals across Chinese Common Crawl snapshots. Top: adjacent-year and five-year Jaccard overlap. Bottom: average yearly entrants in the top-100 lists. F Additional Token Evolution Analyses Following Section 5, this section supplements the main all-token temporal analysis with high- frequency top-100 token views and a top-ksen- sitivity check. We focus on token turnover and persistence in Chinese Common Crawl snapshots, using two complementary signals: Jaccard overlap across time gaps and the average number of yearly entrants in the top-100 list. Figure 14 shows a clear contrast between stable and fast-changing categories. Online Video has the highest adjacent-year overlap and the fewest yearly entrants, indicating that its high-frequency token head is comparatively stable. Online Gam- bling shows the opposite pattern: it has the lowest overlap and the largest number of yearly entrants, indicating rapid replacement of high-frequency to- kens. Anomalous also shows substantial turnover, with many new top-100 tokens entering each year. Adult Content is more nuanced. Its adjacent-year top-token Jaccard remains moderate, but its five- year overlap is much lower, suggesting that this category combines short-term template stability with longer-term lexical replacement. This explains why Adult Content still shows substantial temporal drift even when recurring fragments such as “国 产精品(domestic premium content)”, “精品视 频 (premium videos)”, and “精品国产(domestic premium series)” remain visible across snapshots. Sensitivity to the top-kthreshold.We further re- peat the lexical-evolution analysis with top-10, top- 100, and full category vocabularies to test whether the conclusions depend on the truncation thresh- old. The main pattern is that smallerkemphasizes stability in the lexical head, whereas the full vo- cabulary exposes volatility in the long tail. This effect is especially strong for Online Video, whose adjacent-year Jaccard drops from 0.76 at top-10to 0.60 at top-100and then to 0.31 over the full vocab- ulary. Adult Content and Anomalous show similar declines, from 0.53 to 0.43 to 0.20 and from 0.49 to 0.25 to 0.18, respectively. In contrast, Normal Content remains relatively stable across thresholds, with adjacent-year Jaccard of 0.39 at top-10, 0.39 at top-100, and 0.48 over the full vocabulary. Online Gambling is unstable at every scale, with adjacent- year Jaccard staying low from 0.24 to 0.17 and 0.15. Overall, these ablations motivate the main-text all- token view when the goal is to expose long-tail lexical renewal, while the top-100analysis keeps the high-frequency patterns interpretable. Examples of short-lived and long-lived tokens. • Normal Content. Short-lived examples include “影片名称(film title)”, “世界杯(World Cup)”, and “我欲封天(I Shall Seal the Heavens)”, which spike in a single year; long-lived exam- ples include “中文字幕(Chinese subtitles)”, “联系我们(contact us)”, “有限公司(limited company)”, and “关于我们(about us)”, which persist across most or all snapshots. This cate- gory is comparatively stable overall. • Adult Content. Short-lived examples include “日本一级特黄大片(Japanese explicit-content phrase)”, “亚洲成人(Asian adult)”, “亚洲天 堂 (Asian paradise)”, and “一区二区在线观看 (Zone 1/2 watch online)”; long-lived examples include “国产精品(domestic premium con- tent)”, “精品视频(premium videos)”, “精品国 产 (domestic premium series)”, and “久 久(repeated jiu-jiu template)”, reflecting a sta- ble template-like core mixed with rapid surface turnover. The key pattern is stable templates plus changing surface realizations. •Online Gambling. Short-lived examples in- clude “太阳城(Sun City)”, “乐橙体育 (Lecheng Sports)”, “申博太阳城(Shenbo Sun City)”, and “云顶集团手机版(Genting Group mobile version)”; long-lived examples include “开奖结果(lottery results)”, “双色球(Double Color Ball)”, “娱乐城(casino)”, “免费一级 (free first-level content)”, and “澳门正版资料 (Macau official materials)”, showing recurring gambling-related templates but rapid platform and brand replacement. This is the clearest case of rapid lexical turnover. • Online Video. Short-lived examples include “盗墓笔记第二季(The Lost Tomb Season 2)”, “完美世界txt下载(Perfect World TXT down- load)”, and “网址在线观看(watch online at URL)”; long-lived examples include “在线观 看(watch online)”, “在线播放(stream online)”, “在线视频(online video)”, “免费视频(free video)”, and “免费观看(watch for free)”, con- sistent with the high stability of generic video- access phrases. This is the most stable category. •Anomalous. Short-lived examples include “小 说零(novel zero)”, “一二(Zone 1/Zone 2)”, “五月婷(May Ting)”, and “最高占成(high- est commission share)”; long-lived examples include “一区二区三区(Zones 1/2/3)”, “一区 二区(Zones 1/2)”, “一二三区(Zones 1/2/3)”, “大香蕉(big banana)”, and “卡三卡(card- three-card)”, illustrating both recurring peculiar phrases and rapidly changing surface variants. This category combines contextually irrelevant recurring forms with strong year-to-year drift. •Online Gaming. Short-lived examples include “完美世界前传下载(Perfect World prequel download)”, “国际完美世界下载(Interna- tional Perfect World download)”, and “小说 改编的网页游戏(web games adapted from novels)”; long-lived examples include “破解版 (cracked version)”, “小游戏(mini games)”, “游 戏平台 (game platform)”, “王者荣耀(Honor of Kings)”, and “游戏下载(game download)”, indicating a more persistent gaming vocabulary than Gambling or Anomalous. Its stability is closer to mainstream content than to Gambling or Anomalous. G Token Clouds Following Section 4 and Section 5, this section provides qualitative token-cloud views of high- frequency Chinese tokens in the audited open cor- pora and Chinese Common Crawl snapshots. For each corpus, we read long Chinese tokens (more than 2 Chinese characters), use token-count statis- tics to rank tokens, and visualize the top-2,000 tokens. Font size is determined by count ranking. These token clouds visualize surface forms rather than category ratios, providing qualitative context for Table 1 and Table 2. Figure 15 shows that the six representative broad- web corpus slices differ sharply in their high- frequency surface forms. OSCAR is visually domi- nated by video and adult-access templates, whereas mC4 contains many lottery- or betting-like phrases. HPLT, CulturaX, CWT, and ROOTS show more generic web, institutional, or template-like text. This qualitative contrast aligns with the main re- sult that pollution is widespread but uneven across corpus families, and that different corpora exhibit different dominant pollution profiles. Figure 16 shows a different pattern for curated Chinese corpora. Their top tokens are mostly generic content or web-template phrases, consistent with their much lower pollution ratios in Table 1. However, because token clouds only visualize the highest-count tokens, they should not be read as ev- idence that these corpora are pollution-free. Rather, they show that in cleaner corpora, residual polluted tokens are less visually dominant and often remain in a lower-frequency tail. Figure 17 provides a temporal view of Com- mon Crawl. The dominant high-frequency tokens change across yearly snapshots, with visible shifts among video-access phrases, site templates, and adult-content templates. This supports the tempo- ral analysis in Section 5: Chinese Common Crawl pollution is not only high in aggregate, but also changes in its surface vocabulary over time. OSCARmC4 HPLTCulturaX CWTROOTS Figure 15: Token clouds for six representative broad-web corpus slices. WanJuanMAPCC SkyPileCCI3 WuDao Figure 16: Token clouds for curated Chinese corpus slices. 20212022 20232024 20252026 Figure 17: Token clouds for Chinese Common Crawl snapshots from 2021 to 2026. H Additional Token Tree and Search Evidence Examples Following Section 6, this section provides addi- tional token-tree examples and their corresponding search evidence. These examples expand the main- text discussion in two directions. First, they show how a tree can trace an entire polluted family back to a compact root, making related surface variants reviewable as one group rather than as disconnected strings. Second, they show how pollution can arise compositionally: a longer token can become pol- luted even when the shorter subtokens from which it is formed are normal in isolation. The paired search-evidence figures further show why category labels are assigned with contextual web evidence rather than from token strings alone. Figure 18 illustrates the compositional case with “北京赛车”. The shorter subtokens “北京(Bei- jing)” and “赛车(Car racing)” are normal when in- terpreted separately, but their combination refers to a PK10-style lottery betting term. The descendants then capture common gambling-search variants such as platform, strategy, purchase, and WeChat- group phrases. This example shows that a flat key- word list would either miss the composed polluted expression or over-block its normal components; the tree instead localizes the polluted meaning to the composed token family. Figure 19 shows a second compositional case. Here “菲律宾(Philippines)” and “申博(Apply for a PhD)” can be normal tokens separately, but their combination is associated with an online gambling brand. Descendants such as agent, official-site, lo- gin, and international-casino variants reveal how one composed token expands into a broader pro- motional family. This case is important because it shows that pollution can emerge from ordinary- looking named entities and abbreviations, rather than from an obviously harmful root. Figure 20 and Figure 21 illustrate the polluted- root pattern. In Figure 20, “大道香蕉” functions as a recurring adult-site keyword, and the descen- dants add modifiers such as duration, language, online viewing, and video type. In Figure 21, “浪 小辉” refers to an adult-content entity, and the de- scendants add identity, regional, and content-type variants. In both cases, the root itself is already polluted, so the tree acts as a compact summary of many longer surface forms. This makes the dataset easier to review and helps future cleaning systems generalize beyond exact-string matching. The corresponding search-evidence figures pro- vide the contextual basis for these labels. Figure 22 and Figure 23 show that the gambling labels are supported by search results and webpages con- nected to lottery betting, gambling platforms, and betting-site promotion. Figure 24 and Figure 25 show that the adult-content labels are supported by adult-site results and explicit-content contexts. These examples illustrate a recurring difficulty in Chinese web-pollution auditing: many polluted to- kens are short, euphemistic, or overloaded with be- nign meanings. Web context is therefore necessary for auditable category assignment, while the tree structure makes the resulting decisions traceable across related token variants. Online gambling 北京赛车 A PK10-style online lottery Normal content 北京 Beijing Normal content 赛车 Car racing A polluted online gambling token family rooted at "北京赛车" Representative token: A PK10-style online lottery betting game spread through tutorials, sure-win claims, and group recruitment, a typical gambling promotion term. ② Polluted tokens can emerge from composing normal tokens. Online gamblingNormal + Normal Online gambling 北京赛车PK + PK Online gambling 北京赛车微信群 + WeChat group Online gambling 北京赛车怎么 + how Online gambling 北京赛车怎么买 + buy Online gambling 北京赛车怎么玩 + play Online gambling 北京赛车怎么买稳赢 +sure win Online gambling 北京赛车怎么玩才 + so as to Online gambling 北京赛车怎么玩才稳 + win steadily Figure 18: Token tree rooted at “北京赛车”. The tree shows how two individually normal subtokens, “北京” and “赛车”, compose into an Online Gambling token family. Online gambling 菲律宾申博 An online gambling brand Normal content 菲律宾 Philippines Normal content 申博 Apply for a PhD A polluted online gambling token family rooted at "菲律宾申博" Representative token: Linked to the offshore gambling platform Sunbet and its account- registration sites, a typical gambling promotion term. ② Polluted tokens can emerge from composing normal tokens. Online gamblingNormal + Normal Online gambling 菲律宾申博代理 + agent Online gambling 菲律宾申博亚洲 + Asia Online gambling 菲律宾申博太阳 + sun Online gambling 菲律宾申博太阳城 + city Online gambling 菲律宾申博太阳城官网 + official site Online gambling 菲律宾申博太阳城国际 + international Online gambling 菲律宾申博太阳城游戏 + games Online gambling 菲律宾申博太阳城官网打不开 + can’t open Online gambling 菲律宾申博太阳城国际娱乐 + gaming Online gambling 菲律宾申博太阳城游戏登入 + login Online gambling 菲律宾申博太阳城国际娱乐城 + casino Online gambling 菲律宾申博太阳城游戏登入不了 + fails Figure 19: Token tree rooted at “菲律宾申博”. The tree shows how normal subtokens can compose into a family of Online Gambling tokens. Representative token: "Dadao Xiangjiao" is a recurring spam keyword on Chinese adult-video sites, a token directly linked to sexually explicit content. ① Roots capture minimal recurring polluted subtokens. A polluted adult content token family rooted at "大道香蕉" Adult content 大道香蕉久 + jiu Adult content 大道香蕉中文大在线 + Chinese,online Adult content 一本大道香蕉 + yiben Adult content 一本大道香蕉中文 + Chinese Adult content 一本大道香蕉综合视频 + compilation videos Adult content 一本大道香蕉中文在线 + online Adult content 欧美一本大道香蕉综合视频 + Western Adult content 大道香蕉 A common adult-site keyword Adult content 专区欧美一本大道香蕉综合视频 + section Adult content 国产v亚洲v专区欧美一本大道香 蕉综合视频 + domestic and Asian tags Adult content 一本大道香蕉中文在线视频 + video Adult content 一本大道香蕉中文在线视频观看 + watch Figure 20: Token tree rooted at “大道香蕉”. The root captures a recurring Adult Content subtoken, and descendants show longer adult-site variants. Adult content 浪小辉 A gay porn actor Adult content 猛男浪小辉 A muscular hunk Adult content 中国男同浪小辉 A Chinese gay porn actor Representative token: "Lang Xiaohui" is a performer in Chinese gay adult videos, a token directly linked to sexually explicit content. ① Roots capture minimal recurring polluted subtokens. A polluted adult content token family rooted at "浪小辉" Adult content 浪小辉GAY国产 Domestic gay porn star Adult content 打桩浪小辉 Explicit gay porn content Adult content 猛男浪小辉GAY国产 + domestic GAY Adult content chinese猛男浪小辉 + chinese Adult content CHINESE猛男浪小辉GAY国产 + CHINESE Adult content chinese猛男浪小辉gay国产 + domestic gay Figure 21: Token tree rooted at “浪小辉”. The root captures a recurring Adult Content entity, and descendants show longer surface variants. 部分待验证token A lottery betting game Figure 22: Search evidence for “北京赛车”, showing web context to identify the token family as Online Gambling. 部分待验证token gambling platform Figure 23: Search evidence for “菲律宾申博”, showing web context to identify the token family as Gambling. A common adult-site keyword Figure 24: Search evidence for “大道香蕉”, showing web context to identify the token family as Adult Content. 部分待验证token A gay porn actor Figure 25: Search evidence for “浪小辉”, showing web context used to identify the token family as Adult Content.