Paper deep dive
CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts
Runhan Song, Qiqi Liu, Chuanzhou Pan, Zhenquan Ding, Youquan Xian, Chongru Fan, Lei Cui, Wei Wang, Zhiyu Hao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/17/2026, 4:34:06 AM
Summary
The paper introduces CipherSight, a robust HTTPS website fingerprinting (WF) framework that addresses out-of-distribution (OOD) challenges caused by temporal and geographic drifts. Unlike existing methods relying on raw TCP packet sequences, CipherSight utilizes TLS-record-level attributes and a hierarchical transformer architecture to capture intra-flow and inter-flow dependencies. It employs Masked Record Modeling (MRM) for pretraining and fine-grained record-resource semantic supervision (privileged supervision) during fine-tuning to learn stable, generalizable representations. Experiments demonstrate that CipherSight achieves 95.41% accuracy in closed-world settings and maintains over 90% accuracy under distribution shifts, outperforming baselines.
Entities (10)
Relation Signals (8)
CipherSight → employs → Masked Record Modeling
confidence 95% · CipherSight employs a masked record modeling (MRM) task to capture contextual traffic semantics
CipherSight → uses → TLS Record
confidence 95% · CipherSight learns website representations from TLS records by jointly encoding multiple record-level attributes.
CipherSight → employs → LoRA
confidence 90% · A privileged semantic teacher further transfers resource-level knowledge through LoRA fine-tuning.
CipherSight → handles → Temporal Drift
confidence 90% · It also maintains over 90% accuracy in geographic drift settings... and retains 92.99% under temporal drift
CipherSight → handles → Geographic Drift
confidence 90% · It also maintains over 90% accuracy in geographic drift settings
CipherSight → outperforms → STC-WF
confidence 90% · CipherSight... consistently outperforming all evaluated baselines.
CipherSight → outperforms → VarCNN
confidence 90% · CipherSight achieves 95.41% accuracy... consistently outperforming all evaluated baselines.
Existing Methods → relieson → TCP Packet
confidence 90% · Existing methods primarily learn from raw TCP packet sequences
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:HTTPS website fingerprinting (WF) aims to identify visited websites from metadata observable in encrypted traffic. However, real-world deployments introduce a significant out-of-distribution (OOD) problem caused by temporal and geographic changes, while previously unseen websites are common in open-world scenarios. Existing methods primarily learn from raw TCP packet sequences and struggle to capture stable and generalizable website representations, resulting in performance degradation under practical conditions. We propose CipherSight, a TLS-record-based hierarchical framework for robust HTTPS WF. Unlike existing approaches that rely on TCP packet sequences and are sensitive to transport-layer artifacts, CipherSight learns website representations from TLS records by jointly encoding multiple record-level attributes. It introduces a hierarchical architecture that captures both intra-flow dependencies among TLS records and inter-flow interactions across concurrent flows, enabling the model to exploit structural patterns in HTTPS traffic. Besides, to learn robust representations, CipherSight employs a masked record modeling (MRM) task to capture contextual traffic semantics and leverages fine-grained record-resource annotations as privileged supervision through structure-aware objectives and semantic distillation. Experiments show that CipherSight achieves 95.41% accuracy across more than 2,000 website classes in the closed-world setting and maintains over 90% accuracy under both temporal and geographic drift, consistently outperforming all evaluated baselines.
Tags
Links
- Source: https://arxiv.org/abs/2608.13905v1
- Canonical: https://arxiv.org/abs/2608.13905v1
Trouble viewing inline? Open PDF directly →
Full Text
44,436 characters extracted from source content.
Expand or collapse full text
CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts Runhan Song Qiqi Liu Chuanzhou Pan Zhenquan Ding Youquan Xian Chongru Fan Lei Cui Wei Wang Zhiyu Hao Abstract HTTPS website fingerprinting (WF) aims to identify visited websites from metadata observable in encrypted traffic. However, real-world deployments introduce a significant out-of-distribution (OOD) problem caused by temporal and geographic changes, while previously unseen websites are common in open-world scenarios. Existing methods primarily learn from raw TCP packet sequences and struggle to capture stable and generalizable website representations, resulting in performance degradation under practical conditions. We propose CipherSight, a TLS-record-based hierarchical framework for robust HTTPS WF. Unlike existing approaches that rely on TCP packet sequences and are sensitive to transport-layer artifacts, CipherSight learns website representations from TLS records by jointly encoding multiple record-level attributes. It introduces a hierarchical architecture that captures both intra-flow dependencies among TLS records and inter-flow interactions across concurrent flows, enabling the model to exploit structural patterns in HTTPS traffic. Besides, to learn robust representations, CipherSight employs a masked record modeling (MRM) task to capture contextual traffic semantics and leverages fine-grained record-resource annotations as privileged supervision through structure-aware objectives and semantic distillation. Experiments show that CipherSight achieves 95.41% accuracy across more than 2,000 website classes in the closed-world setting and maintains over 90% accuracy under both temporal and geographic drift, consistently outperforming all evaluated baselines. Introduction Website fingerprinting (WF) aims to identify the websites visited by users from metadata observable in encrypted traffic, without decrypting application payloads. It provides a practical approach for studying residual information leakage in encrypted communications and evaluating the effectiveness of traffic analysis defenses (18). Early WF studies primarily relied on explicit signals, including DNS queries, TLS Server Name Indication (SNI), and destination IP addresses (5). However, these signals are becoming increasingly unreliable due to emerging privacy-preserving mechanisms, e.g., Encrypted Client Hello (ECH) hides SNI, encrypted DNS protocols protect domain resolution, and shared CDN services weaken the direct mapping between IP addresses and websites (22; 13; 12; 14). Consequently, recent HTTPS WF approaches, inspired by Tor-based traffic analysis, exploit ciphertext-observable side channels, ranging from packet-level features, e.g., lengths, directions, and timing patterns (26; 1; 21; 7) to HTTPS-specific flow statistics and higher-layer traffic structures (3; 5; 28; 19). Although existing methods achieve high performance under closed-world settings, real-world deployment remains challenging mainly due to distribution shifts and unseen traffic in open-world scenarios. Website content evolves over time, while routing paths, CDN selection, and localized resources vary across regions, causing substantial shifts in encrypted traffic distributions. We therefore evaluate temporal and geographic drift as OOD generalization settings, together with open-world recognition, which requires detecting websites absent from training (17; 32; 15). There are three challenges that arise in these settings. Challenges First, packet-level representations are highly sensitive to transport-layer variability. TCP segmentation, retransmissions, and dynamic network conditions can introduce substantial variations, causing the same website resources to be mapped into substantially different TCP packet sequences. Such network-induced variations introduce noise into learned website representations and hinder robust generalization. According to our analysis, more than 80% of TCP packet sequences are inconsistent across repeated traffic captures, highlighting their instability under diverse real-world network conditions. Second, modern HTTPS page loads involve multiple concurrent flows. Tor-oriented methods typically flatten a trace into a single packet sequence (26; 1; 21; 7), implicitly overlooking the inherent multi-flow structure of HTTPS traffic. Recent HTTPS-specific approaches have started to incorporate flow context (28; 19). However, simply modeling individual flows remains insufficient. Packets within the same flow share connection-level and transport semantics, while interactions among concurrent flows collectively reveal the organization of webpage loading. Thus, a robust representation should capture these two complementary levels of dependencies. Third, encryption creates a semantic gap between observable traffic patterns and underlying webpage composition. Web resources are serialized into HTTP messages, encrypted into TLS records, and transported as TCP packets. While resource-level properties, such as type, size, and request semantics, more directly characterize webpage composition, encrypted traces expose only indirect manifestations of these properties. Training solely with website-level labels leaves the correspondence between TLS records and resource semantics largely unexplored, preventing the model from leveraging fine-grained resource information for robust WF. Figure 1: Challenges and motivation for CipherSight. Robust HTTPS WF requires reduced sensitivity to transport variability, explicit modeling of dependencies within and across flows, and supervision that connects encrypted records with web-resource structure. To address these challenges, we propose CipherSight, a hierarchical transformer framework for robust HTTPS WF. Unlike existing packet-level representations that are highly affected by transport-layer dynamics, CipherSight models webpage loads at the TLS record level, which provides a more stable abstraction over TCP packets. To capture the hierarchical structure of modern HTTPS traffic, CipherSight introduces an intra-flow encoder using block-diagonal self-attention to learn dependencies among records within each flow and an inter-flow encoder that models interactions across concurrent flows for webpage-level representation learning. Furthermore, CipherSight exploits fine-grained resource structures as privileged supervision, aligning TLS record representations with resource boundaries and attributes to learn richer semantic representations beyond website-level labels. A privileged semantic teacher further transfers resource-level knowledge through LoRA fine-tuning. Privileged components are discarded at inference. Figure 1 summarizes the motivation and corresponding design principles. Extensive experiments on more than 2,000 monitored websites show that CipherSight achieves 95.41% closed-world accuracy and retains 92.99% under temporal drift, exceeding the best-performing baseline by 13.03 percentage points (p). It also maintains over 90% accuracy in geographic drift settings, outperforming the best regional baselines by at least 8.52 p, and achieves 97.64% AUROC in open-world recognition. Ablations further confirm that privileged supervision provides its clearest gains under geographic drift. Contributions Our contributions are summarized as follows: • We design a structure-preserving representation for HTTPS WF that reduces dependence on TCP packetization, bringing it closer to higher-level semantics. • We propose a hierarchical architecture that jointly models multiple TLS-record attributes and explicitly separates intra-flow dependencies from inter-flow interactions, which suppresses transport-specific artifacts without discarding the multi-flow structure of web pages. • We introduce fine-grained privileged supervision based on web-resource semantics for HTTPS WF. CipherSight not only aligns TLS-record states with resource boundaries and attributes, but also transfers resource knowledge through privileged distillation during LoRA fine-tuning. • We evaluate CipherSight on six HTTPS collections covering more than 2,000 websites under closed-world, temporal-drift, geographic-drift, and open-world settings. The results demonstrate gains in closed-world top-1 accuracy, temporal and geographic robustness, and open-world AUROC over representative baselines. Related Work Website Fingerprinting Historically, HTTPS destinations could often be identified directly from hostnames exposed in plaintext DNS queries or the SNI field of TLS ClientHello (2). Visited websites can also be inferred from the sets and sequences of contacted destination IP addresses (11). However, encrypted DNS and ECH increasingly conceal hostname signals, while CDN deployment, shared hosting, DNS-based load balancing, and IP address churn weaken direct domain-to-IP mappings. Consequently, reliable website identification increasingly depends on side-channel patterns that remain observable in encrypted traffic under realistic network environments. Inferring visited websites from encrypted traffic side channels is commonly known as website fingerprinting (WF). WF has been studied most extensively in Tor, an anonymity network designed to conceal communication destinations (8; 16). Tor-oriented methods commonly represent a page load as a flat sequence of observable packet or cell directions, bursts, and timings. Convolutional models such as AWF, DF, and VarCNN learn discriminative representations from directional sequences (24; 26; 1), while TikTok explicitly incorporates timing information (21). Although effective on Tor traces, these flat sequence representations, when transferred to HTTPS traffic, do not explicitly preserve TLS-record boundaries or the concurrent multi-flow structure of modern page loads. Website Fingerprinting over HTTPS HTTPS-native WF methods increasingly move beyond flat TCP packet sequences. H&W derives lightweight fingerprints from HTTP-version-dependent parallel loading, while ADU recovers application data unit length sequences before classification (5; 3). STC-WF models inter-flow spatio-temporal correlations using a graph neural network, and CTX-Aware exploits flow context in encrypted-proxy traffic (28; 19). These studies demonstrate the benefits of higher-layer and flow-aware representations, but do not jointly encode multiple TLS-record attributes with explicit intra-flow and inter-flow dependencies in a unified hierarchy. Semantic information has also been explored for improving WF generalization. STAR aligns encrypted traffic traces with crawl-time semantic profiles for zero-shot website retrieval (4). Related preprints investigate resource-level semantic distillation and semantics-aware traffic augmentation (10; 31). CipherSight differs by treating fine-grained alignment between TLS records and resource spans as privileged information under the LUPI paradigm (29), combining record- and span-level supervision with privileged distillation in a hierarchical ciphertext encoder. Website Fingerprinting under Distribution Shifts High performance in closed-world settings cannot guarantee generalization to real-world conditions. Prior Tor studies show that changes in collection time, network environment, and browsing context, together with unmonitored traffic, can induce distribution shifts that substantially degrade WF performance (16; 6; 15; 32; 7). Temporal and geographic drift alter the input distribution over monitored classes, whereas open-world recognition additionally requires rejecting websites unseen during training (17; 30). Existing robustness studies primarily focus on Tor or isolate a single source of mismatch. We therefore evaluate HTTPS WF under temporal drift, geographic drift, and open-world recognition to assess robustness across complementary deployment conditions. Method Figure 2: Overview of CipherSight. During representation construction, observable TLS records are grouped by flow and tokenized. Block-diagonal intra-flow attention (L1) encodes record tokens within each flow, while full inter-flow attention (L2) aggregates flow summaries for website classification. Training-time privileged structure provides record-resource supervision (P1) and semantic teacher distillation (P2). Overview Figure 2 presents CipherSight, whose representation construction, hierarchical architecture, and privileged structure are motivated by the three challenges identified in Figure 1. First, CipherSight represents traffic at the TLS-record level rather than the TCP-packet level, reducing sensitivity to transport-layer packetization. Ciphertext-observable record attributes are extracted and tokenized, while masked record modeling (MRM) guides pretraining. Second, record tokens are processed by an intra-flow encoder (L1) and an inter-flow encoder (L2), which separately model dependencies within flows and interactions across flows, reflecting the multi-flow organization of HTTPS page loads. Third, training-time privileged supervision connects ciphertext patterns with webpage structure. Record-resource supervision (P1) is used during pretraining and fine-tuning, while a semantic teacher (P2) is introduced during fine-tuning. Representation and Hierarchical Encoding TLS-Record Representation and Tokenization. TLS records are protocol data units carried above TCP. After TCP reassembly, their boundaries and outer headers can be recovered without decryption (23; 9). A TLS record may span multiple TCP segments, while a single segment may contain bytes from multiple records. Consequently, record-level modeling is less directly coupled to TCP segmentation and retransmission than packet-level modeling, although it remains affected by changes in website content and TLS behavior. In the representation construction phase, TLS records extracted from webpage-load trace ri,jr_i,j are grouped into flows fif_i. Categorical features are mapped to learnable embeddings, whereas numerical features combine logarithmic bucket embeddings with linear projections of their log-scaled values. The resulting feature vectors are summed and normalized to produce the record token i,je_i,j. Intra-Flow Encoding (L1). Flattening all records into a single sequence can create direct dependencies between unrelated connections and make the representation sensitive to their collection-specific interleaving. L1 instead encodes records from the same flow to preserve TLS-flow boundaries. For every valid flow fif_i, we prepend an instance of a shared learnable token FCLSiFCLS_i to form (FCLSi,i,1,…,i,ni)(FCLS_i,e_i,1,…,e_i,n_i). All flow sequences are concatenated for batched processing by a shared Transformer. Let ϕ(p)φ(p) denote the flow associated with token position p. L1 applies the attention bias Mpq(1)=0,ϕ(p)=ϕ(q) and p,q are valid,−∞,otherwise.M^(1)_pq= cases0,&φ(p)=φ(q) and p,q are valid,\\ -∞,&otherwise. cases (1) The block-diagonal mask uses each TLS flow as a local context boundary and restricts self-attention to records from that flow. Copies of a shared learnable [FCLS] token aggregate within-flow information into flow summaries iu_i, while i,jh_i,j denotes the contextual representation of each record. Rotary positions are reset within each flow, and cross-flow interaction is deferred to L2. Inter-Flow Encoding (L2). A modern website load typically involves connections from multiple domains, so a single flow cannot represent the complete webpage structure. CipherSight therefore prepends a learnable page token [PCLS] to the valid flow summaries and feeds (PCLS,1,…,m)(PCLS,u_1,…,u_m) to the inter-flow encoder L2. Unlike L1, L2 applies full self-attention across valid flow representations. Using flow summaries as an interface between L1 and L2 avoids all-to-all mixing between individual records from different connections while allowing webpage-level information to be combined across flows. The final hidden state of [PCLS] serves as the student page representation sh_s. A website classifier consisting of a linear layer, GELU activation, dropout, and a final class projection produces the website logits so_s. Pretraining with Record-Resource Supervision Figure 3: Fine-tuning architecture of CipherSight. The LoRA-adapted ciphertext student receives resource-structure supervision (P1), while a privileged semantic teacher (P2) transfers resource-level knowledge through feature and logit distillation, LfeatL_feat and LlogitL_logit. The student and teacher are supervised by LclsSL_cls^S and LclsTL_cls^T, respectively. Only the student is retained for inference. Masked Record Modeling (MRM). Website labels do not directly supervise individual record representations. MRM replaces the complete embeddings of randomly selected valid records with a learnable [MASK] token and requires L1 to reconstruct their direction, length bucket, and outer TLS content type from the remaining within-flow context. The three cross-entropy losses are summed as LMRML_MRM, enabling contextual record modeling without resource annotations. Record-resource Supervision (P1). MRM captures ciphertext context but does not explicitly model the correspondence between TLS records and webpage resources. Prior work has demonstrated that resource semantics can improve the robustness of WF representations (10). Motivated by this observation, we provided novel training-time alignments between website resources and TLS records in P1 to supervise L1 at both record and resource levels. At the record level, P1 predicts begin/inside labels and inner MIME types over aligned records, while unaligned records are ignored by these losses. At the resource level, aligned record states are mean-pooled to predict response attributes and record counts. Algorithm 1 summarizes this two-level supervision and pooling procedure. Algorithm 1 Record-Resource Supervision and Aligned-Record Pooling Input: Record states H, record targets BIY_BI and typeY_type, retained aligned spans Iq\I_q\, and resource attributes resY_res Output: Lseg,Lattr,LcntL_seg,L_attr,L_cnt 1: (^BI,^type)←Hrec()( Y_BI, Y_type)← H_rec(H). 2: Compute LsegL_seg over aligned record positions. 3: Initialize G←∅G← and T←∅T← . 4: for each resource q with a nonempty retained span IqI_q do 5: q←|Iq|−1∑(i,j)∈Iqi,jg_q←|I_q|^-1 _(i,j)∈ I_qh_i,j. 6: Append qg_q to G. 7: Append (yqlen,yqmime,log(1+|Iq|))(y_q^len,y_q^mime, (1+|I_q|)) to T. 8: end for 9: (^len,^mime,^)←Hres(G)( y^len, y^mime, c)← H_res(G). 10: Compute LattrL_attr and LcntL_cnt using T. 11: return Lseg,Lattr,LcntL_seg,L_attr,L_cnt. Pretraining Objective. The pretraining stage objective combines ciphertext-context reconstruction from MRM with resource-structure supervision (P1): LP1 L_P1 =λsegLseg+λattrLattr+λcntLcnt, = _segL_seg+ _attrL_attr+ _cntL_cnt, (2) Lpre L_pre =λMRMLMRM+LP1. = _MRML_MRM+L_P1. MRM learns dependencies among observable records, whereas P1 relates their hidden states to web-resource structure. The pretraining stage directly optimizes the tokenizer and L1, and L2 is introduced for website classification in the fine-tuning stage. Privileged Semantic Fine-Tuning LoRA Adaptation. Pretraining establishes record-level representations, which should not be overwritten by unrestricted task adaptation. We therefore freeze the pretrained backbone and introduce low-rank updates W~=W+(α/r)BA W=W+(α/r)BA into the query, key, value, and output projections of every L1 and L2 attention layer. The LoRA matrices, student classifier, P1 heads, and semantic teacher remain trainable. This allows L2 to learn webpage-level aggregation while limiting changes to the pretraining representation. Privileged Semantic Teacher (P2). Website-label fine-tuning alone may emphasize environment-specific correlations and weaken the semantic structure learned during pretraining. As shown in Figure 3, privileged semantic teacher (P2) supplies website-level semantic guidance using resource attributes available only during training. P2 first constructs resource tokens from resource semantic features. Then a resource semantic encoder summarizes the resource token sequence as rh_r. The teacher fuses rh_r with a stop-gradient copy of the student feature sh_s to produce th_t, and an MLP produces teacher logits to_t. The teacher is trained with website-label cross-entropy and uses LfeatL_feat and LlogitL_logit to distill resource semantics into the student. Formally, let ¯x=N(x) h_x=N(h_x) and x=softmax(x/T)p_x=softmax(o_x/T) for x∈s,tx∈\s,t\. The distillation losses are Lfeat L_feat =MSE(¯s,sg[¯t]), =MSE ( h_s,sg[ h_t] ), (3) Llogit L_logit =T2KL(sg[t]∥s). =T^2KL (sg[p_t]\,\|\,p_s ). Here, N(⋅)N(·) denotes L2 normalization, T is the distillation temperature, and sgsg stops gradients. Model Origin Accuracy Precision Recall Macro-F1 Top-5 VarCNN Tor 92.16% 92.63% 92.02% 91.91% 97.27% ARES Tor 91.22% 91.65% 91.07% 90.85% 97.11% RF Tor 91.09% 91.27% 90.84% 90.63% 97.16% DF Tor 89.85% 90.29% 89.56% 89.36% 96.61% TikTok Tor 89.32% 89.69% 89.04% 88.77% 96.50% TF Tor 86.94% 87.37% 86.62% 86.41% – AWF Tor 62.66% 66.33% 62.40% 61.95% 81.59% H&W HTTPS 86.45% 86.67% 86.23% 85.76% 94.31% STC-WF HTTPS 77.75% 80.00% 77.66% 77.34% 90.65% CTX-Aware HTTPS 08.18% 07.00% 08.14% 05.88% – CipherSight HTTPS 95.41± 0.11% 94.85± 0.14% 95.00± 0.13% 94.67± 0.14% 97.71± 0.02% Table 1: Closed-world performance comparison on 2,008 website classes. CipherSight results are reported as the mean ± standard deviation over five seeds. Fine-Tuning Objective and Inference. Let LclsS=CE(s,y)L_cls^S=CE(o_s,y) and LclsT=CE(t,y)L_cls^T=CE(o_t,y) denote the student and teacher classification losses. The fine-tuning stage minimizes Lft= L_ft= λclsLclsS+λstruct(Lseg+Lattr+Lcnt)+LclsT _clsL_cls^S+ _struct(L_seg+L_attr+L_cnt)+L_cls^T (4) +λlogitLlogit+λfeatLfeat. + _logitL_logit+ _featL_feat. Continuing P1 helps preserve record-resource structure during adaptation, while P2 transfers webpage-level semantics to the ciphertext student. MRM is used only in pretraining. At inference, CipherSight uses only TLS-record metadata, the L1 and L2 encoders, and the website classifier. The MRM, P1, P2, and all resource annotations are removed. Evaluation Datasets and Experimental Setup Collection and Construction. Each automated webpage visit produces a PCAP trace and a corresponding SSL key log. We reconstruct TCP flows and derive the model inputs from ciphertext-observable TLS metadata, including record direction, length, outer content type, timing, and flow membership, without decrypting application payloads. For dataset annotation only, we use the SSL key log with standards-compliant TLS and HTTP parsing to recover the record-resource alignments and resource attributes required by P1 and P2. Datasets. We evaluate six collections spanning four dates and three regions. Closed-world evaluation uses an 8:2 split of us0304, yielding 64,386 training and 15,848 test traces from 2,008 classes. Temporal drift trains on 80% of us0304 and tests us0320, collected 16 days later, over 1,976 shared classes. Geographic drift trains on 80% of us0409 and tests same-day fr0409 and sg0409 over 1,936 shared classes. The held-out 20% of us0409 traces measure performance degradation. Open-world evaluation reuses the closed-world split and adds 213,099 traces from disjoint websites. Metrics. We report top-1 accuracy, macro-F1, and top-5 accuracy where supported. Open-world confidence is maximum softmax probability. With monitored traffic as the positive class, AUROC and AUPR measure known-unknown discrimination, while OSCR jointly reflects monitored classification and unknown rejection across thresholds. Baselines. We compare Tor-oriented methods (24; 26; 1; 21; 27; 25; 7) and HTTPS-native methods (5; 28; 19). Tor baselines use 5,000-packet sequences, truncated or zero-padded, with direction, length, and timing when applicable. Missing preprocessing or open-world evaluation is implemented from reported definitions without altering model architecture. Neural baselines train for 100 epochs, except STC-WF for 300. To favor baselines, we select closed-world checkpoints by peak test accuracy and drift checkpoints on the fixed source-domain 20% holdout before evaluation on complete target sets. Open-world evaluation reuses those checkpoints. CipherSight uses its final 20,000-step checkpoint. Implementation. CipherSight uses PyTorch, 8 NVIDIA GPUs, and BF16 precision. Both stages run for 20,000 steps. CipherSight uses seeds 42–46, and the baselines use seed 42 unless stated otherwise. Closed-World Evaluation As shown in Table 1, CipherSight achieves the best closed-world performance across 2,008 website classes, reaching 95.41% accuracy and 94.67% macro-F1. It outperforms VarCNN, the strongest baseline, by 3.25 and 2.76 percentage points (p) on these metrics, respectively. The standard deviations remain below 0.15 p across five seeds, indicating stable performance across runs. By comparison, the top-5 improvement over VarCNN is only 0.44 p. This contrast indicates that CipherSight primarily improves first-choice discrimination rather than merely retaining the correct website among several candidates. STC-WF warrants a qualified interpretation. A separate run with seed 52 on the same split reaches 84.78% accuracy, while it achieves 92.25% on the us0409 source-domain holdout under the geographic-drift protocol. These results indicate sensitivity to training randomness rather than uniformly weak performance. CTX-Aware achieves only 8.18% accuracy under our adaptation. Its full configuration required more than 128 GB of memory, so we restricted each webpage trace to six flows and used a smaller random forest. Temporal Drift Model Accuracy (%) Macro-F1 (%) VarCNN 67.81 (-24.35) 67.30 (-24.61) ARES 72.84 (-18.38) 71.92 (-18.93) RF 66.07 (-25.02) 65.82 (-24.81) DF 67.92 (-21.93) 66.61 (-22.75) TikTok 66.03 (-23.29) 64.74 (-24.03) TF 67.75 (-19.19) 66.49 (-19.92) AWF 29.10 (-33.56) 29.07 (-32.88) H&W 79.96 (-06.49) 77.92 (-07.84) STC-WF 73.93 (-03.82) 72.50 (-04.83) CipherSight 92.99 (-02.42) 92.07 (-02.60) Table 2: Temporal-drift performance from us0304 to us0320 with a 16-day collection gap. Values in parentheses denote descriptive percentage-point differences from each model’s closed-world result on us0304. Singapore France Model Accuracy (%) Macro-F1 (%) Top-5 (%) Accuracy (%) Macro-F1 (%) Top-5 (%) VarCNN 67.53 (-24.98) 66.20 (-26.24) 83.07 (-14.30) 73.12 (-19.39) 70.70 (-21.74) 85.05 (-12.32) ARES 79.53 (-13.16) 78.15 (-14.48) 89.97 (-07.23) 81.08 (-11.61) 79.11 (-13.52) 90.58 (-06.62) RF 57.69 (-35.12) 57.44 (-35.19) 73.25 (-24.10) 82.07 (-10.74) 80.19 (-12.44) 91.03 (-06.32) DF 74.95 (-17.39) 72.83 (-19.26) 87.20 (-10.10) 76.48 (-15.86) 73.24 (-18.85) 86.41 (-10.89) TikTok 70.00 (-21.98) 67.80 (-23.94) 84.26 (-13.14) 74.51 (-17.47) 71.24 (-20.50) 85.08 (-12.32) TF 74.46 (-16.56) 72.44 (-18.43) – 75.67 (-15.35) 72.70 (-18.17) – AWF 32.09 (-36.74) 31.93 (-36.70) 54.92 (-30.77) 45.99 (-22.84) 45.18 (-23.45) 67.52 (-18.17) H&W 75.94 (-14.82) 72.87 (-17.35) 84.52 (-11.82) 67.56 (-23.20) 62.57 (-27.65) 76.56 (-19.78) STC-WF 60.60 (-31.65) 57.52 (-34.32) 75.12 (-22.60) 54.38 (-37.87) 49.74 (-42.09) 68.96 (-28.76) CipherSight 91.09 (-04.41) 89.99 (-05.21) 95.20 (-02.50) 90.59 (-04.91) 89.19 (-06.01) 95.08 (-02.63) Table 3: Geographic-drift performance from us0409 to Singapore and France. Values in parentheses denote descriptive percentage-point differences from each model’s us0409 source-domain holdout. CipherSight values are means over five seeds. Table 2 reports performance under a temporal distribution shift caused by a 16-day gap in data collection. H&W is the strongest baseline with 79.96% accuracy, while STC-WF reaches 73.93% and slightly outperforms ARES, the strongest Tor-oriented baseline. Although both HTTPS-specific methods trail the leading Tor-oriented baselines in the closed-world setting, their ranking reverses under temporal drift. This result indicates that in-distribution closed-world accuracy is not a reliable proxy for robustness to temporal drift. CipherSight achieves 92.99% accuracy and 92.07% macro-F1, outperforming H&W by 13.03 and 14.15 percentage points (p), respectively. Its descriptive differences from the separate closed-world result are only −2.42-2.42 and −2.60-2.60 p, smaller than those of all baselines. Its accuracy margin over H&W also widens from 8.96 p in the closed-world setting to 13.03 p under temporal drift. Although these differences are not paired source-to-target drops, the pattern is consistent with CipherSight learning representations that remain informative as website traffic evolves. Geographic Drift Table 3 evaluates geographic drift by training on us0409 and testing on the same-day sg0409 and fr0409 collections. This design reduces temporal confounding and helps isolates regional variation. Baseline rankings vary by region. ARES is the strongest baseline in Singapore with 79.53% accuracy, whereas RF leads in France with 82.07%. RF differs by 24.38 percentage points (p) between regions, showing that performance in one region does not reliably predict another. CipherSight achieves the highest accuracy, macro-F1, and top-5 accuracy in both regions, reaching 91.09% accuracy in Singapore and 90.59% in France. It outperforms the strongest regional baselines by 11.56 and 8.52 p, respectively. Its accuracy differs by only 0.50 p between regions, and it exhibits the smallest degradation under geographic drift across all reported metrics. These results suggest that CipherSight retains stable record- and flow-level evidence across regions. Ablation Study Test No-Priv (%) P1 Pre. (%) Full Priv. (%) us0320 92.69 / 91.73 92.99 / 92.07 92.99 / 92.07 us0409 90.46 / 89.02 90.93 / 89.53 91.11 / 89.73 fr0409 89.04 / 87.37 90.15 / 88.73 90.59 / 89.19 sg0409 89.73 / 88.42 90.59 / 89.46 91.09 / 89.99 Table 4: Ablation of privileged supervision under temporal and geographic drift. No-Priv. excludes P1 and P2, P1 Pre. applies P1 only during pretraining, and Full Priv. applies P1 in both stages and introduces P2 during LoRA fine-tuning. Cells report mean accuracy/macro-F1 (%) over five seeds. Table 4 reports the contribution of privileged supervision under distribution shifts. Under temporal drift, P1 Pre. improves accuracy over No-Priv. by 0.30 and 0.47 percentage points (p) on us0320 and us0409, while Full Priv. yields gains of 0.30 and 0.65 p. This result shows that P1 pretraining provides most of the temporal benefit. The contribution is more pronounced under geographic drift. P1 Pre. improves accuracy by 0.86 p on sg0409 and 1.11 p on fr0409, while Full Priv. increases these gains to 1.37 and 1.55 p, with corresponding macro-F1 gains of 1.58 and 1.83 p. Although modest in absolute terms, these gains are approximately 31% as large as CipherSight’s overall accuracy degradation under geographic drift, with a similar ratio of about 30% for macro-F1. Overall, the ablation demonstrates the effectiveness of training-time privileged supervision, particularly in improving geographic robustness. Open-World Evaluation Model AUROC (%) AUPR (%) OSCR (%) VarCNN 89.56 60.50 86.22 ARES 89.93 67.27 86.05 RF 93.04 72.03 87.61 DF 87.06 63.46 83.72 TikTok 86.21 61.49 82.77 TF 82.75 26.95 77.15 AWF 70.11 27.48 52.40 H&W 57.23 07.39 49.42 STC-WF 77.92 35.69 68.67 CipherSight 97.64 88.54 94.81 Table 5: Open-world evaluation results using maximum softmax probability (MSP) as the confidence score. As shown in Table 5, CipherSight consistently outperforms all baselines, achieving 97.64% AUROC, 88.54% AUPR, and 94.81% OSCR. Relative to RF, the strongest baseline, it yields absolute gains of 4.60, 16.51, and 7.20 percentage points, respectively. The substantial improvements in AUPR and OSCR demonstrate more effective rejection of unmonitored samples in the highly imbalanced open-world setting, without compromising monitored-site classification. Limitations More refined data preprocessing methods and sampling-window strategies may further improve model performance (20). CipherSight could be extended to a range of scenarios, including longer-term temporal drift, larger numbers of clients, and HTTP/3 traffic. In addition, privileged supervision requires SSL keys, limiting packet-only reuse, although inference remains ciphertext-only. Conclusion We presented CipherSight, a hierarchical Transformer for robust HTTPS WF under distribution shifts. CipherSight models ciphertext-observable TLS records within and across flows. It uses record-resource supervision and privileged semantic distillation during training while preserving ciphertext-only inference. Across closed-world, temporal, geographic, and open-world evaluations, CipherSight consistently outperformed the baselines. Ablations showed gains from privileged supervision, particularly under geographic drift. Together, these results highlight the value of stable TLS-record representations, explicit inter-flow modeling, and training-time resource semantics for robust encrypted-traffic learning under the evaluated conditions. References Bhat et al. (2019) S. Bhat, D. Lu, A. Kwon, and S. Devadas Var-CNN: a data-efficient website fingerprinting attack based on deep learning. Proceedings on Privacy Enhancing Technologies 2019 (4), p. 292–310. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting, Datasets and Experimental Setup. Chai et al. (2019) Z. Chai, A. Ghafari, and A. Houmansadr On the importance of Encrypted-SNI (ESNI) to censorship circumvention. In 9th USENIX Workshop on Free and Open Communications on the Internet, External Links: Link Cited by: Website Fingerprinting. Chen et al. (2026) Z. Chen, G. Cheng, D. Niu, Y. Zhao, Y. Zhou, and S. Jiang Ultimate encrypted traffic feature engineering: HTTPS encrypted traffic classification using restored application data unit length. IEEE Transactions on Dependable and Secure Computing 23 (1), p. 1290–1307. External Links: Document, Link Cited by: Introduction, Website Fingerprinting over HTTPS. Cheng et al. (2026) Y. Cheng, Y. Zhu, B. Li, X. Deng, Y. Cai, Y. Ren, and Q. Liu STAR: semantic-traffic alignment and retrieval for zero-shot HTTPS website fingerprinting. In Proceedings of the IEEE Conference on Computer Communications (INFOCOM), p. 1–10. External Links: Link Cited by: Website Fingerprinting over HTTPS. Cheng et al. (2025) Y. Cheng, Y. Zhu, B. Li, P. Sun, Y. Ding, X. Deng, and Q. Liu HOLMES & WATSON: a robust and lightweight HTTPS website fingerprinting through HTTP version parallelism. In Proceedings of the ACM Web Conference 2025, p. 1078–1092. Cited by: Introduction, Website Fingerprinting over HTTPS, Datasets and Experimental Setup. Cherubin et al. (2022) G. Cherubin, R. Jansen, and C. Troncoso Online website fingerprinting: evaluating website fingerprinting attacks on Tor in the real world. In Proceedings of the 31st USENIX Security Symposium, p. 753–770. External Links: Link Cited by: Website Fingerprinting under Distribution Shifts. Deng et al. (2023) X. Deng, Q. Yin, Z. Liu, X. Zhao, Q. Li, M. Xu, K. Xu, and J. Wu Robust multi-tab website fingerprinting attacks in the wild. In Proceedings of the 2023 IEEE Symposium on Security and Privacy (SP), p. 1005–1022. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting under Distribution Shifts, Datasets and Experimental Setup. Dingledine et al. (2004) R. Dingledine, N. Mathewson, and P. Syverson Tor: the Second-Generation onion router. In 13th USENIX Security Symposium (USENIX Security 04), San Diego, CA. External Links: Link Cited by: Website Fingerprinting. Eddy (2022) W. Eddy Transmission control protocol (TCP). RFC Technical Report 9293, RFC Editor. External Links: Document, Link Cited by: TLS-Record Representation and Tokenization.. Fan et al. (2026) C. Fan, W. Wang, W. Huang, Z. Ding, J. Shi, L. Cui, Z. Hao, and X. Yun ResAware: cross-environment website fingerprinting via resource-privileged distillation. External Links: 2606.17462, Document, Link Cited by: Website Fingerprinting over HTTPS, Record-resource Supervision (P1).. Hoang et al. (2021) N. P. Hoang, A. A. Niaki, P. Gill, and M. Polychronakis Domain name encryption is not enough: privacy leakage via IP-based website fingerprinting. Proceedings on Privacy Enhancing Technologies 2021 (4), p. 420–440. External Links: Document, Link Cited by: Website Fingerprinting. Hoffman and McManus (2018) P. Hoffman and P. McManus DNS queries over HTTPS (DoH). RFC Technical Report 8484, RFC Editor. Cited by: Introduction. Hu et al. (2016) Z. Hu, L. Zhu, J. Heidemann, A. Mankin, D. Wessels, and P. Hoffman Specification for DNS over transport layer security (TLS). RFC Technical Report 7858, RFC Editor. Cited by: Introduction. Huitema et al. (2022) C. Huitema, S. Dickinson, and A. Mankin DNS over dedicated QUIC connections. RFC Technical Report 9250, RFC Editor. Cited by: Introduction. Jansen et al. (2024) R. Jansen, R. Wails, and A. Johnson Repositioning real-world website fingerprinting on Tor. In Proceedings of the 23rd Workshop on Privacy in the Electronic Society, p. 124–140. External Links: Document, Link Cited by: Introduction, Website Fingerprinting under Distribution Shifts. Juarez et al. (2014) M. Juarez, S. Afroz, G. Acar, C. Diaz, and R. Greenstadt A critical evaluation of website fingerprinting attacks. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, p. 263–274. External Links: Document, Link Cited by: Website Fingerprinting, Website Fingerprinting under Distribution Shifts. Koh et al. (2021) P. W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsubramani, W. Hu, M. Yasunaga, R. L. Phillips, I. Gao, T. Lee, E. David, I. Stavness, W. Guo, B. Earnshaw, I. Haque, S. M. Beery, J. Leskovec, A. Kundaje, E. Pierson, S. Levine, C. Finn, and P. Liang WILDS: a benchmark of in-the-wild distribution shifts. In Proceedings of the 38th International Conference on Machine Learning, p. 5637–5664. External Links: Link Cited by: Introduction, Website Fingerprinting under Distribution Shifts. Li et al. (2018) S. Li, H. Guo, and N. Hopper Measuring information leakage in website fingerprinting attacks and defenses. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, p. 1977–1992. External Links: Document, Link Cited by: Introduction. Ma et al. (2024) X. Ma, J. Qu, M. Shi, B. An, J. Li, X. Luo, J. Zhang, Z. Li, and X. Guan Website fingerprinting on encrypted proxies: a flow-context-aware approach and countermeasures. IEEE/ACM Transactions on Networking 32 (3), p. 1904–1919. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting over HTTPS, Datasets and Experimental Setup. Peng et al. (2025) W. Peng, L. Cui, W. Cai, W. Wang, X. Cui, Z. Hao, and X. Yun Bottom aggregating, top separating: an aggregator and separator network for encrypted traffic understanding. IEEE Trans. Inf. Forensics Secur. 20, p. 1794–1806. External Links: Link, Document Cited by: Limitations. Rahman et al. (2020) M. S. Rahman, P. Sirinam, N. Mathews, K. G. Gangadhara, and M. Wright Tik-Tok: The Utility of Packet Timing in Website Fingerprinting Attacks. Proceedings on Privacy Enhancing Technologies 2020 (3), p. 5–24. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting, Datasets and Experimental Setup. Rescorla et al. (2026) E. Rescorla, K. Oku, N. Sullivan, and C. A. Wood TLS encrypted client hello. RFC Technical Report 9849, RFC Editor. Cited by: Introduction. Rescorla (2018) E. Rescorla The transport layer security (TLS) protocol version 1.3. RFC Technical Report 8446, RFC Editor. External Links: Document Cited by: TLS-Record Representation and Tokenization.. Rimmer et al. (2018) V. Rimmer, D. Preuveneers, M. Juarez, T. Van Goethem, and W. Joosen Automated website fingerprinting through deep learning. In Proceedings of the 2018 Network and Distributed System Security Symposium (NDSS), External Links: Document, Link Cited by: Website Fingerprinting, Datasets and Experimental Setup. Shen et al. (2023) M. Shen, K. Ji, Z. Gao, Q. Li, L. Zhu, and K. Xu Subverting website fingerprinting defenses with robust traffic representation. In 32nd USENIX Security Symposium (USENIX Security 23), p. 607–624. External Links: Link Cited by: Datasets and Experimental Setup. Sirinam et al. (2018) P. Sirinam, M. Imani, M. Juarez, and M. Wright Deep fingerprinting: undermining website fingerprinting defenses with deep learning. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, p. 1928–1943. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting, Datasets and Experimental Setup. Sirinam et al. (2019) P. Sirinam, N. Mathews, M. S. Rahman, and M. Wright Triplet fingerprinting: more practical and portable website fingerprinting with N-shot learning. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, p. 1131–1148. External Links: Document, Link Cited by: Datasets and Experimental Setup. Tan et al. (2024) X. Tan, C. Peng, P. Xie, H. Wang, M. Li, S. Chen, and C. Zou Inter-flow spatio-temporal correlation analysis based website fingerprinting using graph neural network. IEEE Transactions on Information Forensics and Security 19, p. 7619–7632. External Links: Document, Link Cited by: Challenges, Introduction, Website Fingerprinting over HTTPS, Datasets and Experimental Setup. Vapnik and Vashist (2009) V. Vapnik and A. Vashist A new learning paradigm: learning using privileged information. Neural Networks 22 (5–6), p. 544–557. External Links: Document, Link Cited by: Website Fingerprinting over HTTPS. Wang et al. (2023) J. Wang, C. Lan, C. Liu, Y. Ouyang, T. Qin, W. Lu, Y. Chen, W. Zeng, and P. S. Yu Generalizing to unseen domains: a survey on domain generalization. IEEE Transactions on Knowledge and Data Engineering 35 (8), p. 8052–8072. External Links: Document Cited by: Website Fingerprinting under Distribution Shifts. Xian et al. (2026) Y. Xian, X. Zeng, L. Meng, L. Cui, R. Song, W. Wang, Z. Ding, P. Liu, and Z. Hao More than meets the eye: a semantics-aware traffic augmentation framework for generalizable website fingerprinting. External Links: 2605.11402, Document, Link Cited by: Website Fingerprinting over HTTPS. Yuan et al. (2024) X. Yuan, T. Li, L. Li, R. Li, Z. Wang, and X. Luo HSWF: enhancing website fingerprinting attacks on Tor to address real world distribution mismatch. Computer Networks 241, p. 110217. External Links: Document, Link Cited by: Introduction, Website Fingerprinting under Distribution Shifts.