Paper deep dive
Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation
Shuhuan Chen, Xiangyu Zhu, Weisong Zhao, Siran Peng, Tianshuo Zhang, Haoyuan Zhang, Haichao Shi, Xiao-Yu Zhang, Zhen Lei
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/1/2026, 2:09:31 AM
Summary
The paper introduces Private Face Distillation (PFD), a framework for publishing private face recognition (FR) training datasets while preserving privacy. It addresses the 'identity paradox' where identity cues useful for recognition also enable re-identification. PFD uses Orthogonal Geometry Preservation (OGP) to decouple source identity from proxy geometry and Relational Topology Alignment (RTA) to distill proxy images that maintain recognition utility. Experiments show PFD improves downstream FR utility (e.g., +3.94% TAR@FAR=1e-3 on IJB-C) while significantly reducing source-identity linkability compared to baselines like Face Anonymization and Differential Privacy.
Entities (10)
Relation Signals (9)
Private Face Distillation → usescomponent → Orthogonal Geometry Preservation
confidence 95% · It uses Orthogonal Geometry Preservation to construct decoupled proxy identities
Private Face Distillation → usescomponent → Relational Topology Alignment
confidence 95% · and Relational Topology Alignment to preserve identity relations for recognition learning.
Private Face Distillation → evaluatedon → IJB-C
confidence 92% · On IJB-C surveillance, it improves TAR@FAR=1e-3 by 3.94% over the baseline
Relational Topology Alignment → achieves → geometry_preservation
confidence 90% · RTA distills proxy images from random noise by aligning the intra-class and inter-class topology of transformed embeddings.
Orthogonal Geometry Preservation → achieves → identity_decoupling
confidence 90% · OGP resolves this by mapping private identity representations into a proxy space that preserves recognition-useful hyperspherical geometry while reducing direct correspondence to source identity coordinates.
Private Face Distillation → evaluatedon → LFW
confidence 90% · We evaluate PFD on LFW... and our proposed Multi-Scenario Benchmark.
Private Face Distillation → outperforms → Face Anonymization
confidence 90% · PFD achieves stronger utility than the evaluated publication baselines... improving TAR@FAR=1e-3 by 3.94% over the baseline
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: \emph{the identity cues that make released faces useful for recognition supervision are also the cues that make them linkable to real individuals.} A protected face should be decoupled from the original identity, yet still behave as a reliable identity sample for training. Removing these cues too aggressively may destroy the class structure needed for recognition learning, whereas preserving them too faithfully may increase source-identity linkability. We argue that this paradox stems from conflating source-aligned identity semantics with recognition-useful proxy identity geometry. The former should be suppressed to reduce linkage to private individuals, while the latter should be preserved for FR learning. Based on this insight, we propose \textbf{Private Face Distillation}, an identity-decoupling and geometry-preserving framework. It uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning. Experiments across multiple domain-shifted FR scenarios show that Private Face Distillation achieves stronger utility than the evaluated publication baselines. On IJB-C surveillance, it improves $\mathrm{TAR}@\mathrm{FAR}{=}1\text{e-}{3}$ by 3.94\% over the baseline while reducing source-identity linkability. These results suggest that private FR training dataset publication should decouple source-identity correspondence while preserving proxy identity geometry.
Tags
Links
- Source: https://arxiv.org/abs/2607.27764v1
- Canonical: https://arxiv.org/abs/2607.27764v1
Trouble viewing inline? Open PDF directly →
Full Text
47,416 characters extracted from source content.
Expand or collapse full text
Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation Shuhuan Chen1,3, Xiangyu Zhu 2,4, Weisong Zhao5, Siran Peng2,4, Tianshuo Zhang2,4, Haoyuan Zhang2,4, Haichao Shi1, Xiao-Yu Zhang 1, Zhen Lei2,4,6,7 Abstract Publishing private face recognition (FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: the identity cues that make released faces useful for recognition supervision are also the cues that make them linkable to real individuals. A protected face should be decoupled from the original identity, yet still behave as a reliable identity sample for training. Removing these cues too aggressively may destroy the class structure needed for recognition learning, whereas preserving them too faithfully may increase source-identity linkability. We argue that this paradox stems from conflating source-aligned identity semantics with recognition-useful proxy identity geometry. The former should be suppressed to reduce linkage to private individuals, while the latter should be preserved for FR learning. Based on this insight, we propose Private Face Distillation, an identity-decoupling and geometry-preserving framework. It uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning. Experiments across multiple domain-shifted FR scenarios show that Private Face Distillation achieves stronger utility than the evaluated publication baselines. On IJB-C surveillance, it improves TAR@FAR=1e-3TAR@FAR=1e-3 by 3.94% over the baseline while reducing source-identity linkability. These results suggest that private FR training dataset publication should decouple source-identity correspondence while preserving proxy identity geometry. Codes. Introduction High-quality face recognition training datasets are crucial for improving model performance in domain-specific scenarios, yet their publication is increasingly restricted by identity privacy concerns (Zhao et al. 2025a). Unlike generic visual data, facial images directly encode biometric identity, making released datasets vulnerable to unauthorized re-identification, identity tracing, and even deepfake misuse (Meden et al. 2021; Zhang et al. 2025; Peng et al. 2026). Once publicly distributed, private faces may be persistently linked across datasets, applications, or surveillance scenarios, creating risks that are difficult to revoke. To address this issue, Face Dataset Publication (Zhang et al. 2026b) has recently been proposed as a privacy-aware data release paradigm, aiming to transform a private face dataset into a protected counterpart for downstream model training. This work studies its FR-specific setting, where identity cues are required for recognition supervision yet also enable linkage between released faces and source individuals. Figure 1: Performance-linkability trade-off on IJB-C surveillance under private FR training dataset publication. PFD improves downstream utility while reducing source-identity linkability under the evaluated attack protocol. Existing methods typically weaken the association between released faces and private individuals through face anonymization or differential privacy. Face anonymization (FA) modifies or synthesizes facial appearances to reduce their visual correspondence to the original individuals (Kuang et al. 2021; Yuan et al. 2022; Li et al. 2021). While this can reduce apparent identity similarity, it mainly changes who the face looks like, without specifying how protected samples should behave as identity classes in the recognition space. Differential privacy (DP) provides theoretically grounded protection by limiting the influence of individual samples or identities through perturbation or noise injection (Zhu et al. 2020; Meng et al. 2022). However, perturbations that support privacy guarantees also modify the training signals in released data, potentially obscuring subtle identity differences needed for recognition learning. Moreover, Face Dataset Publication often relies on limited private data, making domain-specific generative proxy synthesis difficult. These limitations reveal a fundamental identity paradox when published data are used for FR training: the same identity cues that make protected faces useful for recognition supervision also make them linkable to real individuals. A protected face should reduce correspondence to its source identity while remaining a reliable identity sample for recognition training. Suppressing identity cues too strongly may weaken class consistency and separability, while preserving them too faithfully may increase source-identity linkability. We argue that this paradox reflects not an inherent incompatibility between privacy protection and recognition learning, but a conflation of two forms of identity information. Source-aligned appearance and embedding cues link released faces to real individuals and increase privacy risk. Recognition learning, by contrast, depends on class geometry, including intra-class consistency, inter-class separability, and identity relations, without preserving the original identity correspondence. Private FR training dataset publication should therefore weaken source-identity linkability while preserving proxy identity geometry. Released faces need not reproduce the original individuals, but should retain how proxy classes are organized for recognition learning. Motivated by this objective, we transform private identity representations into protected proxy identities while preserving recognition-useful geometry. We instantiate Private Face Distillation (PFD), an identity-decoupling and geometry-preserving framework for private FR training dataset publication. PFD separates source-aligned identity correspondence from recognition-useful proxy geometry through two designs. Orthogonal Geometry Preservation constructs a transformed space with distinct proxy class targets while preserving hyperspherical identity relations. Relational Topology Alignment transfers this geometry to releasable images by aligning intra-class and inter-class relations, enabling stable, discriminative proxy classes for training. Only proxy images and arbitrary class labels are released, while source data, the orthogonal transformation, and adapted model remain local to the publisher. As shown in Fig. 1, PFD achieves a favorable performance-linkability trade-off, improving downstream FR utility while reducing source-identity linkability under the evaluated attacks. Our contributions are threefold: • We identify the identity paradox in private FR training dataset publication, where source-aligned correspondence increases linkability while proxy identity geometry supports recognition learning. • We introduce Private Face Distillation, which combines OGP and RTA to construct identity-decoupled RGB proxies while preserving recognition-useful geometry. • Across four domain shifts, PFD achieves the highest average utility among the evaluated publication methods while maintaining low source-identity linkability under the tested attacks. Related Work Face Dataset Publication Face Dataset Publication aims to transform a private face dataset into a protected counterpart that can be released for downstream training while reducing identity privacy risks (Zhang et al. 2026b). Existing methods mainly rely on differential privacy or face anonymization. Differential-privacy-based approaches (Chen et al. 2022; Tsai et al. 2025) provide theoretically grounded protection through perturbation or privacy-constrained optimization, but do not explicitly specify what identity supervision should remain for FR adaptation. Face anonymization methods (Barattin et al. 2023; Kuang et al. 2024; Yuan et al. 2024) reduce apparent identity similarity by synthesizing facial appearances, but visual identity replacement does not directly define how protected samples should be organized in the recognition space. In contrast, our work focuses on its FR training-data setting, aiming to reduce source-identity correspondence while preserving proxy identity geometry. Dataset Distillation Dataset distillation (D) aims to synthesize a compact proxy set that can substitute for a large real dataset in downstream training (Wang et al. 2020). Existing methods typically construct such proxies through optimization-based or generation-based routes (Cazenavette et al. 2022; Chen et al. 2025; Li et al. 2025; Chan Santiago et al. 2025; Zhang et al. 2026a). This formulation is naturally relevant to Face Dataset Publication, as both settings rely on proxy data to replace real training data. However, conventional D focuses on training utility and does not require the proxies to protect private identities. Prior work shows that distilled proxies may provide implicit privacy effects (Dong et al. 2022), suggesting their potential for privacy-aware data release. Yet such privacy is only a by-product and does not address the identity paradox, where identity information is both the supervision signal and the privacy risk. Our work adapts D’s proxy-data advantage to private FR training dataset publication by reducing source-identity correspondence while preserving recognition-useful proxy identity geometry. Distinction from PPFR Privacy-preserving face recognition (PPFR) generally protects facial inputs or representations used during recognition (Wang et al. 2023; Mi et al. 2023, 2024; Dai et al. 2025). Private FR training dataset publication instead focuses on releasing protected data for reuse in downstream training. Such data should remain usable without access to the publisher’s private components. PFD only releases labeled RGB proxy images, while the source data, learned transformation, and adapted publisher model remain private. A third party can train an FR model on these proxies using a standard pipeline, and the resulting model directly accepts ordinary RGB test images without protected-domain conversion or architectural modification. We therefore evaluate PFD as a published training resource rather than as a protected recognition system, considering both downstream recognition utility and source-identity linkability in the released proxy set. Figure 2: Overview of PFD. OGP learns an internal transformed identity space by orthogonally mapping raw embeddings, separating original and transformed identities, and preserving hyperspherical geometry. RTA distills proxy images from random noise by aligning the intra-class and inter-class topology of transformed embeddings. Only the distilled proxy images and arbitrary labels are published as a reusable FR training set for joint training with public data. Method Data Release Setting In the private FR training dataset publication setting (Zhang et al. 2026b), we release only protected RGB proxies and arbitrary class labels for downstream FR training. Private source images and internal states, including the source-to-proxy correspondence, adapted model, orthogonal matrix, and paired original and transformed embeddings, are not released. Once distillation and publisher-side evaluation are complete, these intermediate states are no longer needed and can be discarded without affecting the proxy set. The released proxy set requires no publisher-side components. The method and training protocol may be known, and the proxies may be analyzed using auxiliary identity images, public FR models, or off-the-shelf reconstruction models. We focus on whether released images can be linked to source identities while supporting FR training. Disclosure of retained publisher-side information is outside this setting. Private Face Distillation As illustrated in Fig. 2, PFD distills a protected RGB proxy training set by transferring recognition-useful identity geometry from private target-domain data into releasable proxy images. It first learns an internal transformed proxy space through Orthogonal Geometry Preservation (OGP), which retains hyperspherical identity relations while reducing direct correspondence to the original source coordinates. Based on this space, Relational Topology Alignment (RTA) distills proxy images from random noise by aligning their representations with the intra-class and inter-class topology of transformed identity prototypes. Orthogonal Geometry Preservation Directly distilling private embeddings may transfer source-aligned identity semantics into published proxies, while removing identity structure would destroy the angular geometry needed for FR adaptation. OGP resolves this by mapping private identity representations into a proxy space that preserves recognition-useful hyperspherical geometry while reducing direct correspondence to source identity coordinates. Geometry-Preserving Identity Decoupling. Margin-based FR methods (Sun et al. 2020; Boutros et al. 2022; Zhao et al. 2023, 2025b) learn ℓ2 _2-normalized embeddings on a hypersphere ℳ⊂ℝDM ^D, where identity discrimination mainly depends on angular relations. Given a private target-domain image x x, we compute its normalized embedding as e^=normℓ2(ϕθ(x^))∈ℳ e=norm_ _2( _θ( x)) , where ϕθ _θ is a trainable FR backbone initialized from a public pre-trained model and jointly optimized with OGP. To preserve this recognition-useful geometry without exposing the original identity coordinates, OGP maps each target-domain embedding e e into a proxy space through an orthogonal transformation: e^ort=Morte^,Mort=exp(Ω),Ω=A−A⊤, e_ort=M_ort e, M_ort= ( ), =A-A , (1) where A∈ℝD×DA ^D× D is learnable and Ω∈(D) ∈ so(D) is skew-symmetric. The matrix exponential gives Mort∈SO(D)M_ort∈ SO(D), allowing OGP to learn a strictly orthogonal transformation without heuristic orthogonality penalties or post-projection (Lezcano-Casado and Martínez-Rubio 2019). For a column-wise normalized prototype matrix P∈ℝD×NP ^D× N and Port=MortP_ort=M_ortP, orthogonality preserves the within-space topology, Port⊤Port=P⊤P_ort P_ort=P P, whereas source-to-proxy correspondence is measured by P⊤Port=P⊤MortP P_ort=P M_ortP. OGP therefore retains the hyperspherical geometry required for FR while defining transformed coordinate targets for proxy identities. Further analysis of this distinction and its privacy implications is provided in the Supplementary Materials. Discriminative Proxy Identity Learning. Although the orthogonal transformation preserves hyperspherical geometry, it should not collapse to the identity mapping or remain too close to the original identity coordinates. To construct non-trivial proxy identities, we train original and transformed identities jointly in the same FR embedding space. For N private identities, we build a classifier head hθh_θ with 2N2N prototypes, where the first N classes correspond to original identities and the remaining N classes correspond to transformed proxy identities. Given an embedding e e with label y y, its transformed counterpart e^ort e_ort is assigned label y=y^+Ny= y+N. The classification loss is written as ℒclsOGP=ℒCE(hθ([e^,e^ort]),[y^,y]),L_cls^OGP=L_CE (h_θ ([ e, e_ort] ),[ y,y] ), (2) where [⋅,⋅][·,·] denotes batch-wise concatenation and ℒCEL_CE is a margin-based cross-entropy loss. This objective treats original and transformed identities as distinct classes while keeping them in a shared hyperspherical FR space. To further reduce direct correspondence to the original identity coordinates, we impose a repulsion constraint: ℒrep=e^∼ℬ[max(0,|cos(e^ort,e^)|−m)],L_rep=E_ e [ (0,| ( e_ort, e)|-m ) ], (3) where ℬB is the current mini-batch and m is a cosine margin. Meanwhile, to keep the proxy classifier consistent with the orthogonal geometry, we enforce prototype isomorphism: ℒiso=1N∑i=1N[1−cos(Port(i),MortPreal(i))],L_iso= 1N _i=1^N [1- (P_ort^(i),M_ortP_real^(i) ) ], (4) where Preal,Port∈ℝD×NP_real,P_ort ^D× N collect the normalized original and transformed class prototypes of hθh_θ column-wise. The overall OGP objective is ℒOGP=ℒclsOGP+λrepℒrep+λisoℒiso.L_OGP=L_cls^OGP+ _repL_rep+ _isoL_iso. (5) Through this objective, OGP separates source and proxy identities at the class level while preserving their within-space relational geometry. The complete stage therefore defines discriminative proxy coordinate targets rather than treating an orthogonal rotation alone as a privacy mechanism. The resulting backbone ϕθ _θ and classifier head hθh_θ provide the internal proxy embedding space for subsequent distillation, but are not released with the published dataset. Distillation via Relational Topology Alignment RTA converts the proxy identity geometry learned by OGP into releasable proxy images. Following the D paradigm (Yin et al. 2023), each proxy image s is initialized from Gaussian noise and optimized with the adapted backbone ϕθ _θ and classifier head hθh_θ. Rather than directly matching private embeddings or facial appearances, RTA distills relational topology to preserve intra-identity consistency and inter-identity relations. We further use Batch Normalization regularization to align proxy images with the feature statistics of ϕθ _θ: ℒbn=∑l(‖μl(s)−lRM‖2+‖σl2(s)−lRV‖2),L_bn= _l ( \| _l(s)-BN_l^RM \|_2+ \| _l^2(s)-BN_l^RV \|_2 ), (6) where μl(⋅) _l(·) and σl2(⋅) _l^2(·) are the batch mean and variance at layer l, and lRMBN_l^RM and lRVBN_l^RV are the running statistics of ϕθ _θ. We also impose a classification loss: ℒclsRTA=ℒCE(hθ(ϕθ(s)),y),L_cls^RTA=L_CE (h_θ ( _θ(s) ),y ), (7) where y is the internal transformed-class label, consistently remapped to an arbitrary proxy label before release. The loss encourages discriminative proxy classes in the FR space. Relational Topology Alignment. Although ℒbnL_bn and ℒclsRTAL_cls^RTA make proxy images compatible with the adapted FR model ϕθ _θ, they do not preserve the identity geometry learned by OGP. RTA complements these losses by aligning distilled proxies with the relational topology of the transformed proxy identity space. Specifically, stacking private embeddings row-wise as E, we compute ^ort=^Mort⊤ E_ort= EM_ort and apply K-Means per identity to extract K prototypes, where K matches the publication budget. This forms a normalized prototype matrix ∈ℝNK×DP ^NK× D as anchors of proxy identity geometry. Let ∈ℝNK×DE ^NK× D denote the normalized proxy embeddings in prototype order. We define the prototype topology and proxy-to-prototype affinity as ∗=⊤,syn=⊤.A^*=PP , ^syn=EP . (8) Since all embeddings are normalized, each entry is a cosine similarity. We align the target and proxy relation profiles by ℒRTA _RTA =1NK‖(syn−∗)⊙‖1 = 1NK\|(A^syn-A^*) \|_1 (9) +1NK(NK−1)‖(syn−∗)⊙(−)‖1, + 1NK(NK-1)\|(A^syn-A^*) (1-I)\|_1, where I and 1 are the NK×NKNK× NK identity and all-ones matrices, ⊙ is the Hadamard product, and ‖1=∑i,j|Xij|\|X\|_1= _i,j|X_ij| is the entrywise ℓ1 _1 norm. Separate normalization balances diagonal proxy-to-prototype alignment and off-diagonal intra- and inter-identity relations. Although ∗A^* is orthogonally invariant, synA^syn uses the transformed prototypes P, while the adapted backbone and proxy classifier anchor synthesis in the proxy space. The final distillation objective is ℒdistill=ℒclsRTA+λRTAℒRTA+λbnℒbn.L_distill=L_cls^RTA+ _RTAL_RTA+ _bnL_bn. (10) Through this objective, RTA transfers the proxy identity topology required for recognition learning to releasable images. This is a one-time publisher-side optimization, and after distillation, the released images require no private transformation, adapted model, or special inference pipeline. Method Type Venue Cross-Age Surveillance Large-Pose VIS-NIR Avg. AdaFace (Kim et al. 2022) Base CVPR’22 85.02 89.75 88.97 95.00 89.69 Private Data Joint Training Ref. – 92.13 93.81 89.72 97.48 93.29 G2Face (Yang et al. 2024) FA TIFS’24 88.15 92.05 88.30 95.30 90.95 NullFace (Kung et al. 2026) FG’26 88.02 92.36 87.32 95.35 90.76 DP-LoRA (Tsai et al. 2025) DP ICCV’25 85.62 91.41 83.98 92.35 88.34 Metric-DP (Zhang et al. 2026b) TPAMI’26 86.22 89.72 88.17 94.88 89.75 SRe2L (Yin et al. 2023) D NeurIPS’23 85.02 91.78 88.12 94.48 89.85 NCFM (Wang et al. 2025) CVPR’25 84.25 91.64 88.20 95.41 89.88 Ours – 88.42 93.69 89.09 95.43 91.66 Table 1: Comparison of downstream FR adaptation performance (%). “Base” denotes the original pre-trained FR model, while “Ref.” denotes direct adaptation using private images. Best and second-best results among publication methods and the baseline. Figure 3: Source-identity linkability (%). Top: direct attacks on proxy datasets. Bottom: attacks on reconstructed images. Bars show Visual Matching (left) and Rank-1 Re-ID (right) success rates across target-domain scenarios. PFD averages 0.70%, close to the 0.67% random-guessing rate and the 0.96% point estimate of the evaluated DP baseline. Method LFW CFP-FP CPLFW AgeDB CALFW G2Face 99.13 96.43 88.87 93.42 92.80 NullFace 99.25 96.87 89.53 93.95 93.06 DP-LoRA 99.08 95.71 88.32 92.58 92.42 Metric-DP 99.35 96.86 89.60 93.95 93.05 SRe2L 99.33 96.76 89.67 93.60 92.82 NCFM 99.32 96.50 89.27 93.42 92.98 Ours 99.36 96.93 90.07 94.33 93.12 Table 2: Comparison on standard FR benchmarks (%). Experiments Datasets CASIA-WebFace (Yi et al. 2014) is used as the public FR dataset for both pretraining and joint training with the published proxy dataset. We evaluate PFD on LFW (Huang et al. 2008), CFP-FP (Sengupta et al. 2016), CPLFW (Zheng and Deng 2018), AgeDB (Moschoglou et al. 2017), CALFW (Zheng et al. 2017) and our proposed Multi-Scenario Benchmark. Multi-Scenario Benchmark We construct a multi-scenario benchmark to evaluate downstream FR utility and source-identity linkability under private FR training dataset publication. Each scenario includes a private domain set for proxy generation and a disjoint test set for evaluation. For open-set evaluation, private identities are screened against CASIA-WebFace using feature-space similarity to reduce overlap. Each private set contains 150 identities with 50 images per identity, while each published set contains 150 arbitrary proxy classes with 10 images per class under a fixed budget. The source-to-proxy association is retained only for attack evaluation and is not released. Age Shift (Cross-Age). Derived from B3FD (Bešenić et al. 2023), this scenario evaluates cross-age FR adaptation under substantial age-related appearance variations. Evaluation is conducted on a disjoint test set following the standard LFW 1:1 verification protocol, using oldest-image anchors, genuine pairs from the same identity, and impostor pairs from different identities. Resolution Shift (Surveillance). Using the IJB-C dataset (Maze et al. 2018), this scenario focuses on surveillance-domain FR adaptation with low-resolution and noisy imagery. A private surveillance subset is used for proxy distillation, while performance is evaluated on the remaining disjoint split under the official IJB-C protocol. Geometric Shift (Large-Pose). Built upon Multi-PIE (Gross et al. 2008), this scenario emphasizes large-pose FR adaptation under extreme variations. The private split is biased toward challenging non-frontal views, while evaluation compares extreme-pose probes with identity-level cross-pose gallery templates. Modality Shift (VIS-NIR). Using LAMP-HQ (Yu et al. 2021), this scenario studies cross-spectral FR adaptation between VIS and NIR domains. The private split maintains balanced VIS and NIR samples per identity, while evaluation is conducted on a disjoint VIS-NIR verification set following the same 1:1 verification protocol as LFW. Figure 4: Visual comparison across four scenarios. Relative to the displayed D-based (Wang et al. 2025), FA-based (Kung et al. 2026), and DP-based (Zhang et al. 2026b) methods, PFD produces more visually abstract proxies while retaining useful proxy classes for downstream FR adaptation. Figure 5: Representative post-release reconstruction stress test using off-the-shelf Arc2Face (Papantoniou et al. 2024). PFD reconstructions show weaker apparent source correspondence in these examples. Experiment Setting We use IR-50 (Deng et al. 2019) as the backbone and AdaFace (Kim et al. 2022) as the FR classification loss. The public FR model is pre-trained on CASIA-WebFace. All face images are detected, aligned, and cropped to 112×112 following (Kim et al. 2022). FR model optimization is performed using SGD with cosine learning-rate decay and linear warmup, starting from 0.1. The proxy images are optimized from random noise using Adam with a learning rate of 0.25 (Yin et al. 2023). Training details are included in the Supplementary Materials. Evaluation Metrics Downstream FR Utility. We evaluate downstream utility using scenario-specific FR protocols. For the Cross-Age and VIS-NIR scenarios, we report standard 1:1 verification accuracy. For the Surveillance scenario, we follow the official IJB-C protocol and report TAR@FAR=1e-3TAR@FAR=1e-3. For the Large-Pose scenario, we form genuine and impostor scores between extreme-pose probes and identity-level cross-pose gallery templates and report TAR@FAR=1e-3TAR@FAR=1e-3. For LFW-style datasets, we follow their official evaluation protocols. Source-Identity Linkability. To evaluate source-identity linkability under direct matching and reconstruction attacks, we follow prior privacy evaluation protocols (Cherepanova et al. 2021) and report Visual Matching (VisMatch) and Rank-1 Re-Identification (Re-ID). Visual Matching associates each proxy or reconstruction with its nearest private source image, while Rank-1 Re-ID evaluates Top-1 retrieval of its associated source identity from the private gallery. Both report attack success when all 150 balanced source identities are provided as candidates, yielding a random-assignment rate of 1/150=0.67%1/150=0.67\%. The source-to-proxy association is used only for scoring and remains hidden from the attacker. Utility and Source Linkability Recognition Performance. Tab. 1 compares PFD with representative FA-, DP-, and D-based methods, as well as direct private-data joint training, under the same protocol. Existing methods weaken identity cues or alter facial appearance but lack an explicit mechanism for preserving transferable proxy identity geometry under domain shifts. PFD achieves the best performance across all scenarios, improving the average from 89.69% to 91.66%. It is also the only method that consistently surpasses the baseline, indicating stronger preservation of transferable identity structure. The advantage remains in the challenging Large-Pose setting, where PFD improves performance to 89.09% despite severe pose variation. Results show that PFD preserves recognition-useful proxy geometry across domain shifts. Method Components Cross-Age Surveillance Large-Pose VIS-NIR OGP RTA Perf. ↑ VisMatch ↓ Re-ID ↓ Perf. ↑ VisMatch ↓ Re-ID ↓ Perf. ↑ VisMatch ↓ Re-ID ↓ Perf. ↑ VisMatch ↓ Re-ID ↓ Baseline × × 85.02 24.13 28.61 91.78 91.33 93.80 88.12 89.47 93.68 94.48 96.07 96.80 +RTA × ✓ 86.20 46.00 64.85 92.20 96.47 97.63 89.18 96.47 96.52 95.63 99.53 99.39 Ours ✓ ✓ 88.42 0.27 0.41 93.69 0.47 0.72 89.09 0.80 1.03 95.43 0.73 0.76 Table 3: Ablation study of utility and source-identity linkability (%) across four target-domain adaptation scenarios. Performance is measured by Acc/TAR, while linkability is evaluated by VisMatch and Rank-1 Re-ID attack success rates. To evaluate transfer beyond the scenario-specific test sets, Tab. 2 reports results on standard FR benchmarks. We use the Large-Pose-adapted model for LFW, CFP-FP, and CPLFW, and the Cross-Age-adapted model for AgeDB and CALFW, following each benchmark’s dominant variation. PFD obtains the highest reported score on all benchmarks and shows favorable transfer beyond the scenario-specific test sets. Source-Identity Linkability. Fig. 3 reports direct matching and Arc2Face reconstruction attacks using VisMatch and Rank-1 Re-ID. D- and FA-based methods show substantially higher source linkage across scenarios. PFD achieves a mean attack success rate of 0.70%, close to the 0.67% chance rate and the 0.96% DP baseline, while preserving stronger adaptation utility. Fig. 4 and Fig. 5 visualize published proxies and their Arc2Face reconstructions. D- and FA-based outputs retain more source-aligned appearance, whereas PFD produces abstract proxies with weaker apparent source correspondence after reconstruction. These observations agree with the quantitative linkage results in Fig. 3, evaluated under the same matching protocols. Source-Attribute Recoverability. To assess demographic leakage, we use a fixed off-the-shelf attribute predictor (Karkkainen and Joo 2021) to compare race, gender, and age predictions between proxy images and their mapped sources. We report Macro-F1 and balanced accuracy to account for class imbalance. Since the source pseudo-labels span six observed age ranges, the age chance level is 1/6=16.67%1/6=16.67\%. As shown in Tab. 4, average balanced accuracies are 23.53%, 50.31%, and 15.38%, close to chance levels of 25.00%, 50.00%, and 16.67%, indicating limited source-attribute recoverability under this predictor. Scenario Macro-F1 Balanced Acc. Race Gender Age Race Gender Age Cross-Age 22.01 54.60 16.12 21.22 54.73 18.73 Surveillance 25.05 45.80 13.61 25.44 50.86 14.49 Large-Pose 21.53 45.12 15.23 22.48 50.46 15.71 VIS-NIR 17.21 38.86 7.98 24.98 45.20 12.57 Avg. 21.45 46.10 13.24 23.53 50.31 15.38 Table 4: Source-demographic attribute recoverability from released PFD proxies (%). Figure 6: RTA aligns distilled proxies (orange) with transformed prototypes (green) to preserve the transferable relational geometry for downstream FR adaptation. Ablation Study We conduct incremental ablation studies to analyze the individual contributions of OGP and RTA to both target-domain adaptation utility and the reduction of source-identity linkability. Quantitative results are reported in Tab. 3, accompanied by qualitative visualizations for further insight. Figure 7: OGP reduces source-identity matching, yielding abstract proxies (bottom) versus RTA-only (middle). Relational Topology Alignment. Adding RTA consistently improves target-domain adaptation performance over regular distillation. Fig. 6 further shows that, without RTA, the distilled proxies tend to collapse into limited regions of the embedding space. In contrast, RTA preserves both intra- and inter-identity relationships by explicitly aligning relational topology, leading to more discriminative and transferable proxy representations for downstream FR adaptation. Orthogonal Geometry Preservation. OGP is designed to reduce source-identity linkability rather than directly improve adaptation utility. It weakens the correspondence between proxy and source identities while retaining recognition-useful identity geometry. Compared with the RTA-only variant, incorporating OGP reduces both VisMatch and Re-ID attack success from high levels to near chance across all scenarios while maintaining comparable adaptation performance. These results confirm the intended linkability-reduction role of OGP. Fig. 7 provides the corresponding image-level comparison. Conclusion We study the identity paradox in private FR training dataset publication, where recognition-useful cues can also enable linkage to source identities. We propose Private Face Distillation to reduce source-aligned correspondence while preserving proxy identity geometry for recognition learning. Through Orthogonal Geometry Preservation and Relational Topology Alignment, PFD distills an identity-decoupled, geometry-preserving RGB proxy training set. Across domain shifts, the released proxies provide strong downstream utility and near-chance source linkage under the evaluated direct and reconstruction attacks. References S. Barattin, C. Tzelepis, I. Patras, and N. Sebe (2023) Attribute-preserving face dataset anonymization via latent code optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 8001–8010. Cited by: Face Dataset Publication. K. Bešenić, J. Ahlberg, and I. S. Pandžić (2023) Picking out the bad apples: unsupervised biometric data filtering for refined age estimation. The visual computer 39 (1), p. 219–237. Cited by: Age Shift (Cross-Age).. F. Boutros, N. Damer, F. Kirchbuchner, and A. Kuijper (2022) ElasticFace: elastic margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, p. 1578–1587. Cited by: Geometry-Preserving Identity Decoupling.. G. Cazenavette, T. Wang, A. Torralba, A. A. Efros, and J. Zhu (2022) Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, p. 4750–4759. Cited by: Dataset Distillation. J. A. Chan Santiago, P. Tirupattur, G. K. Nayak, G. Liu, and M. Shah (2025) MGD3: mode-guided dataset distillation using diffusion models. In Proceedings of the 42nd International Conference on Machine Learning (ICML), Cited by: Dataset Distillation. J. Chen, C. Yu, C. Kao, T. Pang, and C. Lu (2022) DPGEN: differentially private generative energy-guided network for natural image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 8387–8396. Cited by: Face Dataset Publication. M. Chen, J. Du, B. Huang, Y. Wang, X. Zhang, and W. Wang (2025) Influence-guided diffusion for dataset distillation. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: Dataset Distillation. V. Cherepanova, M. Goldblum, H. Foley, S. Duan, J. P. Dickerson, G. Taylor, and T. Goldstein (2021) LowKey: leveraging adversarial attacks to protect social media users from facial recognition. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: Source-Identity Linkability.. W. Dai, B. Li, N. Dong, G. Bai, and J. S. Dong (2025) FracFace: breaking the visual clues—fractal-based privacy-preserving face recognition. In The Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS), External Links: Link Cited by: Distinction from PPFR. J. Deng, J. Guo, N. Xue, and S. Zafeiriou (2019) ArcFace: additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Experiment Setting. T. Dong, B. Zhao, and L. Lyu (2022) Privacy for free: how does dataset condensation help privacy?. In Proceedings of the 39th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 162, p. 5378–5396. External Links: Link Cited by: Dataset Distillation. R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker (2008) Multi-pie. In 2008 8th IEEE International Conference on Automatic Face & Gesture Recognition (FG), Vol. , p. 1–8. Cited by: Geometric Shift (Large-Pose).. G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller (2008) Labeled faces in the wild: a database forstudying face recognition in unconstrained environments. In Workshop on faces in’Real-Life’Images: detection, alignment, and recognition, Cited by: Datasets. K. Karkkainen and J. Joo (2021) FairFace: face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 1548–1558. Cited by: Source-Attribute Recoverability.. M. Kim, A. K. Jain, and X. Liu (2022) AdaFace: quality adaptive margin for face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 18750–18759. Cited by: Table 1, Experiment Setting. Z. Kuang, H. Liu, J. Yu, A. Tian, L. Wang, J. Fan, and N. Babaguchi (2021) Effective de-identification generative adversarial network for face anonymization. In Proceedings of the 29th ACM International Conference on Multimedia (ACMMM), M ’21, p. 3182–3191. Cited by: Introduction. Z. Kuang, X. Yang, Y. Shen, C. Hu, and J. Yu (2024) Facial identity anonymization via intrinsic and extrinsic attention distraction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 12406–12415. Cited by: Face Dataset Publication. H. Kung, T. Varanka, T. Sim, and N. Sebe (2026) NullFace: training-free localized face anonymization. In 2026 IEEE 20th International Conference on Automatic Face and Gesture Recognition (FG), Vol. , p. 1–10. Cited by: Table 1, Figure 4. M. Lezcano-Casado and D. Martínez-Rubio (2019) Cheap orthogonal constraints in neural networks: a simple parametrization of the orthogonal and unitary group. In Proceedings of the 36th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 97, p. 3794–3803. Cited by: Geometry-Preserving Identity Decoupling.. J. Li, L. Han, R. Chen, H. Zhang, B. Han, L. Wang, and X. Cao (2021) Identity-preserving face anonymization via adaptively facial attributes obfuscation. In Proceedings of the 29th ACM International Conference on Multimedia (ACMMM), M ’21, New York, NY, USA, p. 3891–3899. Cited by: Introduction. W. Li, G. Li, K. Maeda, T. Ogawa, and M. Haseyama (2025) Hyperbolic dataset distillation. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 38, p. 67427–67458. Cited by: Dataset Distillation. B. Maze, J. Adams, J. A. Duncan, N. Kalka, T. Miller, C. Otto, A. K. Jain, W. T. Niggel, J. Anderson, J. Cheney, and P. Grother (2018) IARPA janus benchmark - c: face dataset and protocol. In 2018 International Conference on Biometrics (ICB), Vol. , p. 158–165. Cited by: Resolution Shift (Surveillance).. B. Meden, P. Rot, P. Terhörst, N. Damer, A. Kuijper, W. J. Scheirer, A. Ross, P. Peer, and V. Štruc (2021) Privacy–enhancing face biometrics: a comprehensive survey. IEEE Transactions on Information Forensics and Security (TIFS) 16 (), p. 4147–4183. Cited by: Introduction. Q. Meng, F. Zhou, H. Ren, T. Feng, G. Liu, and Y. Lin (2022) Improving federated learning face recognition via privacy-agnostic clusters. In International Conference on Learning Representations (ICLR), Cited by: Introduction. Y. Mi, Y. Huang, J. Ji, M. Zhao, J. Wu, X. Xu, S. Ding, and S. Zhou (2023) Privacy-preserving face recognition using random frequency components. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 19673–19684. Cited by: Distinction from PPFR. Y. Mi, Z. Zhong, Y. Huang, J. Ji, J. Xu, J. Wang, S. Wang, S. Ding, and S. Zhou (2024) Privacy-preserving face recognition using trainable feature subtraction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 297–307. Cited by: Distinction from PPFR. S. Moschoglou, A. Papaioannou, C. Sagonas, J. Deng, I. Kotsia, and S. Zafeiriou (2017) AgeDB: the first manually collected, in-the-wild age database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Cited by: Datasets. F. P. Papantoniou, A. Lattas, S. Moschoglou, J. Deng, B. Kainz, and S. Zafeiriou (2024) Arc2Face: a foundation model for id-consistent human faces. In Computer Vision – ECCV 2024, p. 241–261. Cited by: Figure 5. S. Peng, H. Zhang, L. Gao, T. Zhang, X. Zhu, B. Li, W. Zhao, and Z. Lei (2026) DiffusionFF: a diffusion-based framework for joint face forgery detection and fine-grained artifact localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 14095–14105. Cited by: Introduction. S. Sengupta, J. Chen, C. Castillo, V. M. Patel, R. Chellappa, and D. W. Jacobs (2016) Frontal to profile face verification in the wild. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), Vol. , p. 1–9. Cited by: Datasets. Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei (2020) Circle loss: a unified perspective of pair similarity optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Geometry-Preserving Identity Decoupling.. Y. Tsai, Y. Li, C. Yu, X. Ren, P. Chen, Z. Chen, and F. Buet-Golfouse (2025) Differentially private fine-tuning of diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 4561–4571. Cited by: Face Dataset Publication, Table 1. S. Wang, Y. Yang, Z. Liu, C. Sun, X. Hu, C. He, and L. Zhang (2025) Dataset distillation with neural characteristic function: a minmax perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 25570–25580. Cited by: Table 1, Figure 4. T. Wang, J. Zhu, A. Torralba, and A. A. Efros (2020) Dataset distillation. External Links: 1811.10959, Link Cited by: Dataset Distillation. Z. Wang, H. Wang, S. Jin, W. Zhang, J. Hu, Y. Wang, P. Sun, W. Yuan, K. Liu, and K. Ren (2023) Privacy-preserving adversarial facial features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 8212–8221. Cited by: Distinction from PPFR. H. Yang, X. Xu, C. Xu, H. Zhang, J. Qin, Y. Wang, P. Heng, and S. He (2024) Gšface: high-fidelity reversible face anonymization via generative and geometric priors. IEEE Transactions on Information Forensics and Security (TIFS) 19 (), p. 8773–8785. Cited by: Table 1. D. Yi, Z. Lei, S. Liao, and S. Z. Li (2014) Learning face representation from scratch. External Links: 1411.7923, Link Cited by: Datasets. Z. Yin, E. Xing, and Z. Shen (2023) Squeeze, recover and relabel: dataset condensation at imagenet scale from a new perspective. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36, p. 73582–73603. Cited by: Distillation via Relational Topology Alignment, Table 1, Experiment Setting. A. Yu, H. Wu, H. Huang, Z. Lei, and R. He (2021) LAMP-hq: a large-scale multi-pose high-quality database and benchmark for nir-vis face recognition. International Journal of Computer Vision (IJCV) 129, p. 1467–1483. Cited by: Modality Shift (VIS-NIR).. L. Yuan, W. Chen, X. Pu, Y. Zhang, H. Li, Y. Zhang, X. Gao, and T. Ebrahimi (2024) PRO-face c: privacy-preserving recognition of obfuscated face via feature compensation. IEEE Transactions on Information Forensics and Security (TIFS) 19 (), p. 4930–4944. Cited by: Face Dataset Publication. L. Yuan, L. Liu, X. Pu, Z. Li, H. Li, and X. Gao (2022) PRO-face: a generic framework for privacy-preserving recognizable obfuscation of face images. In Proceedings of the 30th ACM International Conference on Multimedia (ACMMM), M ’22, New York, NY, USA, p. 1661–1669. Cited by: Introduction. J. Zhang, L. Dai, F. Ye, Z. Chen, P. Li, X. Yang, and B. Sheng (2026a) Dataset distillation via a noise-unconstrained generative model. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (), p. 1–18. Cited by: Dataset Distillation. T. Zhang, L. Gao, S. Peng, X. Zhu, and Z. Lei (2025) DevFD : developmental face forgery detection by learning shared and orthogonal lora subspaces. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 38, p. 4103–4128. Cited by: Introduction. Y. Zhang, J. Ji, T. Wang, R. Zhao, W. Wen, and Y. Xiang (2026b) Make identity indistinguishable: utility-preserving face dataset publication with provable privacy guarantees. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 48 (1), p. 127–139. Cited by: Introduction, Face Dataset Publication, Data Release Setting, Table 1, Figure 4. R. Zhao, Y. Zhang, T. Wang, W. Wen, Y. Xiang, and X. Cao (2025a) Visual content privacy protection: a survey. ACM Comput. Surv. 57 (5). Cited by: Introduction. W. Zhao, X. Zhu, Z. He, X. Zhang, and Z. Lei (2023) Cross-architecture distillation for face recognition. In Proceedings of the 31st ACM International Conference on Multimedia (ACMMM), M ’23, New York, NY, USA, p. 8076–8085. Cited by: Geometry-Preserving Identity Decoupling.. W. Zhao, X. Zhu, H. Shi, X. Zhang, G. Zhao, and Z. Lei (2025b) Global cross-entropy loss for deep face recognition. IEEE Transactions on Image Processing (TIP) 34 (), p. 1672–1685. Cited by: Geometry-Preserving Identity Decoupling.. T. Zheng, W. Deng, and J. Hu (2017) Cross-age lfw: a database for studying cross-age face recognition in unconstrained environments. CoRR abs/1708.08197. External Links: Link Cited by: Datasets. T. Zheng and W. Deng (2018) Cross-pose lfw: a database for studying cross-pose face recognition in unconstrained environments. Beijing University of Posts and Telecommunications, Tech. Rep 5 (7), p. 5. Cited by: Datasets. Y. Zhu, X. Yu, M. Chandraker, and Y. Wang (2020) Private-knn: practical differential privacy for computer vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Introduction.