Paper deep dive
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
Leon Ranke, Wolfgang HĂŒbner, Ronny Hug, Michael Arens, JĂŒrgen Beyerer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/28/2026, 4:37:09 AM
Summary
This paper proposes a compute-aware multi-task protocol to evaluate Privacy-Enhancing Technologies (PETs) in computer vision, arguing that classification accuracy is an insufficient proxy for utility. The authors demonstrate that PETs with similar classification performance can differ significantly in their ability to support other tasks like angle prediction and jigsaw puzzle solving. The study analyzes various transformations including Gaussian blurring, Locally Orderless Images, block-based primitives, and learnable image encryption schemes (Tanaka, E-Tanaka, EtC), highlighting the need for evaluation protocols that assess task-dependent learnability, shortcut leakage, and image-domain obfuscation jointly.
Entities (13)
Relation Signals (9)
Angle Prediction â probes â geometric consistency
confidence 92% · Angle Prediction [12] estimates the relative rotation... This task differs from classification... it requires consistent extraction of orientation-dependent structure
Jigsaw Puzzle Solving â probes â Spatial Compatibility
confidence 92% · Jigsaw Puzzle Solving [40] probes spatial compatibility and local-to-global structural consistency.
Classification Accuracy â isinsufficientproxyfor â Generic Vision Tasks
confidence 90% · As a result, it is too simplistic as a proxy for generic vision tasks.
Gaussian Blurring â istypeof â Irreversible Transformation
confidence 90% · We evaluate controlled irreversible transformations... Gaussian Blurring acts as a low-pass filter... Gaussian blur and LOIs remove information... even when their parameters are known
Animal Species Classification â isusedfor â Evaluation
confidence 90% · For classification, we use the Animal Species Classification dataset [1]... Angle prediction is evaluated on the same dataset.
ResNet-34 â isusedfor â Classification
confidence 90% · We use a ResNet-34 [15] with an N-class classification head.
EtC â usesprimitive â Block Rotation + Flipping
confidence 88% · Encryption then Compression (EtC) [25] scheme (Block Rotation + Flipping, Colour Shuffling, Block Shuffling)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection. This trade-off is commonly evaluated using image classification, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks. Exhaustive downstream-task evaluation, however, is computationally expensive because models must often be trained for each PET transformation and parameter setting. We therefore propose a compute-aware multi-task protocol for evaluating PETs in model training. It combines lightweight proxy tasks that target complementary aspects of visual structure while remaining simple and fast to compute. Across irreversible privacy transformations, key-based block primitives, and learnable image encryption schemes, we demonstrate that PETs with similar classification accuracy can differ substantially on other tasks. The outcomes highlight the need for PET evaluation protocols that move beyond classification-only reporting.
Tags
Links
- Source: https://arxiv.org/abs/2608.27066v1
- Canonical: https://arxiv.org/abs/2608.27066v1
Trouble viewing inline? Open PDF directly â
Full Text
58,570 characters extracted from source content.
Expand or collapse full text
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations Leon Ranke 1,2 , Wolfgang HĂŒbner 1 , Ronny Hug 1 , Michael Arens 1 , and JĂŒrgen Beyerer 1,2 1 Fraunhofer Institute of Optronics, System Technologies and Image Exploitation (IOSB), GutleuthausstraĂe 1, 76275 Ettlingen, Germany 2 Karlsruhe Institute of Technology (KIT) firstname.lastname@iosb.fraunhofer.de https://w.iosb.fraunhofer.de Abstract. Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely processing it, creating a trade-off between task performance and protection. This trade-off is commonly evaluated using image classifica- tion, which primarily captures semantic separability and remains robust despite significant geometric, spatial layout or local boundary alterations. As a result, it is too simplistic as a proxy for generic vision tasks. Exhaus- tive downstream-task evaluation, however, is computationally expensive because models must often be trained for each PET transformation and parameter setting. We therefore propose a compute-aware multi-task pro- tocol for evaluating PETs in model training. It combines lightweight proxy tasks that target complementary aspects of visual structure while remaining simple and fast to compute. Across irreversible privacy trans- formations, key-based block primitives, and learnable image encryption schemes, we demonstrate that PETs with similar classification accuracy can differ substantially on other tasks. The outcomes highlight the need for PET evaluation protocols that move beyond classification-only re- porting. Code is available at: https://github.com/LeonRanke/Task- Dependent-Learnability. Keywords: Privacy-Preserving Machine Learning· Utility Evaluation 1 Introduction Privacy-Preserving Machine Learning (PPML), as applied to computer vision tasks, aims to reduce the exposure of sensitive visual content while retaining sufficient information for effective learning. For representation-altering PETs (Privacy-Enhancing Technologies), such as learnable image encryption [17,34,48] and related image-domain transformations [18,25,47], this creates a fundamen- tal tension: the transformation should obscure human-interpretable information, while the resulting representation must still support downstream learning. arXiv:2608.27066v1 [cs.CV] 27 Aug 2026 2L. Ranke et al. Existing evaluations of representation-altering PETs commonly report clas- sification accuracy as the main measure of utility [18,25,34,48]. This provides a useful but narrow view: classification primarily tests whether class-discriminative statistics remain separable under the transformed data distribution. It does not necessarily probe whether geometric relations, spatial layout, boundary compat- ibility, or orientation-dependent structures are preserved [51]. Indeed, models can maintain high classification accuracy even when spatial structure and high- frequency components are substantially degraded [2,10,16]. Transformed images may therefore support classification while being poorly suited for tasks such as detection, segmentation, tracking, pose estimation, or geometric matching. A natural solution is to evaluate PETs on a broad suite of downstream tasks, but in practice this is often prohibitively expensive. Because transformations alter low-level statistics and spatial structure, pre-training on natural images cannot be assumed to generalise across parametrisations, including encryption keys, design parameters, and scale. Rigorous evaluation would therefore require training from scratch for each transformation family and parameter setting. This motivates a compute-aware intermediate evaluation stage: before investing in expensive downstream pipelines, lightweight diagnostic proxy tasks (e.g. relative angle prediction or jigsaw puzzle-solving) can probe which broad types of visual structure remain learnable. In this paper, learnability is used in an operational, empirical sense. Given a downstream task, a transformation T λ,Îș , 3 a model, and a training protocol, learnability is estimated as the task performance achieved by a model trained from scratch on transformed samplesex = T λ,Îș (x). Learnability is thus not an intrinsic property of the transformation alone, instead it characterises what in- formation a particular learning algorithm can access under finite data and com- putational resources. In this light we distinguish obfuscation from encryption: the former irreversibly alters an image, whereas the latter is invertible given the key Îș (security is not part of this definition). Information may be preserved in principle yet remain inaccessible, or difficult to exploit, to a model that receives no key-dependent information. Our results show that learnability under privacy-motivated image transfor- mations is strongly task-dependent. Transformations that preserve compara- tively high classification accuracy can severely impair tasks that rely on geo- metric consistency and spatial order. Conversely, key-based transformations can introduce deterministic shortcut cues that some tasks exploit, resulting in high performance, while performance on other tasks collapses. By relating these task- specific behaviours to image-domain obfuscation and leakage metrics, we find that no single utility or obfuscation metric fully characterises a transformation. Together, these findings motivate evaluation protocols for privacy-preserving vi- sion that go beyond classification performance, jointly assessing task-dependent learnability, shortcut leakage, and image-domain obfuscation. In summary, our main contributions are: 3 Elements of a transformation family are indexed by the parameter vector λ; Îș denotes a key, if applicable. Beyond Classification: Task-Dependent Learnability3 1. A compute-aware multi-task protocol for evaluating task-dependent learn- ability under privacy-motivated image transformations. 2. An analysis of a controlled set of multi-scale image transformations, span- ning obfuscation, block-based transforms and Learnable Image Encryption schemes [25,34,48], showing how scale and key structure affect learnability. 3. Evidence that classification accuracy, proxy-task utility, and image-domain obfuscation metrics capture distinct and individually incomplete aspects of representation-altering PETs in PPML. 2 Related Work Research on PPML for computer vision has primarily focused on secure com- putation mechanisms [13, 19, 20, 29, 33], privacy-preserving training protocols [7, 42, 50, 56], and data transformations [17, 18, 25, 34, 48]. Comparatively little attention has been devoted to systematic evaluation methodologies that jointly assess image-domain obfuscation and downstream task utility. 2.1 Utility Evaluation in PPML The notion of model utility varies across PETs. Here, utility refers to how well a protected training pipeline preserves the predictive performance of an equivalent plaintext pipeline. A useful distinction can be drawn between (i) functionality- preserving approaches, which aim to reproduce plaintext computations in pro- tected domains; (i) training-oriented PETs, modifying the optimisation process; and (i) representation-altering PETs, that transform images before training. Functionality-preserving PETs include Fully Homomorphic Encryption [11], Secure Multiparty Computation [14] and Trusted Execution Environments [41]. These approaches execute computations in protected domains while reproducing the behaviour of the corresponding plaintext computation. Utility degradation therefore arises mainly from practical constraints, including reduced numerical precision, approximated non-linear functions, and restricted model architectures. Image classification on benchmarks such as MNIST and CIFAR-10 remains fea- sible [13,20,24,39,44,45,53,54], although often at substantial computational and communication overhead, which can hinder scalability. Training-oriented PETs include Differential Privacy (DP) [8] and Federated Learning (FL) [36]. These approaches retain the original input representation while modifying the optimisation or collaboration process. DP-based training introduces a direct privacyâutility trade-off through mechanisms such as gradient clipping and adding calibrated noise [3,42], whereas FL distributes optimisation across multiple data holders, where utility can degrade through noisy updates, heterogeneous client data, or partial client participation [50,56]. Representation-altering PETs include learnable image encryption [17,34,48] and related image-domain transformations [18,25,47]. These approaches trans- form images before training and require models to learn directly from the trans- formed representation. Unlike functionality-preserving PETs, they do not gener- ally preserve arbitrary plaintext computations. Their utility, therefore, depends 4L. Ranke et al. Table 1: Datasets used to evaluate model utility in representation-altering PETs. Representation-altering schemeDataset(s) used to evaluate utility Learnable Image Encryption [48]CIFAR-10 [23] Pixel-Based Image Encryption [47]CIFAR-10 [23], STL-10 [5] Block-wise Scrambled Image Recognition [34] CIFAR-10, CIFAR-100 [23] InstaHide [18] CIFAR-10, CIFAR-100 [23], MNIST [28], ImageNet [6] (subset) Learnable Image Encryption for Vision Transformers [17] CIFAR-10 [23], Tiny-ImageNet [27] on whether the transformation retains sufficient task-relevant information, and it has predominantly been evaluated through downstream classification perfor- mance relative to training on plaintext images (see Tab. 1). This dominance of classification benchmarks creates an evaluation gap for representation-altering PETs. Classification accuracy indicates whether semantic labels remain predictable, but provides limited evidence about other structural properties of the transformed representation. In particular, it does not directly assess whether global orientation, spatial compatibility, local boundary continu- ity, or position-dependent cues are preserved [51]. These properties matter for many downstream vision tasks, yet they are rarely examined because training complete downstream models in transformed domains is computationally ex- pensive. We address this gap by introducing diagnostic proxy tasks that probe complementary structural requirements while remaining inexpensive enough to evaluate across multiple parameter and key settings. The tasks are adapted from self-supervised objectives [12] but serve here as diagnostics of structural learn- ability rather than pre-training objectives. 2.2 Image Security Metrics In contrast to utility evaluation, the image encryption literature provides ex- tensive statistical security evaluation suites [35,43]. Commonly reported metrics quantify image-domain properties associated with leakage or obfuscation, in- cluding perceptual leakage between plain and transformed images (e.g., MSE, PSNR, SSIM [49], LPIPS [52], and MI), encrypted-domain statistics (e.g., en- tropy, adjacent-pixel correlation, spectral flatness, divergence from uniformity), and differential or key-sensitivity properties. Table 2 summarises these metrics. We treat these quantities as indicators of image-domain obfuscation and leak- age rather than as formal privacy guarantees. They measure perceptual, statisti- cal, or differential properties of transformed images, but, by themselves, they do not establish resistance to de-obfuscation, reconstruction, recognition, attribute inference, or membership inference attacks [9, 37, 38, 46, 55]. Nevertheless, they are commonly used in image-encryption evaluations [35, 43] and provide use- Beyond Classification: Task-Dependent Learnability5 Table 2: Image-domain obfuscation and leakage indicators for RGB images. Metrics are computed per channel (R, G, B) and averaged. Intensities are in 0,...,Lâ 1, with L = 256 throughout. Name Abbreviation Range Heuristic target Perceptual Leakage Mean Squared Error MSE[0, (Lâ 1) 2 ] L 2 â1 6 Peak Signal-to-Noise Ratio (dB) PSNR[0,â)10 log 10 6(Lâ1) L+1 Structural Similarity Index Measure [49] SSIM[â1, 1]0 Learned Perceptual Image Patch Similarity [52] LPIPS[0,â)â Mutual Information (bits) MI[0, log 2 (L)]0 Encrypted Image Statistics Chi-Squared Goodness-of-Fit Test Ï 2 [0,â) â€ Ï 2 0.05 (Lâ1) KullbackâLeibler Divergence D KL [0,â)0 Shannon Information Entropy (bits) IE[0, log 2 (L)]log 2 (L) 2D Information Entropy [26] (bits) 2D-IE[0, log 2 (2Lâ1)] log 2 (2Lâ1) Adjacent-pixel correlation C[â1, 1]0 Spectral Flatness Measure SFM[0, 1]1 Differential Analysis Number of Pixels Change Rate (w.r.t. I) NPCR I [0, 1]1â 1 L Unified Average Changing Intensity (w.r.t. I) UACI I [0, 1] L+1 3L Key Analysis Key-space size |K|[1,â)â„ 2 128 Number of Pixels Change Rate (w.r.t. Îș) NPCR Îș [0, 1]1â 1 L Unified Average Changing Intensity (w.r.t. Îș) UACI Îș [0, 1] L+1 3L ful complementary information about how strongly a transformation alters the visible image domain. 3 Methodology The objective of this paper is to characterise how image-domain transformations affect the learnability of different visual structures. To this end, transformations are viewed as operators that remove, rearrange, or statistically alter components of the visual signal. Downstream proxy tasks then act as diagnostic probes that test whether these components remain accessible to the learning algorithm. Transformations are defined as parametrisable operators T λ,Îș :X â e X acting on the plain image spaceX â R HĂWĂ3 , where λ denotes the scale parameter and Îș â K a secret key. For non-keyed obfuscations, Îș = â . Each transformation induces a modified data distribution with altered spatial, geometric, perceptual, or statistical properties. 3.1 Transformation Operators We evaluate controlled irreversible transformations, key-based block primitives, and learnable image encryption schemes [25,34,48]. The first group acts as a di- agnostic baseline that removes specific visual structures in a controlled manner, while the remaining two groups represent transformations used in learnable im- age encryption. Together, they enable a systematic analysis of which structural 6L. Ranke et al. Ï = 2Ï = 8Ï = 20Ï = 50 B = 4B = 8B = 32B = 104 B = 4B = 8B = 32B = 104B = 4B = 8B = 32B = 104 B = 4B = 8B = 32B = 104B = 4B = 8B = 32B = 104 B = 4B = 8B = 32B = 104B = 4B = 8B = 32B = 104 B = 4B = 8B = 32B = 104B = 4B = 8B = 32B = 104 Gaussian Blurring [30]Locally Orderless Images (LOIs) [22] Block ShufflingPixel Shuffling Colour ShufflingBlock Rotation + Flipping NP-Transformation Tanaka [48] E-Tanaka [34]EtC [25] Fig. 1: Qualitative examples of transformation families across increasing scale. components remain learnable under different forms of image-domain alteration. Figure 1 provides qualitative examples of the considered transforms. Gaussian Blurring acts as a low-pass filter, removing high-frequency com- ponents while preserving spatial topology and global intensity relationships. For scale parameter Ï, images are obtained via separable convolution with a Gaus- sian kernel of standard deviation Ï and kernel size k = 2â3Ïâ+1 [30]. Increasing Ï progressively suppresses fine-scale details, isolating the contribution of spatial- frequency bands to downstream task performance. Locally Orderless Images (LOIs) discard spatial ordering within non- overlapping BĂ B windows while approximately preserving marginal intensity distributions inside those windows [22]. As B increases, local correlations are re- moved at progressively larger scales, while global intensity statistics are retained. LOIs, therefore, provide a controlled mechanism for analysing the impact of spa- tial order at different scales on learnability. Beyond Classification: Task-Dependent Learnability7 Block-Based Transformations are invertible operations used in the con- sidered learnable image encryption schemes [25, 34, 48]. Images are partitioned into non-overlapping B Ă B blocks, followed by deterministic, key-based oper- ations applied within or across blocks. Depending on the primitive, these oper- ations either preserve local block content while disrupting global arrangement, or disrupt local adjacency while preserving block-level placement. The anal- ysed primitives are: permutation of spatial block positions (Block Shuffling); block rotations by multiples of 90 ⊠combined with horizontal and vertical flips (Block Rotation + Flipping); permutation of pixel positions within blocks (Pixel Shuffling); permutation of the RGB channels per block (Colour Shuffling); and pixel-wise intensity inversion (Negative-Positive (NP) Transformation). Learnable Image Encryption Schemes combine multiple block-based primitives into structured pipelines. These schemes are invertible given the secret key Îș, but can induce transformed-domain regularities that are difficult to exploit without access to the secret key. We analyse: Tanaka [48] (Pixel Shuffling and NP Transformation); Extended Tanaka (E-Tanaka) [34] (Pixel Shuffling, NP Transformation, Block Shuffling); and the Encryption then Compression (EtC) [25] scheme (Block Rotation + Flipping, Colour Shuffling, Block Shuffling). Gaussian Blurring uses Ï as the scale parameter λ, whereas all other trans- formations use the block size B. The distinction between obfuscations and keyed transformations matters here: Gaussian blur and LOIs remove information from the image even when their parameters are known, whereas block-based transfor- mations and learnable image encryption schemes preserve information. A learner receiving only transformed images may nonetheless be unable to exploit this in- formation effectively, because the transformation can misalign with the modelâs inductive biases. Some PETs address this restriction by implicitly providing the model with key-dependent information (e.g., E-Tanaka [34] provides the model with information related to the block permutation). Moreover, when a fixed key is used across training and testing, deterministic key-dependent regularities may become learnable and act as shortcut cues. 4 Our protocol therefore measures fixed-key empirical learnability under a specified training setup, rather than in- trinsic information preservation or formal privacy. 3.2 Evaluation Tasks The goal is not to approximate the performance of every possible downstream task. Instead, a tractable protocol is introduced that can be repeated across many transformations, scales, and keys to identify which broad types of visual structure remain learnable. The selected tasks act as diagnostic probes: they are inexpensive relative to full detection or segmentation pipelines and require different structural properties of the transformed image distribution. Rather than serving as independent benchmarks, they function as a combined test suite for structural sufficiency under transformation. 4 See Sec. 5.3 for empirical evidence of key-induced shortcut cues. 8L. Ranke et al. Table 3: Overview of investigated proxy tasks, including approximate training time. TaskProbesApprox. Time Classificationsemantic separabilityâ 1.4 h Angle Predictionglobal orientation / geometric consistency â 0.8 h Jigsaw Puzzle Solving spatial compatibility / local structureâ 9.4 h Image Classification predicts a discrete class label y â 0,...,N â 1 given transformed imagesex = T λ,Îș (x). We use a ResNet-34 [15] with an N-class classification head. Since classification provides a sparse supervision signal, a wide range of input variations can be tolerated as long as class-discriminative statistics remain separable. Classification is included because it is the dominant task for evaluating the utility of representation-altering PETs. 5 Angle Prediction [12] estimates the relative rotation âα = α 2 âα 1 between two independently rotated and transformed views of the same image: ex 1 = T λ,Îș (R α 1 · x),ex 2 = T λ,Îș (R α 2 · x).(1) The angles α 1 ,α 2 are sampled randomly from a uniform distribution, with transformations applied after rotation. A CNN processes the two transformed images concatenated along the channel dimension. A regression head predicts (sin(âα), cos(âα)) to avoid discontinuities in angular regression. This task dif- fers from classification in two ways. First, it requires consistent extraction of orientation-dependent structure from two transformed views of the same im- age. Second, the target is continuous, making the task more sensitive to gradual degradation of geometric information. Jigsaw Puzzle Solving [40] probes spatial compatibility and local-to-global structural consistency. Given an image split into a regular kĂk grid, the pieces are permuted according to an unknown permutation Ï. Following the transformer based JPDVT [31], the model recovers Ï by assigning shuffled pieces to their original grid positions. Varying k changes the spatial scale of the task: larger values reduce the context available per piece and increase the solution space. Jigsaw puzzle solving tests whether spatial relationships remain learnable af- ter transformation. Downstream tasks such as detection, segmentation, pose esti- mation, and tracking depend not only on semantic separability but also on spatial layout, boundary compatibility, and consistent relationships between neighbour- ing regions. Jigsaw solving acts as a tractable intermediate probe for spatially exploitable information. Importantly, high jigsaw accuracy does not necessarily imply preservation of natural spatial structure. Under fixed-key transformations, performance may also reflect deterministic cues induced by the transformations, such as position-dependent block patterns or key-specific regularities. Table 3 summarises the proxy tasks investigated. 5 See Sec. 2.1 for a discussion of utility evaluation in the PET literature. Beyond Classification: Task-Dependent Learnability9 4 Experimental Setup All models are trained from scratch exclusively on transformed data e X. They do not receive paired plaintext images, secret keys, or transformation-specific adaptation modules. This enables comparing transformations within a common learning setup and identifying which types of structure remain accessible to standard learners under limited computational resources and data volumes. A predefined set of scale parameters λ is applied consistently during both training and testing across all transformation families. For key-based transfor- mations, a fixed-key setting is evaluated: one secret key Îș is sampled per configu- ration and used for both training and testing. This aligns with common learnable image encryption evaluations, in which models are trained directly in a static transformed domain. However, it also means that models may exploit deter- ministic key-dependent regularities. Parameter choices for each experiment are reported in the supplementary material. Training is performed for a predefined maximum number of epochs with early stopping, providing an upper bound on optimisation cost while mitigating overfitting. For classification, we use the Animal Species Classification dataset [1], which has N = 15 classes, with images resized to 416Ă416. The dataset provides suffi- cient spatial resolution to meaningfully analyse scale-dependent transformations while remaining computationally tractable. Performance is evaluated via top-1 accuracy with a chance level of 1/N. Angle prediction is evaluated on the same dataset. Images are rotated and then cropped to exclude hard image edges that leak rotational information. Per- formance is reported as the mean angular error in degrees. Since the angular error is measured as the shortest circular distance between the predicted and ground-truth relative angle, errors lie in [0 ⊠, 180 ⊠]. For uniformly distributed rotations, the expected error of a random or constant predictor is therefore 90 ⊠. In the puzzle-solving experiments, images from the Animal Species Classifi- cation dataset are resized to 384Ă384 pixels and divided into a regular kĂk grid, where k â4, 6, 8 denotes the puzzle size. For each puzzle size, the transforma- tion block size B is chosen relative to the side length p = 384/k of one puzzle piece. We consider settings in which each puzzle piece contains r Ă r blocks, with r â 0.5, 1, 2, 4, 8. The considered configurations and resulting values are shown in Fig. 2. This construction controls the relationship between the transfor- mation and puzzle scales, avoiding misalignment between transformation blocks and piece boundaries. Performance is measured as the percentage of correctly placed puzzle pieces; the chance level of a kĂ k puzzle is 1/k 2 . 5 Results & Discussion This section analyses how the considered transformations affect task-dependent learnability. We first report classification performance, which aligns with the dominant utility view in prior work, and then compare it with relative-angle prediction and jigsaw puzzle solving. Finally, we relate task performance to image-domain obfuscation metrics. 10L. Ranke et al. image side length 384 piece size 384/k B = 2 p B = p/2 0.5Ă 0.5 setting one block spans 2Ă 2 pieces 2Ă 2 setting four blocks per piece (a) Illustration for k = 4. The red block shows the 0.5Ă0.5 setting, while the blue block shows the 2Ă 2 setting. (b) Resulting transformation block size B, in pixels, for each configura- tion. piece contains rĂ r blocks k 0.5 1 2 4 8 4Ă 4 192 96 48 24 12 6Ă 6 128 64 32 16 8 8Ă 8 96 48 24 12 6 Values assume an input resolution of 384Ă 384. For a puzzle piece of side length p = 384/k, the block size is B = p/r, where rĂ r is the number of BĂ B blocks per piece. Fig. 2: Relationship between the puzzle size k and the transformation block size B. 5.1 Image Classification Figure 3 shows the top-1 classification accuracies as λ increases. The plain base- line reaches 87.00%, while chance performance is 6.67%. Classification remains well above chance for all transformations, even under substantial image-domain alteration. Gaussian blur degrades accuracy gradually, but still achieves 57.10% at Ï = 70. LOIs show a stronger scale effect, decreasing from near-baseline per- formance at small block sizes to 27.75% when correlations are removed globally. Block-based transformations differ according to which structures they dis- rupt. Colour shuffling, block rotation/flipping, and the NP-transformation pre- serve high classification accuracy across most scales, indicating that class dis- criminative information remains accessible despite substantial visual changes. In contrast, pixel shuffling increasingly disrupts local adjacency and reduces accu- racy at larger block sizes. Block shuffling shows the opposite trend: small blocks strongly disrupt the global arrangement, whereas larger blocks preserve more of the layout and therefore recover classification performance. The composed encryption schemes inherit these behaviours. Tanaka degrades with increasing pixel-shuffling scale; EtC improves as larger shuffled blocks pre- serve more global structure, and E-Tanaka remains comparatively low across scales because it combines local and global disruption. Overall, classification captures only semantic separability: high accuracy indicates that semantic labels remain predictable, but it does not imply that geometric or spatial structures required by other tasks are preserved. This does not mean classification is always the more permissive probe. As Sec. 5.2 shows, for some transformations (e.g., Gaussian Blurring), classification degrades faster than angle prediction, so the achieved task utility depends on which structures a transformation removes and which are required for solving the task. Any single task gives an incomplete and potentially misleading picture of utility. Beyond Classification: Task-Dependent Learnability11 010203040506070 Ï 0 20 40 60 80 100 Top-1 Accuracy [%] â Gaussian Blurring 10 0 10 1 10 2 B Locally Orderless Images (LOIs) 10 0 10 1 10 2 B 0 20 40 60 80 100 Top-1 Accuracy [%] â Block-Based transformations Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation 10 0 10 1 10 2 B Learnable Image Encryption Schemes Tanaka E-Tanaka EtC Plain BaselineChance level 1 N Fig. 3: Classification accuracy against transformation strength. Accuracy remains above chance for all transformations, indicating that semantic separability is preserved even under substantial image-domain alteration. 5.2 Angle Prediction Figure 4 reports the mean angular error, with a chance performance of 90 ⊠. The plain baseline achieves 2.48 ⊠. Unlike classification, angle prediction directly tests whether a coherent global reference frame remains accessible. Gaussian blur has a limited effect except for large scales, suggesting that relative orientation can be inferred from coarse global structure and does not require fine, high-frequency detail. LOIs behave differently: performance remains close to the baseline for small blocks but degrades rapidly once correlations at relevant scales are removed, approaching chance in global settings. For key-based primitives, colour shuffling, NP transformation, and block ro- tation/flipping have little effect on angle prediction, while pixel shuffling de- grades at larger block sizes. Block shuffling again shows a scale-dependent inverse trend: small shuffled blocks disrupt global orientation cues, whereas larger blocks preserve enough layout for accurate prediction. Among composed schemes, E- Tanaka substantially impairs angle prediction across all scales, while Tanaka and EtC depend strongly on block size. These results demonstrate that transforma- tions with similar classification accuracy can differ substantially in the geometric information they preserve. 12L. Ranke et al. 010203040506070 Ï 10 1 10 2 Mean Angle Error [ ⊠] â Gaussian Blurring 10 0 10 1 10 2 B Locally Orderless Images (LOIs) 10 1 10 2 B 10 1 10 2 Mean Angle Error [ ⊠] â Block-Based transformations Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation 10 1 10 2 B Learnable Image Encryption Schemes Tanaka E-Tanaka EtC Plain BaselineChance Level 90 ⊠Fig. 4: Mean angular error against transform strength (lower is better). Prediction remains robust under transformations that preserve coarse global structure, such as moderate blur, but degrades strongly when spatial order is removed at larger scales or when local and global disruptions are combined. 5.3 Jigsaw Puzzle Solving Figure 5 reports piece accuracy for k â 4, 6, 8 puzzles. The plain baseline achieves high piece accuracy across all puzzle sizes, with chance performance decreasing from 6.25% for 4Ă 4 puzzles to 1.56% for 8Ă 8 puzzles. LOI strongly reduces jigsaw performance when spatial order is removed at puzzle-relevant scales. When transformation blocks are as large as, or larger than, puzzle pieces, performance approaches chance. As more transformation blocks fit inside each puzzle piece, local spatial information becomes available, and performance improves. This confirms that the task is sensitive to boundary compatibility and local-to-global spatial consistency. The key-based transformations reveal a different effect. Pixel shuffling and several composed schemes remain solvable in the fixed-key setting, even when they disrupt global image structure. This does not indicate preservation of nat- ural spatial compatibility. Instead, the puzzle solver can exploit deterministic transformation-induced regularities, such as key-dependent block or position cues that are consistent across training and testing. Evaluating these settings with a key different from the one used in training causes performance to collapse toward chance (see dashed lines in Fig. 5). Comparing LOI and Pixel Shuffling Beyond Classification: Task-Dependent Learnability13 468 Puzzle Sizek 0 20 40 60 80 100 Piece Accuracy [%] â Locally Orderless Images (LOIs) 468 Puzzle Sizek Block Shuffling 468 Puzzle Sizek Pixel Shuffling 468 Puzzle Sizek 0 20 40 60 80 100 Piece Accuracy [%] â Tanaka 468 Puzzle Sizek E-Tanaka 468 Puzzle Sizek EtC Blocks Per Piece:0.51248Plain BaselineChance level Fig. 5: Jigsaw puzzle-solving accuracy for puzzles of different sizes; dashed coloured curves indicate evaluation with a different key from the training key. Puzzle solving ex- poses spatial placement cues that classification does not capture. LOI and some block- shuffling settings degrade strongly when spatial order is removed at puzzle-relevant scales, whereas several key-based transformations remain highly solvable. confirms that the high performance of the latter is due to key-induced regular- ities. LOI can be viewed as a variant of Pixel Shuffling, with different keys for each block. Attacks on block-scrambling schemes (such as EtC [25]) use puzzle solvers to reconstruct the plain image [4]. The high task utility can therefore also be interpreted as a leakage signal. Block shuffling further shows that performance depends on the alignment between transformation scale and puzzle scale: some settings preserve or expose useful placement cues, while others remove them. Jigsaw solving provides a complementary diagnostic view. Low performance indicates that spatial information required by the puzzle solver is no longer read- ily accessible. High performance may reflect either the preservation of natural spatial structure or shortcut cues introduced by transformations. This distinction highlights the importance of interpreting tasks jointly rather than in isolation. 5.4 Obfuscation-Utility Trade-off Task utility alone is only one aspect of PET evaluation. A transformation that preserves high task performance while remaining perceptually close to the orig- inal provides limited obfuscation; one that strongly alters the image but de- stroys downstream learnability is equally impractical. Evaluation should there- fore jointly consider task utility and image-domain leakage indicators. Figure 6 relates normalised task utility to an LPIPS-based obfuscation score [52], where higher utility indicates better learnability relative to the plain baseline and higher LPIPS indicates lower perceptual similarity to the original. The metrics in Tab. 2 14L. Ranke et al. 0.00.20.40.60.81.0 Normalised LPIPSâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.81.0 Normalised LPIPSâ Angle Prediction 0.00.20.40.60.81.0 Normalised LPIPSâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 6: Obfuscation-utility trade-off. Normalised task utility against normalised LPIPS. For angle prediction, the utility direction is inverted, since lower errors are better. constitute the broader evaluation framework; LPIPS is used here as a represen- tative perceptual indicator. The full PET evaluation should consider additional metrics (see the supplementary material for the remaining metrics in Tab. 2). The trade-off varies substantially across tasks and transformations. Some transformations retain high classification accuracy while offering limited per- ceptual obfuscation. Others increase perceptual distance but degrade geometric or spatial learnability. Fixed-key transformations add a further complication, achieving high perceptual obfuscation while still exposing deterministic cues that can be exploited by puzzle-solving. No single point on the obfuscation axis, therefore, predicts utility across tasks: the same LPIPS level corresponds to very different learnability depending on which structures a task requires. 6 Conclusion We studied learnability under image obfuscation as a task-dependent property. Results show that classification accuracy, the dominant utility measure in PET evaluation, provides only a partial view. Transformations that preserve semantic separability can differ substantially in how well they preserve geometric con- sistency and spatial compatibility. The proposed protocol combines classifica- tion, relative-angle prediction, and jigsaw puzzle solving as lightweight diagnostic probes for complementary structures. Further, we show that fixed-key transfor- mations can introduce deterministic shortcut cues that some tasks exploit. This highlights the need to interpret proxy-task performance jointly rather than in iso- lation. Future evaluations of privacy-preserving vision methods can benefit from moving beyond classification-only reporting to jointly consider task-dependent learnability, shortcut leakage, and image-domain obfuscation. We establish that the three tasks capture distinct structural properties, but not whether they pre- dict full downstream tasks. Training downstream models on a subset of configu- rations is the natural next step and would confirm the proxiesâ practical value. Beyond Classification: Task-Dependent Learnability15 References 1. Aashman, V., Shreshth, J., Khanak, A.: Animal species classification â v3 (2023), https://w.kaggle.com/datasets/utkarshsaxenadn/animal-image- classification-dataset/data 2. Brendel, W., Bethge, M.: Approximating CNNs with bag-of-local-features models works surprisingly well on ImageNet. In: International Conference on Learning Representations (2019) 3. Bu, Z., Dong, J., Long, Q., Su, W.J.: Deep learning with gaussian differential privacy. Harvard Data Science Review 2(23) (2020). https://doi.org/10.1162/ 99608f92.cfc5d25 4. Chuman, T., Kiya, H.: A jigsaw puzzle solver-based attack on image encryption using vision transformer for privacy-preserving DNNs. Information 14(6) (2023) 5. Coates, A., Ng, A., Lee, H.: An analysis of single-layer networks in unsupervised feature learning. In: Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings (2011) 6. Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A large- scale hierarchical image database. In: IEEE Conference on Computer Vision and Pattern Recognition (2009) 7. Dong, J., Durfee, D., Rogers, R.: Optimal differential privacy composition for ex- ponential mechanisms. In: International Conference on Machine Learning. PMLR (2020) 8. Dwork, C.: Differential privacy. In: International Colloquium on Automata, Lan- guages and Programming (2006) 9. Geiping, J., Bauermeister, H., Dröge, H., Moeller, M.: Inverting gradients: How easy is it to break privacy in federated learning? In: Advances in Neural Information Processing Systems. vol. 33 (2020) 10. Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F.A., Brendel, W.: ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In: International Conference on Learning Representations (2018) 11. Gentry, C.: Fully homomorphic encryption using ideal lattices. In: Proceedings of the 41st Annual ACM Symposium on Theory of Computing. Association for Computing Machinery (2009). https://doi.org/10.1145/1536414.1536440 12. Gidaris, S., Singh, P., Komodakis, N.: Unsupervised representation learning by pre- dicting image rotations. In: International Conference on Learning Representations (2018) 13. Gilad-Bachrach, R., Dowlin, N., Laine, K., Lauter, K., Naehrig, M., Wernsing, J.: CryptoNets: Applying neural networks to encrypted data with high throughput and accuracy. In: Proceedings of the 33rd International Conference on Machine Learning. vol. 48. PMLR (2016) 14. Goldreich, O., Micali, S., Wigderson, A.: How to play any mental game, or a com- pleteness theorem for protocols with honest majority. In: Providing sound founda- tions for cryptography: on the work of Shafi Goldwasser and Silvio Micali. ACM (2019). https://doi.org/10.1145/3335741.3335755 15. He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016) 16. Hermann, K., Chen, T., Kornblith, S.: The origins and prevalence of texture bias in convolutional neural networks. In: Advances in Neural Information Processing Systems. vol. 33 (2020) 16L. Ranke et al. 17. Hirose, M., Imaizumi, S., Kiya, H.: Learnable image encryption without key management for privacy-preserving vision transformers. IEEE Access 13 (2025). https://doi.org/10.1109/ACCESS.2025.3635235 18. Huang, Y., Song, Z., Li, K., Arora, S.: InstaHide: Instance-hiding schemes for private distributed learning. In: Proceedings of the 37th International Conference on Machine Learning. vol. 119. PMLR (2020) 19. Jarin, I., Eshete, B.: PRICURE: Privacy-preserving collaborative inference in a multi-party setting. In: Proceedings of the 2021 ACM Workshop on Security and Privacy Analytics. Association for Computing Machinery (2021). https://doi. org/10.1145/3445970.3451156 20. Juvekar, C., Vaikuntanathan, V., Chandrakasan, A.: GAZELLE: A low-latency framework for secure neural network inference. In: Proceedings of the 27th USENIX Security Symposium. USENIX Association (2018) 21. Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: Interna- tional Conference on Learning Representations (2015) 22. Koenderink, J.J., Van Doorn, A.J.: The structure of locally orderless images. In- ternational Journal of Computer Vision 31 (1999) 23. Krizhevsky, A.: Learning multiple layers of features from tiny images. Tech. rep., University of Toronto (2009) 24. Kucur, E.N., Buyuktanir, T., Ugurelli, M., Yildiz, K.: Privacy-preserving machine learning techniques: Cryptographic approaches, challenges, and future directions. Applied Sciences 16 (2026). https://doi.org/10.3390/app16010277 25. Kurihara, K., Shiota, S., Kiya, H.: An encryption-then-compression system for the JPEG standard. In: 2015 Picture Coding Symposium (PCS). IEEE (2015) 26. Larkin, K.G.: Reflections on shannon information: In search of a natural information-entropy for images. arXiv preprint arXiv:1609.01117 (2016) 27. Le, Y., Yang, X., et al.: Tiny ImageNet visual recognition challenge. CS 231N 7 (2015) 28. LeCun, Y., Bottou, L., Bengio, Y., Haffner, P.: Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 (1998) 29. Lee, J., Lee, E., Lee, J.W., Kim, Y., Kim, Y.S., No, J.S.: Precise approximation of convolutional neural networks for homomorphically encrypted data. IEEE Access 11 (2023). https://doi.org/10.1109/ACCESS.2023.3287564 30. Lindeberg, T.: Scale-space theory: A basic tool for analyzing structures at different scales. Journal of Applied Statistics 21 (1994) 31. Liu, J., Teshome, W., Ghimire, S., Sznaier, M., Camps, O.: Solving masked jig- saw puzzles with diffusion vision transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 32. Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019) 33. Lou, Q., Feng, B., Charles Fox, G., Jiang, L.: Glyph: Fast and accurately train- ing deep neural networks on encrypted data. In: Advances in Neural Information Processing Systems. vol. 33. Curran Associates, Inc. (2020) 34. Madono, K., Tanaka, M., Onishi, M., Ogawa, T.: Block-wise scrambled image recognition using an adaptation network. arXiv preprint arXiv:2001.07761 (2020) 35. Mahalakshmi, K., Nagarajan, S.: Comprehensive review and analysis of image en- cryption techniques. IEEE Access 13 (2025). https://doi.org/10.1109/ACCESS. 2025.3578158 36. McMahan, B., Moore, E., Ramage, D., Hampson, S., Arcas, B.A.y.: Communication-efficient learning of deep networks from decentralized data. In: Beyond Classification: Task-Dependent Learnability17 Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS). vol. 54. PMLR (2017) 37. McPherson, R., Shokri, R., Shmatikov, V.: Defeating image obfuscation with deep learning. arXiv preprint arXiv:1609.00408 (2016) 38. Melis, L., Song, C., De Cristofaro, E., Shmatikov, V.: Exploiting unintended fea- ture leakage in collaborative learning. In: 2019 IEEE symposium on Security and Privacy (SP). IEEE (2019) 39. Mohassel, P., Zhang, Y.: SecureML: A system for scalable privacy-preserving ma- chine learning. In: 2017 IEEE Symposium on Security and Privacy (SP). IEEE (2017). https://doi.org/10.1109/SP.2017.12 40. Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: European Conference on Computer Vision. Springer (2016) 41. Ohrimenko, O., Schuster, F., Fournet, C., Mehta, A., Nowozin, S., Vaswani, K., Costa, M.: Oblivious multi-party machine learning on trusted processors. In: Pro- ceedings of the 25th USENIX Conference on Security Symposium. USENIX Asso- ciation (2016) 42. Phuong, T.T., et al.: Differentially private stochastic gradient descent via com- pression and memorization. Journal of Systems Architecture 135 (2023). https: //doi.org/10.1016/j.sysarc.2022.102819 43. SaberiKamarposhti, M., Ghorbani, A., Yadollahi, M.: A comprehensive survey on image encryption: Taxonomy, challenges, and future directions. Chaos, Solitons & Fractals 178 (2024). https://doi.org/10.1016/j.chaos.2023.114361 44. Schneider, T., Wang, H.C., Yalame, H.: HE-SecureNet: An efficient and usable framework for model training via homomorphic encryption. In: Proceedings of the 24th Workshop on Privacy in the Electronic Society (2025). https://doi.org/10. 1145/3733802.3764063 45. Shafran, A., Segev, G., Peleg, S., Hoshen, Y.: Crypto-oriented neural architecture design. In: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE (2021). https://doi.org/10.1109/ICASSP39728. 2021.9413592 46. Shokri, R., Stronati, M., Song, C., Shmatikov, V.: Membership inference attacks against machine learning models. In: 2017 IEEE symposium on Security and Pri- vacy (SP). IEEE (2017) 47. Sirichotedumrong, W., Kinoshita, Y., Kiya, H.: Pixel-based image encryption with- out key management for privacy-preserving deep neural networks. IEEE Access 7 (2019). https://doi.org/10.1109/ACCESS.2019.2959017 48. Tanaka, M.: Learnable image encryption. In: 2018 IEEE International Conference on Consumer Electronics-Taiwan (ICCE-TW) (2018). https://doi.org/10.1109/ ICCE-China.2018.8448772 49. Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Process- ing 13(4) (2004) 50. Wen, J., Zhang, Z., Lan, Y., Cui, Z., Cai, J., Zhang, W.: A survey on federated learning: Challenges and applications. International Journal of Machine Learning and Cybernetics 14(2) (2023). https://doi.org/10.1007/s13042-022-01647-y 51. Zamir, A.R., Sax, A., Shen, W., Guibas, L.J., Malik, J., Savarese, S.: Taskonomy: Disentangling task transfer learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018) 52. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018) 18L. Ranke et al. 53. Zhang, Y., Zheng, M., Shang, Y., Chen, X., Lou, Q.: HEPrune: Fast private training of deep neural networks with encrypted data pruning. In: Advances in Neural Information Processing Systems. vol. 37. Curran Associates, Inc. (2024). https: //doi.org/10.52202/079017-1616 54. Zhou, I., Tofigh, F., Piccardi, M., Abolhasan, M., Franklin, D., Lipman, J.: Secure multi-party computation for machine learning: A survey. IEEE Access 12 (2024). https://doi.org/10.1109/ACCESS.2024.3388992 55. Zhu, L., Liu, Z., Han, S.: Deep leakage from gradients. In: Advances in Neural Information Processing Systems. vol. 32 (2019) 56. Zuo, X., Luopan, Y., Han, R., Zhang, Q., Liu, C.H., Wang, G., Chen, L.Y.: FedViT: Federated continual learning of vision transformer at the edge. Future Generation Computer Systems 154 (2024). https://doi.org/10.1016/j.future.2023.11. 038 Beyond Classification: Task-Dependent Learnability19 Table 4: Hyperparameters for Animal Species Classification. Hyperparameter Value DatasetAnimal Species Classification dataset [1] Number of classes15 Data augmentation None Model architecture ResNet-34 [15] Initial weightsRandom (trained from scratch) Input resolution3Ă 416Ă 416 Loss functionCross-entropy OptimizerAdam [21] Learning rate10 â4 Batch size64 Max. epochs100 Early stoppingpatience 5 on validation accuracy Learning-rate schedule Multistep at epochs [20, 50] with Îł = 0.1 Checkpoint selection Best validation accuracy Evaluation metricTop-1 classification accuracy Random seed73 FrameworkPyTorch 2.11.0+cu126 Gaussian blur Ï 1, 2, 4, 6, 8, 10, 14, 16, 20, 30, 40, 50, 60, 70 Block size B1, 2, 8, 16, 26, 32, 104, 208, 416 Hardware1Ă NVIDIA GeForce RTX 4090 Supplementary Material Hyperparameter Settings Tables 4 to 6 list the hyperparameters used for classification, relative angle pre- diction, and jigsaw puzzle solving, respectively. Unless stated otherwise, PyTorch defaults are used. Block sizes B are chosen as divisors of the respective input size, so that blocks tile each image exactly. All experiments use a single random seed per configuration. Obfuscation-Utility Trade-off Figures 7 to 16 show the trade-off between task utility and the remaining image- domain obfuscation metrics complementing the LPIPS trade-off in the main pa- per. In each figure, colour indicates the transformation family and marker shape the individual transformation; each point corresponds to one transformationâ scale configuration. Task utility is normalised per task as u = performanceâ chance baselineâ chance , where baseline denotes the plain-baseline performance and chance the chance level (1/N for classification, 90 ⊠mean angular error for angle prediction, 1/k 2 20L. Ranke et al. Table 5: Hyperparameters for Relative Angle Prediction. Hyperparameter Value DatasetAnimal Species Classification dataset [1] Angle samplingα 1 ,α 2 âŒU[â180 ⊠, 180 ⊠] Data augmentation None Model architecture CNN (2Ă conv, max-pool, 4Ă linear) Initial weightsRandom (trained from scratch) Input resolution3Ă 304Ă 304 Loss functionCosine similarity OptimizerAdam [21] Learning rate10 â4 Batch size32 Max. epochs50 Early stoppingpatience 5 on validation loss Learning-rate schedule Reduce on plateau, factor 0.5, patience 3 Checkpoint selection Lowest validation loss Evaluation metricMean angular error in degrees Random seed42 FrameworkPyTorch 2.11.0+cu126 Gaussian blur Ï 1, 2, 4, 6, 8, 10, 14, 16, 20, 30, 40, 50, 60, 70 Block size B4, 8, 16, 38, 76, 152 Hardware1Ă NVIDIA GeForce RTX 4090 for jigsaw solving). For angle prediction, whose raw metric decreases with qual- ity, the error is inverted before normalisation so that higher u uniformly indicates higher utility. A value of u = 1 corresponds to plaintext-level performance and u = 0 to chance; values above 1 occur where transformed-domain performance exceeds the identity baseline. For keyed transformations, we attribute these val- ues to deterministic key-induced shortcut cues. However, values above 1 also occur for non-keyed transformations (notably angle prediction under mild Gaus- sian blur and small-block LOIs). In these cases the effect is explained by the transformation mildly aiding the specific task and model (i.e., low-pass filter- ing easing orientation estimation for a small CNN), rather than by shortcut memorisation. The obfuscation metrics are minâmax normalised per metric across all config- urations, with direction-aware inversion so that higher values uniformly indicate stronger obfuscation. Differential and key-sensitivity metrics (NPCR, UACI,|K|) are omitted from these plots, as they characterise a transformationâs sensitivity rather than a per-configuration obfuscation level. Beyond Classification: Task-Dependent Learnability21 Table 6: Hyperparameters for Jigsaw Puzzle Solving. Hyperparameter Value DatasetAnimal Species Classification dataset [1] Puzzle sizes k4, 6, 8 PermutationsSampled uniformly at random during training Data augmentation None Model architecture JPDVT [31] Diffusion timesteps T 1000 Model patch size16 Initial weightsRandom (trained from scratch) Input resolution3Ă 384Ă 384 Loss functionCosine similarity OptimizerAdamW [32] Learning rate10 â4 Batch size64 Max. epochs100 Early stoppingpatience 5 on validation loss Learning-rate schedule Multistep at epochs [20, 50, 75] with Îł = 0.1 Checkpoint selection Lowest validation loss Evaluation metricPiece accuracy Random seed42 FrameworkPyTorch 2.11.0+cu126 Block size BCoupled to puzzle size k Hardware2Ă NVIDIA GeForce RTX 4090 0.00.20.40.60.81.0 Normalised MSEâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.81.0 Normalised MSEâ Angle Prediction 0.00.20.40.60.81.0 Normalised MSEâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 7: Obfuscationâutility trade-off: normalised task utility against normalised MSE. 22L. Ranke et al. 0.00.20.40.60.8 Normalised PSNRâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.8 Normalised PSNRâ Angle Prediction 0.00.20.40.60.8 Normalised PSNRâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 8: Obfuscationâutility trade-off: normalised task utility against normalised PSNR. 0.00.20.40.60.81.0 Normalised SSIMâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.81.0 Normalised SSIMâ Angle Prediction 0.00.20.40.60.81.0 Normalised SSIMâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 9: Obfuscationâutility trade-off: normalised task utility against normalised SSIM. 0.20.40.60.81.0 Normalised MIâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.20.40.60.81.0 Normalised MIâ Angle Prediction 0.20.40.60.81.0 Normalised MIâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 10: Obfuscationâutility trade-off: normalised task utility against normalised MI. Beyond Classification: Task-Dependent Learnability23 0.00.20.40.60.81.0 Normalised Chi-2â 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.81.0 Normalised Chi-2â Angle Prediction 0.00.20.40.60.81.0 Normalised Chi-2â Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 11: Obfuscationâutility trade-off: normalised task utility against normalised Ï 2 . 0.00.20.40.60.8 Normalised KL-Divâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.8 Normalised KL-Divâ Angle Prediction 0.00.20.40.60.8 Normalised KL-Divâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 12: Obfuscationâutility trade-off: normalised task utility against normalised D KL . 0.00.20.40.60.8 Normalised IEâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.8 Normalised IEâ Angle Prediction 0.00.20.40.60.8 Normalised IEâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 13: Obfuscationâutility trade-off: normalised task utility against normalised IE. 24L. Ranke et al. 0.20.40.60.8 Normalised IE2Dâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.20.40.60.8 Normalised IE2Dâ Angle Prediction 0.20.40.60.8 Normalised IE2Dâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 14: Obfuscationâutility trade-off: normalised task utility against normalised 2D- IE. 0.00.20.40.60.81.0 Normalised Câ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.20.40.60.81.0 Normalised Câ Angle Prediction 0.00.20.40.60.81.0 Normalised Câ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 15: Obfuscationâutility trade-off: normalised task utility against normalised C. 0.00.10.20.30.40.5 Normalised SFMâ 0.0 0.2 0.4 0.6 0.8 1.0 Normalised Task Utility â Classification 0.00.10.20.30.40.5 Normalised SFMâ Angle Prediction 0.00.10.20.30.40.5 Normalised SFMâ Jigsaw Puzzle Solving Gaussian BlurringLocally Orderless Images (LOIs)Block-Based transformationsLearnable Image Encryption Schemes Gaussian Blurring Locally Orderless Images (LOIs) Block Shuffling Block Rotation + Flipping Pixel Shuffling Colour Shuffling NP-Transformation Tanaka E-Tanaka EtC Fig. 16: Obfuscationâutility trade-off: normalised task utility against normalised SFM.