Paper deep dive
Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils
Karel Becerra, Boris Mederos, Dean Snow, RamĂłn A. Mollineda
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizability, and subjective feature engineering. This study presents an uncertainty-aware deep learning framework for sex attribution in prehistoric hand stencils that explicitly models, propagates, and aggregates uncertainty throughout the analytical pipeline. The methodology combines dual image processing, dual contour extraction, structured silhouette augmentation, model architectural diversity, and ensemble-based decision aggregation. The pipeline generates twelve plausible silhouette realizations per stencil to capture boundary uncertainties, which are processed by two ensembles of ten deep neural networks each (EfficientNet-B3 and MobileViT-S) trained on 14,036 contemporary hand samples. Furthermore, a triangulated validation scheme integrates ensemble predictions with unsupervised 2D latent-space manifold mapping (UMAP + k-NN) and explainable AI spatial attributions (LayerCAM) to ensure anatomical consistency. On contemporary data, ensemble models achieve strong classification performance, with accuracies exceeding 88% in older age groups. When applied to prehistoric stencils, the framework produces both sex predictions and confidence measures of internal agreement, enabling the distinction between morphologically stable and ambiguous cases. Convergence across ensemble predictions, latent-space structure, and interpretability analyses shows that uncertainty can become a measurable component of archaeological inference, enabling robust and reproducible decoding of ancient rock art.
Tags
Links
- Source: https://arxiv.org/abs/2608.14539v1
- Canonical: https://arxiv.org/abs/2608.14539v1
Trouble viewing inline? Open PDF directly â
Full Text
107,405 characters extracted from source content.
Expand or collapse full text
Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils Karel Becerra a , Boris Mederos b , Dean Snow c , RamĂłn A. Mollineda d,1 a Data Science and AI Division, Azyri, Miami, Florida, USA b Departamento de FĂsica y MatemĂĄticas, Instituto de IngenierĂa y TecnologĂa, Universidad AutĂłnoma de Ciudad JuĂĄrez, Ciudad JuĂĄrez, Mexico c Professor Emeritus of Anthropology at Pennsylvania State University, Pennsylvania, USA d Institute of New Imaging Technologies, Universitat Jaume I, CastellĂł de la Plana, Spain Abstract Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a chal- lenging problem due to the absence of ground truth, population differences between contemporary and pre- historic groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizability, and subjective feature engineering. This study presents an uncertainty-aware deep learning framework for sex attribution in prehis- toric hand stencils that explicitly models, propagates, and aggregates uncertainty throughout the analytical pipeline. The methodology combines dual image processing, dual contour extraction, structured silhouette augmentation, model architectural diversity, and ensemble-based decision aggregation. The pipeline generates twelve plausible silhouette realizations per stencil to capture boundary uncertainties, which are processed by two ensembles of ten deep neural networks each (EfficientNet-B3 and MobileViT-S) trained on 14,036 con- temporary hand samples. Furthermore, a triangulated validation scheme integrates ensemble predictions with unsupervised 2D latent-space manifold mapping (UMAP + k-N) and explainable AI spatial attributions (LayerCAM) to ensure anatomical consistency. On contemporary data, ensemble models achieve strong classi- fication performance, with accuracies exceeding 88% in older age groups. When applied to prehistoric stencils, the framework produces both sex predictions and confidence measures of internal agreement, enabling the distinction between morphologically stable and ambiguous cases. The convergence observed across ensemble predictions, latent-space structure, and interpretability analyses suggests that uncertainty can be transformed into an explicit and measurable component of archaeological inference, providing a robust, transparent, and reproducible approach to decoding ancient rock art. Keywords: Prehistoric hand stencils, hand silhouettes, sex classification, deep learning, uncertainty management, explainable AI 1. Introduction Echoes of ancient hands in cave art can be found worldwide, with hand stencilsânegative handprints created by spraying pigment around the handâbeing the most common form (Snow 2006; Lundborg 2014). â Corresponding author: mollined@uji.es Email addresses: karel.becerra@azyri.com (Karel Becerra), boris.mederos@uacj.mx (Boris Mederos), drs17@psu.edu (Dean Snow), mollined@uji.es (RamĂłn A. Mollineda) They provide some of the earliest evidence of sym- bolic behavior and artistic expression, that some re- searchers have related to personal signatures, ritual marks, or protective symbols to ward off danger (Janssens 1957; Permana, Pojoh, et al. 2017). As suggested by (FernĂĄndez-Navarro, Fidalgo-Casares, et al. 2025), widespread participation in rock art may indicate group cohesion, as well as an intention to transmit cultural values across generations. Recent research has leveraged advanced, non-des- tructive techniques for analyzing and measuring ar- arXiv:2608.14539v1 [cs.CV] 14 Aug 2026 chaeological hands (Snow 2013). To mitigate the in- accuracies and optical distortions associated with 2D photographs and in-situ measurements, photogram- metric 3D models of archaeological hand stencils have been used to generate 2D orthophotos, on which land- marks and semilandmarks were subsequently identi- fied for morphometric analysis (FernĂĄndez-Navarro, CamarĂłs, et al. 2022; FernĂĄndez-Navarro, Fidalgo- Casares, et al. 2025). While most studies have focused on estimating the sex of the artists behind Palae- olithic cave art (Wang et al. 2010; Snow 2013), re- cent groundbreaking work has expanded this focus to include age prediction (FernĂĄndez-Navarro, Fidalgo- Casares, et al. 2025). Strong evidence suggests that both sexes, including children and subadults, were ac- tively involved in these artistic practices, challenging the long-standing belief that they were primarily cre- ated by adult males. The biological evidence for sexual dimorphism in human hands is extensive and well-established, en- compassing differences in hand size, shape, bone struc- ture, dermatoglyphic patterns, and morphometric traits (Snow 2013; Karakostis et al. 2015; Bindurani et al. 2016; Yune et al. 2019; PolcerovĂĄ et al. 2023). Stud- ies consistently show that males have larger and more robust handsâwith broader palms and thicker finger bonesâthan females, with these differences emerging in early childhood (by age 3) and becoming more pro- nounced during development (Karakostis et al. 2015; Bindurani et al. 2016). Skeletal analyses reveal that the proximal phalanges are significantly larger in men, with the greatest dimorphism observed in the thumb (Karakostis et al. 2015). Hand proportions, such as the hand index (breadth-to-length ratio), are also re- liably higher in males and have proven useful for sex estimation across populations, albeit with moderate accuracy (Bindurani et al. 2016; Suleiman et al. 2024). Dermatoglyphic studies demonstrate consistent sex- based differences in ridge counts, particularly on the radial sides of the thumb and index finger, which may reflect prenatal developmental influences (Pol- cerovĂĄ et al. 2023). Geometric morphometric analy- ses and recent advances in machine learning further support the presence of subtle but detectable shape and proportion differences that can be leveraged for sex classification, even in cases where such differences are imperceptible to human observers (KrĂĄlĂk et al. 2014; Yune et al. 2019; FernĂĄndez-Navarro, Garate- Maidagan, et al. 2024). Collectively, these findings provide robust multidisciplinary support for the pres- ence of sexual dimorphism in human hands. The analysis of prehistoric handprints to deter- mine the sex of Paleolithic cave artists has attracted growing interest in archaeology, offering insights into gender roles. However, while some studies have shown promising results, critical methodological and popula- tion-based challenges remain. Traditional morphome- tric methods for sexing prehistoric handprints, based on size, angles, and digit ratios, are limited by signif- icant hand size overlap between sexes and ages, poor generalizability across populations, and low accuracy (often around 60â65%) (Snow 2006; Snow 2013; Jowa- heer and Agnihotri 2011; Galeta et al. 2014). Man- ning index (2D:4D ratio) also show substantial within- and between-population variation (McIntyre 2006; J. T. Manning et al. 2007), reducing reliability unless pop- ulation parameters are known (Nelson, J. Manning, et al. 2006). Another challenge, noted by (Gunn 2006), is that stencil images created from a same hand can vary in size measurements by up to 5 m. Unsur- prisingly, earlier studies aimed at estimating the sex of prehistoric hand stencils have yielded notably con- flicting results (Nelson, Hall, et al. 2017). Traditional approaches rely on manually extracted handcrafted features, such as geometric measurements and digit ratios, which are heavily dependent on sub- jective domain expertise, often fail to generalize across datasets, capture only predefined low-level traits, and scale poorly. Some of the aforementioned limitations were addressed by an early automated approach (Wang et al. 2010) combining image processing techniques with a machine learning algorithm (Support Vector Machines) to extract a number of hand features. A more recent work (FernĂĄndez-Navarro, Fidalgo-Casares, et al. 2025) also automates the location of landmarks, the calculation of engineered features, and decision- making using discriminant analysis. Nonetheless, sub- jectivity persists in the selection of human-designed features, while more abstract morphological patterns remains unexplored. Recent work (Mollineda et al. 2025; FernĂĄndez- Navarro, GonzĂĄlez-Marfil, et al. 2025) introduces deep learning techniques for sex prediction from prehistoric handprints. In the absence of annotated ancient data, deep neural networks (DNNs) are trained on contem- porary hand silhouettes and then transferred to Pale- olithic samples. By learning hierarchical representa- tions directly from images, these models avoid manual feature engineering and capture complex, data-driven patterns that may enhance accuracy and robustness 2 across domains. The underlying hypothesis is that hand morphology encodes high-level, sex-informative features that remain sufficiently invariant across age and long-term evolutionary change. While DNNs offer significant promise, the anal- ysis of prehistoric hand stencils remains affected by a range of archaeological and technical uncertainties. Pigment erosion, incomplete preservation, surface ir- regularities, uneven illumination, and poorly defined stencil boundaries can hinder both automated seg- mentation and manual contour delineation, with the latter being particularly sensitive to observer judg- ment and image noise. These challenges are illus- trated in Figure 1 using examples from El Castillo Cave (Snow 2006). Representations based on photogrammetric 3D mo- dels and 2D orthophotos can reduce viewpoint- and scale-related distortions, but they do not eliminate these inherent sources of uncertainty. Moreover, un- certainty may arise during 3D reconstruction and or- thophoto generation: reconstruction errors can prop- agate to the final representation, while orthoprojec- tion depends on the reconstructed surface geometry and the choice and orientation of the projection plane. Orthoprojection does not generally preserve intrinsic distances on curved or irregular cave surfaces, and the resulting distortions may vary across the same hand stencil according to local surface orientation. Although such distortions are small on nearly pla- nar surfaces, they can introduce relevant geometric uncertainty into morphometric measurements on ir- regular ones. This issue is particularly relevant to sex-attribution methods, for which small variations in finger dimensions, hand proportions, or landmark locations can affect the extracted descriptors and sub- sequent classification. Finally, the absence of biolog- ical ground truth for Palaeolithic individuals repre- sents an intrinsic source of uncertainty that cannot be resolved through geometric preprocessing alone. Overall, these limitations, ranging from the ab- sence of biological ground truth to surface degrada- tion, underscore the need for approaches that explic- itly account for uncertainty to support robust inter- pretations. Accordingly, methods that generate diver- sity through plausible geometric realizations of mul- tiple source representations, identify patterns across multiple hypotheses, quantify uncertainty, and pro- vide interpretable evidence for their decisions appear well suited to this task. This study proposes a multi-layered, uncertainty- aware inference methodology that models, propagates, and aggregates variation hierarchically across a deep learning pipeline. Building upon the classifier en- semble and spatial attribution-based interpretability method proposed in (Mollineda et al. 2025), the present work substantially expands the previous framework in both methodological scope and uncertainty manage- ment. The extended framework systematically incor- porates variability across all stages of the analysis: source-level variation through dual image processing, annotation-level variation through dual contour ex- traction, morphological variation through structured silhouette generation, uncertainty quantification at multiple stages, an additional interpretability tool ba- sed on latent-space analysis and manifold mapping, and a triangulated validation scheme designed to iden- tify convergent evidence across predictions, interpre- tability findings, and uncertainty measures. By prop- agating controlled variability throughout the pipeline, the proposed methodology goes beyond the previous ensemble-based approach, deriving final classifications from the convergence of multiple, highly structured yet individually imperfect observations. Nine hand-stencil images are considered as case studies to demonstrate the practical applicability of the proposed methodology and its value as a tool to support archaeological interpretation. The cases are analyzed jointly through their sex predictions, as- sociated confidence measures, latent-space organiza- tion, and spatial attribution patterns, providing com- plementary evidence for interpreting the model out- puts. Importantly, the nine cases encompass a broad range of outcomes, from highly consistent predictions supported by multiple sources of evidence to cases characterized by substantial uncertainty or disagree- ment. This diversity enables a detailed examination of different analytical scenarios and illustrates how the methodology can support the interpretation of both well-supported and ambiguous cases. These nine images are not intended to provide a representative sample of the morphological variability of Upper Palae- olithic hand stencils; rather, their sole purpose is to demonstrate the practical value and analytical capa- bilities of the proposed methodology. The main contributions of this work are summa- rized below: âą A novel uncertainty-informed computational fra- mework for sex attribution in prehistoric hand stencils, which explicitly models, propagates, and 3 (a) img_26(b) img_28(c) img_31(d) img_38 Figure 1: Four prehistoric hand images from El Castillo cave (Snow 2006). aggregates uncertainty across all stages of the analytical pipeline. âą A structured silhouette representation strategy that generates multiple plausible contour real- izations through dual annotation, morphologi- cal perturbations, and multi-channel composi- tions, capturing annotation and morphological uncertainty. âą A triangulated validation scheme integrating en- semble predictions, latent-space analysis (UMAP + k-N), and aggregated interpretability maps (LayerCAM), providing convergent evidence and enabling the diagnosis of epistemic stability. The remainder of this manuscript is organized as follows. Section 2 reviews previous research on sex es- timation from hand morphology and prehistoric hand stencils, with particular attention to the evolution from traditional morphometric approaches to contem- porary machine learning and deep learning method- ologies. Section 3 introduces the contemporary and prehistoric datasets and the proposed inference frame- work, detailing the silhouette extraction and augmen- tation procedures, the ensemble-based classification strategy, and the post-hoc analysis techniques used for validation and interpretation. Section 4 presents the experimental protocol adopted to benchmark the deep learning models on contemporary hand silhou- ettes and to deploy the framework on prehistoric hand stencils, including results obtained on both data do- mains. Section 5 discusses the significance of the find- ings, the strengths and limitations of the proposed approach, and practical considerations for archaeo- logical interpretation. Finally, Section 6 summarizes the main conclusions and outlines potential directions for future research. 2. Related work The analysis of hand stencils in rock art for sex es- timation has evolved substantially, progressing from traditional anthropological measurements to more ad- vanced computational approaches. In contrast to most reviews in this field, this section adopts a distinct perspective by focusing on four key dimensions: the nature of the hand data (Paleolithic stencils, simu- lated stencils, or contemporary), the feature extrac- tion approach (manually handcrafted features, auto- matically extracted handcrafted features, or automat- ically learned features), the decision rule and the ex- perimental protocol. Structuring the reviewed studies along these dimensions enables a clearer categoriza- tion and helps reveal common trends in hand sten- cil analysis. Table 1 summarizes the selected studies, chosen for their experimental designs combining hand feature extraction and automatic sex classification. One of the earliest and foundational studies was conducted by (Snow 2006), who introduced a two- stage analytical approach based on predictive discrim- inant analysis. Using a modern reference sample of 111 adult individuals of European descent, the first stage applied Fisherâs linear discriminant functions to overall hand length and finger lengths (D2âD5) to separate adult males from individuals with smaller hands, a group that in archaeological contexts may include both adult females and subadults. Because absolute size alone cannot reliably distinguish women from young boys, the second stage applied the same discriminant framework to digit ratios (D2:D4 and D2:D5) to separate adult females from subadult males. The predictive equations, learned from the contempo- rary training set, were then applied to a small sample of six Paleolithic stencils, yielding one of the first de- mographic interpretations of cave artists. Expanding on this, (Wang et al. 2010) introduced an automated machine learning pipeline using con- 4 ReferenceTraining dataTest data Features extraction Decision ruleProtocol (FernĂĄndez- Navarro, GonzĂĄlez-Marfil, et al. 2025) ContemporaryPaleolithic Automatically learned DNNCross-domain (FernĂĄndez- Navarro, Fidalgo-Casares, et al. 2025) Contemporary Simulated, Paleolithic Automatically handcrafted LDACross-domain (Mollineda et al. 2025) ContemporaryPaleolithic Automatically learned DNNCross-domain (Rabazo- RodrĂguez et al. 2017) ContemporaryPaleolithic Manually handcrafted LDACross-domain (Nelson, Hall, et al. 2017) ContemporaryContemporary Automatically handcrafted Fisherâs LDALeave-one-out (Permana, Arifin, et al. 2015). ContemporaryPaleolithic Manually handcrafted Nearest class mean Cross-domain (Galeta et al. 2014) ContemporaryContemporary Manually handcrafted Fisherâs LDA Leave-one-out, Cross-domain (Wang et al. 2010)ContemporaryPaleolithic Automatically handcrafted SVMCross-domain (Snow 2006)ContemporaryPaleolithic Manually handcrafted Fisherâs LDA Cross-domain evaluation Table 1: Studies on sex attribution in prehistoric hand stencils, categorized by the type of hand data used (simulated stencils, Paleolithic stencils, or contemporary hand images). Abbreviations: LDA = Linear Discriminant Analysis; DNN = Deep Neural Network; SVM = Support Vector Machine. temporary hand images. Their methodology involved transforming the images into the HSV color space and applying K-means clustering to isolate hand con- tours. By identifying key points of interest, specif- ically fingertips and valleys, they extracted geomet- ric features including finger lengths, finger widths, palm dimensions, and the well-known D2:D4 ratio. To reduce the effect of scale differences across images, these measurements were normalized by the length of the middle finger. The middle finger was omitted af- ter normalization due to its constant value, and the thumb was excluded due to unreliable measurements. Finally, a Support Vector Machine (SVM) classifier, trained on modern data represented as 33-dimensional feature vectors, was applied to prehistoric handprints from French caves. This study constitutes one of the earliest applications of machine learning to sex esti- mation in archaeological handprints. The issue of population variability and method- ological reliability was directly addressed by (Galeta et al. 2014) using an independent modern French data- set of 100 right-hand prints from the University of Bordeaux. In a first analysis, linear discriminant func- tions based on hand and finger lengths (D2âD5) along- side the D2:D4 and D2:D5 ratios were determined directly on the French dataset using a leave-one-out cross-validation scheme. Second, to evaluate general- izability across populations, they applied pre-existing discriminant functions trained on a U.S. sample by (Snow 2006) directly to the French dataset. The model trained on U.S. data exhibited poor generalization; differences in average hand size between the U.S. and French populations led to heavily biased classifica- tions. Consequently, the authors concluded that mod- ern morphometric references are unreliable for esti- mating the sex of prehistoric artists. Concurrently, other researchers investigated the application of these methods across diverse geograph- ical contexts. (Permana, Arifin, et al. 2015) employed the manually extracted D2:D4 digit ratio on both con- temporary individualsâ181 adults and juveniles from three villagesâand hand stencils from Pettakere Cave in Indonesia. Their approach used a simple decision rule based on comparison with class means, apply- ing a rule derived from contemporary data directly to archaeological samples. In a similar vein, (Rabazo- RodrĂguez et al. 2017) examined sexual dimorphism in hand stencils from El Castillo Cave (Spain). Their 5 study relied on a contemporary dataset of 77 hand stencils, from which three key features were extracted: overall hand length, index finger length, and ring fin- ger length. A discriminant analysis model was trained on this modern dataset and directly applied to classify 21 Paleolithic stencils. A significant shift towards geometric morphomet- rics is evident in the work of (Nelson, Hall, et al. 2017). This study relied on a contemporary dataset of 132 hand stencils, capturing hand morphology through 19 two-dimensional landmarks. The authors applied Generalized Procrustes Analysis for shape alignment, followed by Principal Component Analysis for dimen- sionality reduction, and Fisherâs Linear Discriminant Analysis for classification. Their experimental proto- col relied on cross-validation to estimate performance, due to the absence of an independent test set. A comprehensive framework proposed by (FernĂĄndez- Navarro, Fidalgo-Casares, et al. 2025) leverages a di- verse, multi-source dataset to enhance the classifica- tion of ancient hand stencils. This dataset includes contemporary scans from 546 individuals (categorized by sex and age), 53 experimental stencils created in a limestone quarry using Paleolithic techniques (ochre and plant tubes), and 124 authentic Paleolithic sten- cils from eight Iberian caves. The methodology uti- lizes 32 anatomical landmarks to capture hand mor- phology. To isolate shape from external variables, Procrustes coordinates were computed to remove the effects of size, position, and orientation. Hand size was quantified via centroid size, while pairwise shape differences and variability were assessed using Pro- crustes distances and Principal Component Analysis, respectively. By applying discriminant analysis to these handcrafted features, the framework provides a robust and generalizable model for characterizing the authors of prehistoric rock art. An alternative approach was explored by (FernĂĄndez- Navarro, GonzĂĄlez-Marfil, et al. 2025), who developed a supervised deep learning pipeline utilizing hand sil- houettes. Their multi-source dataset comprised 247 contemporary hand scans, 11,076 images from the public 11K Hands database (representing 190 unique subjects with a female-to-male ratio of approximately 1.8:1), 61 experimental stencils produced on limestone by modern adults, and 45 Paleolithic stencils. Due to the complexity of the rock backgrounds, the exper- imental and Paleolithic stencils were manually seg- mented using Procreate 1 ; conversely, silhouettes from the contemporary data were extracted automatically. To achieve this automation, a ResU-Net model (Di- akogiannis et al. 2020) was initially trained on the 247 scanned hands and then applied to the 11K Hands dataset. To refine the model, 120 high-quality pre- dictions from the 11K Hands collection were manu- ally selected, added to the original training set, and used to retrain the architecture from scratch. After filtering for motion blur or shape-altering artifacts, 5,556 high-quality silhouettes were successfully ex- tracted from the 11K Hands database. These were combined with the 247 scanned silhouettes and 61 experimental stencils to form the training and eval- uation sets. The authors assessed several pretrained convolutional neural networks across ten independent train-validation-test splits. EfficientNetV2-S emerged as the most stable model and was subsequently de- ployed to predict the biological sex of the 45 archaeo- logical stencils from Spanish caves. Beyond improving generalization, the combination of scanned and ex- perimental data served to narrow the domain gap be- tween modern and ancient samples; nonetheless, the modelâs outputs remain susceptible to variations in manual segmentation and the documented bias to- ward classifying subadult hands as female. Existing literature has relied largely on handcrafted features, ranging from linear measurements and digit ratios to landmark-based geometric morphometrics. Recent studies have begun to move beyond this para- digm by using deep learning models trained on con- temporary hand data and evaluated on archaeolog- ical samples. In particular, (Mollineda et al. 2025; FernĂĄndez-Navarro, GonzĂĄlez-Marfil, et al. 2025) ex- plore data-driven pipelines where sex-discriminant fea- tures are automatically learned from hand silhouettes, moving away from hand-engineered linear measure- ments. This shift represents a key step toward more objective and flexible frameworks for the analysis of prehistoric hand stencils, while highlighting the need for robust cross-domain validation when models trained on modern data are applied to the Paleolithic field. Despite these advancements, significant challenges persist. A fundamental limitation is the lack of defini- tive ground truth, as the biological sex of prehistoric stencils must be inferred from contemporary reference models rather than directly observed. Furthermore, 1 Procreate is a professional digital illustration app designed for iPad (https://procreate.com). 6 while the classification pipeline utilizes hand silhou- ettes, those derived from Paleolithic art are inher- ently products of manual delineation. This process is susceptible to inter-observer variability and sub- jective bias, further compounded by factors such as pigment diffusion, rock texture, and varying states of preservation, Consequently, subtle inaccuracies in the traced contours may propagate through the model, amplifying the uncertainty of the final predictions. An additional challenge stems from the cross-domain setting, where models trained on modern or experi- mental data are deployed on archaeological samples. Moreover, discrepancies in pose, acquisition condi- tions, and population-specific characteristics can hin- der the generalization of learned features. Collec- tively, these challenges underscore the necessity for ro- bust frameworks that account for segmentation vari- ability, dataset representativeness, and domain adap- tation, while moving beyond simple accuracy to pro- vide reliable measures of predictive uncertainty. 3. Materials and methods 3.1. Overview The proposed study aims to infer biological sex from prehistoric cave hand stencils through a silhou- ette-based computer vision pipeline, which is orga- nized into three main phases: (i) silhouette extraction and augmentation, (i) ensemble-based classification of silhouettes, and (i) post-hoc analysis. Figure 2 illustrates this computer vision pipeline. 3.2. Datasets Experiments were conducted using two data sources: a large set of 14,036 annotated X-ray images from contemporary subjects, used to train DNNs, and a specialized collection of nine prehistoric hand-stencil images, used as case studies to illustrate the practi- cal value of the proposed methodology. Samples from both sources were represented as binary hand silhou- ettes, enabling their integration within a unified ex- perimental framework. This standardized representa- tion supported the development of a DNN-based sex classifier trained on contemporary annotated silhou- ettes, which was then applied to the prehistoric data to generate sex-related hypotheses. 3.2.1. Contemporary data This study relies on the RSNA Bone Age Chal- lenge dataset (Larson et al. 2018; Halabi et al. 2019), Table 2: Summary of RSNA Bone Age Challenge dataset. females males total training set5,778 6,833 12,611 validation set 652773 1,425 test set100100200 which comprises 14,036 X-ray images of the left hand from pediatric patients aged between 1 month and 19 years. The data were collected from two insti- tutions: Lucile Packard Childrenâs Hospital at Stan- ford University (2,983 images) and Childrenâs Hos- pital Colorado (11,053 images) (Larson et al. 2018). The dataset was randomly partitioned into 12,611 samples (approximately 90%) for training and 1,425 samples (approximately 10%) for validation. In addi- tion, a separate test set of 200 hand radiographs (100 male and 100 female) was constructed from the pic- ture archiving and communication systems at Stan- ford, ensuring no overlap with the training and val- idation studies. Each radiograph is annotated with skeletal (bone) age and patient sex based on the as- sociated clinical radiology report. Silhouettes were obtained using the Segment Anything Model (SAM) (Kirillov et al. 2023) with a zero-prompt segmentation approach (Mollineda et al. 2025). Table 2 summarizes the sample counts and class distribution, while Fig. 3 presents representative test cases across both sexes and age groups. Figure 4 illustrates the class-conditional bone age distributions for the training, validation, and test sets. While the male and female classes exhibit unimodal and bimodal patterns, respectively, the highest con- centration of samples lies between 10 and 16 years of age for both sexes. This age range corresponds to the attainment of skeletal and structural maturity of the hand, where sexually dimorphic traits become more stable and pronounced. The inclusion of a wide range of ages and morphological variability is expected to enhance the robustness and generalization capacity of the learned models. 3.2.2. Prehistoric data A total of nine prehistoric hand stencils from two distinct sources are used as independent case stud- ies. The first subset comprises four samples from El Castillo cave (Spain), illustrated in Figure 1, which were selected for their superior contrast and well-defined edge morphology. These characteristics facilitate highly 7 Figure 2: Uncertainty-aware inference system: quantifying sexual dimorphism in prehistoric cave hand stencils. Generated using Gemini Notebook from authorsâ prompts and methodology. reliable manual contouring during the silhouette ex- traction process. The second subset consists of five samples previously analyzed by (Mollineda et al. 2025), originating from at least four well-known sites (Cos- quer, Chauvet, Pech Merle, and El Castillo). These images, shown in Figure 5, were chosen because earlier studies proposed initial hypotheses regarding their sex attribution. Re-examining these cases under the sub- stantially expanded uncertainty-aware framework in- troduced in this study is expected to provide more ro- bust and reliable assessments, while offering a broader analysis of the evidence supporting each prediction. In addition, the use of nine hand stencils enables a detailed, case-by-case examination of the framework and illustrates how its different analytical components can support archaeological interpretation under con- ditions of substantial uncertainty. 3.3. Methodology 3.3.1. Overview The proposed methodology adopts a multi-layered, uncertainty-aware inference design to address key chal- lenges in the analysis of Paleolithic hand stencils, par- ticularly the lack of ground-truth information on the artistsâ biological sex. Rather than attempting to eliminate uncertainty, the framework models, prop- agates, diversifies, and hierarchically aggregates un- certainty across the entire analytical pipeline. The uncertainty management strategies are summarized below across the three operational phases: Phase 1 Silhouette extraction and augmentation Source-level uncertainty At the archaeolog- ical and image-acquisition level, uncertainty arises from the absence of biological ground truth, surface degradation, pigment ero- sion, illumination variability, and histori- cal interpretative disagreement. To mit- igate photometric ambiguity, each stencil is processed along two parallel paths: the original image and a tonally enhanced ver- sion. This dual representation avoids com- mitting to a single visual interpretation of degraded cave art. Annotation-level uncertainty Manual contour extraction introduces boundary ambiguity and annotator bias. Instead of assuming a single definitive silhouette, the framework derives two independent contours from the original and enhanced images, yielding two plausible hand silhouettes. These are treated as alternative representations within a range of admissible boundaries rather than as ex- act geometries. 8 Figure 3: Representative samples of contemporary test X-ray images (top row) and their corresponding extracted silhouettes (bottom row). Each column displays a unique case identified by its 4-digit ID above, with associated sex and age metadata provided below. Figure 4: Bone age distribution by sex across the training, validation, and test subsets of the RSNA Bone Age Challenge dataset. Source: (Mollineda et al. 2025). Reproduced with permission. Morphological-level uncertainty To explic- itly model contour variability, structured shape perturbations are generated through binary operators: intersection (AND), union (OR), arithmetic averaging, and a morpho- logical mid-set interpolation. In addition, three-channel compositional variants are built to account for representation dependency in convolutional architectures. In total, 12 silhouette representations are produced per stencil, expanding the morphological hypothesis space prior to classification. Phase 2 Ensemble-based classification of silhouettes Model-level uncertainty Epistemic and sto- chastic uncertainties are addressed through architectural diversity and model replica- tion. Two DNN architectures (Efficient- Net-B3 and MobileViT-S), each trained ac- ross 10 independent runs with distinct ran- dom initializations, process all 12 silhou- ette variants, yielding a total of 120 prob- abilistic predictions per architecture (240 per stencil). This strategy intentionally amplifies variability prior to aggregation. Hierarchical decision aggregation For each architecture, predictions are aggregated us- ing both sum-based and max-based fusion strategies. Consensus class labels are paired with a silhouette support rate, defined as the proportion of silhouette variants con- sistent with the predicted class. This met- 9 (a) Cosquer(b) Chauvet(c) Unknown(d) Pech Merle(e) El Castillo Figure 5: Five prehistoric hand images analyzed in the exploratory study (Mollineda et al. 2025). Labels identify the source caves. ric serves as a morphological agreement in- dex across contour realizations. Phase 3 Post-hoc analysis Latent-space structural validation To assess whether sex-specific structure is intrinsic to the latent space learned by one the DNN architectures, latent feature vectors are pro- jected into two-dimensional (2D) space us- ing an unsupervised manifold learning tech- nique. A non-parametric k-nearest neigh- bors (k-N) classifier is then applied in the 2D embedding space. Agreement between ensemble predictions and 2D space classi- fication indicates structural robustness un- der dimensionality reduction. Interpretability-level aggregation Interpre- tability is assessed using attribution maps generated across silhouette variants and Ef- ficientNet model instances. To suppress explanation noise and model-specific arti- facts, 120 maps are aggregated via the geo- metric median, yielding a robust prototype explanation per stencil. In summary, the framework follows five key prin- ciples: (i) multiplicity before commitment, (i) struc- tured perturbation, (i) architectural diversification, (iv) hierarchical aggregation, and (v) explicit quantifi- cation of internal agreement. The pipeline operates as a controlled uncertainty propagation and consensus system rather than as a single deterministic rule. The next subsections provide a detailed descrip- tion of the three phases outlined above. 3.3.2. Phase 1: silhouette extraction and augmenta- tion This phase comprises a multi-stage pipeline (see Fig. 6) encompassing two preprocessing operations: silhouette extraction and silhouette augmentation. Sil- houette extraction transforms raw hand stencils into normalized, binary silhouettes, while silhouette aug- mentation manages manual tracing uncertainty by gen- erating anatomically consistent, locally perturbed con- tour variants through a series of controlled spatial and pixel-wise transformations. Silhouette extraction. The silhouette extrac- tion process consists of several key steps: âą Adjustment of tonal image parameters (so- urce-level uncertainty management). Cave art images typically present substantial chal- lenges for automated segmentation due to un- even illumination, pigment degradation, surface irregularities, and background textures. The original image undergoes parametric adjustment through multiple enhancement areas: exposure, brightness, contrast, highlights, shadows, and sharpness. The methodology then branches into two separate paths, applying silhouette outlin- ing and extraction to both the original and en- hanced images to mitigate uncertainties arising from degraded image quality. âą Manual contouring (annotation-level un- certainty management). Hand silhouette ex- traction was achieved by manually contouring both original and enhanced cave hand images to isolate hand morphology. This dual-image approach (see Fig. 7) establishes a robust foun- dation for mitigating subjective annotation bias 10 Figure 6: Silhouette extraction and augmentation pipeline (phase 1). From left to right: manual contouring is performed on both original and enhanced hand stencils, followed by the extraction of two primary silhouettes and their subsequent geometric normalization. The data augmentation phase comprises four binary operations (AND, OR, AVG, and morphological mid-set) between the extracted silhouettes. Finally, 3D composition operators combine the six resulting variants (two primary + four binary augmentations) into three-channel representations. Generated using Gemini Notebook from authorsâ prompts and methodology. while providing complementary perspectives on ambiguous boundaries. By processing these two versions, the pipeline also compensates for the degraded quality and irregular pigmentation of ancient rock art. âą Silhouette extraction. The manual contours derived from the original and enhanced cave im- ages feed into a hand silhouette extraction stage via contour masking. This process transforms outlined hand regions into clean binary hand masks (silhouettes) while suppressing background information. By decoupling hand morphology from photographic variability and surface irreg- ularities, this approach enables standardized sha- pe analysis in subsequent phases (see Fig. 6). This process comprises three primary steps: i) background suppression, i) binary silhouette ex- traction, and i) denoising and smoothing. Back- ground suppression isolates the hand by zeroing out all background pixels. Binary silhouette ex- traction maps all hand pixels to white, yielding a binary image. Denoising and smoothing mit- igate impulse noise via median filtering, while morphological operations eliminate artifacts, fill holes, and regularize contour boundaries. âą Geometric normalization. Manual axis align- ment standardizes hand poses through rotation, scaling, and translation, ensuring consistency while strictly preserving anatomical morphol- ogy (see geometric in Fig. 6). Silhouette augmentation. Silhouette augmen- tation operations, defined from the two normalized hand silhouettes, explicitly account for both morpho- logical and representational uncertainty. Motivated by the uncertainty of manually con- touring degraded, low-contrast images, silhouette aug- mentation acknowledges that contour placement is in- fluenced by noise, surface deterioration, and subjec- tive judgment; therefore, extracted silhouettes should be interpreted as plausible instances within a broader space of admissible shapes rather than precise ge- ometries. From the two manually contoured silhou- ettes, augmentation generates locally perturbed vari- ants via pixel-wise, morphological, and stacking op- erators, simulating other plausible contour variations without compromising global anatomy. From a learning perspective, each augmented sil- houette represents a noisy observation (or a stochastic 11 (a) Manual contouring(b) Original hand stencil(c) Enhanced hand stencil Figure 7: Manual contouring applied to original versus enhanced cave hand stencil images. realization) of the underlying hand morphology. By evaluating all variants with ensembles of silhouette- based sex classifiers and fusing their weak predictions, the approach promotes statistical consensus through aggregation: consistent morphological cues are rein- forced across variants, whereas spurious or annotation- specific artifacts are averaged out. This strategy en- ables the final prediction to emerge from the conver- gence of multiple imperfect observations, yielding de- cisions that are less sensitive to local contour inac- curacies and more representative of the latent hand shape supported by the available visual evidence. Let S 1 and S 2 be the two primary extracted binary silhouettes from the original and the enhanced hand images, respectively, such that S 1 , S 2 â 0, 1 HĂW . Ten new augmented variants were generated from four binary and six 3D composition operators. Binary operators. Silhouette augmentation is partially achieved using a set of binary operators de- tailed below: âą Pixel-wise AND (logic): S and = S 1 â© S 2 . The AND operation yields an intersection silhouette, representing a high-confidence region where both silhouettes agree. It can be understood as a lower bound on the true hand shape, and re- ferred to as âconsensus silhouetteâ. âą Pixel-wise OR (logic): S or = S 1 âȘS 2 . The OR operation produces a union silhouette, forming a comprehensive shape envelope that includes all pixels from either silhouette. It could be an upper bound on the true hand shape, and re- ferred to as âcomposite silhouetteâ. âą Average (probabilistic): S avg = 1 2 (S 1 + S 2 ). The average silhouette requires careful consider- ation since it transitions from binary to continuous- valued representation (soft silhouette). It can be interpreted as the probability that a pixel belongs to the true silhouette; thus, it can be called âprobabilistic silhouetteâ. âą Morphological mid-set (geometric): Let S and and S or be the inner and the outer silhouettes, such that S and â S or . The morphological mid- set is S mid = p â S or | d(p, S and ) †d(p, S c or ), where d denotes Euclidean distance. Unlike sta- tistical approaches, this geometric operation per- forms a distance-based shape interpolation, where each point in the output is closer (or equidis- tant) to the inner silhouette than to the outer boundary. The result is a geometrically inter- mediate shape that captures continuous transi- tions between contours. 3D composition operators. The silhouette- based sex classification DNNs were originally designed to process RGB images with three input channels. To adapt these architectures for grayscale or binary in- puts, these 2D images were broadcast across three channels to construct a data structure equivalent to standard RGB data. This allowed for the full retrain- ing of the DNNs on the new image domain. Conse- quently, the analysis of the original silhouettes (S 1 , S 2 ), as well as their augmented versions (S and , S or , S avg , S mid ), required replicating each silhouette across three channels prior to being presented to the models. 12 This architectural requirement motivated the ex- ploration of alternative three-channel inputs, cons- tructed by stacking various silhouette variantsâderived from the two primary hand silhouettes and their four augmentationsâinto composite 3D tensors. These stacked 3D compositions were then used as inputs to the pre-trained DNNs. Within this pilot study, the following three-channel compositions were constructed and evaluated: âą S 12A = (S 1 , S 2 , S and ) âą S 12O = (S 1 , S 2 , S or ) âą S 12V = (S 1 , S 2 , S avg ) âą S 12M = (S 1 , S 2 , S mid ) âą S AOV = (S and , S or , S avg ) âą S AOM = (S and , S or , S mid ) In summary, for each cave image of a prehistoric handprint, a total of 12 silhouette representations were generated: S 1 , S 2 , S and , S or , S avg , S mid , S 12A , S 12O , S 12V , S 12M , S AOV , and S AOM . 3.3.3. Phase 2: silhouette classification (sex predic- tion via score aggregation) In the absence of biological ground truth, sex pre- diction from cave hand stencils is formulated as a cross-domain classification task. Models are trained on a source domain with sufficient labeled samples (RSNA Bone Age Challenge dataset) and subsequently applied to a target domain lacking annotations (pre- historic hand stencils). Furthermore, the scarcity of high-quality samples in the target domain precludes the reliable use of domain adaptation techniques to mitigate distributional discrepancies between contem- porary and prehistoric hand data. Together, the lack of ground truth, the inability to reduce distributional shifts between source and target domains, and the ambiguity and subjectivity inher- ent in manually tracing prehistoric hand contours in- troduce significant uncertainty into individual model predictions. To address this issue, this section proposes a clas- sification strategy based on aggregating multiple weak (i.e., individually unreliable) predictions generated by ensembles of models applied to all silhouette variants. This approach integrates both representational diver- sity (across silhouette variants) and model diversity (across architectures and instances) into a consensus decision obtained through the fusion of numerous in- dividual predictions. As a result, it enhances robust- ness to outliers and yields more stable classification outcomes. Specifically, two ensembles of 10 model instances each are constructed from two complementary deep neural network architectures: EfficientNet-B3 (Tan and Le 2019) and MobileViT-S (Mehta and Raste- gari 2022). The former provides a favorable perfor- manceâcomplexity trade-off for capturing fine-grained geometric details, whereas the latter employs a hy- brid convolutionalâtransformer design to combine lo- cal feature extraction with long-range contextual mod- eling. This architectural diversity helps reduce sys- tematic bias and promotes predictive variability within the ensemble. Given a cave image, 10 EfficientNet-B3 and 10 MobileViT-S model instances process 12 distinct sil- houette representations (2 manually contoured and 10 augmented variants), yielding 120 predictions per ar- chitecture (240 predictions per image in total). Each prediction is expressed as a probability distribution over two classes: Female and Male. For each archi- tecture, predictions from the 10 independent model instances are aggregated using two strategies: âą Sum aggregation: Computes the cumulative posterior probability for each class across all ensemble members, favoring the class with the highest overall support. âą Max aggregation: Selects the maximum pre- dicted probability for each class across all mod- els, prioritizing the most confident individual prediction. To ensure comparability, all model instances share a consistent standard training configuration: a batch size of 24, an initial learning rate of 0.001, and 100 training epochs. The Adam optimizer is employed for its adaptive convergence properties. Training is guided by a cross-entropy loss with equal class weights, ensuring a balanced objective for binary classifica- tion. This configuration achieves stable convergence while maintaining strong generalization performance on the contemporary validation subset. To further enhance generalization and robustness to viewpoint and alignment variations, data augmentation is ap- plied using moderate geometric transformations: ro- tation (±10 ⊠), width shifts of up to 10%, and height 13 shifts of up to 10%. These transformations preserve the semantic integrity of the silhouettes while simu- lating realistic stochastic distortions. Since comprehensive benchmarking for sex classi- fication falls outside the scope of this work, explicit hyperparameter optimization was deliberately omit- ted in favor of developing an accessible, inference frame- work rather than maximizing predictive performance. This design choice helps ensure that the methodology remains user-friendly for non-specialists. Therefore, more thorough hyperparameter tuning represents a clear avenue for further improvements. 3.3.4. Phase 3: post-hoc analysis and interpretation The final phase of the methodology focuses on post hoc interpretability and analysis, aiming to sub- stantiate the modelâs decision-making process. Specif- ically, this stage aims to provide insight into the inter- nal mechanisms of the deep learning ensemble by ex- amining both the geometric structure of feature repre- sentations in latent space and spatial attribution map- ping. These post-hoc techniques foster model trans- parency and provide a rigorous means to explain and justify sex predictions through anatomically coherent evidence. Latent-Space Analysis and Manifold Map- ping. For each trained EfficientNet-B3 instance in the ensemble, the model automatically learns 1,536 high-level, sex-discriminant features. To make this high-dimensional data human-readable, these features are projected into a two-dimensional (2D) space us- ing UMAP (Uniform Manifold Approximation and Projection) (McInnes et al. 2018). This technique learns a nonlinear manifold in an unsupervised man- ner, meaning it reorganizes the data based on struc- tural similarities without initially knowing the class labels. This dimensionality reduction serves two key purposes. First, it enables the visual assessment of the location of prehistoric hand stencils within the distribution of modern training samples. Second, it supports a secondary, non-parametric sex classifica- tion directly in the 2D space using the k-N rule. The 12 silhouette variants of the nine Paleolithic hand stencils are projected into each learned 2D space, yielding 108 projections per model (12 variants Ă 9 cave images). This provides a robust test set com- posed exclusively of samples designed to mitigate un- certainties associated with manual delineation. Fig- ure 8 illustrates a representative 2D projection gener- Figure 8: Interpretable 2D manifold generated by one of the ten EfficientNet-B3 models, using only the three silhouette variants S and , S or , S mid per stencil to reduce visual clutter. ated by one of the ten EfficientNet-B3 instances, using only three silhouette representations (S and , S or , and S mid ) per stencil to avoid visual clutter. For each variant, a complementary classification is derived using a k-N decision rule, utilizing the manifold-projected contemporary training silhouettes as reference data. These individual predictions are first combined via majority voting across variants to produce an instance-level consensus for each cave im- age. In a second aggregation step, majority voting across all ensemble instances (all manifolds) deter- mines the final classification consensus. Model interpretability via spatial attribu- tion. Spatial attribution maps aim to evaluate whether model decisions rely on coherent and anatomically meaningful regions of the hand silhouette. The result- ing heatmaps emphasize areas such as finger propor- tions and palm width that contribute to the predicted sex. For each prehistoric hand image, attribution maps were generated via LayerCAM (Jiang et al. 2021) across the 12 silhouette representations and the 10 indepen- dent EfficientNet-B3 instances, yielding 120 maps per image. Therefore, these maps capture variability aris- ing from contour uncertainty and model stochasticity. They were subsequently combined using the geomet- ric median method (Weiszfeld 1937), building a sin- gle robust prototype attribution map per cave image, which emphasizes spatial patterns while suppressing outliers. Results are analyzed in Sect. 4.2.3. 14 4. Experiments This section is organized into two subsections to bridge the gap between controlled model validation and practical archaeological application. The first subsection evaluates the ensembleâs performance on contemporary hand silhouettes, establishing a per- formance baseline within the source domain where ground-truth labels are certain. The second subsec- tion details the deployment of the inference frame- work on prehistoric hand stencils, the target domain of this study. By benchmarking the models on mod- ern samples first, we provide a transparent and jus- tified foundation for interpreting classification results and post-hoc analysis on the Paleolithic data. 4.1. Contemporary data: model benchmarking on mod- ern silhouettes 4.1.1. Experimental design A contemporary dataset of hand silhouettes with sex annotations was derived from the RSNA Bone Age Challenge, which contains 14,036 left-hand X- ray images from pediatric patients aged between 1 month and 19 years. Silhouettes were extracted using the Segment Anything Model (SAM) (Kirillov et al. 2023) through the zero-prompt segmentation strat- egy proposed in (Mollineda et al. 2025). This ap- proach achieved a silhouette detection sensitivity ex- ceeding 99% on the training and validation subsets, and 98% on the test subset. Accordingly, the num- bers of modern hand silhouettes used to train, vali- date, and test the DNN models are 12,509, 1,416, and 196, respectively. The broad age range is expected to facilitate the learning of robust sexual dimorphism features across all stages of skeletal development. Im- portantly, the highest concentration of samples oc- curs between 10 and 16 years for both sexes, a stage at which bone maturity is largely attained, thereby providing highly informative patterns for sex-related morphological differentiation. All images were resized to a 400Ă 300 resolution and normalized to the [0, 1] range. The EfficientNet-B3 and MobileViT-S architectures were implemented in PyTorch (Paszke et al. 2019) us- ing the Torchvision and timm libraries, respectively, and initialized with ImageNet-pretrained weights (Deng et al. 2009). Each model was adapted with a sin- gle fully connected output layer to predict the class- conditional probability distribution for each input sam- ple. While initialized with pretrained weights to pro- vide a meaningful starting point, all models were fully retrained on the hand silhouette images. Hyperpa- rameter values remained constant across all classifi- cation tasks, as specified in Sect. 3.3.3. For each architecture, ten independent model in- stances were trained, sharing identical designs, initial weights, and hyperparameters. For each model in- stance, the training subset is used to optimize the DNN parameters over 100 epochs, the validation sub- set to select the optimal model checkpoint (i.e., the epoch yielding the highest validation accuracy), and the test subset to conduct a final independent evalu- ation of the selected model. The primary source of diversity across the ten individual model was the stochasticity in optimiza- tion introduced by the random ordering of training samples. Given the non-convex nature of deep learn- ing loss landscapes, varying the sample sequence dic- tates the optimization trajectory through the high- dimensional weight space. This path-dependency causes models with identical configurations to converge to- ward different local minima and develop unique deci- sion boundaries, thereby fostering the predictive di- versity essential for a robust ensemble. The resulting ensembles (one per architecture) ag- gregated the predictions of their ten constituent mod- els using both sum and max rules. Classification per- formance was evaluated based on the overall accuracy (success rate). Training and inference were conducted on an on- premise server equipped with an AMD Ryzenâą 9 7950X CPU, 64 GB of DDR5 RAM, and an NVIDIA GeForce RTX 4090 GPU (24 GB). The software environment included Ubuntu 22.04.4 LTS, Python 3.10.16, and PyTorch 2.5.1. 4.1.2. Sex prediction results on contemporary data The benchmarking of the models on modern sil- houettes acts as a critical baseline to validate the en- sembleâs predictive performance within a source do- main where ground-truth sex labels are certain. The results in Table 3 demonstrate that both the Efficient- Net-B3 and MobileViT-S architectures, and particu- larly their respective ensembles, exhibit a high capac- ity for sex prediction, achieving substantial accuracy across the contemporary test set. This performance is particularly pronounced in the [6â12) and [12â19] age groups, where the EfficientNet-B3 ensemble using sum aggregation achieves accuracies of up to 85.4% 15 Table 3: Performance of EfficientNet-B3 and MobileViT-S on contemporary data, featuring age-stratified results for the test set and overall results for training and validation sets. Results are reported as mean accuracy ± Standard Error of the Mean (SEM) for individual model instances, alongside ensemble accuracy achieved via sum and max aggregation rules. Square brackets indicate inclusion of the boundary values, whereas parentheses indicate their exclusion. ArchitectureModel Training ValidationTest data datadataAll ages[0â6) y[6â12) y [12â19] y EfficientNet-B3 Mean±SEM82.6± 0.7 75.8± 0.3 81.7± 0.4 68.1± 1.9 81.5± 1.0 85.3± 0.6 Sum aggregation85.477.7 85.7 76.2 85.4 88.4 Max aggregation 85.9 78.184.7 76.284.387.2 MobileViT-S Mean±SEM80.3± 1.4 74.8± 0.7 80.6± 1.7 66.7± 4.1 79.4± 1.9 85.2± 1.6 Sum aggregation84.277.082.766.782.087.2 Max aggregation84.677.883.261.983.1 88.4 and 88.4%, respectively. A clear trend is observed whereby classification accuracy improves with increasing age of the sub- jects. This pattern appears to reflect two main fac- tors. First, the dataset exhibits a higher sample den- sity between 10 and 16 years, leading to greater rep- resentation in the older strata ([6â12) and [12â19]), which likely supports more robust learning of sex- discriminative features than in the youngest cohort [0â6). Second, morphological manifestations of sexual dimorphism become progressively more pronounced with age and are further consolidated during adoles- cence, yielding more distinct and informative skele- tal patterns for model discrimination. This finding provides evidence for the modelsâ ability to reliably classify adult hands. Results also highlight the effectiveness of the en- semble approach, as these configurations consistently outperform the mean performance of individual model instances. This improvement is particularly pronounced for EfficientNet-B3 with both aggregation rules in the [0â6) age stratum, where the ensemble accuracy (76.2%) substantially exceeds the average accuracy of individ- ual instances (68.1%± 1.9%). Across all datasets and test age strata, EfficientNet- B3 achieves higher average performance than Mobile- ViT-S for individual models, and this advantage ex- tends to most ensemble configurations under both ag- gregation rules. This behavior may be attributed to the greater architectural capacity of EfficientNet-B3 (both in depth and parameterization) which enables the capture of more subtle morphological patterns. In addition, MobileViT architectures may be more sensitive to specific hyperparameter settings than the highly optimized and robust EfficientNet framework, particularly given that no hyperparameter tuning was conducted in this study. Nevertheless, these observa- tions should be interpreted with caution and do not support broad claims regarding the general superior- ity of one architecture over another, as the analysis is restricted to specific model variants applied to a single, specialized task. Figure 9 presents the confusion matrices, detail- ing class-wise correct and incorrect predictions across age-stratified test data for both ensembles (architec- tures) under the sum-aggregation rule. They show a clear and coherent pattern across age groups and ar- chitectures, with classification performance improv- ing with age and both ensembles exhibiting similar behavior under sum-based aggregation. Predictions for the Male class are consistently more stable, while those for the Female class display greater variability, particularly in the youngest age group. This differ- ence becomes less pronounced in older strata, where both classes are recognized with higher accuracy. Im- portantly, although Male predictions tend to achieve higher recall, this does not imply that Female predic- tions are intrinsically more reliable. Rather, the re- sults suggest that Female classifications may require stronger or more distinctive evidence, leading to fewer correct detections overall. Consequently, these pat- terns should be interpreted as reflecting differences in class-specific model sensitivity rather than straight- forward differences in prediction confidence. 4.2. Prehistoric data: model application and inter- pretability on cave hand stencils This section presents the application of the infer- ence and analysis pipeline described in 3.3 to the Pa- leolithic hand stencils. It consists of three parts, each reporting results obtained by aggregating evidence across all silhouette variants and model instances: 16 Figure 9: Confusion matrices summarizing class-wise correct and incorrect predictions across age-stratified test data for both ensembles (architectures) under the sum-aggregation rule. Sex prediction results on prehistoric data. The two ensembles, each comprising 10 model instances of EfficientNet-B3 and MobileViT-S trained on contem- porary hand silhouettes from the RSNA Bone Age Challenge dataset, are applied to the 12 silhouette representations of the nine Paleolithic hand stencils. The ensemble strategy is explained below. Post-hoc latent-space analysis and manifold map- ping. Both contemporary and Paleolithic hand sil- houettes are projected into 2D manifolds derived from EfficientNet-B3 latent spaces using UMAP. Each sil- houette variant is classified via a k-N rule based on the distribution of projected contemporary training samples. Predictions are aggregated across variants and EfficientNet-B3 instances to provide a comple- mentary sex classification for each stencil. Post-hoc interpretability analysis via spatial attribution. For each prehistoric hand image, vi- sual attribution maps are generated using LayerCAM across the 12 silhouette variants and the 10 Efficient- Net-B3 instances, yielding 120 maps per image. These maps are aggregated into a single, robust consensus attribution map for each cave image, highlighting spa- tial patterns relevant to sex prediction. 4.2.1. Sex prediction results on prehistoric data This section provides the necessary details to un- derstand the ensemble strategies and presents and dis- cusses the aggregated results for sex prediction of the hand stencils. Section 3.3.3 introduced the general methodology for sex prediction via score aggregation. In summary, two ensembles of 10 model instances each process the 12 silhouette variants of every hand stencil, yielding 120 predictions per ensemble. For each architecture, these predictions are aggregated using both sum- and max-based strategies to produce robust predictions. Each prediction is expressed as a probability distri- bution over two classes, Female and Male. At the first level of aggregation, the outputs across the 10 ensemble instances are combined for each in- dividual silhouette variant. Under sum aggregation, cumulative posterior probabilities are computed for each class, while under max aggregation, the maxi- mum probability per class is selected. The consen- sus class for each silhouette variant is defined as the one with the highest aggregated score. Additionally, 17 a classification margin is computed as the difference between the class ensemble scores (normalized for the sum rule). The sign of this margin indicates the ensembleâs preferred class for that variant, while its magnitude, the ensembleâs confidence. At the second aggregation level, the predictions obtained from the 12 silhouette variants are combined to determine the final classification of the original hand stencil. The final prediction is assigned to the majority class among the 12 variants and is eval- uated using two confidence metrics: the Silhouette Support Rate (SSR) and the Average Classification Margin (ACM). The SSR quantifies the proportion of silhouette variants supporting the final consensus class. Given the two-class system without an absten- tion mechanism, the SSR ranges from 0.5 (maximum ambiguity) to 1 (complete consensus across all vari- ants). The ACM is computed by first averaging the signed classification margins across all variants, and then taking the absolute value of that mean as a fine- grained confidence measure ranging from 0 (minimum certainty) to 1 (maximum confidence). Table 4 summarizes sex prediction results for the nine Paleolithic hand stencils using classifier ensem- bles based on two DNN architectures (EfficientNet-B3 and MobileViT-S), under both sum- and max-based fusion strategies. Final predictions are compared with prior or existing hypotheses derived from earlier sta- tistical or morphometric approaches. The results indicate that ensembles produce stable and coherent predictions across architectures and ag- gregation strategies, with most cases showing agree- ment between sum and max rules within the same model. This internal consistency suggests that the ag- gregation process effectively reduces variability across the 120 individual predictions per stencil, yielding ro- bust consensus outputs. In particular, cases with high SSR (i.e., SSR â 1) and high ACM (i.e., ACM â„ 0.5) can be interpreted as high-confidence predictions, reflecting strong agreement across both silhouette vari- ants and ensemble members. This level of certainty is exemplified by the samples img_26 (Female), img_28 (Female), Cosquer cave (Female), and El Castillo 25 (Male), all of which exhibit complete agreement among the 12 silhouette variants (SSR = 1) across both deep learning ensembles and both aggregation strate- gies, while also demonstrating strong model confi- dence (ACM > 0.5). A broader pattern emerges when comparing the two DNN architectures: EfficientNet-B3 typically yields more stable predictions than MobileViT-S, frequently accompanied by higher confidence measures. This trend is consistent with earlier observations on con- temporary data, where EfficientNet-B3 demonstrated superior average performance, suggesting that its la- tent representations are more effective at capturing discriminative features for this task. Nonetheless, Mo- bileViT-S maintains a high level of classification con- sistency, further validating the overall robustness and cross-model reliability of the ensemble approach. This evidence further supports the decision to conduct the post-hoc analyses using the EfficientNet-B3 ensemble. When comparing the ensemble predictions with prior hypotheses, the results reveal a heterogeneous pattern of agreement and divergence. Several cases are consistent with earlier morphometric or statisti- cal interpretations, particularly those historically re- garded as less ambiguous. For example, img_26, and Cosquer were unanimously classified as Female with high SSR and ACM values, aligning perfectly with ex- isting archaeological hypotheses. Similarly, El Castillo 25 was strongly classified as Male across all configura- tions, in complete agreement with prior estimations. In contrast, strong divergences were identified in two other El Castillo cases (img_28 and img_38), where all ensembles produced highly consistent and confi- dent predictions that differed from prior hypotheses. Cases like img_31 and Pech Merle remain complex, as the models split between Male and Female predic- tions, mirroring the lack of consensus found in tradi- tional morphometric or statistical approaches. Confidence measures (SSR and ACM) play a key role in interpreting these results. Cases with lower SSR or ACM values tend to correspond to histori- cally ambiguous or controversial stencils, suggesting that the ensemble appropriately reflects underlying uncertainty rather than forcing overconfident deci- sions. Conversely, high-confidence cases indicate strong structural consistency across silhouette variants and model instances. Overall, the results suggest that while AI can pro- vide highly confident sex attributions for well-defined stencils, cases with low SSR and ACM values indicate inherent morphological ambiguity in the prehistoric data. The combination of aggregation strategies and confidence metrics provides a nuanced view of the re- sults, allowing both agreement and uncertainty to be explicitly quantified. 18 Table 4: Sex attribution results for nine Upper Paleolithic hand stencils using EfficientNet-B3 and MobileViT-S ensembles, including fusion strategies, silhouette supportrates (SSR), average classification margins (ACM), and comparison with prior archaeological hypotheses. F and M denote the Female and Male classes, respectively. Cases from El Castillo cave (selected for their high contrast and well-defined edges) Cave Image* Prior / Existing Hypotheses EfficientNet-B3 MobileVit-S (Snow 2006) Sum aggregation Max aggregation Sum aggregation Max aggregation Step 1 Step 2 Inference Class SSR/ACM Class SSR/ACM Class SSR/ACM Class SSR/ACM img_26 Strong Female Strong Female Adult Female F 1/0.66 F 1/0.56 F 1/0.71 F 1/0.55 img_28 Weak Female Weak Male Adolescent Male F 1/0.71 F 1/0.63 F 1/0.77 F 1/0.57 img_31 Strong Female Weak Male Adolescent Male M 0.83/0.18 M 0.92/0.19 F 0.83/0.05 F 0.92/0.05 img_38 Strong Female Female Adult Female M 1/0.26 M 1/0.32 M 1/0.24 M 1/0.24 * Case names are consistent with the ID numbers shown in the original images. Cases from the work (Mollineda et al. 2025) Cave Image* Prior / Existing Hypotheses EfficientNet-B3 MobileVit-S Sum aggregation Max aggregation Sum aggregation Max aggregation (Wang et al. 2010) (Snow 2013) (Mollineda et al. 2025) Class SSR/ACM Class SSR/ACM Class SSR/ACM Class SSR/ACM Cosquer Female N/A Female F 1/0.75 F 1/0.72 F 1/0.81 F 1/0.66 Chauvet Male N/A Strong Male F 1/0.29 F 1/0.33 F 0.75/0.08 M 0.58/0.04 Unknown Male N/A Strong Male M 0.83/0.09 M 0.67/0.04 M 0.92/0.17 F 0.58/0.03 Pech Merle Male Female Weak Male M 0.92/0.15 M 0.50/0.05 F 0.83/0.17 F 0.83/0.13 El Castillo 25 Male Male Strong Male M 1/0.68 M 1/0.62 M 1/0.79 M 1/0.75 * Case names are consistent with those used in (Mollineda et al. 2025). 19 Figure 10: Mean SSR (support for the predicted class across all silhouette variants) as a function of k. A sensitivity analysis was conducted over the range k â [1, 50). Since SSR remains consistently above 0.5, the k-N predictions are shown to be robust and largely insensitive to the choice of k. 4.2.2. Post-hoc latent-space analysis and manifold map- ping Each projected silhouette variant is classified us- ing the k-N rule based on the distribution of contem- porary training samples embedded within 2D mani- folds. The resulting predictions are aggregated, first across silhouette variants and then across 2D map- pings, yielding a complementary sex classification hy- pothesis for each stencil. Consistent with the confi- dence measures defined in 4.2.1, the SSR is computed for each hand stencil within each 2D manifold, and the resulting SSR values are averaged across the ten mappings. Accordingly, the SSR ranges from 0.5, in- dicating maximum ambiguity, to 1, indicating com- plete consensus across all silhouette variants. Rather than selecting an arbitrary value for k, a sensitivity analysis was conducted across the range k â [1, 50), leveraging the high density of the pro- jected contemporary data. For each value of k, the mean SSR was calculated across all 2D mappings to generate a curve representing SSR behavior as a func- tion of k for each hand stencil (see Figure 10). This analysis revealed that the k-N aggregation predic- tions are invariant to changes in k (SSR initially scales with k before stabilizing significantly above 0.5), de- monstrating the robustness of the class-conditional densities within the mappings and aggregation ap- proach based on k-N. As illustrated in Figure 10, the SSR curves stabilize for k â„ 20; then, the average SSR across the interval [20, 50) was adopted as the fi- nal confidence measure for the k-N classification of each stencil. Table 5 summarize the sex classification results for the nine cave hand stencils, including both the predic- tions obtained using the k-N rule on the 2D manifold mappings and those produced by the EfficientNet-B3 ensembles. This comparison aims to evaluate the de- gree of convergence between the highly interpretable 2D representations and the DNN models, which ex- hibit strong generalization capabilities. Overall, there is strong concordance between the k-N classifications and the EfficientNet-B3 ensem- ble predictions. In all nine cases, the predicted sex label obtained from the k-N classifier matches the label reported by the ensemble. This consistency sug- gests that the class structure encoded in the high- dimensional latent spaces is largely preserved in the reduced two-dimensional manifold and remains sepa- rable using a simple distance-based method. Four cases show particularly strong support val- ues across both averaging schemes (embedding-space averaging and fusion-strategy averaging), indicating robust local clustering in the 2D space. These include img_26 and img_28 from El Castillo Cave (Female, 0.966â1.0 support), the Cosquer case from Cosquer Cave (Female, 0.989â1.0 support), and El Castillo 25 (Male, 0.981â1.0 support). In these cases, the k-N classifier identifies highly discriminant neighborhoods in the 2D projection, consistent with the decisive en- semble predictions. The classification for these four hand stencils is illustrated in Figure 8. The three points (silhouettes) corresponding to each case are po- sitioned consistently outside the region of confusion between the two classes. Moderate support values are observed in img_31 (Male, 0.697â0.87), img_38 (Male, 0.821â1.0), the Chauvet case (Female, 0.656â1.0), the unknown case (Male, 0.793â0.75), and the Pech Merle case (Male, 0.823â0.71). Although their support levels are lower than in the most robust examples, the predicted class remains stable across both approaches. The reduced support values indicate less compact clustering of the 12 silhouette variants in the 2D space, but do not result in label inversion relative to the ensemble out- put. In contrast to the higher-confidence cases, as il- lustrated in Figure 8, the points associated with these five hands are located much closer to the inter-class overlap region, thereby explaining the lack of consen- sus among the different classification rules. Notably, no discrepancies were observed between the k-N classifier applied to the 2D manifolds and the EfficientNet-B3 ensemble predictions. While mod- 20 Table 5: Aggregated k-N classification performance within 2D manifolds across the 12 silhouette variants and the 10 EfficientNet- B3 instances. k-NNEfficientNet-B3 Cave Image Predicted Class SSR + Predicted Class SSR/ACM â img_26Female0.982Female1/0.61 img_28Female0.966Female1/0.67 img_31Male0.697Male0.87/0.18 img_38Male0.821Male1/0.29 CosquerFemale0.989Female1/0.73 ChauvetFemale0.656Female1/0.31 UnknownMale0.793Male0.75/0.06 Pech MerleMale0.823Male0.71/0.10 El Castillo 25Male0.981Male1/0.65 + Average of the mean SSRs (computed across all 2D manifolds) over the interval k â [20, 50). * Average SSR and ACM values across both fusion strategies for the EfficientNet-B3 ensemble (derived from Table 4). Figure 11: Robust LayerCAM attribution maps obtained by geometric median aggregation across 12 silhouette variants and 10 EfficientNet model instances. est variations in support metrics occurred, they did not alter the final class assignments. This alignment suggests that the sex-specific structure captured within the high-dimensional latent representations learned by EfficientNet-B3 instances is robust enough to re- main discriminable even after aggressive unsupervised dimensionality reduction and to be recoverable through a simple neighborhood-based decision rule. 4.2.3. Post-hoc interpretability analysis via spatial at- tribution To assess whether the predicted sex labels were supported by coherent and anatomically meaningful visual cues, 120 LayerCAM attribution maps (12 sil- houette variantsĂ 10 EfficientNet-B3 instances) were generated for each stencil and aggregated using the geometric median. This aggregation suppresses contour- specific artifacts and model stochasticity, yielding a robust prototype explanation per hand. The nine re- sulting maps, which are shown in Fig. 11, represent stable spatial patterns across both silhouette uncer- tainty and model replication. Across all nine cases, relevance is consistently con- centrated along the fingers and interdigital spaces, particularly around middle and proximal phalanges and the metacarpal heads, while the central palm and wrist regions show minimal activation. This spatial selectivity suggests that the models rely primarily on finger geometry and relative spacing (features con- sistent with sexually dimorphic morphology) rather than on global hand size or background structure. Table 6 brings together predictive, latent-space, and interpretability indicators for each case. A clear distinction emerges within the four fe- male predictions. A first group composed of img_26, img_28, and the Cosquer case shares a highly consis- tent relevance structure characterized by distributed activation across the central phalangeal regions of the index, middle and ring fingers, with consistent empha- sis on interdigital spacing and some minor activation on the thumb. These maps are spatially compact and anatomically organized. Importantly, all three cases exhibit maximal or near-maximal silhouette support in Table 4 and strong k-N support in the 2D man- ifolds in Table 5, indicating both predictive stabil- ity and compact latent-space clustering. In contrast, the Chauvet stencil exhibits a more diffuse attribution pattern, irregularly distributed across the middle and proximal phalanges of the index and middle fingers. This case corresponds to the lowest k-N support among female classifications in Table 5 and to one of the largest discrepancies between prior archaeological hypotheses and deep learning predictions (Table 4). 21 Table 6: Case-by-case interpretive analysis integrating spatial attribution patterns and aggregation model confidence measures. Stencil Sex Attribution Pattern Spatial Ensemble k-N Interpretive Assessment Coherence SSR SSR img_26 F Broad activation across the index, middle and ring fingerswith interdigital emphasis. High 1.0 High Morphologically stable female pattern consistent withmaximal ensemble agreement. img_28 F Focal activation in the central ring finger with interdigitalemphasis. High 1.0 High Structurally compact female pattern consistent withmaximal ensemble agreement. img_31 M Dispersed activation mainly in the proximal regions of thelittle and index fingers. Moderate 0.83â0.92 Low A dispersed activation pattern reflects architecturaldisagreement and smallclassification margins. img_38 M Sharp, focal activation in proximal middle-ring regions. High 1.0 Moderate Internally confident classification despite reversalof prior interpretation. Cosquer F Broad activation across the index, middle and ring fingerswith interdigital sensitivity. High 1.0 High Strong agreement across models, silhouette variants andprior hypotheses; stableembedded representation. Chauvet F Diffuse activation in the mid region of the index fingers. Low 0.58â1 Moderate Structurally ambiguous female case mirrors discrepancy withprior hypotheses andarchitectural disagreement. Unknown M Sharp, focal activation at palmâfinger junction. Moderate 0.58â0.92 Moderate Stable but not compact; moderate support acrossdecision approaches. Pech Merle M Weakly focused activation at palmâfinger junction. Moderate 0.50â0.92 Moderate Morphologically ambiguous case; mirrors historicalinterpretive discrepancy andarchitectural disagreement. El Castillo 25 M Focal activation at palmâfinger junction. High 1.0 High Strongly stable male pattern with maximal ensemble andlatent-space agreement. 22 The spatial dispersion observed in the Chauvet map therefore aligns with its reduced structural support and interpretive ambiguity. The five male-classified cases generally exhibit less focal patterns, with strong activation around the me- tacarpophalangeal joints of fingers, and occasionally extending toward the ulnar side of the palm below the little finger. Among them, img_38 and El Castillo 25 (both with maximal support in Table 4) present well-defined spatial clusters. Their compact attri- bution patterns mirror their strong ensemble agree- ment and high latent-space support in Table 5. In contrast, img_31 and Pech Merle show broader acti- vation zones consistent with intermediate or reduced support values (Table 4). Notably, Pech Merle, char- acterized by interpretive disagreement, exhibits one of the most diffuse patterns among male cases, par- alleling its lower silhouette support and architectural divergence. As in the Chauvet example, spatial dis- persion appears to reflect underlying morphological or representational ambiguity rather than classification instability alone. 5. Discussion The proposed methodology was structured as a multi-layered uncertainty management system. In- stead of treating uncertainty as noise to be suppressed, it is modeled and propagated across the full analyti- cal pipeline before being resolved through hierarchical aggregation. Uncertainty is addressed at the source level (dual image processing), annotation level (dual contour extraction), morphological level (structured silhouette perturbations), representational level (mul- tiple three-channel compositions), model level (archi- tectural diversity and stochastic replication), and de- cision level (ensemble fusion with silhouette support rates). Each stage introduces controlled variability prior to aggregation, ensuring that final predictions emerge from the convergence of multiple imperfect but systematically structured observations. The decision-making process incorporates criteria beyond simple probability estimates. Structural vali- dation was conducted using an unsupervised UMAP projection followed by non-parametric k-N classifi- cation to test whether class separability persists un- der aggressive dimensionality reduction. Addition- ally, spatial attribution maps are aggregated via the geometric median to ensure explanation stability across contour variants and model instances. The interpretabil- ity results provide independent support for the inter- nal coherence of the framework. Because each Layer- CAM map aggregates 120 individual attribution in- stances, the resulting spatial patterns reflect relevance structures that are stable under both contour uncer- tainty and stochastic model variation. Across cases, activation patterns consistently tar- get anatomically meaningful regions of the hand, par- ticularly the fingers, interdigital spaces, proximal and middle phalanges, and metacarpal heads, while as- signing comparatively little relevance to the palm and wrist. This observation is noteworthy because several of these structures have previously been identified as important sources of sexual dimorphism in skeletal studies. For example, the proximal phalanges have been reported to exhibit significant sex-related dif- ferences, with the thumb showing particularly pro- nounced dimorphism. The correspondence between these anthropologically established markers and the regions highlighted by the attribution maps suggests that the classifiers are relying on biologically plausi- ble morphological cues rather than on spurious im- age characteristics or global size-related proxies. Al- though attribution maps cannot establish causal rela- tionships between anatomical traits and classification outcomes, the observed spatial consistency provides additional evidence that the models capture meaning- ful aspects of hand morphology associated with sexual differentiation. The convergence of these activation patterns across multiple silhouette variants and en- semble members further suggests that the identified anatomical regions are robust to contour uncertainty and model stochasticity, thereby strengthening confi- dence in the interpretability of the resulting predic- tions. Moreover, a qualitative correspondence emerges between attribution coherence, ensemble confidence rates, and latent-space separability. Cases exhibit- ing maximal internal consensus (SSR â 1) display sharply localized and anatomically organized activa- tion patterns. By contrast, cases associated with ar- chitectural disagreement or reduced silhouette sup- port show more diffuse and spatially heterogeneous relevance distributions. This alignment indicates that spatial dispersion in the geometric median maps may function as a visual correlate of epistemic ambiguity rather than mere saliency noise. The layered design thus transforms uncertainty from a liability into an analyzable signal. Agreement across predictive aggre- gation, latent-space validation, and attribution stabil- 23 ity provides convergent evidence for morphologically stable classifications, whereas divergence across these layers highlights structurally ambiguous cases. The framework therefore functions not merely as a pre- dictive system, but as a structured mechanism for diagnosing stability in the absence of archaeological ground truth. From a practical perspective, the proposed frame- work is intended to support archaeological interpre- tation through the joint analysis of predictions, con- fidence measures, latent-space organization, and at- tribution patterns, rather than through binary sex labels alone. Hand stencils exhibiting strong con- sistency across ensemble configurations, high silhou- ette support rates, large classification margins, con- sistent positioning within the latent-space manifolds, and agreement between ensemble and k-N classifica- tions may be considered relatively robust candidates for sex attribution. In contrast, cases characterized by reduced support rates, small classification margins, proximity to the overlap region between classes in the latent-space projections, or disagreement among an- alytical components should be interpreted as inher- ently ambiguous. Such cases are not necessarily clas- sification failures; instead, they may reflect genuinely intermediate morphologies, insufficient visual infor- mation, or limitations of the available reference pop- ulation. Accordingly, the confidence indicators pro- duced by the framework should be treated as archae- ological evidence in their own right. High-confidence predictions can contribute to broader demographic in- terpretations, whereas low-confidence cases may be more appropriately reported as indeterminate rather than being forced into categorical assignments. This perspective shifts the focus from obtaining definitive classifications toward evaluating the strength and re- liability of the available evidence, promoting more transparent and reproducible interpretations of pre- historic hand stencils. A fundamental limitation of this study arises from the difficulty of establishing how closely contempo- rary reference populations resemble the Upper Pale- olithic populations to which the framework is ulti- mately applied. Given the considerable temporal and evolutionary separation between these groups, sub- stantial differences in the distribution of sex-related hand morphology are plausible, yet their magnitude remains unknown. All supervised learning approaches for prehistoric hand stencil analysis, whether based on morphometric measurements, engineered features, or deep learning, necessarily rely on modern reference data because no ground-truth sex labels exist for ar- chaeological hand stencils. Consequently, the models assume that at least part of the morphological signa- tures of sexual dimorphism observed in present-day human populations remain sufficiently stable across long temporal scales. While this assumption is sup- ported by the persistence of broad anatomical differ- ences between sexes, it cannot be directly verified for Upper Paleolithic populations. Evolutionary, demo- graphic, nutritional, developmental, and population- specific factors may have influenced hand morphology over thousands of years, potentially altering the dis- tribution of sex-related traits. Previous studies have already demonstrated that predictive functions devel- oped in one contemporary population may general- ize poorly to another due to differences in average hand morphology and population structure. There- fore, even highly consistent predictions should not be interpreted as direct determinations of biological sex, but rather as probabilistic inferences relative to the morphological patterns learned from the available contemporary reference population. Although future methodologies based on unlabeled target-domain data may help narrow the gap between contemporary and prehistoric populations, the feasibility of such approa- ches remains strongly conditioned by the limited num- ber, quality, and representativeness of currently avail- able archaeological samples. Consequently, the pro- posed framework should be understood as a tool for generating evidence-based hypotheses and quantify- ing their internal consistency, rather than as a mech- anism for obtaining definitive sex attributions. This framework could be further advanced by en- hancing its ability to represent input variability and quantify uncertainty within the predictive process. Moving toward a more continuous and expressive rep- resentation of plausible hand shapes would enable a richer characterization of morphological variability and a more robust assessment of prediction stability under structured perturbations. In parallel, incorporating mechanisms to capture uncertainty in model predic- tions would shift from point estimates to predictive distributions, enabling a more systematic quantifica- tion of confidence and its integration into result in- terpretation. More broadly, virtually all stages of the analytical pipeline offer opportunities for systematic scaling to improve uncertainty modeling. From early image pro- cessing and manual contour delineationâwhere inde- 24 pendent realizations can be incorporatedâeach stage can be expanded to propagate variability forward. The cumulative combination of outputs across stages enables a more comprehensive exploration of plausible interpretations, strengthening the robustness of final predictions. Together, these directions point toward a more fully probabilistic formulation of the frame- work, enhancing its ability to disentangle data-driven ambiguity from model uncertainty, particularly in the absence of reliable ground truth. 6. Conclusions This study introduced a multi-layered uncertainty- aware framework for sex attribution in prehistoric hand stencils. Rather than treating uncertainty as an unde- sirable artifact to be eliminated, the proposed method- ology models, propagates, and aggregates uncertainty across multiple levels of analysis, including image ac- quisition, manual annotation, silhouette representa- tion, model training, and decision making. The re- sulting framework combines controlled variability with hierarchical aggregation, allowing final classifications to emerge from the convergence of multiple comple- mentary observations rather than from a single deter- ministic prediction. Experimental results on contemporary hand sil- houettes proved that ensemble-based deep learning models can reliably capture sexually dimorphic mor- phological patterns, particularly in age groups where such traits are more strongly expressed. When trans- ferred to prehistoric hand stencils, the framework pro- duced stable predictions for several cases while si- multaneously identifying others as intrinsically am- biguous through reduced silhouette support and lower classification margins. These confidence indicators of- fer valuable guidance in mitigating the overinterpre- tation of uncertain archaeological evidence. A key contribution of the study is the triangula- tion of evidence through three complementary per- spectives: ensemble-based classification, latent-space validation through unsupervised learning of interpre- table 2D manifold, and aggregated LayerCAM ex- planations. The strong agreement observed among these components indicates that the inferred classifi- cations are supported by consistent latent representa- tions and anatomically meaningful visual cues. Con- versely, cases exhibiting disagreement across these anal- yses reveal regions of epistemic uncertainty that war- rant cautious interpretation. More broadly, the proposed framework demon- strates that uncertainty itself can be treated as an in- formative signal. Agreement across predictive, struc- tural, and interpretability analyses provides conver- gent evidence for morphologically stable classifications, whereas divergence highlights cases where the avail- able evidence remains insufficient for definitive attri- bution. In this sense, the methodology functions not only as a predictive system but also as a tool for as- sessing the reliability of archaeological inferences in the absence of ground truth. Future work should move the framework toward more explicitly probabilistic formulations, strength- ening its capacity to represent input variability and to quantify uncertainty in predictions. These develop- ments may help disentangle morphological ambiguity from model-related uncertainty. Acknowledgments The authors would like to express their sincere gratitude to Professor Roberto Ontañón Peredo for kindly providing the scaled images of the El Castillo cave hand stencils used in Figures 1 and 5, as well as in part of Figure 7a. These images were instrumental in both the development of the proposed methodology and the presentation of the results. The authors also acknowledge the partial finan- cial support provided by grant AIA2025-163919-C54, funded by MICIU/AEI/10.13039/501100011033 (Spain). Declaration of competing interest The authors declare that there is no conflict of interest that could affect the independence of the re- search reported in this paper. Data availability To promote transparency, reproducibility, and in- dependent verification of the reported results, the data- sets generated in this study are publicly available. 2 . The source code developed for the classification pipeline and the visualization of the experimental results is also publicly available. 3 2 Thedatageneratedinthis workisavailablethroughthislink: https://w.kaggle.com/datasets/karelbecerra/decoding- the-past-in-prehistoric-hand-stencils 3 The code written in this work is available through this link: https://github.com/karelbecerra/hand-paintings-in-rock-art 25 Ethics approval Not applicable. Declaration of generative AI and AI-assisted technologies in the manuscript preparation pro- cess During the preparation of this work, the authors used generative artificial intelligence tools, specifically ChatGPT (GPT-5.5) developed by OpenAI, to im- prove the readability and clarity of selected sections of the manuscript, and Gemini Notebook to generate in- fographics. All scientific ideas, analyses, and interpre- tations were conceived and developed independently by the authors. The authors carefully reviewed and edited all AI-assisted content and assume full respon- sibility for the accuracy, integrity, and conclusions of the published article. References Bindurani, MK, AN Kavyashree, and LP Subhash (2016). âSexual dimorphism on metric valuation of hand dimensionsâ. In: International Journal of Anatomy and Research 4.2, p. 2212â2215. doi: http://dx.doi.org/10.16965/ijar.2016.180. Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei (2009). âImageNet: A large-scale hierarchical image databaseâ. In: 2009 IEEE Con- ference on Computer Vision and Pattern Recog- nition, p. 248â255. doi: https://doi.org/10. 1109/CVPR.2009.5206848. Diakogiannis, Foivos I, Francois Waldner, Peter Cac- cetta, and Chen Wu (2020). âResUNet-a: A deep learning framework for semantic segmentation of remotely sensed dataâ. In: ISPRS Journal of Pho- togrammetry and Remote Sensing 162, p. 94â114. doi: https://doi.org/10.1016/j.isprsjprs. 2020.01.013. FernĂĄndez-Navarro, V, E CamarĂłs, and D Garate (2022). âVisualizing childhood in Upper Palaeolithic soci- eties: Experimental and archaeological approach to artistsâ age estimation through cave art hand stencilsâ. In: Journal of Archaeological Science 140, p. 105574. doi: https://doi.org/10.1016/j. jas.2022.105574. FernĂĄndez-Navarro, V, D Fidalgo-Casares, D GarcĂa- MartĂnez, and D Garate-Maidagan (2025). âDe- coding Palaeolithic Hand Stencils: Age and Sex Identification Through Geometric Morphometricsâ. In: Journal of Archaeological Method and Theory 32.1, p. 24. doi: https://doi.org/10.1007/ s10816-025-09693-w. FernĂĄndez-Navarro, V, D Garate-Maidagan, and D GarcĂa-MartĂnez (2024). âOntogeny and sexual di- morphism in the human hands through a 2D ge- ometric morphometrics approachâ. In: American Journal of Biological Anthropology 185.1, e25001. doi: https://doi.org/10.1002/ajpa.25001. FernĂĄndez-Navarro, V, A GonzĂĄlez-Marfil, I Arganda- Carreras, and D Garate-Maidagan (2025). âRec- ognizing past shapes: sex differentiation through deep learning on European Upper Palaeolithic hand stencilsâ. In: Digital Applications in Archaeology and Cultural Heritage, e00453. doi: https://doi. org/10.1016/j.daach.2025.e00453. Galeta, Patrik, Jaroslav Bruzek, and Martina LĂĄzniÄkovĂĄ- GaletovĂĄ (2014). âIs sex estimation from hand- prints in prehistoric cave art reliable? A view from biological and forensic anthropologyâ. In: Journal of Archaeological Science 45, p. 141â149. doi: http://dx.doi.org/10.1016/j.jas.2014. 01.028. Gunn, Robert G (2006). âHand sizes in rock art: in- terpreting the measurements of hand stencils and printsâ. In: Rock Art Research: The Journal of the Australian Rock Art Research Association (AURA) 23.1, p. 97â112. Halabi, Safwan S, Luciano M Prevedello, Jayashree Kalpathy-Cramer, Artem B Mamonov, Alexan- der Bilbily, Mark Cicero, Ian Pan, Lucas AraĂșjo Pereira, Rafael Teixeira Sousa, Nitamar Abdala, et al. (2019). âThe RSNA pediatric bone age ma- chine learning challengeâ. In: Radiology 290.2, p. 498â 503. doi: http://dx.doi.org/10.1148/radiol. 2018180736. Janssens, Paul A (1957). âMedical views on prehis- toric representations of human handsâ. In: Medical History 1.4, p. 318â322. doi: http://dx.doi. org/10.1017/S0025727300021499. Jiang, Peng-Tao, Chang-Bin Zhang, Qibin Hou, Ming- Ming Cheng, and Yunchao Wei (2021). âLayer- CAM: Exploring Hierarchical Class Activation Maps for Localizationâ. In: IEEE Transactions on Image Processing 30, p. 5875â5888. issn: 1057-7149. doi: https://doi.org/10.1109/TIP.2021.3089943. 26 Jowaheer, Vandna and Arun Kumar Agnihotri (2011). âSex identification on the basis of hand and foot measurements in Indo-Mauritian population - A model based approachâ. In: Journal of Forensic and Legal Medicine 18.4, p. 173â176. doi: https: //doi.org/10.1016/j.jflm.2011.02.007. Karakostis, FA, E Zorba, and K Moraitis (2015). âSex- ual dimorphism of proximal hand phalangesâ. In: International Journal of Osteoarchaeology 25.5, p. 733â 742. doi: http://dx.doi.org/10.1002/oa.2340. Kirillov, Alexander, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr DollĂĄr, and Ross Girshick (2023). âSeg- ment Anythingâ. In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), p. 3992â 4003. doi: https://doi.org/10.1109/ICCV51070. 2023.00371. KrĂĄlĂk, Miroslav, Stanislav Katina, and Petra UrbanovĂĄ (2014). âDistal part of the human hand: study of form variability and sexual dimorphism using ge- ometric morphometricsâ. In: Anthropologia integra 5.2, p. 7â25. doi: http://dx.doi.org/10.5817/ AI2014-2-7. Larson, David B, Matthew C Chen, Matthew P Lun- gren, Safwan S Halabi, Nicholas V Stence, and Curtis P Langlotz (2018). âPerformance of a deep- learning neural network model in assessing skele- tal maturity on pediatric hand radiographsâ. In: Radiology 287.1, p. 313â322. doi: http://dx. doi.org/10.1148/radiol.2017170236. Lundborg, Göran (2014). âThe Hand and the Brain: From Lucyâs Thumb to the Thought-Controlled Robotic Handâ. In: Springer. Chap. Handprints from the Past, p. 41â48. doi: 10.1007/978- 1-4471-5334-4_5. Manning, John T, Andrew JG Churchill, and Michael Peters (2007). âThe effects of sex, ethnicity, and sexual orientation on self-measured digit ratio (2D: 4D)â. In: Archives of sexual behavior 36.2, p. 223â 233. doi: https://doi.org/10.1007/s10508- 007-9171-6. McInnes, Leland, John Healy, and James Melville (2018). âUMAP: Uniform Manifold Approximation and Projection for Dimension Reductionâ. In: arXiv preprint. doi: https : / / doi . org / 10 . 48550 / arXiv.1802.03426. McIntyre, Matthew H (2006). âThe use of digit ratios as markers for perinatal androgen actionâ. In: Re- productive biology and endocrinology 4, p. 1â9. doi: http://dx.doi.org/10.1186/1477-7827- 4-10. Mehta, Sachin and Mohammad Rastegari (2022). âMo- bileViT: Light-weight, General-purpose, and Mobile- friendly Vision Transformerâ. In: arXiv preprint. doi: https://doi.org/10.48550/arXiv.2110. 02178. Mollineda, RamĂłn A, Karel Becerra, and Boris Mederos (2025). âSex classification from hand X-ray images in pediatric patients: How zero-shot Segment Any- thing Model (SAM) can improve medical image analysisâ. In: Computers in Biology and Medicine 197, p. 111060. issn: 0010-4825. doi: https:// doi.org/10.1016/j.compbiomed.2025.111060. url: https://w.sciencedirect.com/science/ article/pii/S001048252501412X. Nelson, Emma, Jason Hall, Patrick Randolph-Quinney, and Anthony Sinclair (2017). âBeyond size: The potential of a geometric morphometric analysis of shape and form for the assessment of sex in hand stencils in rock artâ. In: Journal of Archaeological Science 78, p. 202â213. doi: https://doi.org/ 10.1016/j.jas.2016.11.001. Nelson, Emma, John Manning, and Anthony Sinclair (2006). âNews Using the length of the 2nd to 4th digit ratio (2D: 4D) to sex cave art hand stencils: Factors to considerâ. In: Before Farming 2006.1, p. 1â7. Paszke, Adam, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala (2019). âPyTorch: An Im- perative Style, High-Performance Deep Learning Libraryâ. In: arXiv preprint. doi: https://doi. org/10.48550/arXiv.1912.01703. Permana, R, K Arifin, and I Pojoh (2015). âThe Ap- plication of Digit Ratio 2D:4D in Predicting Male- Female Hands on Prehistoric Cave Hand Stencils in Indonesiaâ. In: Asian Journal of Humanities and Social Studies (ISSN: 2321â2799) 3.05. doi: https://doi.org/10.24203/ajhss.v3i5.3249. Permana, R, Ingrid HE Pojoh, and Karina Arifin (2017). âMabedda Bola ritual in South Sulawesi; The rela- tionship between handprints in traditional house and hand stencils in prehistoric cavesâ. In: Wa- cana, Journal of the Humanities of Indonesia 18.3, 27 p. 6. doi: https://doi.org/10.17510/wacana. v18i3.633. PolcerovĂĄ, Lenka, Richard L Jantz, Miroslav KrĂĄlĂk, MĂĄria ChovancovĂĄ, and Martin Äuta (2023). âSex differences in radioulnar contrasts of the finger ridge counts across 21 human population sam- plesâ. In: Annals of human biology 50.1, p. 370â 389. doi: http://dx.doi.org/10.1080/03014460. 2023.2247970. Rabazo-RodrĂguez, Ana MarĂa, Mario Modesto-Mata, LucĂa Bermejo, and Marcos GarcĂa-DĂez (2017). âNew data on the sexual dimorphism of the hand stencils in El Castillo Cave (Cantabria, Spain)â. In: Journal of Archaeological Science: Reports 14, p. 374â381. Snow, Dean R (2006). âSexual dimorphism in Upper Palaeolithic hand stencilsâ. In: Antiquity 80.308, p. 390â404. doi: http://dx.doi.org/10.1017/ S0003598X00093704. â (2013). âSexual dimorphism in European Upper Paleolithic cave artâ. In: American Antiquity 78.4, p. 746â761. doi: http://dx.doi.org/10.7183/ 0002-7316.78.4.746. Suleiman, Muritala Odidi, B Danborno, SA Musa, JA Timbuak, AO Yusuf, and HO Suleiman (2024). âSex estimation and sexual dimorphism analysis through hand anthropometry: Insights from a cross- sectional studyâ. In: Forensic Science International: Reports 10, p. 100374. doi: https://doi.org/ 10.1016/j.fsir.2024.100374. Tan, Mingxing and Quoc Le (2019). âEfficientNet: Re- thinking Model Scaling for Convolutional Neural Networksâ. In: arXiv preprint. doi: https://doi. org/10.48550/arXiv.1905.11946. Wang, James Z, Weina Ge, Dean R Snow, Prasen- jit Mitra, and C Lee Giles (2010). âDetermining the sexual identities of prehistoric cave artists us- ing digitized handprints: a machine learning ap- proachâ. In: Proceedings of the 18th ACM interna- tional conference on Multimedia, p. 1325â1332. doi: http://dx.doi.org/10.1145/1873951. 1874214. Weiszfeld, Endre (1937). âSur le point pour lequel la somme des distances de n points donnĂ©s est min- imumâ. In: Tohoku Mathematical Journal, First Series 43, p. 355â386. Yune, Sehyo, Hyunkwang Lee, Myeongchan Kim, Sha- hein H Tajmir, Michael S Gee, and Synho Do (2019). âBeyond human perception: sexual dimor- phism in hand and wrist radiographs is discernible by a deep learning modelâ. In: Journal of Digital Imaging 32.4, p. 665â671. doi: http://dx.doi. org/10.1007/s10278-018-0148-x. 28