Paper deep dive
3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment
Alam Noor, Luis Almeida, Mohamed Daoudi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/16/2026, 3:45:51 AM
Summary
This paper introduces the 3D Sheep Pain Facial Expression System (3D-SPFES), a novel monocular depth-aware geometric graph neural network for assessing sheep pain. Unlike previous 2D methods, 3D-SPFES integrates facial landmarks (ears, eyes, nose) into 3D Euclidean space using VideoDepthAnything from single RGB camera inputs. It employs a Weighted Geometric Graph Neural Network (WG-GNN) with geometry-aware message passing and scaled dot-product attention to analyze spatial relationships and surface normals. The system achieves a validation accuracy of 73.2% and a test accuracy of 78.33%±2.2%, utilizing bilateral detection averaging and a Normalized Pain Score (NPS) derived from the Sheep Pain Facial Expression Scale (SPFES).
Entities (10)
Relation Signals (8)
Luis Almeida → affiliatedwith → CISTER Research Center
confidence 99% · Luis Almeida... CISTER Research Center
Alam Noor → affiliatedwith → CISTER Research Center
confidence 99% · Alam Noor... CISTER Research Center
Mohamed Daoudi → affiliatedwith → CISTER Research Center
confidence 99% · Mohamed Daoudi... CISTER Research Center
3D-SPFES → isbasedon → SPFES
confidence 95% · intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES)
3D-SPFES → uses → WG-GNN
confidence 95% · This paper presents the 3D Sheep Pain Facial Expression System (3D-SPFES)... A Weighted Geometric Graph Neural Network (WG-GNN) studies this graph
3D-SPFES → uses → VideoDepthAnything
confidence 95% · estimated from a single RGB camera by using VideoDepthAnything
WG-GNN → employs → scaled dot-product attention
confidence 93% · enhanced by a scaled dot-product attention method that selectively enhances anatomically relevant inter-landmark messages.
3D-SPFES → achievesaccuracy → 78.33%
confidence 92% · 3D-SPFES WG-GNN system model achieved a validation accuracy of 73.2% and an accuracy of 78.33%±2.2%
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES). This paper presents the \textbf{3D Sheep Pain Facial Expression System (3D-SPFES)}, a novel, monocular depth-aware geometric graph neural network system that integrates each SPFES facial landmark, such as the ears, eyes, and nose, into 3D Euclidean space estimated from a single RGB camera by using VideoDepthAnything, thus preventing the need for specialized depth hardware. Each landmark node includes a feature vector containing its 3D spatial coordinates, estimated surface normal, and facial attribute class embedding. Edges linked to nodes are assigned weights based on an aggregate metric that combines both Euclidean distance and surface co-planarity in a 3D space. A Weighted Geometric Graph Neural Network (WG-GNN) studies this graph using $\mathcal{K} = 3$ geometry-aware message-passing layers enhanced by a scaled dot-product attention method that selectively enhances anatomically relevant inter-landmark messages. The resultant node embeddings are combined into $\mathcal{O} = 3$ pain-level clusters and integrated into a Normalized Pain Score (NPS) within the range of $[0, 100%]$ a confidence-weighted, SPFES-derived scoring method.
Tags
Links
- Source: https://arxiv.org/abs/2608.11050v1
- Canonical: https://arxiv.org/abs/2608.11050v1
Trouble viewing inline? Open PDF directly →
Full Text
61,495 characters extracted from source content.
Expand or collapse full text
3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment Alam Noor, Luis Almeida and Mohamed Daoudi CISTER Research Center, Porto, Portugal Faculty of Engineering University of Porto, Portugal Univ. Lille, CNRS, Centrale Lille, UMR 9189 CRIStAL, Lille, F-59000, France IMT Nord Europe, Institut Mines-Télécom, Univ. Lille, Centre for Digital Systems, F-59000 Lille, France Thanks: This work was not supported by any organization Abstract Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES). This paper presents the 3D Sheep Pain Facial Expression System (3D-SPFES), a novel, monocular depth-aware geometric graph neural network system that integrates each SPFES facial landmark, such as the ears, eyes, and nose, into 3D Euclidean space estimated from a single RGB camera by using VideoDepthAnything, thus preventing the need for specialized depth hardware. Each landmark node includes a feature vector containing its 3D spatial coordinates, estimated surface normal, and facial attribute class embedding. Edges linked to nodes are assigned weights based on an aggregate metric that combines both Euclidean distance and surface co-planarity in a 3D space. A Weighted Geometric Graph Neural Network (WG-GNN) studies this graph using =3K=3 geometry-aware message-passing layers enhanced by a scaled dot-product attention method that selectively enhances anatomically relevant inter-landmark messages. The resultant node embeddings are combined into =3O=3 pain-level clusters and integrated into a Normalized Pain Score (NPS) within the range of [0,100%][0,100\%] a confidence-weighted, SPFES-derived scoring method. We additionally propose bilateral detection averaging, which integrates left and right instances of symmetric face features into a singular SPFES-consistent representative node, enhancing test accuracy by 6.7% compared to single-detection baselines. In a particular sheep facial landmark dataset, 3D-SPFES WG-GNN system model achieved a validation accuracy of 73.2% and an accuracy of 78.33%±2.2%78.33\%± 2.2\%. The proposed system can be deployed on any standard RGB camera platform, including UAVs and mobile devices, without the need for specialist sensing hardware. I Introduction In agriculture, livestock management is important for economic productivity and ethical animal care. Sheep herds are integral across multiple cultural and farming economies globally [20, 15]. A key strategy of effective herd management is the regular monitoring of animal health, with pain assessment recognized as a crucial indicator of welfare condition. Undiagnosed pain in sheep may give rise to systemic infection, decreased productivity, and considerable welfare loss. Therefore, early, non-invasive pain detection possesses significant clinical and societal benefits. A viable approach for non-invasive pain evaluation utilizes the Sheep Pain Facial Expression Scale (SPFES) [13], a validated clinical protocol that measures pain by assessing observable variations in 3 bilateral facial areas: ear position, eye expression, and nose shape. The SPFES encodes nine distinct facial expressions throughout these areas, each associated with a pain intensity score pi∈0,1,2p_i∈\0,1,2\ (indicating no pain, mild pain, or severe pain). In contrast to whole-body pose estimation, the SPFES provides a direct, anatomically-based assessment of pain level that can be obtained from standardized camera image data collected by mobile devices, such manned aerial vehicles (UAVs) or static farm cameras [18]. Previous studies on automated sheep health assessment has mainly relied on full facial detection and pose estimation employing deep learning frameworks [6, 5, 19, 11, 9, 26, 23]. Although these methods have encouraging detection abilities, algorithms consider the face as a singular input and fail to effectively utilize the spatial and semantic links among distinct facial landmarks. Therefore, they are unable to correctly represent the SPFES scoring approach, which requires individual pain evaluation per landmark, followed by overall assessment across landmarks. Moreover, all present methods function within a 2D image framework, neglecting the 3D anatomical structure of the sheep’s face, which is essential for distinguishing facial structures that may appear analogous in projection yet differ in spatial orientation, such as the posterior and bilateral regions of the ear or the nose bridge in relation to the jaw point. To address these challenges, our previous work [14] introduced a 2D Weighted Graph Neural Network (WGNN) that represents SPFES face landmarks as graph nodes and their pain-level correlations as edges, with a cluster classification accuracy of 92.71%. However, the proposed method is limited by mainly two challenges: (i) it performed primarily within the 2D image plane, disregarding depth and surface spatial orientation data, and (i) it depended on a singular camera viewpoint without any 3D geometric details, thereby reducing the predictive ability to perform edge weighting for anatomically adjacent yet structurally different facial areas. This paper proposes the 3D Sheep Pain Facial Expression System (3D-SPFES), a novel WG-GNN system model that addresses both limitations by integrating each facial landmark into 3D Euclidean space, determined by a single RGB camera utilizing a monocular depth estimation neural network. The building of a 3D graph provides geometrically accurate edge weighting, surface-normal-aware feature initialization, and equilibrium-variant message distribution. Additionally, the depth details are extracted from VideoDepthAnything [4], which results in alleviating the reliance on specialized RGB-D hardware like the Intel RealSense D435. The proposed system model is easily scalable on any standardized RGB camera computing system, which includes UAVs and mobile devices. The main scientific contribution over our previous WGNN system model [14] are the WG-GNN system model. Initially, we present a 3D graph representation of SPFES landmarks, wherein each node possesses spatial coordinates i∈ℝ3x_i ^3, surface normals i∈ℝ3d_i ^3, and embeddings for facial part features, while edge weights represent both metric proximity and normal alignment within 3D space. Secondly, we introduce a geometric attention method into the message-passing layers, allowing the network to precisely enhance messages from anatomically connected landmark pairs. Third, we present bilateral detection averaging, which aggregates left and right patterns of symmetrical face features into a singular representative node that conforms with the SPFES standard. Fourth, we use a 5-local-to-global model analysis method that aggregates the softmax predictions from five separately trained models, resulting in a Cohen’s κ of 0.473 (indicating strong concurrence) on the sheep facial expression dataset, which is a 163% increase over the single-model baseline. So, this work is the extension of our previous conference paper [14]. We propose the following 4 novel minimal contributions for sheep facial expressions of pain assessment. The main contributions of this work are as follows: • 3D facial graph representation with monocular depth. We propose a novel 3D-SPFES graph =(,ℰ,,)G=(N,E,O,X) in which each SPFES landmark node is embedded in metric 3D space using VideoDepthAnything, requiring no specialized hardware. Node features consist of 3D coordinates, projected surface normals, and embeddings of facial part categories, which provides the GNN with precise geometric structure lacking in 2D implementations. • Geometric message passing with edge-attention weighting. We propose a Weighted Geometric GNN (WG-GNN) layer that computes geometry-aware messages ij=ωij⋅ϕm(ψi,ψj,δij,i⊤j)m_ij= _ij· _m( _i, _j, _ij,d_i d_j), where ωij _ij encodes both 3D proximity and surface co-planarity and ϕm _m is a learned edge network. A scaled dot-product attention system αij _ij provides data-driven adaptation of inter-landmark message intensity levels. The analyzed attention pattern shows that the eyes node performs as the anatomical center of the face graph, corresponding with the clinical SPFES standard. • Bilateral detection averaging for SPFES-consistent landmark representation. We propose a systematic aggregation method that unifies all corresponding detections (left and right ear; left and right eye) per image, derives a bilateral centroid, and provides the maximum pain score from both instances. This optimizes test accuracy by 6.7% compared to single-detection baselines. • 3D-SPFES WG-GNN evaluation with global model design. We transform local models to a single global model with four 3D-domain augmentations, resulting in an accuracy of 78.33%±2.2%78.33\%± 2.2\%. A 5-local model integrated into a single model that estimates the softmax predictions of all local models obtains a test accuracy of 73.22% with Cohen’s κ=0.473κ=0.473 (which means mild convergence), a mild Pain F1 score of 0.580.58, and Ear κ=0.527κ=0.527, representing the most substantial improvement in the process, without additional training. The remainder of this paper is structured as follows. Section I reviews related work on sheep facial analysis, graph neural networks for behavioural biometrics, and monocular depth estimation. Section I presents the complete 3D-SPFES WG-GNN system model. Section IV describes the experimental evaluation including cross-validation, ensemble analysis, and ablation study. Section V concludes the paper and discusses directions for future work. I Related Work In the past, researchers have studied the recognition and expression of animal faces to identify their behaviors and illnesses. These animals, particularly sheep, exhibit behavior more akin to that of humans [15]. These behaviors are the result of sheep’s natural ability to evade potential threats and problems, as well as their human similarity, such as group living and the ability to recognize their mother and siblings, even after a long period of separation. These two factors have a significant impact on sheep’s behavior and responses in a variety of situations. The literature studies several deep learning approaches for sheep face classification and detection. In particular, Hao et al. [7] studied the sheep facial expression using the Single Shot MultiBox Detector (SSD) algorithm to detect the whole face of sheep and define its expressions. Ying et al. [6] presented the detection of sheep breed using DT-YOLOv5 and were capable of recognizing facial features. Zishuo et al. [5] introduced a procedure that involves two phases: identification and classification. A detection network determines if each sheep’s activity is normal, physiological, or disruptive, while traditional networks multi-scale feature aggregation, attention mechanism, and depthwise convolution module are used to let the network balance model size and detection accuracy. Farah et al. [19] presented single as well as seven-layer convolutional neural network (CNN) models with the help of centroids by UAVs to identify the sheep. Furthermore, after fine-tuning the pre-trained models, they assembled the FCN and defined models to address recall and precision issues, while Li et al. [11] combined CNNs and vision transformers to recognise the sheep faces by extracting feature representations. Hitelman et al. [9] studied face detection and classification using a biometric identification model. They applied faster R-CNN to detect the sheep’s face in an image, while pretrained models were used to identify the face into seven distinct classification models. Zhang et al. [26] used the YOLOv4 model to identify sheep and added the convolutional block attention module (CBAM) to make the extraction of model features more stable with the help of the biometric system. In another work, Zhang et al. [25] used YOLOv7 with multiple attention mechanisms to detect the sheep faces. Additionally, the same authors [24] used the YOLOv7-tiny for the multiple sheep abnormal behaviors using sheep whole body structure from a video camera. Similarly to the previous works, Bati et al. [2] studied the YOLOv5 model with SORT algorithm to detect the whole body of the sheep to identify and track the animal behaviors. Zhang et al. [28] also introduced YOLOv5 for the detection of the whole face of the sheep, while Ayub et al. [1] used YOLOv5s to classify the corresponding activity state in active and non-active for the sheep body. Zhang et al. [27] also introduced the data called multi-view sheep face images and applied the vision transform model to recognize the sheep faces with a heat map. Kelly et al. [10] presented a sheep activity dataset to analyze the behaviors of the sheep with deep learning pre-trained models. Xue et al. [22] introduced a sheep face orientation recognition algorithm to different face orientations and a feature point-matching and reconstructing the sheep face. However, the algorithm requires images from different angles to accurately recreate the sheep face. Pang et al. [17] studied the feature extraction of sheep faces using an attention residual module to aggregate the features received from the input of a spherical camera to capture image data. However, the model relies on classification, making it challenging to identify sheep faces for its application. Cai et al. [3] used convolutional neural networks for the sheep face recognition for disease prevention, while Xinyu et al. [21] presented the dilated convolutional attention module for gender identification using sheep faces as a binary classification model. The existing models, as illustrated in the literature, focus primarily on the sheep’s entire face or body. With such an approach, classifying the sheep’s facial pain is very difficult. To classify the pain, we must examine every part of the sheep’s entire face. Furthermore, the detection models only detect facial parts expression, requiring the clustering of painful parts and their separation from non-painful ones to define and combine the total face pain. Therefore, our study follows the novel sheep facial landmark dataset, proposing the use of WGNN to determine the total pain of the sheep faces. I System Model I-A Graph Formulation and 3D Spatial Representation The proposed 3D-SPFES WG-GNN system model improves upon 2D facial network analysis by integrating each facial part expression (FPE) node into 3D Euclidean space, effectively and precisely representing the anatomical geometry of the sheep face. Depth-aware perception, derived from a VideoDepthAnything vision reconstruction, assigns a spatial coordinate triplet to each identified face landmark. The comprehensive 3D-SPFES graph is shown as in (1), =(,ℰ,,),G=(N,\ E,\ O,\ X), (1) where N represents the set of nodes containing 9 facial expressions (eyes, ears, nose, cheeks, lips, jaw), ℰE represent the set of directed edges that encode spatial and pain-level relationships across the nodes, O represents the set of output state labels corresponding to pain clusters; and ∈ℝ||×3X ^|N|× 3 be the coordinate matrix where the i-th row i=(xi,yi,zi)⊤x_i=(x_i,y_i,z_i) denotes the 3D centroid of a facial part i within the camera frame. The depth dimension ziz_i is essential: it clarifies facial parts that occupy equivalent 2D image orientations (e.g., medial cortex versus lateral cortex of the eye) and offers a more robust structural prior for graph edge generation than plain pixel proximity. A prerecorded diagnosis label is utilized on the RGB channel to obtain per-part pain intensity labels pi∈0,1,2p_i∈\0,1,2\ for all 9 face parts =[p1,p2,…,p9]P=[p_1,p_2,…,p_9], where 0 represents no pain, 1 denotes mild pain, and 2 reflects severe pain. A 3D back-projection process is translated from each identified bounding box centroid from image coordinates to the 3D coordinate system utilizing the depth map of VideoDepthAnything and the camera’s intrinsic matrix K as shown in (2), i=zi⋅−1[uivi1],x_i=z_i·K^-1 bmatrixu_i\\ v_i\\ 1 bmatrix, (2) where (ui,vi)(u_i,v_i) represents the pixel centroid of the i-th face part and ziz_i represents the equivalent depth value obtained from the aligned depth frame. The surface normal i∈ℝ3d_i ^3 at each facial category, derived from the local point cluster proximity by principal component analysis (PCA) applied to the k-nearest depth points, represents the direction of the facial surface in that part. Parts on the flat nose bridge will have normals closely aligned with the camera’s visual orientation, whereas ear tips will display laterally oriented normals; this geometric differentiation is utilized in edge weighting. I-B 3D Node Feature Initialization Each node n∈n contains a feature vector ψ that performs as the input value to the geometric graph neural network. In the first 2D variant, this feature vector is limited to the scaling pain level and a categorical representation of facial-part types. The 3D node extension enhances this representation using the node’s 3D spatial coordinates and surface normal, effectively connecting the graph with explicit geometric contextual information. The in-depth feature vector for face expression part i is shown as (3), ψi(0)=[pi,FPTi,i⊤,i⊤]∈ℝ1+C+3+3, _i^(0)= [\ p_i,\ FPT_i,\ x_i ,\ d_i \ ] ^1+C+3+3, (3) where pi∈0,1,2p_i∈\0,1,2\ represent the pain intensity, FPTi∈ℝCFPT_i ^C show a one-hot or learned embedding of the facial part type (e.g., left eye, right ear, nose tip) with C categories, i=(xi,yi,zi)⊤∈ℝ3x_i=(x_i,y_i,z_i) ^3 represents the 3D spatial position, and i=(dix,diy,diz)⊤∈ℝ3d_i=(d_ix,d_iy,d_iz) ^3 indicates the unit surface normal vector. The complete initial feature dimensionality corresponds to 1+C+61+C+6. Adding ix_i directly into the feature vector, instead of exclusively applying it for edge construction, allows the graph to distinguish position-dependent pain correlations. For instance, a pain score of 2 at the face might exhibit significant clinical repercussions with respect to the equivalent score at the orbital region, with the spatial coordinate providing a distinguishing parameter for the learned transformation matrices. The surface normal id_i adds nodes with proximate 3D coordinates but divergent anatomical orientations, such as the inner and outer jaw borders. I-C 3D Edge Construction with Geometric Weighting In the early 2D WGNN system model [14], edges are assigned according to facial morphology proximity and pain-score symmetry. The 3D WG-GNN system combines the qualitative proximity criterion with a quantitative Euclidean distance threshold and incorporates a normalization factor into the edge weight, resulting in a geometrically accurate connectivity structure. An edge e=(n,w)∈ℰe=(n,w) forms between nodes n and w if and only if both of the conditions δnw=‖n−w‖2≤τ _nw=\|x_n-x_w\|_2≤τ and |pn−pw|≤ϵ|p_n-p_w|≤ε are satisfied simultaneously. While, the initial state requires anatomical proximity in 3D space with a threshold τ, providing that only physically neighboring facial parts are considered as structurally correlated. The second prerequisite specifies pain-level regularity with tolerance ϵ∈0,1,2ε∈\0,1,2\, categorizing facial parts showing an analogous degree of pain. Both conditions must be at the same time observed: two anatomically distinct areas (e.g., left ear and right ear) are unlikely to be linked despite identical pain scores, and two adjacent areas with considerably different pain scores will not be connected even if they are in close proximity. After the construction of an edge, its scaling weight ωnw _nw assesses the degree of strength of the relationship by integrating a radial basis function (RBF) of 3D space with a normal-alignment attribute, as shown in (4), ωnw=exp(−δnw22σ2)⏟proximity term⋅n⊤w+12⏟normal alignment term. _nw= \! (- _nw^22σ^2 )_proximity term· d_n d_w+12_normal alignment term. (4) The proximity factor is a Gaussian kernel centered at zero distance, constrained by bandwidth σ: edges between precisely positioned nodes (small δnw _nw) acquire weights equivalent to 1.0, but edges adjacent to the threshold τ are attributed weights that gradually decrease toward 0. The norm alignment metric is the normalized cosine similarity of surface normals, transformed from [−1,+1][-1,+1] to [0,1][0,1]: co-planar facial parts (parallel normals, n⊤w≈1d_n d_w≈ 1) attain correlation scores close to orthogonally oriented facial parts receive scores near 0.5, and anti-parallel facial parts (physically infeasible for a convex surface but incorporating a for robustness) receive scores near 0. The multiplicative coupling of these two factors shows that a structurally key edge implies both proximity and co-planarity, which is effective in preventing defective connections across anatomical edges, such as the transition from jaw to ear. I-D Optimal 3D Parse Graph Inference For each facial part type (FPT), a parse graph g=(g,ℰg,g)g=(N_g,E_g,O_g) is inferred as a sub-graph of G such that g⊆N_g and ℰg⊆ℰE_g . The 3D formulation augments the inference objective with the coordinate matrix X, so that the optimal parse graph g∗g^* is defined as shown in (5), g∗=argmaxgPd(g∣g,ℰg,ψ,)⏟labelling probability⋅Pd(g,ℰg∣ψ,,)⏟structural probabilityg^*= _g\ P_d\! (O_g _g,\ E_g,\ ψ,\ X )_labelling probability· P_d\! (N_g,\ E_g ψ,\ X,\ G )_structural probability (5) where, Pd(g∣g,ℰg,ψ,)P_d(O_g _g,E_g,ψ,X), is the labeling probability: given the graph structure and all node/edge features, including 3D positions. While, Pd(g,ℰg∣ψ,,)P_d(N_g,E_g ψ,X,G) represents the structural probability, assuming the feature-augmented 3D graph G, what is the likelihood of this specific sub-graph g serving as an explanation for the observed facial expression? By conditioning both factors on X, the inference simultaneously assesses the degree to which the proposed parse graph is geometrically consistent with the sheep’s true facial anatomy. For instance, parse graphs that attribute the same pain cluster to anatomically disparate facial parts exclusively due to their similar pain scores. I-E Geometric Message Passing with Attention The key computational unit is a weighted geometric graph neural network (WG-GNN) that evaluates the 3D-SPFES graph via K successive message-passing layers. Compared to 2D GNNs that define edge weights as scaling matrices multiplying a predefined aggregation function, the proposed system model employs a learned edge network ϕm(⋅) _m(·)—realized as a multi-layer perceptron (MLP), which sequentially analyzes node features, inter-node distance, and normal alignment to generate geometry-aware messages. The message from adjacent node j to node i at layer k is calculated as shown in (6), ij(k)=ωij⋅ϕm(ψi(k−1),ψj(k−1),δij,i⊤j)m_ij^(k)= _ij· _m\! ( _i^(k-1),\ _j^(k-1),\ _ij,\ d_i d_j ) (6) where ωij _ij represents the precomputed geometric edge weight, whereas ϕm:ℝ2dψ+2→ℝdh _m:R^2d_ψ+2 ^d_h transforms the concatenated feature pair, scalar distance, and scalar normal alignment into a hidden-dimensional message vector. The scalar δij _ij and the dot product i⊤jd_i d_j can be given as explicit inputs to ϕm _m to allow the system to learn message transformations that are dependent on distance and orientation, surpassing the information encoded solely by the static weight ωij _ij. A scaled dot-product attention method is applied to enable the system to assign differential weights to communication from various neighbors based on learned feature compatibility, rather than relying exclusively on predetermined geometric weights. Projections for queries and keys are derived from the existing features of each node as shown in (7), αij(k)=softmaxj((q(k)ψi(k−1))⊤(k(k)ψj(k−1))dh), _ij^(k)=softmax_j\! ( (W_q^(k)\, _i^(k-1) ) (W_k^(k)\, _j^(k-1) ) d_h ), (7) where q(k),k(k)∈ℝdh×dψW_q^(k),W_k^(k) ^d_h× d_ψ are the query and key projection matrices for layer k, and the softmax is computed over all neighbors j∈(i)j (i). The 1/dh1/ d_h scaling factor prevents the dot products from growing large in magnitude as dhd_h increases, which would push the softmax into regions of extremely small gradients. While the aggregate message at the node i in the layer k is the attention-weighted sum of all incoming geometric messages as given in (8), ℳi(k)=∑j∈(i)αij(k)⋅ij(k)M_i^(k)= _j (i) _ij^(k)·m_ij^(k) (8) This formulation integrates the structural prior shown by ωij _ij, which adjusts the magnitude of each message, with the data-driven attention αij(k) _ij^(k), which selectively enhances or reduces messages based on learned feature compatibility, resulting in more in-depth aggregation than either process separately provides. The node feature is subsequently updated using a linear transformation following by a non-linear activation as given in (), ψi(k)=ℱ(σ)((k)ℳi(k)+(k)), _i^(k)=F(σ)\! (W^(k)\,M_i^(k)+b^(k) ), (9) where (k)∈ℝdψ×dhW^(k) ^d_ψ× d_h is the layer-k weight matrix, (k)∈ℝdψb^(k) ^d_ψ is the bias vector, and ℱ(σ)F(σ) denotes the ReLU activation function. After K such layers, the final node embeddings ψi()i∈\ _i^(K)\_i are passed to the clustering and pain-scoring module. I-F 3D Cluster-Based Pain Score Computation In addition to message passing, the GNN gives O cluster assignments (1,2,…,o)(C_1,C_2,…,C_o) that group the nine facial part expressions based on their pain level and geometric contextual information. The total pain score for cluster jC_j is estimated as a weighted average, with each face part’s determination adjusted by its SPFES-derived impact weight ωi _i and a depth-confidence weight ρi _i as shown in (10), j=∑i∈jωi⋅pi⋅ρi∑i∈jωi,S_j= _i _j _i· p_i· _i _i _j _i, (10) where ρi=‖i‖∈[0,1] _i=\|d_i\|∈[0,1] is the surface normal magnitude, which serves as a proxy for depth estimation confidence: nodes whose local point cloud neighborhood yielded a stable, well-conditioned normal estimate (high ‖i‖\|d_i\|) contribute with full weight, while nodes with poorly estimated normals (e.g., occluded ear tips or near-boundary pixels) are down-weighted automatically. This confidence-aware formulation is a direct consequence of the 3D representation and has no analog like in the WGNN 2D model. While the total pain score across all clusters is then obtained as a weighted sum where cluster-level weights wjw_j reflect the relative importance of each cluster according to the SPFES standards as given in (11), p=∑j=1owj⋅jT_p= _j=1^ow_j·S_j (11) Finally, the Normalized Pain Score (NPS) maps pT_p to the interval [0,100%][0,100\%] by dividing by the maximum achievable pain score maxT_ that is the value obtained when all facial parts exhibit pi=2p_i=2 with full depth confidence is given in (12), =(pmax)×100.NPS= ( T_pT_ )× 100. (12) I-G Joint Loss Function with Geometric Regularization The proposed WG-GNN system model is trained in an end-to-end way by minimizing a joint loss that integrates cluster-classification cross-entropy with a geometric regularization factor. Let p∈Δ⟨o⟩U_p∈ o represent the true one-hot probability distribution across o clusters, and let q∈Δ⟨o⟩V_q∈ o represent the predicted softmax distribution. The cross-entropy loss function is defined as in (13), ℒCE=−∑i=1op(i)logq(i)L_CE=- _i=1^oU_p(i)\, \,V_q(i) (13) A geometric regularization term ℒgeoL_geo is added to enable resemblance in the final node embeddings of geometrically proximate nodes, such as those linked by highly weighted edges, to be similar in the learned feature space. Nodes connected by a robust edge (high ωnw _nw) should possess similarities of final embeddings, whereas nodes linked by a weak edge (low ωnw _nw, approaching the threshold τ) are not required to exhibit similarity as given in (14), ℒgeo=∑(n,w)∈ℰωnw−1⋅‖ψn()−ψw()‖22L_geo= _(n,w) _nw^-1· \| _n^(K)- _w^(K) \|_2^2 (14) The inverse weighting ωnw−1 _nw^-1 shows that the penalty for dissimilar embeddings is greatest for the most geometrically significant edges (high ωnw _nw, low ωnw−1 _nw^-1). It is important to observe that we minimize a ωnw−1 _nw^-1 from above to avoid causing instability in the case of very weak edges. The total training loss (ℒL) is shown in (15), ℒ=ℒCE+λℒgeoL=L_CE+λ\,L_geo (15) where λ>0λ>0 is a hyperparameter controlling the trade-off between discriminative classification accuracy and geometric smoothness of the learned embedding space. While λ is tuned on a validation set, values in the range [0.01,0.1][0.01,0.1] are recommended as a starting point. IV Experimental Work We conducted experiments on the sheep facial landmarks dataset developed using SPFES parameters [13] to assess sheep facial pain using the proposed 3D-SPFES WG-GNN system model. The dataset is divided into training, testing, and validation images. We meticulously labelled each image of the nose, ears, and eyes with bounding boxes to detect the level of pain, with bilateral annotations for both the left and right ears and both left and right eyes. We trained the model using PyTorch on a GPU-equipped workstation with a GeForce RTX A3000. The ResNet-18 visual backbone was initialized with ImageNet pretrained weights and fine-tuned end-to-end with a learning rate of 5×10−55× 10^-5, while the GNN head used a learning rate of 5×10−45× 10^-4. We utilize the AdamW optimizer with weight decay of 5×10−45× 10^-4 and employ a class-weighted cross-entropy loss with label smoothing (ϵ=0.1ε=0.1) combined with geometric regularization and visual consistency losses. Every model was trained for 200 epochs per fold using stratified 5-fold cross-validation to ensure robust performance estimation on this limited dataset. IV-1 Dataset We developed the sheep facial landmarks dataset with bounding boxes using high-resolution images from well-known sources according to the SPFES standard [13]. These sources are Mendeley11 1 http://dx.doi.org/10.17632/y5sm4smnfr.5 and the study carried out in [16, 14]. We use the scale to evaluate expression in three bilateral facial parts: ears positioning, eyes constraint, and nose posture. We evaluate these part expressions based on the presence or absence of abnormal expression, categorizing them as not present (pain 0), slightly present (pain 1), or substantial (pain 2), as shown in Fig. 1. Since the SPFES protocol scores bilateral features as single indicators, our pipeline automatically averages left and right detection centroids and takes the maximum pain score across both instances during preprocessing. Fig. 1: Sheep facial landmarks with bounding boxes for the evaluations in which Ear-Flat (EF), Ear-Flipped (EFP), Ear-Rotated (ER), Eyes-FullyOpen (EyE), Eyes-NotClassifiable (EyN), Eyes-PartlyClose (EyPC), Nose-ExtendedV (NExV), Nose-ShallowU (NSU), Nose-NotClassifiable (NNC), and Nose-ShallowV (NSV) are defined to classify different pain levels and their corresponding positions. We conducted 4 domain-specific augmenting methods during training to avoid overfitting on this constrained dataset: (i) 3D position jitter (Gaussian noise with σ=0.05σ=0.05 m on back-projected coordinates), (i) horizontal axis flipping to reproduce variations in head orientation, (i) depth scale jitter (multiplicative noise of ±20%± 20\% on the z-coordinate), and (iv) facial part feature dropout (with a probability of 0.30.3, randomly eliminating the visual features of one part) to improve cognitive resiliency toward partial occlusion. These augmentation algorithms connect exclusively to the 3D graph domain and enhance the normal image-level augmentations used during GNN detector training. IV-2 Model Configuration The proposed 3D-SPFES WG-GNN consists of a shared ResNet-18 visual backbone that extracts 256-dimensional feature vectors from cropped facial regions, along with geometric features (facial part type, 3D position, surface normal, and detection confidence) to build a 266-dimensional multi-modal node feature. The WG-GNN performs on the 3D graph with =3K=3 message-passing layers, each with a hidden dimension of dhidden=64d_hidden=64. Graph edges are given weights utilizing a radial basis function (RBF) kernel with parameters σ=1.5σ=1.5 m and τ=2.0τ=2.0 m. The model is trained end-to-end utilizing AdamW with differential learning rates (5×10−55× 10^-5 for the backbone and 5×10−45× 10^-4 for the GNN head), using class-weighted cross-entropy loss with label smoothing, coupled with auxiliary geometric and visual consistency regularization parameters. IV-3 Sheep Facial Depth Map Fig. 2 shows some of the predictions of VideoDepthAnything to analyze the pain level and assessment of sheep faces. Each panel displays the original RGB image with alongside the corresponding depth map estimated by VideoDepthAnything. The depth maps provide the 3D spatial context that enables geometric edge weighting in the proposed WG-GNN. Fig. 2: Sheep facial expression alongside the corresponding monocular depth map (right) estimated by VideoDepthAnything. IV-4 3D Depth Estimation and Graph Construction Figure 3 shows the approach in which the proposed 3D-SPFES WG-GNN system model encodes the entire sheep face in a three-dimensional graph network structure. The monocular depth map created by VideoDepthAnything [4] is analyzed at uniform grid intervals, with each sampled depth point behaving as a mesh node connected to its spatial neighbors by slender edges. This makes 3D facial topology that preserves the anatomical structure of the sheep’s face. The 3 SPFES facial landmarks (ear, eyes, and nose) are shown as key graph nodes on this surface, interconnected by learned attention edges (α=1.00α=1.00). This graphic shows that the proposed system accurately constructs and analyzes over a 3D graph in metric Euclidean space, as in contrast to existing within the 2D image plane [14]. Fig. 3: A full face’s depth map rendered as a 3D graph network structure. Each sampled depth point forms a mesh node (thin edges), creating a 3D face topology. Figure 4 shows 3D depth terrain representations for various test images, applying three rendering processes. The rainbow depth surface (a) provides a height-mapped color map to represent 3D facial geometry, with the nose and jaw region surfacing as the salient pinnacle and the ears forming bilateral elevated shapes. The golden terrain (b) shows the graph node positions as higher points on the depth surface, presenting a concise view of the spatial distance between facial landmarks. The pain-zone surface (c) integrates the original RGB texture with pain-level coloring determined by proximity to each analyzed facial feature, resulting in a visual pain heatmap overlaying the 3D facial geometry. The visualizations confirm that VideoDepthAnything generates geometrically deep maps that precisely represent the unique 3D geometry of the sheep’s face. (a) (b) (c) Fig. 4: 3D depth terrain representations of accurate test images: (a) Rainbow depth surface for a frontal pain sheep (NPS = 25.0%), showing the special facial geometry with mild ear physical discomfort and severe pain in the nose and eyes. Golden terrain, including graph nodes, rises at face landmark regions. (c) Pain-zone surface for a mild-pain case (NPS =25.0%=25.0\%), with the surface colored according to relative proximity to each pain-classified facial area (orange == mild pain). IV-5 Evaluation Results Figure 5 shows the evaluation results of the proposed 3D-SPFES WG-GNN system model. Each figure shows associated facial landmarks with colored bounding boxes (orange for ears, green for eyes, and magenta for nose), expression class labels, and a triangle graph construction connecting all 3 facial morphology parts. The graph edges are shown with 3D depth-sensitive thickness, where closer facial features give thicker edges, visually representing the spatial depth correlations extracted from VideoDepthAnything. The graph nodes are described by their predicted pain cluster (green indicating no to low pain, orange signifying mild pain), with the NPS displayed at the bottom of each image. These examples are not utilized in training and show the model’s ability to generalize across several sheep breeds, head orientations (frontal and lateral), age categories (including lambs), and operational environments. Fig. 5: Evaluation results of the proposed 3D-SPFES WG-GNN system model on accurately identified test images. Each image presents identified facial landmarks along with bounding boxes, expression labels, and a 3D depth-aware graph structure linking the ears, eyes, and nose. Edge thickness is correlated with depth proximity. Node colors represent anticipated pain clusters: green (no to low pain) and orange (mild pain). IV-6 Global Model Facial Pain Assessment We analyze the proposed 3D-SPFES WG-GNN system model, applying both a single 3D-SPFES WG-GNN local model and a 5-local-to-global 3D-SPFES WG-GNN model integration approach that averages the softmax probability distributions of all five local models. Table I presents a thorough comparison. The 3D-SPFES WG-GNN global model gives a significant gain in all consensus evaluation metrics without requiring further training. Cohen’s κ rises from 0.180 (slight consistency) to 0.473 (moderate consistency), indicating a 163% relative increase. The ear part shows significant enhancement, increasing from κ=0.098κ=0.098 to κ=0.527κ=0.527, but the eyes part improves from a negative κ=−0.053κ=-0.053 to a positive κ=0.470κ=0.470. The improvements result from each 3D-SPFES WG-GNN local model learning different decision boundaries due to varying training partitions and augmentation variability, with the averaging of the five softmax distributions mitigating individual prediction errors at the No Pain / Mild Pain boundary. Method Acc.(%) W-F1 κ MCC M-F1 Local 76.67 64.5 0.180 0.214 35.5 Global 78.33 78.0 0.473 0.486 49.3 Δ +1.7 +13.5 +0.293 +0.272 +13.8 TABLE I: 3D-SPFES WG-GNN local model vs. 3D-SPFES WG-GNN global model on the held-out test set. The detailed breakdown of the 3D-SPFES WG-GNN global model is presented in Table I, which gives details on the accuracy of each individual portion and pain level. 80% of the accuracy is achieved in the ear and eye areas, while 75% is achieved in the nose. The recall rate for no pain is 95.2%, and the diagnosis of mild pain is made with a precision of 78 Subset Correct / Total Accuracy (%) Per Facial Part Ear 16 / 20 80.0 Eyes 16 / 20 80.0 Nose 15 / 20 75.0 Overall 47 / 60 78.33 Per Pain Level No Pain (0) 40 / 42 95.2 Mild (1) 7 / 15 46.7 Severe (2) 9 / 12 40.01 TABLE I: 3D-SPFES WG-GNN global test accuracy by facial part and pain level. Table I presents the full per-class classification report for the 3D-SPFES WG-GNN global model across all 3 facial parts. Part Pr0 R0 Pr1 R1 κ MCC Overall 0.85 0.95 0.78 0.47 0.473 0.486 Ear 0.87 0.93 0.75 0.60 0.527 0.531 Eyes 0.82 1.00 1.00 0.40 0.470 0.517 Nose 0.87 0.93 0.67 0.40 0.422 0.430 TABLE I: 3D-SPFES WG-GNN global model classification report. Precision (Pr) and recall (R) of (Pr0/R0) = No Pain; (Pr1/R1) = Mild Pain. IV-7 Discriminative Performance Table IV focuses on the Area Under the Curve (AUC) and Average Precision (AP) metrics. The nose area demonstrates superior discriminative performance for both no pain (AUC =0.881=0.881, AP =0.940=0.940) and mild pain (AUC =0.733=0.733, AP =0.637=0.637), affirming its status as the most informative facial indication for pain evaluation. The Ear Mild Pain AP of 0.642 shows the positive effect of bilateral averaging in detecting ear-position variations correlated with pain. The eyes’ mild AP of 0.347 correlates with the diagnostic difficulties of discriminating partial eye closure from a fully opened condition. Part AUC AP P0 P1 P2 P0 P1 P2 Overall 0.735 0.656 0.404 0.860 0.507 0.055 Ear 0.631 0.760 0.316 0.758 0.642 0.071 Eyes 0.738 0.453 0.579 0.897 0.347 0.111 Nose 0.881 0.733 0.316 0.940 0.637 0.071 TABLE IV: ROC-AUC and Average Precision per facial part and pain class. P0 = No Pain, P1 = Mild, P2 = Severe. IV-8 Per-Expression Accuracy A classification accuracy for each of the 10 SPFE expression categories seen in the test set. No-pain expressions prevail, with ears flat at 93%, eyes fully open at 81%, and noses shallow and U-shaped at 92%. The Nose-NotClassifiable class shows 60% accuracy after the use of the bilateral detection fix, a result absent in previous single-detection processing test results. Pain-indicative expressions show decreases in individual accuracy due to the limits on sample sizes (30–40 test datasets each); however, the 3D-SPFES WG-GNN global model achieved a non-zero mild recall at the aggregate level (47%) as multiple low-confidence correct predictions across expressions contribute when softmax probabilities are averaged. IV-9 GNN Graph Structure and Edge Attention Figure 6 shows the averaged attention weight matrix αij _ij and the geometric weight matrix ωij _ij across 20 test images. The attention pattern shows that the Eyes node represents the focal point of the facial graph: Ear→ , Eyes→ , Eyes→ , and Nose→ all exhibit α=1.0α=1.0, whereas Ear↔ connections are inactive. This hub-and-spoke layout corresponds with the clinical SPFES observation that eye expressions correlate with both auricular and nose pain symptoms. The geometric weights ω=0.013ω=0.013 validate that the RBF bandwidth σ=1.5σ=1.5 m provides significant distance-dependent edge weighting using VideoDepthAnything’s metric depth scale. Fig. 6: Learned attention weights αij _ij (left) and geometric weights ωij _ij (right) averaged over 20 test images. The Eyes node serves as the anatomical hub, receiving and distributing messages between the ear and nose. IV-10 Normalised Pain Score (NPS) Analysis This experiment evaluates the use of the proposed 3D-SPFES WG-GNN global model to compute the Normalized Pain Score (NPS), a persistent pain rating within the range of [0,100%][0,100\%], derived from the weighted aggregation of cluster pain ratings from all three face parts. Figure 7 presents the distribution of NPS defined by the highest whole-face pain level. Table V contains the NPS details. No-pain images focus primarily at an NPS of approximately 0%, whereas mild-pain images range from 8% to 17%, with a median of 13.5%. The only severe-pain image returns an NPS of 11.1%, positioning it inside the mild-pain range due to a model mistake. The substantial divergence between no-pain and mild-pain NPS values indicates that the NPS offers a clinically important continuous assessment of pain intensity. Fig. 7: NPS distribution by maximum whole-face pain level (left) and raw pain score distribution per facial part (right). No-pain images cluster at NPS ≈0%≈ 0\%; mild-pain images span 8–17%. Pain Level n Mean Std Min Max No Pain (0) 14 0.0% 0.0% 0.0% 0.0% Mild (1) 5 13.5% 3.2% 8.3% 16.7% Severe (2) 1 11.1% — 11.1% 11.1% Overall 20 3.89% 6.3% 0.0% 20.0% TABLE V: Normalised Pain Score (NPS) statistics on the sheep images. IV-11 Model Calibration Figure 8 presents the reliability graphs for each pain cluster classification. The No Pain calibration curve has a nearly monotonic trend proximal to the ideal calibration diagonal, signifying that when the model forecasts No Pain with a certain confidence, that confidence is accurately correlated with the actual result. The mild pain calibration exhibits a logical trend at intermediate probability (0.3–0.5). The calibration of severe pain is nearly 0%, correlating with the model’s performance, as we found no severe samples from the no-pain class during testing on a specific batch of sheep images. Fig. 8: Calibration curves (reliability diagrams) for No Pain, Mild Pain, and Severe Pain. The dashed diagonal represents perfect calibration. IV-12 Ablation Study Table VI presents the ablation study showing the impact of each design possibility over seven incremental model designs. We develop such setups by systematically integrating system model sections into the baseline, thereby isolating the impact of each alterations. ID Key Change Acc. κ M-F1 C1 Circular baseline 100.0 1.000 1.00 C2 Correct task 68.3 0.062 0.20 C3 + Weights [1:4:12] 38.3 0.044 0.43 C4 + Weights [1:2:4] 65.0 0.007 0.18 C5 + Reduced model 65.0 0.007 0.18 C6 + Local Models + bilateral 76.7 0.180 0.22 C7 + Gloabl 78.3 0.473 0.58 TABLE VI: Ablation study of D-SPFES WG- GNN global model progressive impact of design decisions on test performance. Acc. = accuracy (%), κ = Cohen’s Kappa, M-F1 = Mild Pain F1-score. The ablation shows 3 major processes. The change from C1 to C2 shows the complexity of the 3D-SPFES WG-GNN system model by eliminating the circular pain-score leakage (where the pain score served as both an input feature and the cluster label), resulting in a decrease in accuracy from an apparent 100% to an actual 68.3%. The shift from C5 to C6 represents the integrated impact of local 3D-SPFES WG-GNN models and bilateral detection averaging, resulting in an accuracy enhancement of 11.7 percentage points and an insignificant recall gain from 13.3% to 46.7%. The transition from C6 to C7 shows that the global 3D-SPFES WG-GNN model approach results in the most significant Kappa increase (+0.293), enhancing the model from slight agreement (κ=0.180κ=0.180) to mild agreement (κ=0.473κ=0.473) and nearly tripling the mild pain F1-score from 0.22 to 0.58 without retraining. IV-13 State-of-the-art (SOTA) Comparison We compared the efficacy of 3D-SPFES against the state-of-the-art models analyzed in the prior study [15] and other baseline comparisons, as shown in Table VII. The comparison consists of the initial 2D WGNN model [15] and the proposed 3D-SPFES, accompanied by a global analysis. The reported 92.71% accuracy of the 2D WGNN resulted from a symmetrical task design where the pain score contributed as both an input feature and a prediction goal, hence imposing certain limits on the results. The proposed 3D-SPFES eliminates this limitation by excluding the pain score from the input features and using the total face pain level using a depth map as the cluster label for every facial part, hence enhancing accuracy and reliability. Model Train Acc. Test Acc. κ SVM [12] 71.55% 62.08% — CNN [3] 79.15% 78.56% — CCVT [11] 85.45% 83.33% — EfficientNet [8] 89.33% 86.60% — 2D-WGNN [15]† 92.71% 91.96% 1.000† 3D-SPFES WG-GNN (ours) 78.33± 2.2% 73.22% 0.473 TABLE VII: Comparison with SOTA models. †The 2D-WGNN result used a circular task formulation (pain score in input == prediction target), resulting in certain limits. The 3D-SPFES result uses a corrected, non-trivial task formulation with a depth map and global modeling and Cohen’s κ as the primary agreement metric. Direct numerical comparison between the 3D-SPFES WG-GNN system model and prior state-of-the-art models is not straightforward due to differing working formulations. The prior models performed the per-node classification using the individual part’s pain score as the ground truth input, whereas the proposed 3D-SPFES WG-GNN system model predicts the overall facial depth pain level from 3D geometry and visual features using depth map for each part, representing a significantly more challenging and practical task. The Cohen’s κ=0.473κ=0.473 (indicating mild agreement) obtained by the 3D-SPFES WG-GNN global model represents the initial robust, non-circular assessment of GNN-based sheep pain evaluation on the sheep facial expression dataset. IV-14 Limitations of Study The efficacy of the proposed 3D-SPFES WG-GNN system model is constrained by 3 main limitations. The severe pain category is significantly limited in the dataset, with insufficient test samples, limiting the meaningful identification of pain level 2. Adding to the dataset to a minimum of 450 verified severe-pain samples for each facial part represents the most notable gain. Secondly, the NPS for the individual severe-pain test image (11.1%) falls within the mild-pain spectrum (8.3–16.7%), which shows that the cluster-weight parameterization has not yet achieved a strictly monotonic severity hierarchy. Therefore, a recalibration against validated veterinary pain evaluations is an imminent priority. Third, although the geometric weights ω=0.013ω=0.013 show significant enhancement due to model performance, they are still considerably lower than the learned attention weights (α=1.0α=1.0), showing that the integration of camera intrinsic calibration and dense surface reconstruction could further enhance the geometric aspect of the message-passing mechanism. Additionally, we plan on obtaining additional data from diverse environments featuring various sheep breeds, colors, and sizes to ensure generalizability and to enhance the SPFES evaluation framework with lamb-specific facial features. V Conclusion This study presented the integration of a 3D Weighted Geometric GNN (WG-GNN) model with a VideoDepthAnything monocular depth estimator, aimed at identifying clusters of facial expressions and assigning a Normalized Pain Score (NPS) to the facial expressions of sheep without the necessity for specialized depth-sensing devices. We proposed the framework to detect facial landmarks and estimate their 3D coordinates via monocular depth estimation. We presented bilateral detection averaging to effectively integrate symmetric facial parts and a WG-GNN using geometric message passing and attention mechanisms to form pain-level clusters. The obtained attention validated that the eye node acts as the anatomical center of the facial graph. A global method attained a held-out accuracy of 78.33%, with Cohen’s κ=0.473κ=0.473 (indicating mild agreement) and MCC =0.486=0.486, while each of the 3 facial regions independently achieved κ>0.42κ>0.42. The nose area reportedly had the highest discriminative performance (AUC = 0.881, AP = 0.940). The NPS attained complete distinction between pain-free and mildly painful lambs, validating its clinical usefulness. Ablation research using 6 setups validated the distinct impact of each design approach, with the global model alone achieving a 163% increase in Kappa. The proposed system model is scalable with any standard RGB camera platform and is relevant to a wider range of facial behavioral biometrics, including livestock pain assessment and clinical pain monitoring. References [1] M. Y. Ayub, A. Hussain, M. F. U. Hassan, B. Khan, F. A. Khan, D. Al-Jumeily, and W. Khan (2023) A non-restraining sheep activity detection and surveillance using deep machine learning. In 2023 16th International Conference on Developments in eSystems Engineering (DeSE), Vol. , p. 66–72. External Links: Document Cited by: §I. [2] C. T. Bati and G. Ser (2024) Improved sheep identification and tracking algorithm based on yolov5 + sort methods. Signal, Image and Video Processing 18 (10), p. 6683–6694. External Links: Document, Link, ISSN 1863-1711 Cited by: §I. [3] Z. Cai, M. Chen, R. Jing, and Y. Zhang (2024) Sheep face detection and disease prevention based on convolutional neural network. In Proceedings of the 2023 7th International Conference on Electronic Information Technology and Computer Engineering, EITCE ’23, New York, NY, USA, p. 1501–1505. External Links: ISBN 9798400708305, Link, Document Cited by: §I, TABLE VII. [4] S. Chen, H. Guo, S. Zhu, F. Zhang, Z. Huang, J. Feng, and B. Kang (2025) Video depth anything: consistent depth estimation for super-long videos. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 22831–22840. External Links: Document Cited by: §I, §IV-4. [5] Z. Gu, H. Zhang, Z. He, and K. Niu (2023) A two-stage recognition method based on deep learning for sheep behavior. Computers and Electronics in Agriculture 212, p. 108143. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I. [6] Y. Guo, Z. Yu, Z. Hou, W. Zhang, and G. Qi (2023) Sheep face image dataset and dt-yolov5s for sheep breed recognition. Computers and Electronics in Agriculture 211, p. 108027. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I. [7] M. Hao, Q. Sun, C. Xuan, X. Zhang, M. Zhao, and S. Song (2024) Lightweight small-tailed han sheep facial recognition based on improved ssd algorithm. Agriculture 14 (3). External Links: Link, ISSN 2077-0472, Document Cited by: §I. [8] G. M. S. Himel, Md. M. Islam, and M. Rahaman (2024) Utilizing efficientnet for sheep breed identification in low-resolution images. Systems and Soft Computing 6, p. 200093. External Links: ISSN 2772-9419, Document, Link Cited by: TABLE VII. [9] A. Hitelman, Y. Edan, A. Godo, R. Berenstein, J. Lepar, and I. Halachmi (2022) Biometric identification of sheep via a machine-vision system. Computers and Electronics in Agriculture 194, p. 106713. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I. [10] N. A. Kelly, B. M. Khan, M. Y. Ayub, A. J. Hussain, K. Dajani, Y. Hou, and W. Khan (2024) Video dataset of sheep activity for animal behavioral analysis via deep learning. Data in Brief 52, p. 110027. External Links: ISSN 2352-3409, Document, Link Cited by: §I. [11] X. Li, Y. Xiang, and S. Li (2023) Combining convolutional and vision transformer structures for sheep face recognition. Computers and Electronics in Agriculture 205, p. 107651. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I, TABLE VII. [12] Y. Lu, M. Mahmoud, and P. Robinson (2017) Estimating sheep pain level using facial action unit detection. In 2017 12th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2017), Vol. , p. 394–399. External Links: Document Cited by: TABLE VII. [13] K. M. McLennan, C. J.B. Rebelo, M. J. Corke, M. A. Holmes, M. C. Leach, and F. Constantino-Casas (2016) Development of a facial expression scale using footrot and mastitis as models of pain in sheep. Applied Animal Behaviour Science 176, p. 19–26. External Links: ISSN 0168-1591, Document, Link Cited by: §I, §IV-1, §IV. [14] A. Noor, L. Almeida, M. Daoudi, K. Li, and E. Tovar (2025) Sheep facial pain assessment under weighted graph neural networks. In 2025 IEEE 19th International Conference on Automatic Face and Gesture Recognition (FG), Vol. , p. 1–9. External Links: Document Cited by: §I, §I, §I, §I-C, §IV-1, §IV-4. [15] A. Noor, M. J. Corke, and E. Tovar (2023) Sheep health behavior analysis in machine learning: a short comprehensive survey. Smart Agricultural Technology 6, p. 100366. External Links: ISSN 2772-3755, Document, Link Cited by: §I, §I, §IV-13, TABLE VII. [16] A. Noor, Y. Zhao, A. Koubaa, L. Wu, R. Khan, and F. Y.O. Abdalla (2020) Automated sheep facial expression classification using deep transfer learning. Computers and Electronics in Agriculture 175, p. 105528. External Links: ISSN 0168-1699, Document, Link Cited by: §IV-1. [17] Y. Pang, W. Yu, Y. Zhang, C. Xuan, and P. Wu (2023) An attentional residual feature fusion mechanism for sheep face recognition. Scientific Reports 13 (1), p. 17128. External Links: Document, Link, ISSN 2045-2322 Cited by: §I. [18] N. Sarantinoudis, G. Arampatzis, K. P. Valavanis, and N. Tsourveloudis (2023) Unmanned aerial vehicles and livestock management: an application in western crete. In 2023 International Conference on Unmanned Aircraft Systems (ICUAS), Vol. , p. 159–166. External Links: Document Cited by: §I. [19] F. Sarwar, A. Griffin, S. U. Rehman, and T. Pasang (2021) Detecting sheep in uav images. Computers and Electronics in Agriculture 187, p. 106219. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I. [20] A. E. O. Smith, C. Doidge, T. Knific, F. Lovatt, and J. Kaler (2024) The tales of contradiction: a thematic analysis of british sheep farmers’ perceptions of managing sheep scab in their flocks. Preventive Veterinary Medicine 227, p. 106194. External Links: ISSN 0167-5877, Document, Link Cited by: §I. [21] Z. Xinyu, T. Zhenzhen, Y. Wei, L. Lei, and W. Jihua (2023) DCAM-net: sheep gender identification network based on dilated convolutional attention module. In 2023 13th International Conference on Information Technology in Medicine and Education (ITME), Vol. , p. 283–287. External Links: Document Cited by: §I. [22] J. Xue, Z. Hou, C. Xuan, Y. Ma, Q. Sun, X. Zhang, and L. Zhong (2024) A sheep identification method based on three-dimensional sheep face reconstruction and feature point matching. Animals 14 (13). External Links: Link, ISSN 2076-2615, Document Cited by: §I. [23] X. Yan, C. Yang, B. Shi, L. Ao, and S. Ma (2024) A goat facial recognition approach based on an enhanced yolov8s and keypoint affine transformation. In 2024 5th International Conference on Computer Vision, Image and Deep Learning (CVIDL), Vol. , p. 565–571. External Links: Document Cited by: §I. [24] H. Zhang, Y. Ma, X. Wang, R. Mao, and M. Wang (2023) Lightweight real-time detection model for multi-sheep abnormal behaviour based on yolov7-tiny. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vol. , p. 4191–4196. External Links: Document Cited by: §I. [25] X. Zhang, C. Xuan, Y. Ma, H. Liu, and J. Xue (2024) Lightweight model-based sheep face recognition via face image recording channel. Journal of Animal Science 102, p. skae066. External Links: ISSN 1525-3163, Document, Link, https://academic.oup.com/jas/article-pdf/doi/10.1093/jas/skae066/58661179/skae066.pdf Cited by: §I. [26] X. Zhang, C. Xuan, Y. Ma, H. Su, and M. Zhang (2022) Biometric facial identification using attention module optimized yolov4 for sheep. Computers and Electronics in Agriculture 203, p. 107452. External Links: ISSN 0168-1699, Document, Link Cited by: §I, §I. [27] X. Zhang, C. Xuan, Y. Ma, Z. Tang, and X. Gao (2024) An efficient method for multi-view sheep face recognition. Engineering Applications of Artificial Intelligence 134, p. 108697. External Links: ISSN 0952-1976, Document, Link Cited by: §I. [28] X. Zhang, C. Xuan, J. Xue, B. Chen, and Y. Ma (2023) LSR-yolo: a high-precision, lightweight model for sheep face recognition on the mobile end. Animals 13 (11). External Links: Link, ISSN 2076-2615, Document Cited by: §I.