Paper deep dive
A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/16/2026, 2:54:02 AM
Summary
This paper introduces a benchmark dataset for fine-grained isolated handshape recognition in sign language, grounded in the Hamburg Notation System (HamNoSys). The dataset consists of 144,000 RGB images collected from 15 participants across 160 handshape classes. The authors evaluate four baseline models (ResNet-18, ViT-B/16, Graph Convolutional Network, and XGBoost) using both subject-dependent and leave-one-subject-out (LOSO) protocols, highlighting the challenge of generalizing to unseen participants.
Entities (9)
Relation Signals (8)
Proposed Dataset â hasclasses â 160 handshape classes
confidence 98% · 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart
Proposed Dataset â hassize â 144,000 images
confidence 98% · A balanced dataset of 144,000 RGB images was collected
HamNoSys â definesclassesfor â Proposed Dataset
confidence 95% · A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart.
Proposed Dataset â usedinevaluation â LOSO
confidence 95% · a 15-fold leave-one-subject-out (LOSO) protocol were used
Graph Convolutional Network â evaluatedon â Proposed Dataset
confidence 90% · a graph convolutional network and XGBoost were evaluated from hand landmarks
XGBoost â evaluatedon â Proposed Dataset
confidence 90% · a graph convolutional network and XGBoost were evaluated from hand landmarks
ViT-B16 â evaluatedon â Proposed Dataset
confidence 90% · ResNet-18 and ViT-B/16 were evaluated as appearance-based models
ResNet-18 â evaluatedon â Proposed Dataset
confidence 90% · ResNet-18 and ViT-B/16 were evaluated as appearance-based models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.
Tags
Links
- Source: https://arxiv.org/abs/2608.10588v1
- Canonical: https://arxiv.org/abs/2608.10588v1
Trouble viewing inline? Open PDF directly â
Full Text
52,862 characters extracted from source content.
Expand or collapse full text
A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language Ushnish Sarkar 1,2* , Suvajit Patra 3â , Bhaswar Chattopadhyay 1,2â , Pranab Singha Roy 1â , Tapas Samanta 1,2â 1* Computer and Informatics Group, Variable Energy Cyclotron Centre, Kolkata, 700064, India. 2 Homi Bhabha National Institute, Mumbai, 400094, India. 3 Ramakrishna Mission Vivekananda Educational and Research Institute, Belur, 711202, India. *Corresponding author(s). E-mail(s): u.sarkar@vecc.gov.in; â These authors contributed equally to this work. Abstract Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a handshape dataset grounded in the language- independent Hamburg Notation System (HamNoSys) and baseline models for handshape recognition evaluated on the same. Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 as appearance-based models were evaluated on this dataset, while a graph convolutional network and XGBoost were evaluated on the hand landmarks of the images from this dataset. Both a class- stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent baselines established reproducible reference performance across all four model families, whereas LOSO evaluation showed some reduction when recognition was required to generalise to unseen participants. Additional analysis has been shown to highlight the presence of visually close handshapes that may possibly confuse the standard models . Conclusion: The documented acquisition, curation, and complementary evaluation protocols provide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies. Keywords: Sign language, HamNoSys, handshape recognition, dataset, leave-one-subject-out evaluation 1 Introduction More than 1.5 billion people worldwide experi- ence some degree of hearing loss [1], while a national survey in the United States estimated that approximately 2.8% of adults use a sign language [2]. Accessible computational tools for 1 arXiv:2608.10588v1 [cs.CV] 11 Aug 2026 sign-language transcription and recognition are therefore relevant to a substantial and diverse population. Sign languages are natural languages with their own phonological, morphological, and syntactic structures [3]. At the sublexical level, signs can be analysed through contrastive man- ual parameters such as handshape, orientation, location, and movement, together with linguis- tically relevant non-manual components [3, 4]. Among the manual parameters, handshape pro- vides an important source of lexical contrast and is consequently relevant to computational sign-language recognition and translation [5]. Lin- guistically grounded representation, resource con- struction, and evaluation have therefore remained central concerns in sign-language technology [6, 7]. The visualâgestural modality presents a funda- mental representational difficulty. Video preserves the multidimensional and temporally organised signal, but it does not by itself provide an explicit, discrete, and searchable description of the artic- ulatory form. Glosses can be used to identify lexical sign types or meanings, but differences in handshape, orientation, location, movement, and signer-specific realisation are not encoded by them. When physical sign form, phonetic vari- ation, phonological contrast, or cross-linguistic similarity is to be investigated, a form-based tran- scription system is therefore required. Through such a system, observable sublexical properties can be represented consistently, sign variants can be compared, corpus annotations can be searched, and findings can be evaluated across datasets and sign languages [8, 9]. However, no universally accepted sign-language equivalent of the Interna- tional Phonetic Alphabet has yet been established [9]. Several transcription systems have been devel- oped at different levels of descriptive abstrac- tion. Stokoe notation introduced a parameter- based phonological representation of handshape, location, and movement, principally for Amer- ican Sign Language [4]. The LiddellâJohnson MovementâHold model subsequently represented signs through temporally ordered movement and hold segments [10]. Prosodic Model Handshape Coding provides a theoretically motivated phono- logical representation of selected fingers and joint configuration [11], whereas Sign Language Pho- netic Annotation provides a more anatomically detailed description of individual fingers, joints, thumb configuration, and articulator contact [9, 12, 13]. SignWriting serves primarily as a graph- ical or orthographic representation and has also been employed as an intermediate representa- tion in computational translation [14]. These sys- tems differ in purpose, descriptive coverage, and granularity and are therefore not interchange- able [15, 16]. For the present handshape-centred task, HamNoSys offers greater articulatory detail and broader cross-linguistic applicability than the original Stokoe system, while avoiding the exten- sive joint-by-joint description required by Sign Language Phonetic Annotation [13, 15]. Unlike a notation restricted to a particular phonological model, it can also be used as a mostly phonetic description of observable manual form [17, 18]. Most importantly for dataset construction, the official, non-exhaustive HamNoSys 4 Handshapes Chart provides publicly illustrated reference forms from which a bounded and visually reproducible class inventory can be defined [19]. Reported inconsistencies in the use of HamNoSys anno- tations further support the use of labels tied directly to fixed chart illustrations rather than unconstrained transcription [20]. Existing visual resources support several related tasks, including fingerspelling recognition, corpus-derived hand- shape classification, and synthetic handshape recognition. However, their differing objectives leave a need for a real-image benchmark that com- bines a broad, transcription-defined handshape inventory with systematic evaluation on unseen participants. The supporting comparison with existing resources is provided in Section 2. This need was addressed through the construction of a balanced, HamNoSys-grounded image dataset and its evaluation under complementary subject- dependent and leave-one-subject-out protocols. The class definition, acquisition procedure, and dataset composition are presented in Section 3, while the baseline models and matched-model experiments on external datasets are presented in Section 4. 2 Related Work Fine-grained handshape inventories distinguish selected fingers, joint configuration, thumb behaviour, inter-finger relations, and contact [11, 21, 22]. Segmental descriptions additionally dis- tinguish static articulation from its temporal 2 organisation within a sign [12]. Because neigh- bouring handshapes may differ in only one of these properties, fine-grained recognition requires a broader inventory than that provided by language-specific fingerspelling alphabets. Existing visual resources represent several related but distinct tasks. Continuous-sign cor- pora provide weak or auxiliary handshape labels [23, 24]; LSWH100 contains synthetic images organised using SignWriting-derived categories [25]; and fingerspelling datasets cover restricted language-specific alphabets or digits [26â29]. Other resources represent complete lexical signs [30] or derive hand-pose categories through feature-based clustering [31]. Their differences in visual data, class inventory, and linguis- tic scope are summarised in Table 1. Within the reviewed literature, no resource jointly pro- vides an approximately balanced collection of real isolated-hand images, a broad class inven- tory grounded in a language-independent phonetic notation, and evaluation under both participant- overlapping and participant-independent pro- tocols. This combination defines the specific resource gap addressed in the present study. Ham- NoSys has otherwise been applied primarily at the sign level in corpora, multilingual lexicons, dic- tionary transcription, cross-linguistic comparison, and automatic motion generation [18, 32â37]. As summarised in Table 2, these applications encode lexical citation forms, corpus units, or motion sequences rather than balanced image collections organised by isolated handshape. The HamNoSys 4 Handshapes Chart is itself an illustrated refer- ence inventory rather than an image dataset; its use for defining the proposed classes is described in Section 3. More broadly, gesture-recognition research has addressed humanâcomputer interaction, vir- tual environments, rehabilitation, and sign- language processing [38]. General-purpose gesture resources, however, are not necessarily organised as linguistically defined handshape inventories. Complementary representation families have been adopted for handshape and gesture recogni- tion. RGB images have been processed using con- volutional and transformer-based architectures [39, 40], while hand landmarks have been rep- resented as anatomical graphs or fixed-length feature vectors [41â44]. Representative baselines from these families are evaluated in the present study rather than an exhaustive set of architec- tures. Evaluation design is particularly important for multi-participant data. Participant-overlapping splits measure recognition when the same par- ticipants may occur across partitions, whereas participant-independent protocols assess general- isation to unseen individuals [45]. Both settings are therefore reported, with leave-one-subject-out evaluation used for the systematic participant- independent assessment described in Section 4. 3 Dataset Construction This section defines the handshape class inven- tory and subsequently describes the acquisition protocol, collection software, participants, dataset organisation, and evaluation splits. 3.1 HamNoSys Handshape Inventory and Class Definition HamNoSys Version 2.0 encodes the manual com- ponents of a sign through handshape, orientation, location, and movement [17]. Version 4 addition- ally supports optional non-manual specifications [18]. Handshapes are constructed composition- ally from a basic form and modifiers describ- ing finger selection, bending, thumb position, individual-finger configuration, and intermediate forms. Transitions between handshapes are rep- resented as actions rather than separate dynamic handshape primitives [18]. Because HamNoSys does not define a finite handshape inventory, the official, non-exhaustive HamNoSys 4 Handshapes Chart was used to establish a bounded and reproducible class set [19]. Each distinct illustrated hand model in Fig. 1 was treated as one class; blank cells and cells containing only symbols or cross-references were excluded. This procedure produced 160 classes. The illustrations occupy six Selection rows and four Thumb-opposition rows. The two Thumb- opposition rows containing no illustrationsâTwo Fingers (spread), others in fist position and Four Fingers (spread)âwere excluded. Table 3 reports the class counts for the ten populated rows, while Table 4 groups the same 160 classes by chart column: 106 Selection classes and 54 Thumb- opposition classes. All dataset codes are author- defined and are not official HamNoSys symbols. 3 Table 1 Representative visual handshape and sign-language resources. ResourceVisual dataClasses Label basis or scope Deep Hand [23]> 1 million weakly labelled real frames 60Corpus- and lexicon-derived handshapes PHOENIX14T-HS [24] Continuous-sign videos60Auxiliary handshape labels associated with DGS glosses LSWH100 [25]144,000 synthetic images100SignWriting-derived Libras handshapes ASL Fingerspelling Dataset A [27] 65,774 real RGB images from five users 24Static ASL alphabet Other fingerspelling resources [26, 28, 29] Static real or augmented images 24â41ASL, JSL, or Danish letters and digits LSA64 [30]3,200 videos64Complete Argentinian Sign Language lexical signs Kajiyama et al. [31]Hand images from sign-language data Data- derived Finger-shape and palm-orientation clusters Proposed dataset144,000 real RGB images from 15 participants 160Chart-defined HamNoSys handshapes Table 2 Documented applications of HamNoSys in sign-language resources. ResourceSign language(s)Scale or primary unitFunction of HamNoSys DGS Corpus and GLex [18, 32, 33] German Sign Language Corpus utterances and technical lexical signs Corpus-linked annotation and citation-form description DICTA-SIGN resources [34, 37] BSL, DGS, GSL, and LSF Approximately 1,000 signs per language Citation-form representation and cross-linguistic comparison Corpus-based Dictionary of PJM [35] Polish Sign Language 3,476 lexical signsCitation-form transcription Motion-generation dataset [36] Spanish Sign Language 754 signs and 6,786 videos Intermediate representation for automatic motion generation Table 3 Populated chart selections used as dataset categories. Chart sectionCategoryCode Classes SelectionFistF12 SelectionOne FingerOF20 SelectionTwo Fingers (nonspread)TFN17 SelectionTwo Fingers (spread)TFS19 SelectionFlathand (Four Fingers nonspread)FFN16 SelectionFour Fingers (spread)FFS22 Thumb opposition One Finger, others in fist positionOFO14 Thumb opposition Two Fingers (nonspread), others in fist positionTFO12 Thumb opposition Four Fingers (nonspread)FFO11 Thumb opposition One Finger, others extended (spread)OFOE17 Total160 The resulting inventory is restricted to the forms illustrated in the non-exhaustive chart and therefore does not cover every handshape express- ible in HamNoSys. Extension beyond these 160 classes would require expert definition and valida- tion. 3.2 Participants and Ethics Fifteen university students aged 23â25 years par- ticipated in the data collection. Right-hand dom- inance was reported by fourteen participants and left-hand dominance by one. For each class, par- ticipants examined and reproduced the displayed 4 Fig. 1 The official, non-exhaustive HamNoSys 4 Handshapes Chart [19]. Each distinct drawn hand model is treated as one dataset class. Table 4 Number of drawn models under each chart column group. Chart sectionColumn groupCode Classes SelectionSelected Fingers ExtendedSFE24 SelectionSelected Fingers FlattenedSFF17 SelectionSelected Fingers BentSFB24 SelectionSelected Fingers HookedSFH21 SelectionDerivation ExamplesDE20 Thumb opposition Fingertip-Thumbtip Opposition w/fingers roundedFTR16 Thumb opposition Fingertip-Thumbtip Opposition w/fingers flattenedFTF16 Thumb opposition Fingertip-Thumbtip Opposition w/Hitchhikerâs fingersFTH4 Thumb opposition Fingertip Thumbâs Interphalangeal-joint oppositionFTI4 Thumb opposition Fingertip Thumbâs Metacarpophalangeal-joint oppositionFTM4 Thumb opposition Other Derivation ExamplesDE10 Total160 reference handshape. Finger selection, bending, thumb position, and contact were verified against the reference before recording by an operator with sign-language experience. Participants received an honorarium for their time. The applicable ethical oversight and consent procedures are reported in the Statements and Declarations. 5 3.3 Dataset recording Two interfaces were provided by the acquisition application (Fig. 2). The target HamNoSys hand- shape and live camera view were displayed on the actor panel, while the reference, incoming stream, and recording controls were displayed on the oper- ator panel. The recording duration was set to 10 seconds. Recordings were made indoors against a con- stant dull-white background using a tripod- mounted Logitech Brio RGB camera at 30 frames/s and 640Ă 480 pixels. The acquisition program was run on an Intel Core i7 Windows workstation, while the session was managed from a separate workstation (Fig. 3). No instrumented gloves or body-mounted sen- sors were used. The recording area was kept rea- sonably uncluttered to limit irrelevant background variation and to support later hand localisation. Participants were free to adjust their upper-body posture and hand position within the camera view. The only required actions were to form the target handshape and perform the prescribed rotation of the hand. Every class defined in Section 3.1 was per- formed by each participant. After the reference- guided verification described above, a 10-second clip was recorded while the participant main- tained the handshape and slowly rotated the hand about two approximately orthogonal axes. This efficiently introduced viewpoint, apparent overlap, and self-occlusion variation without intentionally changing the class. Each 10-second clip recorded at 30 frames/s yielded 300 RGB frames. To reduce temporal redundancy while retaining samples throughout the hand rotation, every fifth frame was selected, producing 60 images per clip. Par- ticipant, HamNoSys class, source-video, frame- index, and dominant-hand metadata were stored for each image, yielding 144,000 labelled images. MediaPipe Hands [41], configured as described in Section 4, localised the metadata-defined dom- inant hand in 139,199 images. These images formed the common modelling subset for all four baselines, thereby preventing representation- dependent sample selection. The remaining 4,801 images were retained in the complete dataset. The pipeline is shown in Fig. 4. 3.4 Dataset Organisation and Naming Convention The dataset is organised hierarchically to preserve the provenance of every image and its correspon- dence with the chart-defined handshape inventory. The root directory contains 15 participant folders, labelled s0001 to s0015. Each participant folder is divided into the ten populated HamNoSys cat- egories defined in Table 3. Within each category, images are grouped by handshape subclass and chart variation. A class folder follows the naming convention subclass_categoryvariation, where subclass identifies the relevant chart-column group, category denotes the dataset-specific selection code, and variation indexes the corresponding drawn hand model within that combination. For example, SFE_OF2 denotes the second Selected Fingers Extended form in the One Finger selection category, whereas DE_F4 denotes the fourth Derivation Example associated with the Fist category. Individual image files retain their source-video identifier and frame index using the convention video_frame_frameindex.png. Thus, a path such as s0001/OF/SFE_OF2/v0000014_frame_00000.png identifies participant s0001, the One Fin- ger selection category, the second Selected Fin- gers Extended class, source video v0000014, and extracted frame 00000. Similarly, s0001/F/DE_F4/v0000111_frame_00000.png corresponds to the fourth derivation-example class under the Fist selection. This naming scheme makes each image trace- able to its participant, selection category, chart- defined handshape class, source recording, and frame position. A supplementary class-mapping file provides the complete correspondence between the 160 dataset class identifiers and the illus- trated hand models in the official HamNoSys 4 Handshapes Chart. 6 Fig. 2 Custom data-acquisition software: (a) actor interface displaying the target handshape and live participant view, and (b) operator/controller interface used for verification and recording control. Fig. 3 Controlled classroom recording environment showing the participant, operator, acquisition computer, and tripod- mounted RGB camera from three viewpoints. 3.5 Dataset Characteristics and Class Distribution The complete dataset and common modelling sub- set are summarised in Table 5. The 144,000 images correspond to 900 per class and 60 per class per participant. The modelling subset remains approximately uniform (Fig. 5), ranging from 812 images for DE_F3 to 900 for FTR_TFO1 and SFB_TFS3. Within-class variation in viewpoint, appar- ent orientation, hand position, and self-occlusion was introduced by the video-based acquisition procedure. Figure 6 presents four frames from each of two representative recordings, with one participant , enacting different variations of a handshape, shown per row. 3.6 Evaluation Protocols and Data Splits Two complementary protocols are defined over the 139,199-image modelling subset. A class-stratified Table 5 Principal characteristics of the HamNoSys handshape dataset. CharacteristicValue Total images144,000 Images used for modelling139,199 (96.66%) Images excluded from modelling4,801 (3.34%) Handshape classes160 Participants15 Image resolution 640Ă 480 pixels Image modalityRGB Mean images per class in total subset 900 Mean images per class in mod- elling subset 869.99 Minimum images in a class in modelling subset 812 Maximum images in a class in modelling subset 900 frame-level split is used to provide a subject- dependent reference comparable with conven- tional image-classification benchmarks, including 7 Fig. 4 Dataset construction and common modelling-subset selection pipeline. 14080120160 Class rank after sorting by image count 810 830 850 870 890 900 Images per class Minimum = 812 Maximum = 900 (a) Rank-ordered class sizes Mean = 869.99 160 classes 139,199 images Mean 869.99 Median 871 Range 812900 (b) Distribution summary 820840860880900 Images per class 0 5 10 15 20 25 Number of classes Class Balance of the Modelling Subset Fig. 5 Class balance of the 139,199-image modelling subset. (a) Rank-ordered image counts for the 160 handshape classes, with the minimum and maximum class sizes highlighted. The dashed line indicates the mean class size. (b) Boxplot and histogram summarizing the distribution of per-class image counts. Class sizes range from 812 to 900 images, with a mean of 869.99 images per class. datasets for which signer identities are unavail- able. Because participant identity was preserved during collection, the stricter assessment of gen- eralisation to unseen participants is provided by a 15-fold leave-one-subject-out (LOSO) protocol. 3.6.1 Subject-Dependent Protocol For the subject-dependent evaluation, images were partitioned at frame level using a class- stratified random 70:15:15 split. Temporal redun- dancy was reduced before this split by retaining every fifth frame, as described in Section 3.3. The resulting counts are reported in Table 6. Because 8 Fig. 6 Representative dataset frames from two participants. Each row contains four frames from one participant, illustrating variation in viewpoint, hand position, articulation, and appearance. Faces are blurred to protect participant identity. Table 6 Subject-dependent frame-level split of the 139,199-image modelling subset. Partition Images Percentage Approx. per class Training 97,37069.95%609 Validation 20,80614.95%130 Test21,02315.10%131 Total 139,199100%870 all participants may occur in every partition, recognition under seen-participant conditions is measured by this protocol; unseen-participant generalisation is assessed separately by LOSO. 3.6.2 Subject-Independent Protocol (LOSO) Subject-independent performance was evalu- ated using 15-fold leave-one-subject-out cross- validation, as summarised in Table 7. In each fold, one participant was reserved exclusively for testing, while images from the remaining 14 par- ticipants were divided into class-stratified training and validation partitions using an 85:15 ratio. Identical partitions were used for all four base- lines, and no data from the held-out participant were used for model fitting, validation, checkpoint selection, or other training-stage decisions. Each participant served as the test subject once, and performance was reported as the mean and sample standard deviation across the 15 folds. Table 7 Structure of each leave-one-subject-out fold. Exact image counts vary with the number of usable images contributed by the held-out subject. ComponentDefinition Test setAll usable images from one held-out subject Development setAll usable images from the remaining 14 subjects Training partition85% of the development images, selected using class-stratified sampling Validation partition15% of the development images, selected using the same class-stratified split Number of folds15, with each subject held out exactly once Subject overlapNone between the test subject and the training or validation partitions 4 Experiments Baseline recognition performance is established for the proposed dataset under the subject- dependent and leave-one-subject-out protocols defined in Section 3.6. Four models were eval- uated: two appearance-based models operating on RGB hand crops and two landmark-based models operating on MediaPipe hand keypoints. The comparisons are intended to characterise the 9 benchmark rather than propose a new recogni- tion architecture.Figure 7 presents the two com- plementary evaluation protocols and the model variants used under each protocol. External context is provided by two datasets. LSWH100 [25] contains 144,000 synthetic images in 100 SignWriting-derived classes with predefined training, validation, and test partitions; model output layers were changed to 100 classes and results computed on the predefined test set. ASL Fingerspelling Dataset A [27] contains 65,774 real RGB observations from five users covering 24 static ASL letters (excluding dynamic J and Z). Output layers were changed to 24 classes. A class-stratified 70:15:15 subject-dependent split and five-fold LOSO were used, with 15% of each remaining-user development set reserved for vali- dation where required. These are matched-model references, not direct rankings of dataset quality: image ori- gin, class definition and count, participant com- position, and acquisition protocol differ across datasets. 4.1 Training Configurations and Baseline Models RGB Input based models. For the appearance-based models, MediaPipe Hands [41] was applied in static-image mode to identify the participantâs metadata-defined domi- nant hand. The hand bounding box was expanded by 15 pixels on each side, subject to the image boundaries, and the resulting crop was resized to 224Ă 224 pixels. For ImageNet-pretrained ResNet-18 [39], the classifier was replaced by a 160-class linear head and only the final residual stage and new head were fine-tuned. For ImageNet-pretrained ViT- B/16 [40], the head was replaced by dropout (0.3) and a 160-class linear layer; the final two transformer blocks, encoder normalisation, and head were fine-tuned. Horizontal flipping, rota- tion up to 15 ⊠, colour jitter, and random erasing were used for ResNet-18 augmentation. Affine and perspective transformations and stronger photo- metric augmentation were additionally used for ViT-B/16. Evaluation images were resized and ImageNet-normalised. Landmark Input based models. MediaPipe Hands was also used to extract 21 dominant-hand landmarks from each image, which were preprocessed as in [43]. Detected left hands were mirrored to a common canoni- cal orientation. The coordinates were translated so that the wrist lay at the origin, rotated into a hand-centred coordinate frame, and scaled by the maximum landmark distance. A joint-angle fea- ture, normalised by Ï, was computed for the 15 internal finger joints; the wrist and fingertips were assigned an angle of zero. Each node was therefore represented by (x,y,z,Ξ), giving a 21Ă 4 feature matrix. The anatomical hand connections were used as edges by the graph convolutional network, with self-connections introduced during graph convolu- tion. Five graph-convolutional layers with output dimensions [512, 480, 448, 416, 352], GELU activa- tions, residual connections, batch normalisation, dropout of 0.1, global mean pooling, and a 160- class linear output layer were included. An 84-dimensional vector formed from each landmarkâs x, y, and z coordinates and joint angle was used by the XGBoost baseline [44]: 21Ă (x,y,z,Ξ). For XGBoost, 1,200 boosting rounds were used with maximum depth 10, minimum child weight 3, learning rate 0.02180, gamma 0.01248, subsam- ple ratio 0.6762, column-sampling ratio 0.7482, â 2 regularisation 0.01875, â 1 regularisation 0.09311, and histogram-based tree construction. 4.2 Evaluation Measures The subject-dependent results for the pro- posed dataset and the matched-model results for LSWH100 and ASL Fingerspelling Dataset A are reported in Table 9. Because each partition may contain frames from all participants, this protocol measures recognition under participant overlap and provides a seen-participant reference. It is not interpreted as an estimate of generalisation to unseen participants. 10 Evaluation Design and Baseline Model Variants ATWO COMPLEMENTARY EVALUATION PROTOCOLS 139,199-image modelling subset 160 handshape classes · 15 participants Class labels and participant identities retained Stratify by classGroup by participant SUBJECT-DEPENDENT REFERENCE Class-stratified frame-level split TRAIN 97,370 · 69.95% VALIDATION 20,806 · 14.95% TEST 21,023 · 15.10% Conventional image-classification reference; comparable with datasets for which participant identities are unavailable UNSEEN-SUBJECT GENERALISATION 15-fold leave-one-subject-out (LOSO) MODEL DEVELOPMENT Remaining 14 participants TEST 1 unseen participant Each participant is held out once; performance is reported across folds to assess generalisation to unseen participants BMODEL VARIANTS APPLIED UNDER BOTH PROTOCOLS Same four baseline models in each protocol RGB images ResNet-18 ViT-B/16 LANDMARK coordinates XGBoost GCN Fig. 7 Overview of the evaluation design and baseline model variants. Table 8 Training configuration of the neural baseline models. SettingResNet-18ViT-B/16GCN PretrainingImageNetImageNetNone OptimizerAdamAdamWAdamW Learning rate 5Ă 10 â5 1Ă 10 â4 6.451Ă 10 â4 Weight decay 1Ă 10 â4 5Ă 10 â4 3.285Ă 10 â5 Batch size6432128 Maximum epochs6060150 Early-stopping patience8610 Learning-rate scheduleReduceLROnPlateau CosineAnnealingLR ReduceLROnPlateau Label smoothing0.10.1None Input 224Ă 224 RGB 224Ă 224 RGB 21Ă 4 graph 4.3 Results 4.3.1 Subject-Dependent Performance and External Reference The subject-dependent results for the proposed dataset and the matched-model LSWH100 and ASL Fingerspelling Dataset A references are reported in Table 9. Because frames from all par- ticipants may occur in each partition, these values provide a seen-signer reference rather than an estimate of performance on an unseen signer. On the proposed 160-class dataset, the highest top-1, top-3, and top-5 accuracies were achieved by ViT-B/16, at 86.20%, 95.99%, and 97.79%, respectively. ResNet-18 achieved a comparable top-1 accuracy of 84.72%. The higher top-1 accu- racies of the two RGB-based models relative to the 11 Table 9 Within-dataset test performance on the proposed HamNoSys dataset and ASL Fingerspelling Dataset A under subject-dependent evaluation, and on LSWH100 using its predefined split (%). AccuracyWeighted averageMacro average Model@1@3@5PRF 1 PRF 1 Proposed HamNoSys dataset (160 classes) ResNet-1884.7294.6296.7885.0784.7284.7485.0984.7284.75 ViT-B/1686.2095.9997.7986.5086.2086.1986.5186.1986.20 GCN72.4488.4492.7472.6272.4472.3872.6072.4172.36 XGBoost69.5785.2390.0669.7669.5769.5569.7469.5469.53 LSWH100 (100 classes) ResNet-1886.7097.1798.6787.1186.7086.6987.1186.7086.69 ViT-B/1681.5595.8598.0082.7081.5581.5382.7081.5581.53 GCN79.5494.8296.7980.1179.5479.5279.8079.5479.36 XGBoost72.2690.5094.7672.9672.2672.2972.7172.2372.14 ASL Fingerspelling Dataset A (24 classes) ResNet-1899.94100.00100.0099.9499.9499.9499.9499.9499.94 ViT-B/1699.72100.00100.0099.7299.7299.7299.7299.7199.71 GCN97.2199.0999.3597.2597.2197.2197.1796.9897.05 XGBoost97.6699.1799.4597.6997.6697.6797.4197.5297.46 P = precision; R = recall. The external datasets have different class inventories and acquisition conditions; their results therefore provide context rather than direct estimates of relative dataset quality. landmark-based baselines are consistent with fine- grained distinctions benefiting from appearance information that is not completely retained by the 21-point landmark representation. Among the landmark-based models, GCN exceeded XGBoost by 2.87 percentage points in top-1 accuracy, sug- gesting an advantage from explicitly representing the anatomical connectivity of the hand. For every model, the weighted and macro scores were closely aligned. This agreement is con- sistent with the approximately balanced class dis- tribution and indicates that the aggregate results were not dominated by a small number of larger classes. The highest LSWH100 top-1 accuracy was obtained by ResNet-18 at 86.70%. Relative to the proposed dataset, the change in top-1 accuracy ranged fromâ4.65 to +7.10 percentage points and differed across models. This change in model ordering, together with the different image origins and 100- versus 160-class inventories, prevents a direct ranking of dataset quality from these within-dataset results. Top-1 accuracy above 97% was obtained by all four models on ASL Fingerspelling Dataset A. These near-ceiling participant-overlapping results must be interpreted in relation to its smaller 24- class alphabetic inventory and different acquisi- tion conditions. The external evaluations therefore establish matched-model reference points rather than direct measures of the relative quality of the three datasets. Because ViT-B/16 achieved the strongest subject-dependent performance on the proposed dataset, it was selected for the sub- sequent confusion analysis. The broad-category confusion matrix in Figure 8 is strongly concen- trated along the diagonal, indicating that the ten broad handshape categories were generally sep- arated successfully. The remaining off-diagonal predictions occurred primarily between related categories and motivated a finer target-level anal- ysis. As shown in Figure 1, several target classes differ only in subtle properties such as finger selec- tion, bending, thumb position, or contact. The five target-level pairs with the largest bidirectional confusion counts are shown in Figure 9. For each pair, the pair error rate was calculated as the total number of errors in both directions divided by the combined support of the two classes. 12 Fig. 8 Broad-category confusion matrix for ViT-B/16 under subject-dependent evaluation. Rows denote true categories and columns denote predicted categories; cell values are test-sample counts. 4.3.2 Leave-One-Subject-Out Performance The LOSO results for the proposed dataset and ASL Fingerspelling Dataset A are reported in Table 10. LOSO evaluation could not be conducted on LSWH100 because signer-identity information is not provided with that dataset. On the proposed dataset, numerically close mean top-1 accuracies were obtained by ResNet- 18 and ViT-B/16, at 45.38% and 45.22%, respec- tively. Although its top-1 accuracy was lower, GCN achieved the highest top-3 and top-5 accura- cies, at 69.23% and 78.58%. This result indicates that the landmark graph frequently retained the correct class among its leading predictions even when the top-ranked prediction was incorrect. Relative to the participant-overlapping evalu- ation, the ResNet-18 and ViT-B/16 top-1 accu- racies decreased by 39.34 and 40.98 percent- age points, respectively. The fold standard devi- ations also demonstrate substantial variation 13 0510152025 Number of errors 18 10 15 15 9 A predicted as B 1 FTR_OFOE2FTR_OFOE3 12.26% 2 FTI_TFO1FTR_TFO1 11.99% 3 SFB_FFS1SFB_FFS3 11.92% 4 FTR_OFO2FTR_OFO3 10.94% 5 DE_F5 SFE_F1 9.36% Class A and illustration Handshape pair Class B and illustration Pair error 0510152025 Number of errors 14 22 16 14 16 B predicted as A Most Frequent Target-Level Confusions Pairs are ranked by the sum of errors in both directions. Fig. 9 Five most frequently confused target-level handshape pairs for ViT-B/16 under subject-dependent evaluation. HamNoSys chart illustrations are displayed alongside the corresponding class codes. The left and right bars report directional misclassification counts, while the pair error rate represents the total bidirectional errors relative to the combined support of the two classes. among held-out participants. These results iden- tify unseen-participant generalisation as the prin- cipal challenge of the proposed 160-class bench- mark rather than indicating a failure of within- participant handshape recognition. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% for ViT-B/16 to 87.40% for ResNet-18. These val- ues are not directly comparable with those of the proposed dataset because ASL Dataset A con- tains only 24 classes and was collected under different conditions. In contrast, the proposed inventory contains 160 fine-grained classes, includ- ing several visually similar handshapes that differ in limited articulatory properties. The confu- sion pairs in Figure 9 illustrate this fine-grained separation problem. Improved representations of local finger articulation and greater robustness to inter-participant variation are therefore important directions for further modelling. Per-participant top-1 accuracies for the pro- posed dataset are shown in Figure 10. 5 Conclusion A balanced handshape dataset grounded in the official HamNoSys 4 Handshapes Chart was pre- sented, comprising 144,000 RGB images from 15 participants across 160 illustrated classes. RGB appearance baselines were provided by ResNet- 18 and ViT-B/16, while hand-landmark baselines were provided by a GCN and XGBoost. Com- plementary references for seen-participant and participant-disjoint recognition were provided by the subject-dependent and LOSO protocols. A substantial effect of evaluation protocol on recognition performance was observed, and the 160-class inventory remained challenging under participant-disjoint testing. Matched-model evaluations were also conducted on synthetic LSWH100 and real ASL Dataset A. The dataset and baselines are intended to support phonology- grounded research and the development of tran- scription, recognition, and translation tools across sign languages, including under-resourced settings where dictionaries may be available but labelled datasets remain scarce. Limitations. Few limitations should be noted. Data were col- lected from 15 university students in a controlled indoor environment using a single RGB camera, and broader demographic, environmental, and sensor variability was therefore not represented. The class inventory was restricted to the 160 static, single-hand forms illustrated in the non- exhaustive HamNoSys chart; dynamic transitions, two-handed configurations, orientation, location, 14 Table 10 Leave-one-subject-out performance on the proposed HamNoSys dataset and ASL Fingerspelling Dataset A (%). Values are reported as mean± sample standard deviation across 15 held-out subjects for the proposed dataset and five held-out users for ASL Dataset A. ModelAccuracy@1Accuracy@3Accuracy@5 Proposed HamNoSys dataset (15 subjects) ResNet-18 45.38± 7.48 67.72± 8.80 75.74± 8.35 ViT-B/16 45.22± 6.92 68.47± 7.75 76.49± 7.18 GCN 43.49± 7.75 69.23± 9.15 78.58± 8.25 XGBoost 39.66± 6.87 64.29± 8.75 74.29± 8.22 ASL Fingerspelling Dataset A (5 users) ResNet-18 87.40± 4.50 96.32± 1.21 97.95± 0.68 ViT-B/16 82.20± 5.36 95.23± 1.56 97.44± 0.73 GCN 84.22± 4.03 95.86± 1.31 97.75± 0.81 XGBoost 86.14± 3.81 96.00± 0.96 97.76± 0.49 Weighted averageMacro average ModelPRF 1 PRF 1 Proposed HamNoSys dataset (15 subjects) ResNet-18 46.80± 7.77 45.38± 7.48 43.92± 7.56 46.76± 7.81 45.39± 7.47 43.90± 7.57 ViT-B/16 47.64± 7.29 45.22± 6.92 43.63± 7.07 47.56± 7.31 45.22± 6.89 43.59± 7.05 GCN 44.15± 8.00 43.49± 7.75 42.14± 7.61 44.10± 8.03 43.49± 7.73 42.11± 7.62 XGBoost 40.44± 6.97 39.66± 6.87 38.49± 6.72 40.39± 7.00 39.67± 6.86 38.47± 6.73 ASL Fingerspelling Dataset A (5 users) ResNet-18 88.75± 3.95 87.40± 4.50 87.02± 4.52 88.74± 3.82 87.43± 4.68 87.03± 4.55 ViT-B/16 84.48± 4.70 82.20± 5.36 81.35± 5.60 84.45± 4.70 82.08± 5.60 81.24± 5.74 GCN 86.16± 3.36 84.22± 4.03 83.97± 3.94 83.00± 6.75 81.56± 5.36 80.85± 6.31 XGBoost 88.08± 2.88 86.14± 3.81 85.71± 3.70 84.64± 6.05 83.95± 5.22 82.72± 6.23 P = precision; R = recall. movement, and non-manual components were not included. Reference correspondence was verified by an operator with sign-language experience, but further annotation validation by HamNoSys specialists may strengthen the resource. Finally, only four baseline model families were evaluated. More diverse participants, less-controlled acquisi- tion settings, and advanced fine-grained models should be investigated in future work. Statements and Declarations Ethics approval. The study was conducted collaboratively by the Variable Energy Cyclotron Centre and Ramakrishna Mission Vivekananda Educational and Research Institute under the ethical and administrative procedures mutually established by the participating organisations. All procedures involving human participants were conducted in accordance with the applicable insti- tutional requirements. Consent to participate. Written informed consent was obtained from all participants before data collection. Participation was voluntary, and the study procedure and intended research use of the recordings were explained before the recording sessions. Consent for publication. Written consent was obtained for the publication of participant images. All participant faces shown in this article were anonymised before publication. Data availability. The dataset will be made available after publication upon reason- able request to the corresponding author at u.sarkar@vecc.gov.in. Access will be subject to the conditions established by the collaborating institutions. 15 Fig. 10 Top-1 accuracy for each held-out signer under LOSO evaluation. Dashed lines indicate the respective across-signer means. References [1] World Health Organization. Deafness. WHO Facts in Pictures (2024). URL https://w w.who.int/news-room/facts-in-pictures/de tail/deafness. Published 1 February 2024; accessed 21 July 2026. [2] Mitchell, R. E. & Young, T. A. How many people use sign language? a national health survey-based estimate. The Journal of Deaf Studies and Deaf Education 28, 1â6 (2023). [3] Sandler, W. & Lillo-Martin, D. Sign Lan- guage and Linguistic Universals (Cambridge University Press, Cambridge, UK, 2006). [4] Stokoe, W. C. Sign Language Structure: An Outline of the Visual Communication Sys- tems of the American Deaf. No. 8 in Studies in Linguistics: Occasional Papers (Depart- ment of Anthropology and Linguistics, Uni- versity of Buffalo, Buffalo, NY, 1960). [5] Rastgoo, R., Kiani, K. & Escalera, S. Sign language recognition: A deep sur- vey. Expert Systems with Applications 164, 113794 (2021). [6] Bragg, D. et al. Sign language recognition, generation, and translation: An interdisci- plinary perspective. In Proceedings of the 21st International ACM SIGACCESS Con- ference on Computers and Accessibility, 16â 31 (Association for Computing Machinery, New York, NY, USA, 2019). [7] De Coster, M., Shterionov, D., Van Her- reweghe, M. & Dambre, J. Machine transla- tion from signed to spoken languages: State of the art and challenges. Universal Access in the Information Society 23, 1305â1331 (2024). [8] Garcia, B. & Sallandre, M.-A. Transcrip- tion systems for sign languages: A sketch of the different graphical representations of sign language and their characteristics. In MĂŒller, C. et al. (eds.) BodyâLanguageâ Communication: An International Handbook on Multimodality in Human Interaction, 1125â1138 (De Gruyter Mouton, Berlin, Ger- many, 2013). [9] Tkachman, O., Hall, K. C., Xavier, A. & Gick, B. Sign language phonetic annotation meets phonological CorpusTools: Towards a sign language toolset for phonetic notation and phonological analysis. Proceedings of the Annual Meetings on Phonology 3 (2016). [10] Liddell, S. K. & Johnson, R. E. American sign language: The phonological base. Sign Language Studies 195â277 (1989). 16 [11] Eccarius, P. & Brentari, D. Handshape coding made easier: A theoretically based notation for phonological transcription. Sign Language & Linguistics 11, 69â101 (2008). [12] Johnson, R. E. & Liddell, S. K. Toward a phonetic representation of signs, i: Sequen- tiality and contrast. Sign Language Studies 11, 241â274 (2011). [13] Hall, K. C., Mackie, S., Fry, M. & Tkachman, O. SLPAnnotator: Tools for implement- ing sign language phonetic annotation. In Proceedings of Interspeech 2017, 2083â2087 (2017). [14] Jiang, Z., Moryossef, A., MĂŒller, M. & Ebling, S. Machine translation between spoken languages and signed languages rep- resented in SignWriting. In Findings of the Association for Computational Linguis- tics: EACL 2023, 1706â1724 (Association for Computational Linguistics, Dubrovnik, Croatia, 2023). URL https://aclanthology.o rg/2023.findings-eacl.127/. [15] Hochgesang, J. A. Using design principles to consider representation of the hand in some notation systems. Sign Language Studies 14, 488â542 (2014). [16] Dhanjal, A. S. & Singh, W. Comparative analysis of sign language notation systems for indian sign language. In 2019 Second International Conference on Advanced Com- putational and Communication Paradigms (ICACCP), 1â6 (IEEE, Gangtok, India, 2019). [17] Prillwitz, S., Leven, R., Zienert, H., Hanke, T. & Henning, J. HamNoSys Version 2.0: Hamburg Notation System for Sign Lan- guages: An Introductory Guide, vol. 5 of International Studies on Sign Language and Communication of the Deaf (Signum, Ham- burg, Germany, 1989). [18] Hanke, T. HamNoSysârepresenting sign language data in language resources and language processing contexts. In Stre- iter, O. & Vettori, C. (eds.) Proceedings of the LREC2004 Workshop on the Rep- resentation and Processing of Sign Lan- guages: From SignWriting to Image Pro- cessing. Information Techniques and Their Implications for Teaching, Documentation and Communication, 1â6 (European Lan- guage Resources Association (ELRA), Lis- bon, Portugal, 2004). URL https://w.sign -lang.uni-hamburg.de/lrec/pub/04001.html. [19] Hanke, T. HamNoSys 4 handshapes chart. DGS-Korpus Project, University of Hamburg (2010). URL https://w.sign-lang.un i-hamburg.de/dgs-korpus/files/inhalt_pd f/HamNoSys_Handshapes.pdf. Dated 10 June 2010; drawings by Heiko Zienert, Olga Jeziorski, and Andreas HanĂ. [20] Ferlin, M. et al. Quantifying inconsistencies in the hamburg sign language notation sys- tem. Expert Systems with Applications 256, 124911 (2024). [21] Brentari, D. & Eccarius, P. Handshape contrasts in sign language phonology. In Brentari, D. (ed.) Sign Languages, Cam- bridge Language Surveys, 284â311 (Cam- bridge University Press, Cambridge, UK, 2010). [22] Brentari, D., Coppola, M., Cho, P. W. & Senghas, A. Handshape complexity as a pre- cursor to phonology: Variation, emergence, and acquisition. Language Acquisition 24, 283â306 (2017). [23] Koller, O., Ney, H. & Bowden, R. Deep hand: How to train a cnn on 1 million hand images when your data is continuous and weakly labelled. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3793â3802 (IEEE, 2016). [24] Zhang, X. & Duh, K. Handshape-aware sign language recognition: Extended datasets and exploration of handshape-inclusive methods. In Findings of the Association for Computa- tional Linguistics: EMNLP 2023, 2993â3002 (Association for Computational Linguistics, Singapore, 2023). URL https://aclanthology .org/2023.findings-emnlp.198/. 17 [25] Lobo-Neto, V. C. & Pedrini, H. LSWH100: A handshape dataset for brazilian sign language (Libras) using SignWriting. Data in Brief 56, 110780 (2024). [26] DataMunge. Sign language mnist. Kaggle dataset (2017). URL https://w.kaggle.c om/datasets/datamunge/sign-language-m nist. C0: Public Domain; accessed 21 July 2026. [27] Pugeault, N. & Bowden, R. Spelling it out: Real-time ASL fingerspelling recogni- tion. In 2011 IEEE International Confer- ence on Computer Vision Workshops (ICCV Workshops), 1114â1119 (IEEE, Barcelona, Spain, 2011). [28] Hosoe, H., Sako, S. & Kwolek, B. Recogni- tion of jsl finger spelling using convolutional neural networks. In 2017 Fifteenth IAPR International Conference on Machine Vision Applications (MVA), 85â88 (IEEE, Nagoya, Japan, 2017). [29] Jain, S. ADDSL: Hand gesture detec- tion and sign language recognition on anno- tated danish sign language. arXiv preprint arXiv:2305.09736 (2023). URL https://arxi v.org/abs/2305.09736. [30] Ronchetti, F., Quiroga, F., Estrebou, C. A., Lanzarini, L. C. & Rosete, A. LSA64: An argentinian sign language dataset. In Pro- ceedings of the XXII Congreso Argentino de Ciencias de la ComputaciĂłn (CACIC 2016), 794â803 (Red de Universidades con Carreras en InformĂĄtica (RedUNCI), 2016). URL ht tp://sedici.unlp.edu.ar/handle/10915/5676 4. [31] Kajiyama, T., Endo, R., Kaneko, H., Sano, M. & Shishikui, Y. Sign language image dataset with a hand pose type attribute. In 2022 IEEE International Symposium on Broadband Multimedia Systems and Broad- casting (BMSB), 1â4 (IEEE, 2022). [32] Hanke, T., Schulder, M., Konrad, R. & Jahn, E. Extending the public DGS cor- pus in size and depth. In Proceedings of the LREC2020 9th Workshop on the Repre- sentation and Processing of Sign Languages: Sign Language Resources in the Service of the Language Community, Technological Chal- lenges and Application Perspectives, 75â82 (European Language Resources Association (ELRA), Marseille, France, 2020). URL http s://aclanthology.org/2020.signlang-1.12/. [33] Konrad, R. et al. (eds.) FachgebĂ€rdenlexikon Gesundheit und Pflege (Signum, Seedorf, Germany, 2007). URL http://w.sign-lan g.uni-hamburg.de/glex. [34] Matthes, S. et al. DICTA-SIGNâbuilding a multilingual sign language corpus. In Proceedings of the LREC2012 5th Work- shop on the Representation and Processing of Sign Languages: Interactions between Cor- pus and Lexicon, 117â122 (European Lan- guage Resources Association (ELRA), Istan- bul, Turkey, 2012). URL https://w.sign-l ang.uni-hamburg.de/lrec/pub/12016.html. [35] Ćacheta, J., Czajkowska-Kisil, M., Linde-Usiekniewicz, J. & Rutkowski, P. (eds.) Korpusowy SĆownik Polskiego JÄzyka Migowego/Corpus-based Dictio- nary of Polish Sign Language (Faculty of Polish Studies, University of War- saw, Warsaw, Poland, 2016).URL https://w.slownikpjm.uw.edu.pl/en. [36] Villa-Monedero, M., Gil-MartĂn, M., SĂĄez- Trigueros, D., Pomirski, A. & San-Segundo, R. Sign language dataset for automatic motion generation. Journal of Imaging 9, 262 (2023). [37] Varanasi, A. B., Sinha, M. & Dasgupta, T. Cross-linguistic phonological similarity anal- ysis in sign languages using HamNoSys. In Proceedings of the Workshop on Sign Lan- guage Processing (WSLP), 51â66 (Associa- tion for Computational Linguistics, IIT Bom- bay, Mumbai, India, 2025). URL https: //aclanthology.org/2025.wslp-main.9/. [38] Mitra, S. & Acharya, T. Gesture recognition: A survey. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 37, 311â324 (2007). 18 [39] He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770â778 (IEEE, 2016). [40] Dosovitskiy, A. et al. An image is worth 16x16 words: Transformers for image recogni- tion at scale. In International Conference on Learning Representations (2021). URL https: //openreview.net/forum?id=YicbFdNTTy. [41] Zhang, F. et al. MediaPipe Hands: On-device real-time hand tracking. arXiv preprint arXiv:2006.10214 (2020). URL https://arxi v.org/abs/2006.10214. [42] Kipf, T. N. & Welling, M. Semi-supervised classification with graph convolutional net- works. In International Conference on Learn- ing Representations (2017). URL https://op enreview.net/forum?id=SJU4ayYgl. [43] Sarkar, U., Chakraborti, A., Samanta, T., Pal, S. & Das, A. Enhancing asl recognition with gcns and successive residual connec- tions. In Palaiahnakote, S. et al. (eds.) Pattern Recognition. ICPR 2024 Interna- tional Workshops and Challenges, vol. 15616 of Lecture Notes in Computer Science, 3â16 (Springer Nature Switzerland, Cham, 2025). [44] Chen, T. & Guestrin, C. XGBoost: A scal- able tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Con- ference on Knowledge Discovery and Data Mining, 785â794 (Association for Computing Machinery, New York, NY, USA, 2016). [45] Aly, S. & Aly, W. DeepArSLR: A novel signer-independent deep learning framework for isolated arabic sign language gestures recognition. IEEE Access 8, 83199â83212 (2020). 19