Paper deep dive
Pediatric Bone Age Prediction Using Deep Learning
Al Zadid Sultan Bin Habib, Md. Ekramul Islam, Md Asif Bin Syed, Md Younus Ahamed, Tanpia Tasnim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/21/2026, 4:17:49 AM
Summary
This paper presents a deep learning approach for pediatric bone age prediction using hand X-ray images from the RSNA dataset. The authors utilize EfficientNet models (B0 and B4) and a novel variant, EfficientNetB4 with Additive Attention (EN-AA), to predict bone age. The study demonstrates that EfficientNetB4 and EN-AA outperform EfficientNetB0 in accuracy, leveraging transfer learning and data augmentation to handle image preprocessing and feature extraction effectively.
Entities (7)
Relation Signals (7)
RSNA Bone Age Dataset → usedby → EfficientNetB0
confidence 95% · The method utilizes over 12,000 X-ray images from the RSNA bone age dataset... This work uses two variations of the EfficientNet model (B0 and B4)
RSNA Bone Age Dataset → usedby → EfficientNetB4
confidence 95% · The method utilizes over 12,000 X-ray images from the RSNA bone age dataset... This work uses two variations of the EfficientNet model (B0 and B4)
RSNA Bone Age Dataset → usedby → EN-AA
confidence 95% · The method utilizes over 12,000 X-ray images from the RSNA bone age dataset... This work uses two variations of the EfficientNet model (B0 and B4)
EN-AA → uses → Additive Attention
confidence 95% · EfficientNetB4 is also finetuned with the Additive Attention mechanism.
EfficientNetB4 → outperforms → EfficientNetB0
confidence 90% · EfficientNetB4 and EfficientNetB4 with Additive Attention (EN-AA) successfully predicted the bone ages more accurately... and their performance was better than the EfficientNetB0.
EN-AA → outperforms → EfficientNetB0
confidence 90% · EfficientNetB4 and EfficientNetB4 with Additive Attention (EN-AA) successfully predicted the bone ages more accurately... and their performance was better than the EfficientNetB0.
Bone Age Prediction → diagnoses → Endocrine Disorders
confidence 85% · Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders and provide insight into a child's growth and development. However, conventional bone age prediction methods are often labor-intensive and require specialized radiological expertise. This paper presents a Deep Learning (DL)-based approach to pediatric bone age prediction using EfficientNet with Additive Attention, a state-of-the-art neural network architecture for image classification and regression tasks. The method utilizes over 12,000 X-ray images from the RSNA bone age dataset. It involves image preprocessing, transforming them into three-channel images, and training a Convolutional Neural Network (CNN) to automatically learn the features of hand bone images. This approach provides a more effective and accurate solution for predicting bone age, which is critical in diagnosing pediatric endocrine diseases. This work uses two variations of the EfficientNet model (B0 and B4), where EfficientNetB4 is also finetuned with the Additive Attention mechanism. These three models predict the age for the original age, and their comparison is shown in curves. The predicted ages depict that in most cases, EfficientNetB4 and EfficientNetB4 with Additive Attention (EN-AA) successfully predicted the bone ages more accurately regarding the original age, and their performance was better than the EfficientNetB0. Specific performance metrics are provided to underscore this improvement. Learning curves for training and validation loss confirm effective learning without overfitting or underfitting, further validating our approach's efficacy in pediatric endocrine disease diagnosis.
Tags
Links
- Source: https://arxiv.org/abs/2607.16936v1
- Canonical: https://arxiv.org/abs/2607.16936v1
Trouble viewing inline? Open PDF directly →
Full Text
20,677 characters extracted from source content.
Expand or collapse full text
2023 26th International Conference on Computer and Information Technology (ICCIT) 13–15 December 2023, Cox’s Bazar, Bangladesh Pediatric Bone Age Prediction Using Deep Learning Al Zadid Sultan Bin Habib 1 , Md. Ekramul Islam 2 , Md Asif Bin Syed 3 , Md Younus Ahamed 4 , Tanpia Tasnim 5 1,4 Lane Department of Computer Science and Electrical Engineering, West Virginia University, Morgantown, WV 26506, USA 2 Department of Computer Science & Engineering, Stamford University Bangladesh, Dhaka-1217, Bangladesh 3 Department of Industrial and Management Systems Engineering, West Virginia University, Morgantown, WV 26506, USA 5 Department of Computer Science and Engineering, Green University of Bangladesh, Narayanganj-1461, Dhaka, Bangladesh Email: ah00069@mix.wvu.edu 1 , eislam706@gmail.com 2 , ms00110@mix.wvu.edu 3 , ma00087@mix.wvu.edu 4 , tanpia@cse.green.edu.bd 5 Abstract—Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders and provide insight into a child’s growth and development. However, conventional bone age prediction methods are often labor-intensive and require specialized radiological expertise. This paper presents a Deep Learning (DL)-based approach to pediatric bone age prediction using EfficientNet with Additive Attention, a state-of-the-art neural network architecture for image classification and regression tasks. The method utilizes over 12,000 X-ray images from the RSNA bone age dataset. It involves image preprocessing, transforming them into three- channel images, and training a Convolutional Neural Network (CNN) to automatically learn the features of hand bone images. This approach provides a more effective and accurate solution for predicting bone age, which is critical in diagnosing pedi- atric endocrine diseases. This work uses two variations of the EfficientNet model (B0 and B4), where EfficientNetB4 is also finetuned with the Additive Attention mechanism. These three models predict the age for the original age, and their comparison is shown in curves. The predicted ages depict that in most cases, EfficientNetB4 and EfficientNetB4 with Additive Attention (EN-A) successfully predicted the bone ages more accurately regarding the original age, and their performance was better than the EfficientNetB0. Specific performance metrics are provided to underscore this improvement. Learning curves for training and validation loss confirm effective learning without overfitting or underfitting, further validating our approach’s efficacy in pediatric endocrine disease diagnosis. Index Terms—RSNA Bone Age, Deep Learning, EfficientNet, Additive Attention, Pediatrics, Image Regression. I. INTRODUCTION In the rapidly evolving fields of Computer Science (CS), technology, and medical science, numerous cross-disciplinary advancementshavesignificantlyimpactedhealthcare. Traditionally, medical procedures like X-ray photo scanning and analysis were performed manually. However, with the advent of computer image processing technologies, these tasks have become more efficient and accurate. In medical practice, age is a crucial measure of human growth, but bone age is often a more accurate indicator of biological maturity, as detailed in [1]. Typically, an X-ray of the left wrist is used to assess bone age, which is essential in evaluating adolescent development, screening for genetic disorders, and talent assessment. Traditionally, most X-ray analyses Accepted at ICCIT 2023. This version is the author-prepared manuscript. The final published version appeared in IEEE Xplore. involve manual comparison of a patient’s radiograph with an age-specific atlas or using a bone-specific scoring system. However, these manual methods are time-consuming, subject to limitations, and prone to errors. Numerous domestic and international initiatives have emerged focusing on X-ray bone age image learning, reflect- ing the growing intersection of medical imaging and CS. Advanced software solutions such as BoneXpert [2] have employed Computer Vision (CV) techniques to reconstruct hand bone contours. However, challenges still need to be solved, particularly with low-quality images. The advent of Ar- tificial Intelligence (AI) technology, alongside advancements in computer and graphics technology, offers new avenues for tackling these interdisciplinary challenges. The field of bone age assessment has progressively shifted from manual to automated evaluations, thanks to AI and DL innovations. In this context, our work utilizes the variations of the EfficientNet model with the Additive Attention mechanism [3], showcasing its application in automated and precise bone age assessment. This network optimizes the network depth, width, and image resolution balance to achieve ideal performance. EfficientNet surpasses traditional models in size-to-accuracy ratio [4]. A separable CNN standardizes and normalizes the input for processing X-ray images. This involves segmenting the hand area for image registration, focusing on the central region of interest, and rotating key points to isolate the hand [5]. Utilizing the DeepLabv3 plus [6] and MobileNetV1 [7] ar- chitectures, the network employs separable convolutions as its core unit. We apply various rotations, scaling, and cropping techniques to enhance the dataset and prevent underfitting. Previous research has yet to extensively explore bone age estimation in hand X-ray images using the EfficientNet model. We aim to establish a rapid and efficient method for this task, leveraging the latest advancements in DL models. For further details, readers are directed to [7]–[13] and related references. This study leverages DL and image processing techniques for predicting bone age from over 12,000 children’s hand bone X-ray images in the RSNA bone age dataset [14]. The approach includes three key steps: (1) Preprocessing the X- ray images: size normalization, noise reduction, and histogram equalization, to standardize the images for DL. (2) Feature extraction using the EfficientNet model, focusing on extracting arXiv:2607.16936v1 [cs.CV] 18 Jul 2026 Fig. 1: Sample X-ray bone images. relevant features from hand bone images. (3) Applying a CNN for DL to automatically extract features and accurately determine bone age. The key contributions of this work are summarized as follows: ■ Investigated the application of pre-trained EfficientNet models in the context of bone age prediction, demon- strating their effectiveness in a medical imaging domain. ■ Conducted a comparative analysis of various EfficientNet model variants, providing insights into their performance in bone age estimation, a critical task in pediatric health- care. ■ Innovatively adapted the Additive Attention mechanism, traditionally used in NLP, to the EfficientNetB4 model, creating an enhanced EN-A model for the specific demands of image regression tasks in medical imaging. The subsequent sections of the paper are organized in the following manner: Section I outlines the methodological framework utilized in this work. At the same time, Section I offers a comprehensive explanation of the obtained results. The paper is concluded in Section IV by outlining the limitations and future research directions. I. METHODOLOGICAL FRAMEWORK A. Preprocessing The dataset for our study, sourced from RSNA 2017 [14], consists of over 12,000 X-ray images showcasing children’s hand bones, including a sample depicted in Fig: 1. Accompa- nying each image are relevant details like age and gender. A notable characteristic of this dataset is the variation in image resolution and size, with the dimensions of hand bone images generally ranging from 800 to 2200 pixels in length and width, as identified through statistical analysis. The RSNA X-ray images were divided into two sets for model training and evaluation: 70% for training and validation and 30% reserved for testing. Fig. 2: Histogram of the bone ages before and after normal- ization. Fig. 3: EfficientNetB0 architecture. B. EfficientNet Model EfficientNetB0, the most compact model within the Effi- cientNet family [3], distinguishes itself with only 5.3 mil- lion parameters, far fewer than larger models like ResNet50 or InceptionV3. This efficiency makes it ideal for compu- tational resource conservation. The architecture comprises seven blocks, each integrating convolutions, batch normaliza- tion, and activation functions, designed for an input size of 224x224 pixels, aligning with the ImageNet dataset standards. A notable feature of EfficientNetB0 is its compound scaling technique, balancing depth, width, and resolution, enhancing its performance in diverse CV applications, including image classification, object detection, and segmentation. Regarding effectiveness, EfficientNetB0 attains a top-1 accuracy of 76.3% and a top-5 accuracy of 93.2% on the ImageNet dataset [15], making it competitive among small-scale models. EfficientNetB0, part of the EfficientNet series, is optimized for CV tasks using a modular design that combines multi- ple building blocks and subblocks. Building blocks feature convolutional layers with a bottleneck structure for efficiency, followed by Squeeze-and-Excitation (SE) modules for feature recalibration. Subblocks, which are more straightforward and used in early layers, lack the SE module and have fewer filters. The models incorporate a compound scaling method, adjusting filters, depth, and resolution for balanced archi- tectural complexity. This design approach, coupled with a Fig. 4: Modules in EfficientNet architectures. Fig. 5: Sub-blocks in EfficientNet architectures. standardized stem structure comprising convolutional layers, batch normalization, and max pooling, enables EfficientNetB0 to process low-level features efficiently before advancing to complex tasks. Its stem’s consistent design facilitates scala- bility and comparative analysis within the EfficientNet family, maintaining both efficiency and effectiveness. C. EN-A 1) Image Model: The initial segment of the model employs an EfficientNetB4, pre-trained on ImageNet, for extracting image features. It processes an input tensor with dimensions Fig. 6: Additive Attention architecture. (256, 256, 1), outputting a tensor sized (8, 8, 1792) after the concluding convolutional layer. This tensor undergoes flattening through a flattened layer, resulting in a vector of dimensions (114688). Subsequently, this vector is fed into a sequence of three dense layers containing 128, 64, and 32 neurons, respectively, and employing the ReLU activation function. The final output from this image model is a tensor with dimensions (32). Fig. 7: EfficientNetB4 architecture. 2) Gender Model: The second component of the model is dedicated to processing gender data through a straightforward dense network. It begins with an input tensor sized (1) and channels it through two dense layers. The first layer comprises 64 neurons and the second 32, utilizing the ReLU activation function. The output from this gender-focused model is a tensor with dimensions (32). 3) Concatenate: The output tensors from both the image and gender models are merged using the Concatenate layer in Keras. This operation forms a combined tensor with a shape of (64), effectively integrating the distinct feature sets derived from each model into a single, unified tensor for subsequent processing steps. 4) Attention Mechanism: The concatenated tensor is pro- cessed through an Additive Attention mechanism employing the previously defined additive attention function. The resul- tant output from this attention mechanism is a weighted sum of the concatenated tensor, effectively focusing on specific features within it [16]. This approach allows the model to se- lectively emphasize essential elements in the tensor, enhancing its overall interpretative capability. 5) Dense Layers: The weighted sum tensor undergoes further processing by passing through two dense layers. The first layer contains 32 neurons, and the second comprises 16, utilizing the ReLU activation function. This arrangement allows for additional refinement and extraction of features from the tensor. 6) Output Layer: The model’s final output is generated by a single dense layer of one neuron equipped with a linear activation function. This layer predicts the target variable, translating the processed features into the outcome. 7) Model Definition: The model is built using Keras’ Model class, combining image and gender input layers into a single output layer. This structure effectively merges distinct data sources for a holistic prediction. I. RESULTS ANALYSIS The study focuses on predicting bone age from hand X- ray images of children using the EfficientNet variants and EN-A model. The primary data source is the RSNA Bone Age Assessment dataset comprising hand X-ray images with corresponding bone ages. Employing transfer learning, the pre-trained EfficientNetB0 and EfficientNetB4 models were finetuned on this dataset. Data augmentation techniques like rotation, zooming, and flipping were utilized to enhance the model’s efficacy, effectively expanding the dataset size. Further experimentation involved adjusting hyperparameters and refining data augmentation techniques to optimize model performance. The outcomes affirm the efficacy of using a pre- trained EfficientNet model coupled with data augmentation for bone age prediction. The age predictions made by different EfficientNet models are illustrated in Figures 8, 9, and 10. These figures represent the results from EfficientNetB0, Effi- cientNetB4, and EN-A, respectively. Additionally, Figures 11 and 12 display the learning curves for EfficientNetB4 and EN-A. These demonstrate a decline in training and validation loss over epochs, indicating effective learning and model optimization. A closer analysis of individual predictions reveals variations in performance. From Fig. 13, the predicted ages for EfficientNetB4 model can be observed with respect to the original age. For instance, EfficientNetB4 estimated the ages as 15, 12, 12, and 13 months against the original ages of 17, 14, 11, and 13 months, respectively. Similar outcomes can be observed for EN-A model in Fig. 14. EN-A predictions for ages 14, 5, 13, and 11 months were 11, 4, 15, and 11 months, respectively. Fig. 15 represent the predicted ages by EfficientNetB0 model. The model predicted the ages as 115.9, 113.8, 126, 114.7, 115.3, 121.7, 117.6, and 117.8 months correspondingly for 96, 150, 132, 60, 15, 138, 210, and 24 months. These individual case studies highlight the nuanced performance differences between models. It clearly noticeable that EfficientNetB4 and EN-A models were superior com- pared to the EfficientNetB0 model. Figures 15, 13, and 14 further elucidate these differences by showcasing bone images with both predicted and original ages for EfficientNetB0, EfficientNetB4, and EN-A models, respectively. These visual representations underscore the varying degrees of accuracy each model variant achieves. The study demonstrates the potential of using pre-trained EfficientNet models, enhanced by data augmentation strategies, for accurately predicting bone age from X-ray images. The comparative analysis of different EfficientNet models, including those enhanced with Additive Attention, provides valuable insights into the capabilities and limitations of these approaches in the context of bone age assessment. Fig. 8: Predicted age by EfficientNetB0. Fig. 9: Predicted age by EfficientNetB4. IV. CONCLUSIONS In this study, we employed the EfficientNet models for the critical task of bone age estimation, which is integral to assessing human development and aiding in disease diagnosis. The process involved vital steps such as detailed data prepro- cessing: resizing, normalization, and enhancement to reduce data bias and expand the training set. EfficientNetB0 was pivotal in extracting intricate image features. These features, combined with gender data, were then processed through a fully connected (FC) layer, culminating in the bone age predic- tion. Using a pre-trained model was instrumental in achieving rapid convergence and enhancing the overall predictive accu- racy. Additionally, the study delved into the capabilities of Fig. 10: Predicted age by EN-A. Fig. 11: Learning curves for EfficientNetB4. Fig. 12: Learning curves for EN-A. Fig. 13: Output images with predicted ages by EfficientNetB4. Fig. 14: Output images with predicted ages by EN-A. Fig. 15: Output images with predicted ages by EfficientNetB0. EfficientNetB4 and its variant with Additive Attention (EN- A). Findings indicated that EfficientNetB4 outperformed the B0 model, with the Additive Attention variant demonstrating comparable proficiency. The study identifies two primary avenues for advancement: deploying advanced versions of EfficientNet coupled with specialized attention modules to more precisely target and analyze regions of interest in X- ray images and the extension of training epochs to refine the models further. These future endeavors aim to improve the accuracy of bone age predictions and prepare these models for broader clinical application and eventual market integration. REFERENCES [1] V. Gilsanz and O. Ratib, Hand Bone Age: A Digital Atlas of Skeletal Maturity. Berlin, Heidelberg: Springer, 2005. [2] H. H. Thodberg, S. Kreiborg, A. Juul, and K. D. Pedersen, “The bonexpert method for automated determination of skeletal maturity,” IEEE Transactions on Medical Imaging, vol. 28, no. 1, p. 52–66, 2009. [3] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for con- volutional neural networks,” in International conference on machine learning. PMLR, 2019, p. 6105–6114. [4] V. Kajla, A. Gupta, and A. Khatak, “Analysis of x-ray images with image processing techniques: A review,” in 2018 4th International Conference on Computing Communication and Automation (ICCCA), 2018, p. 1–4. [5] V. I. Iglovikov, A. Rakhlin, A. A. Kalinin, and A. A. Shvets, “Paediatric bone age assessment using deep convolutional neural networks,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th International Workshop, ML-CDS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, Proceedings 4. Springer, 2018, p. 300–308. [6] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder- decoder with atrous separable convolution for semantic image segmen- tation,” in Proceedings of the European conference on computer vision (ECCV), 2018, p. 801–818. [7] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convo- lutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017. [8] Z. Wenxiang, “Automatic evaluation of bone age based on x-ray images,” Ph.D. dissertation, Doctoral dissertation) Go to reference in article, 2018. [9] X. Pan, Y. Zhao, H. Chen, D. Wei, C. Zhao, and Z. Wei, “Fully automated bone age assessment on large-scale hand x-ray dataset,” International journal of biomedical imaging, vol. 2020, 2020. [10] A. M. Mughal, N. Hassan, and A. Ahmed, “Bone age assessment methods: a critical review,” Pakistan journal of medical sciences, vol. 30, no. 1, p. 211, 2014. [11] C. Spampinato, S. Palazzo, D. Giordano, M. Aldinucci, and R. Leonardi, “Deep learning for automated skeletal bone age assessment in x-ray images,” Medical image analysis, vol. 36, p. 41–51, 2017. [12] S. Koitka, A. Demircioglu, M. S. Kim, C. M. Friedrich, and F. Nensa, “Ossification area localization in pediatric hand radiographs using deep neural networks for object detection,” PloS one, vol. 13, no. 11, p. e0207496, 2018. [13] V. De Sanctis, S. Di Maio, A. T. Soliman, G. Raiola, R. Elalaily, and G. Millimaggi, “Hand x-ray in pediatric endocrinology: Skeletal age assessment and beyond,” Indian journal of endocrinology and metabolism, vol. 18, no. Suppl 1, p. S63, 2014. [14] “Rsna bone age — kaggle,” https://w.kaggle.com/datasets/kmader/rsna- bone-age, (Accessed on 04/15/2023). [15] M. Tan and Q. Le, “Efficientnetv2: Smaller models and faster training,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, p. 10 096–10 106. [Online]. Available: https://proceedings.mlr.press/v139/tan21a.html [16] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014.