Paper deep dive
A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities
Rafael Muñoz-Terol, Jesús Peral, Sandra Amador, David Gil
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/20/2026, 4:00:19 AM
Summary
This systematic review evaluates 55 studies published between 2017 and 2023 regarding the application of machine learning (ML) techniques for the diagnosis and treatment of Autism Spectrum Disorder (ASD). The review finds that supervised learning methods are the most frequently used due to their alignment with diagnostic classification needs, while deep learning and hybrid methods are emerging. Key challenges include the need for complex data integration (genetic, clinical, wearable data) and interdisciplinary collaboration to improve diagnostic accuracy and treatment outcomes.
Entities (11)
Relation Signals (8)
Machine Learning → appliedto → Autism Spectrum Disorder
confidence 100% · This systematic review evaluates 55 studies from 2017 to 2023 on the application of machine learning (ML) techniques to ASD.
Supervised Learning → dominates → Machine Learning
confidence 95% · Supervised learning methods dominate, as they align well with ASD diagnostic needs
Brain Data → mostfrequentlyused → Machine Learning
confidence 95% · brain data are the most frequently utilized data type in these studies (45 %)
Hybrid Learning → combines → Deep Learning
confidence 90% · hybrid methods, where unsupervised, deep learning, and fuzzy logic could be included
Hybrid Learning → combines → Unsupervised Learning
confidence 90% · hybrid methods, where unsupervised, deep learning, and fuzzy logic could be included
Wearable Devices → enables → Continuous Monitoring
confidence 90% · incorporating innovative data sources, like wearable devices and biometric sensors, could enable continuous and non-intrusive monitoring
Deep Learning → expandingrole → Machine Learning
confidence 90% · however, the role of deep learning is expanding with greater data availability.
Support Vector Machine → usedfor →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. As the causes of ASD remain unclear, identifying relevant features and hidden correlations is crucial for early diagnosis. This systematic review evaluates 55 studies from 2017 to 2023 on the application of machine learning (ML) techniques to ASD. The primary objective is to examine recent ML applications in ASD research, identifying trends, techniques, and datasets that enhance diagnosis and treatment. Supervised learning methods dominate, as they align well with ASD diagnostic needs; however, the role of deep learning is expanding with greater data availability. Emerging techniques based on hybrid methods, where unsupervised, deep learning, and fuzzy logic could be included, will be interesting to observe in the future. The review highlights key challenges and opportunities, particularly the need for models that can integrate complex data -such as genetic and clinical information- to improve diagnostic accuracy and treatment outcomes. Additionally, incorporating innovative data sources, like wearable devices and biometric sensors, could enable continuous and non-intrusive monitoring, providing a more holistic understanding of ASD. Findings emphasize that addressing current challenges requires interdisciplinary collaboration and expanded datasets tailored to ASD. Future ML models will benefit from broader multimodal data integration, enabling researchers to more comprehensively address the complexities of ASD.
Tags
Links
- Source: https://arxiv.org/abs/2608.18188v1
- Canonical: https://arxiv.org/abs/2608.18188v1
Trouble viewing inline? Open PDF directly →
Full Text
75,501 characters extracted from source content.
Expand or collapse full text
Research article A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities Rafael Mu ̃ noz-Terol a,* , Jesús Peral a , Sandra Amador b , David Gil c a Lucentia Research Group, Department of Software and Computing Systems, University of Alicante, 03690 Alicante, Spain b U.I. for Computer Research, Alicante, Spain c Department of Computer Science Technology and Computation, University of Alicante, 03690 Alicante, Spain ARTICLE INFO Keywords: Autism spectrum disorder (ASD) Machine learning techniques Supervised learning Unsupervised learning Hybrid learning Deep learning ABSTRACT Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. As the causes of ASD remain unclear, identifying relevant features and hidden correlations is crucial for early diagnosis. This systematic review evaluates 55 studies from 2017 to 2023 on the application of machine learning (ML) techniques to ASD. The primary objective is to examine recent ML applications in ASD research, identifying trends, techniques, and datasets that enhance diagnosis and treatment. Supervised learning methods dominate, as they align well with ASD diagnostic needs; however, the role of deep learning is expanding with greater data availability. Emerging techniques based on hybrid methods, where unsupervised, deep learning, and fuzzy logic could be included, will be interesting to observe in the future. The review highlights key challenges and opportunities, particularly the need for models that can integrate complex data —such as genetic and clinical information— to improve diagnostic accuracy and treatment outcomes. Additionally, incorporating innovative data sour- ces, like wearable devices and biometric sensors, could enable continuous and non-intrusive monitoring, providing a more holistic understanding of ASD. Findings emphasize that address- ing current challenges requires interdisciplinary collaboration and expanded datasets tailored to ASD. Future ML models will benefit from broader multimodal data integration, enabling re- searchers to more comprehensively address the complexities of ASD. 1. Introduction Over the past few years, research on autism spectrum disorder (ASD) has been one of the most important topics in computational psychiatry. ASD is a developmental disability characterised by impaired social communication and social interactions, in addition to restricted and constant patterns of conduct (Diagnostic and Statistical Manual of Mental Disorders 1 , DSM-5). Therefore, ASD is a divergent disability related to the appearance of symptoms and severity, risk factors, study, and treatment responses. Regarding computational psychiatry and ASD research frameworks, various machine learning (ML) approaches have been used [1–3], which *Corresponding author. E-mail addresses: rafamt@dlsi.ua.es(R. Mu ̃ noz-Terol), jperal@dlsi.ua.es(J. Peral), saandra.amador@gmail.com(S. Amador), david.gil@ua.es (D. Gil). 1 https://w.psychiatry.org/Psychiatrists/Practice/DSM/Updates-to-DSM/Coding-Updates/2021-Coding-Updates(visited on November 20, 2025). Contents lists available at ScienceDirect Heliyon journal homepage: w.cell.com/heliyon https://doi.org/10.1016/j.heliyon.2025.e44359 Received 8 March 2023; Received in revised form 24 November 2025; Accepted 8 December 2025 Heliyon 12 (2026) e44359 Available online 3 January 2026 2405-8440/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the C BY-NC-ND license ( http://creativecommons.org/licenses/by-nc-nd/4.0/ ). have demonstrated better results than the knowledge-based approaches. Therefore, in these children psychiatry-applied framework, the best doctors’ decisions will enhance the well-being of patients with ASD. This fact implies that the application of ML approaches to the ASD problem is a research trend in detection and management tasks. In this way, ML algorithms are traditionally classified into two different categories according to their learning methods: unsupervised and supervised learning. Thus, as noted by Lloyd et al. [4], in supervised learning, the machine assumes the role from a set of training examples, whereas in unsupervised learning, the machine attempts to detect hidden structures in unlabelled data. Moreover, there is a third category of ML classification algorithms for studying the ASD problem, namely, hybrid algorithms that combine ML techniques and knowledge-based resources. Various ML-based approaches have been implemented in ASD detection and observation frameworks. Crippa et al. [5] conducted an evidence-of-concept study to determine whether an easy higher-limb gesture could be used in categorising low-functioning infants (aged 2–4 years) with ASD by developing a supervised ML method that correctly identifies 15 kindergarten infants with ASD from 15 commonly developing infants through the kinematic analysis of a simple reach-to-drop task. In addition, Duda et al. [6] trained and tested three pairs of ML models on the full 65-element Social Responsiveness Scale result sheets from a set of individuals with either ASD or attention deficit hyperactivity disorder (ADHD), whose scores justified the suitability of this method for distinguishing ASD and ADHD with high accuracy. Bone et al. [7] utilized ML to infer ASD instrument algorithms and enhance commonly used ASD screening and diagnostic tools by applying support vector machine (SVM) approaches. Usta et al. [8] tested the performance of four ML tech- niques: naive Bayes, generalized linear model, logistic regression, and decision trees (DTs). They concluded that these ML models indicate that several others are more remarkable in terms of predictive information and management of the treatment of children with ASD. Rudovic et al. [9] used the latest approaches in deep learning to propose a personalised ML framework for the automatic discrimination of the emotional behaviours of children. Furthermore, it includes the arrangement during robot-assisted autism treatment exhibiting the viability of robot perception of sweetie and commitment in infants with autism and has inferences in the design process of future autism treatments. Thabtah and Peebles [10] proposed a new ML method that in addition distinguishes autistic features of situations and checkpoints and provides users with knowledge rules that can be utilized by domain professionals to recognize the purposes after the classification. Thus, empirical scores from three data sets show that rule-based machine learning provides classifiers with better predictive accuracy, harmonic mean, specificity, and sensitivity than other ML approaches. Tariq et al. [11] theorised that ML tests performed on home videos could rate an assay without compromising accuracy. Consequently, they tested item-level records using a pair of standard analytical tools to develop ML classifiers that enhanced spareness, explainability, and accuracy. They eventually checked whether features from the improved models could be considered by blinded amateur raters from 3-min home videos of infants with and without ASD for rapid and correct ML autistic classification. Liu et al. [12] applied an ML method to detect an eye motion dataset from a face recognition task to categorise infants with and without ASD. Thus, they analysed the performance of the model by considering its sensitivity, accuracy, and specificity in categorising ASD; the scores confirmed the utility of the ML algorithm that manages the face-scanning patterns for identifying infants with ASD. ML hybrid algorithms combine multiple ML techniques or models to develop a new approach that leverages the strengths of each individual component. These algorithms aim to enhance the performance, improve accuracy, and tackle specific challenges that a single technique cannot adequately resolve. In recent years, the research community has explored different approaches for developing hybrid algorithms. These methods include ensemble methods [13], feature selection/extraction combinations [14], and model stacking [15]. Ensemble methods involve a combination of multiple base models, such as DTs, neural networks, and SVMs in the prediction process. In contrast, feature selection/extraction combinations mix different subsets of features or data representations to improve the final prediction or decision. Finally, model stacking trains multiple models on the same dataset and combines their predictions using another model, often referred to as a meta-model or blender. Hybrid-based approaches have also been applied to the ASD problem. Alharbi et al. [16] presented an expert system to evaluate autism under unpredictability that used trusted knowledge-based inference methods with the evidential reasoning technique, where the knowledge base of the system was obtained by utilizing real data from affected people, and by considering doctors’ opinions. Alam et al. [17] proposed an Internet-of-Things belief knowledge base-based hybrid system. This intelligent approach can instinctively identify the signs and symptom data of different autistic infants in real-time and categorised them. The belief-rule-base subsystem included knowledge description variables to accomplish this goal. Thus, infants with autism were categorised based on the signs and symptoms identified by pervasive sensing components. Mugzach et al. [18] developed an ontology that allowed data integration and reasoning with patient data to categorise subjects, and deduce new rules based on ASD and related neurodevelopmental disorders based on this categorization. Gong et al. [19] demonstrated a method to prognosticate autism propensity genes, where genes were first extracted from biomedical literature, and some autism liability genes were then perceived as kernels by the superior understanding. Persisting genes were predicted by developing alliance rules between the kernels and candidates. We can summarise that hybrid learning techniques developed in recent years have shown great potential to improve diagnostic accuracy and personalize treatments in the context of ASD. Unlike traditional methods, hybrid models combine multiple machine learning techniques, such as integrating neural networks with rule-based systems or optimization algorithms. This approach provides greater flexibility and adaptability in the classification and analysis of complex data, enabling a more detailed understanding of patterns associated with autism. At the same time, deep learning has gained traction due to its ability to recognize highly complex patterns in large volumes of multimodal data. Advanced neural network models and deep learning architectures, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs), could be integrated into future studies to improve the detection of the early ASD markers. Implementing transfer learning techniques, for example, would allow models to leverage knowledge from similar domains to enhance accuracy in limited ASD datasets. Additionally, explainable AI (XAI) holds the potential to provide interpretability in these models, helping medical professionals understand and justify algorithm-based decisions. These developments present a promising path for future research, combining the precision of deep learning with the clarity and adaptability needed in a medical R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 2 context. As can be observed from all the studies presented, since the 2010s a series of works have laid the foundations for the application of ML techniques for the detection, diagnosis, and treatment of ASD [5–7,12,13,16,20]. However, after analysing the previous state-of-the-art studies, we believe that a current survey that reviews ML techniques, their suitability and performance, as well as the creation and accessibility of datasets in autism is needed. Consequently, this systematic review aims to summarise the current ML-based approaches that have been applied to the ASD problem in order to address the research questions described in Table 1. We have focused on the last seven years, from 2017 to 2023, to provide an overview of the most modern trends in this area. First, the exploration strategies utilized to retrieve the relevant literature published over the last seven years are outlined. The various studies that describe the predominant techniques, venues, databases, and various performance metrics are also explained. Subsequently, the limitations, challenges, and metrics of all the analysed ML algorithms are described in the Discussion section. Finally, the conclusions and future research directions are presented. 2. Methods and materials The next two subsections describe the search strategies, the research questions and the criteria for choosing the studies included in this review. 2.1.Search strategy and research questions In this review, a search strategy was developed to retrieve studies related to ASD. The search task was performed using major database collections, including ScienceDirect, Scopus, and Multidisciplinary Digital Publishing Institute (MDPI), with the keywords “autism spectrum disorder”, “autism”, “ASD”, and “machine learning”. The review questions (RQs) used in this review are classified into eight items (Table 1). 2.2.Study selection This section outlines the methodology employed to systematically select, review, and classify studies on the application of ML techniques to ASD. To retrieve relevant research works, manuscripts were selected based on well-defined inclusion and exclusion criteria: 1. Inclusion Criteria: •Studies published between 2017 and 2023, capturing recent advancements and trends in ML applications to autism research, reflecting the latest technological innovations and developments in the field. •Manuscripts written in English to ensure consistency in interpretation and to align with the scope of our review. •Research works that specifically addressed ASD, autism, and the application of ML methods, with a focus on diagnosis, inter- vention, or behavioural analysis. 2. Exclusion Criteria: •Manuscripts that did not directly address ASD or apply ML to ASD-specific challenges. •Studies comprised of abstracts, conference proceedings, review articles, tutorials, or non-peer-reviewed materials. •Manuscripts focusing on ML applications outside of the ASD context, even if they used similar techniques in other domains. The process began by identifying studies through systematic keyword searches in ScienceDirect, Scopus, and MDPI databases, resulting in 163 initial publications (83 from ScienceDirect and 80 from Scopus and MDPI) that met the basic inclusion criteria. After duplicate removal, 92 records remained for initial screening. Records that were abstracts, conference proceedings, review articles, tutorials, non-English publications, or unrelated to the topic (n =37) were then excluded, resulting in 55 final studies. Table 1 Research questions that have been applied to the research studies. IDQuestionRationale/Motivation RQ1Where and when were the previous studies published?To determine the quality of scientific contributions. RQ2Which countries are the authors of the contributions based in?To identify the geographic distribution of contributions. RQ3Which type of ML approach is the most frequently used?To identify the most frequently investigated types of approaches (including combined methods) proposed in literature. RQ4Which ML techniques are the most frequently used?To identify the most frequently adopted techniques to develop ensemble approaches in ASD diagnosis in literature. RQ5Which data types are used most frequently in the selected studies?To identify the most frequently investigated data types. RQ6How many patients do the selected studies examine?To determine the number of patients used for ASD studies. RQ7Do they normally develop their own databases in their studies? If not, which public databases are the most frequently used? To identify the most commonly used public databases and determine the number of studies that created their own databases. RQ8Are ML techniques a good choice for managing ASD?To analyse the good performance of ML techniques according to the different features proposed in the literature. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 3 Each retrieved manuscript was independently reviewed by at least two researchers experienced in computation and data science (ML). The abstracts, methods, and results were carefully analysed to ensure that only studies of high relevance and methodological soundness were included. Fig. 1presents the PRISMA flowchart of our selection process [21,22], with all inclusion and exclusion steps outlined. Utilizing the PRISMA methodology enhances review quality and consistency, which is particularly beneficial for systematic reviews in medical and clinical research. This methodology ensures that our review focuses exclusively on high-quality research addressing ML applications in the ASD field, capturing significant trends, techniques, and datasets relevant to current and future advancements. 3. Results The following eight subsections describe the study conducted to answer each of the research questions regarding the review framework. 3.1.RQ1: Where and when were the previous studies published? These studies were published as high-quality contributions in international journals over the past seven years (between 2017 and 2023). All these international journals are ranked in the Journal Citation Report according to their impact factors that place them in one of the four quartiles: Q1, Q2, Q3 or Q4. Moreover, according to the journal-publishing countries listed, various journals from the Netherlands published the largest number of the research works (27,27 %), as shown in Table 2. Fig. 1.PRISMA flowchart that shows the process applied to the manuscripts considered in this study. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 4 3.2.RQ2: Which countries are the authors of the contributions based in? Table 3lists the number of contributions from each country considering two aspects. First, the column labelled as “corre- sponding_author_contributions” ranks the countries according to the country to which the corresponding author's institution belongs. Second, the column labelled as “remaining_author_contributions” ranks the countries according to the country to which the remaining authors' institutions belong. 3.3.RQ3: Which type of ML approach is the most frequently used? Table 4presents the different ML approaches used in all the studies. Interestingly, supervised learning approaches were used most frequently and obtained the highest percentage in detection tasks for individuals on the autism spectrum (Fig. 2). This was expected because the classification was known and supervised. However, it is important to consider the increase in the number of hybrid algorithms, which in many cases complement supervised learning via clustering techniques through a separate phase-by-phase process. As can be seen, most of the work relies on supervised methods since classification (autistic diagnosis) is a clear objective. The strengths of these techniques are thereby that they are effective for classification and prediction when labelled data are available, allowing for greater accuracy in pattern identification. Therefore, they are useful in the early detection of ASD through the identifi- cation of specific markers and the classification of symptoms associated with ASD. However, even though this is a strength, it can become a limitation when the dataset is large or when some records are not labelled where unsupervised techniques may be more appropriate. Their weaknesses are therefore the dependence on the quality and quantity of available labelled data, which may be limited in ASD clinical settings, as well as the risk of overfitting if training data do not adequately represent the diversity of the autistic spectrum. Alternatively, unsupervised techniques have the strengths of allowing the discovery of underlying patterns in the data without requiring prior labelling, which is useful for comprehensive data exploration, facilitating the identification of unique developmental profiles within the autistic spectrum. Weaknesses are that interpretation of results can be challenging due to lack of labels, which could limit direct clinical application. They are sensitive to data quality and the presence of noise, which may affect the accuracy of results. They may have applicability in helping to identify subgroups within the autism spectrum, which can guide per- sonalised interventions and therapeutic approaches. In many cases, combining both methods (supervised and unsupervised) can lead to hybrid techniques that leverage the strengths of both strategies. The algorithms listed as “Others” refer to statistical techniques or ensembles of comparative algorithms. Algorithms that are not included in the other categories are classified in this category. 3.4.RQ4: Which ML techniques are the most frequently used? Table 5lists the ML algorithms used in these studies and their corresponding year intervals (2017–2018, 2019–2020, and 2021–2023). Fig. 3shows that “classical” algorithms were used most frequently. However, it should be noted that in the coming years, new algorithms utilizing cutting-edge methodologies, such as deep learning, could become the most frequently used. The algorithms listed as “Others” include those that are not categorised in the other categories. 3.5.RQ5: Which data types are used most frequently in the selected studies? Table 6lists the different data types that were used in the selected studies. Fig. 4shows that brain data are the most frequently utilized data type in these studies (45 %), followed by clinical data (25 %) and eye tracking (18 %). In contrast, the ballistocardiogram (BCG), prosody, and phenotype were used only once. Moreover, there are three special cases in which the same study used different data types: specifically, the cases that use brain data and eye tracking [53], clinical data and eye tracking [29], and a combination of brain data, eye tracking, and facial recognition [76]. Table 2 Countries contributing to journal publications detailing where research studies related to RQ1 were published. CountryStudies% Studies Netherlands[23–37]27,27 % USA[38–49]21,81 % UK[10,50–59]20 % Ireland[60–65]10,91 % Switzerland[66–71]10,91 % India[72]1,82 % Iran[73]1,82 % Sweden[74]1,82 % Poland[75]1,82 % Egypt[76]1,82 % R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 5 Table 3 Countries of the corresponding and remaining authors' institutions regarding RQ2. Countrycorresponding_author_contributionsremaining_author_contributions USA[25,29,30,33,38,41–43,46,47,51,61,73][23–25,29,30,33,38,39,41–43,46,47,51,61,66,73] China[32,39,40,50,52–54,60,64,74,76][24,32,39,40,50,52–54,60,64,65,74,76] India[24,28,49,70][24,28,49,66,68,70,72] Spain[31,36,67][31,36,44,67] Italy[26,56,57][26,56,57] Bangladesh[48,71,72][48,71,72] United Kingdom[10,62][10,27,34,35,45,68] Saudi Arabia[68,75][62,71,75] Germany[59,69][59,69] Brazil[35,45][35,45] Turkey[27,34][74] Australia[37][37,43,48,71] Canada[23][23,44,46] Egypt[55][41,55] Finland[44][23] Portugal[63][63] Singapore[66][54] Iraq[58][58] United Arab Emirates–[58,62,70] Malaysia–[58,68] Austria–[59] Iran–[67] Japan–[66] Jordan–[66] Pakistan–[48] Sweden–[59] Taiwan–[66] Tunisia–[34] Table 4 ML method types used in the research studies regarding RQ3. ML algorithmsStudies Supervised Learning[10,23,25,27,29,32–34,38–40,42,44–47,49–51,53,55–57,59,63,64,66,67,69,71–75] Unsupervised Learning[35,60,61] Hybrid[24,26,28,30,31,43,48,52,58,68,70,76] Others[36,37,41,54,62] Fig. 2.Machine learning approaches used in the research studies regarding RQ3. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 6 Table 5 Classification of ML Algorithms, subtypes, years interval, and its application in the research studies: extended information to RQ4. ML Algorithm typeAlgorithmYears interval and # of studiesStudies Supervised LearningSVM2017–2018 (6)[35,43,44,46,56,63] 2019–2020 (16)[23–25,27,32–34,40,45,50–53,64,66, 67] 2021–2023 (9)[37,38,47–49,58,65,70,72] Random Forest (RF)2017–2018 (4)[35,42,56,57] 2019–2020 (10)[28–30,37,40,49,51,67,73,74] 2021–2023 (5)[48,58,59,69,72] MLP (Neural Network)2017–2018 (1)[35] 2019–2020 (5)[25,31,40,64,67] 2021–2023 (4)[49,58,65,68] Decision Trees (DT)2017–2018 (1)[62] 2019–2020 (9)[10,27–29,40,51,55,64,67] 2021–2023 (6)[48,49,58,65,69,70] k-nearest neighbours (KNN)2017–2018 (2)[57,75] 2019–2020 (5)[27,28,33,51,66] 2021–2023 (4)[48,49,58,72] Bayes2017–2018 (1)[57] 2019–2020 (5)[28,39,51,64,67] 2021–2023 (5)[37,48,49,58,76] Deep Learning, convolutional, autoencoders2017–2018 (1)[35] 2019–2020 (6)[26,29,31,33,45,50] 2021–2023 (1)[70] Long short-term memory (LSTM)2019–2020 (1)[32] Logistic Regression2017–2018 (3)[44,56,57] 2019–2020 (4)[26–28,64] 2021–2023 (6)[37,48,49,58,65,72] Boosting & Classifiers, Bagging2019–2020 (2)[10,27] 2021–2023 (5)[37,48,49,58,65] Modified Grasshopper Optimization Algorithm (MGOA)2019–2020 (1)[28] Feature selection/PCA2019–2020 (5)[31,32,53,66,67] Probabilistic neural network2019–2020 (1)[66] Linear Discriminant Analysis (LDA)2021–2023 (1)[48] Unsupervised Learningk-means Clustering2017–2018 (1)[61] 2019–2020 (9)[24,52,60] OthersVoting, ensembles classification2019–2020 (1)[54] Statistics2019–2020 (1)[41] Kernel Extreme Learning Machine2021–2023 (1)[36] Fig. 3.ML algorithms used in the research studies in relation to RQ4. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 7 3.6.RQ6: How many patients do the selected studies examine? Table 7classifies studies according to the number of patients that were examined in these studies. Fig. 5shows the number of patients (sample sizes) selected by the studies to develop the performance tasks. In most of the studies, the sample sizes were 51–100 and 1001–2000 (20 % for each sample). However, for a few studies, the sample sizes were 201–300 (5.45 %). After analysing the patient sampling, an attempt was made to correlate the ML algorithms with the datasets used to reduce biases and obtain greater generalisation and extrapolation using the ML algorithms. For this purpose, correlations between RQ4 and RQ6 were sought, and it can be concluded that, regarding deep knowledge, no specific algorithm produces a greater bias with the data. This conclusion should be verified in future studies with larger volumes of data. 3.7.RQ7: Do they normally develop their own databases in their studies? If not, which public databases are the most frequently used? Table 8lists the studies that developed their own databases and those that used public databases. Considering the selected manuscripts, 42 % of them developed their own database, whereas 58 % used public databases (Fig. 6). Moreover [69,75], employed both strategies. The research works shown in Fig. 7used the following public databases: Autism Brain Imaging Data Exchange (ABIDE), Virulence Factor Database (VFDB), Sequence Read Archive database, Ambiente Di Ricerca Interdisciplinare Per L'Analisi Di Neuroimmagini Nell'Autismo (ARIANNA) database of the ARIANNA project, University of California Irvine (UCI) Machine Learning Repository, Karolinska Directed Emotional Faces (KDEF), Autism Genetic Resource Exchange (AGRE), Human Connectome Project (HCP), Na- tional Database for Autism Research (NDAR), NimStim Face Stimulus Set, Hartwell Autism Research, Technology Initiative (iHART), Figshare data repository, the ASD-Net database by the German research consortium, Kaggle, Autism Diagnostic Observation Schedule (ADOS-G/ADOS-2), and the National Insurance Institute (NII) dataset. The most frequently used public database was ABIDE (with a frequency of 31.42 %), followed by the UCI Machine Learning Repository. Table 9lists the public databases used in the different studies. Table 6 Data types used in the research studies with respect to RQ5. Data TypeStudies Brain data[23,24,26,27,30,31,34,35,38,39,42–46,53,54,56,57,60,63,64,66,75,76] Clinical data[10,28,29,37,48,49,55,58,59,62,65,67,71,72] Eye tracking[29,32,36,50–53,68,70,76] Genome[33,40,73,74] BCG[25] Prosody[47] Phenotypes[61] Facial recognition[41,69,76] Fig. 4.Data types used in the research studies in relation to RQ5. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 8 3.8.RQ8: Are ML techniques a good choice for managing ASD? The performance of the set of studies that applied ML techniques was analysed by considering the following five features: best ML method, accuracy (%), sensitivity (%), specificity (%) and area under the curve (AUC) (Fig. 8). Therefore, within the scope of the best ML method, any given ML method could be better because of the great diversity of ML methods that were highlighted as the best in the different environments considered by these studies. Thus, various supervised and unsupervised ML methods were adjudged to be the best methods for this set of studies. Moreover, a subset of these studies introduced deep learning methods as a novelty in this ASD research. Considering the accuracy (%), several research works have an accuracy of more than 90 %, indicating that ML-based ap- proaches were the best for the ASD problem. Consequently, the sensitivity (%) scores obtained by these ML methods were as good as the previous accuracy (%) measures, including a small subset with a score of 100 %. It is well known that when a wide spectrum of approaches have different parameters, sensitivity analyses for model calibration must be performed to demonstrate their influences on the results. Thus, Asheghi et al. [77] demonstrated that parameters ranked using sensitivity analysis can be more reliable because they cover more uncertainties. In addition, within the framework of the specificity (%) measure, the high scores obtained by most ML-based approaches (including scores of 100 %) also indicate that the application of ML methodology to the ASD research problem demonstrates the best performance in this computational biomedical domain. Finally, in terms of the AUC, the scores obtained by these ML methods also demonstrated their good performance for ASD research. Table 7 Number of examined patients in the research studies concerning RQ6. Number of patientsStudies 0–50[23,25,33,41,51,57,60,63,64,75] 51–100[36–38,46,50,53,66,68,70,74,76] 101–200[27,30,34,42–44,47,52] 201–300[29,32,40] 301–1000[24,31,35,39,48,71] 1001–2000[10,28,35,45,49,54,56,59,62,67,72] >2000[26,55,58,61,65,69,73] Fig. 5.Patient samples that were considered in the research studies related to RQ6. Table 8 Studies that developed their own database or used a public database: extended information to RQ7. DatabaseStudies Own database[25,30,32,37,38,41,42,46,47,49–53,57,60,61,63–66,69,75,76] Public database[10,23,24,26–29,31,33–36,39,40,43–45,48,54–56,58,59,62,67–75] R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 9 4. Discussion The application of ML techniques to various research problems improves system performance in different areas. Regarding the ASD problem, research on the use of ML techniques presents challenges due to the large number of variables related to the computational process, which include the specific subdomain corpus, feature selection, learning task, and combination of ML techniques. An appropriate choice of these variables implies the best system performance in this ASD framework, where early diagnosis and treatment by doctors supported by these ML systems will condition the patient's quality of life. Therefore, this study aims to synthesise the literature review process of ML techniques that have been applied to the ASD area during the last seven years (between 2017 and 2023). A total of 55 studies using different ML techniques were identified and reviewed. In the following subsections, we summarise all the RQs analysed and link them to the metrics, challenges and limitations of the proposals studied. Fig. 6.Database creation in relation to RQ7. The research works (%) that developed their own database vs. those that used an existing database are shown. Fig. 7.Public databases used in the research studies regarding RQ7. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 10 4.1.Metrics (RQ1 and RQ2) In these studies, journals from the Netherlands published the largest number of research works, followed by the USA and the UK (RQ1). Regarding the countries to which the contributing authors of the studies belonged (RQ2), the USA led this score, closely fol- lowed by China, whereas other countries were far behind. These results are unsurprising because of the considerable annual invest- ment in Research and Development 2 of both countries, which is significantly greater than that of other countries were studies in this field were conducted. Consequently, the large number of high-quality researchers from both countries is justified. 4.2.Challenges (RQ3, RQ4, and RQ7) As discussed in the previous RQ3 section, the supervised approach was the most frequently used ML approach. This outcome was expected, as in many cases, the goal was to diagnose autism. Therefore, in RQ4, the most commonly used ML techniques were those that employed supervised learning methods, such as SVM, RF, RL, ANN, and DT. The recent deployment of techniques based on new technological developments that extend from deep learning should also be highlighted. The latest convolutional network algorithms in their different architectures are expected to provide tools for one of the future challenges in which heterogeneous variables can be combined, from commonly measured ones found in databases (as discussed in RQ7) to new biomarker data trends. As described in this paper, deep learning is one of the most recent areas of research and the quantification of uncertainly remains a major problem in this field. Thus, various studies were conducted in recent years [78], where a novel approach known as the automated random deactivating Table 9 Public databases used in the research studies related to RQ7. DatabaseStudies ABIDE[24,26,27,31,34,35,39,44,45,54,56] UCI Machine Learning Repository[10,28,48,62,67,72] NDAR[29,55] VFDB[74] Sequence Read Archive[40] ARIANNA[56] KDEF[51] AGRE[73] HCP[45] NimStim Face Stimulus Set[23] iHART[73] Figshare[36,68,70] ASD-Net[69] Kaggle[48,71] ADOS-G/ADOS-2[59] Fig. 8.Evaluation metrics used in the research studies with respect to RQ8. 2 https://asia.nikkei.com/Business/Science/China-passes-US-as-world-s-top-researcher-showing-its-R-D-might(visited on April 16, 2024). R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 11 connective weights approach (ARDCW) was presented. The study was validated using contour maps of the predicted error for different dropouts, accuracy metrics, and success rates, and by comparing with Monte Carlo dropout and quantile regression. As previously mentioned, the cross-cutting of applicable artificial intelligence methods, leading to the analysis of DNA biomarkers for predicting autism, is worth highlighting [33]. The challenges of data science, as well as the opportunities it offers in strategic partnerships, lead to work such as reported in Ref. [27], where crowdsourcing was leveraged by organising a Kaggle competition to build a pool of ML pipelines for diagnosing neurological disorders. Furthermore, it has been applied to the diagnosis of ASD using cortical morphological networks derived from T1-weighted magnetic resonance imaging. In Ref. [28], the authors used the MGOA Algorithm which can detect ASD in all age groups. Although it is not one of the most common algorithms, it has the potential to explore and exploit the search space effectively. In the work of [79], authors identify clinician subjectivity and lack of data-driven deci- sion-making as key barriers affecting treatment quality. To address these, the application of two ML algorithms is proposed to recommend and personalize treatment goals Applied behaviour analysis (ABA) for 29 study participants with ASD. The authors of [80] present a novel strategy for the early diagnosis of ASD through a prediction model combining RF-CART and RFID3. The test results are very promising as the proposed model outperforms state-of-the-art methods both in the AQ10 dataset and in real-world applications regarding reliability, sensitivity, susceptibility, clarity, and false positive rate. All the challenges discussed in this section will be addressed in future research. With regard to the types of data utilized in state-of-the-art studies, structured data is typically employed. However, the challenges and opportunities presented by unstructured data are undeniable. Therapeutic records often provide unstructured information, particularly in the form of free-text clinical notes, which contain significant information not captured in structured records. A similar approach is illustrated in the study by Peng et al. [81], where Natural Language Processing (NLP) tools are evaluated using ASD as a case study. CLAMP, cTAKES, and MetaMap are employed to analyse 544 full-text articles and 20,408 PubMed abstracts to extract ASD-related terms. The analysis protocols utilized in this study have broader applicability and can be extended to other neuropsy- chiatric or neurodevelopmental disorders lacking well-defined terminology sets to describe their phenotypic presentations. It is crucial to emphasize that the development of efficient and effective ML models for ASD involves close collaboration among computer scientists, clinicians, psychologists, and other stakeholders. Computer scientists contribute their expertise by developing robust algorithms and computational techniques for analysing large datasets related to ASD, while clinicians provide valuable insights into the clinical manifestations of the disorder and the specific challenges faced by individuals with ASD and their families. Psy- chologists contribute their understanding of cognitive and behavioural aspects of ASD, helping to guide the selection of relevant features and outcome measures for the ML models. Additionally, input from other stakeholders, such as educators, caregivers, and individuals with ASD themselves, ensures that the models are designed with a holistic understanding of the needs and perspectives of the ASD community. By leveraging the diverse expertise and perspectives of these stakeholders, the primary objective is to develop ML models that are not only technically robust but also clinically meaningful and relevant to the real-world challenges faced by individuals with ASD. In terms of privacy and security considerations, it is important to acknowledge that studies of this nature involve handling sensitive data. It is essential to ensure the protection of participants in medical research, particularly when dealing with individuals with neurodevelopmental disorders. Data collection in this context should be conducted ethically and respecting the privacy of individuals. Informed consent must be obtained from participants or, when applicable, from family members, and the research should be approved by ethical committees of institutions/universities. Data should be stored and processed securely to protect the confidentiality of sensitive medical and behavioural information. It is essential to implement robust security measures to prevent unauthorized access and to maintain data integrity and confidentiality throughout the research process. Additionally, transparency in the use of ML al- gorithms must be considered, and results should be interpreted and communicated responsibly and ethically, taking into account the potential implications for the health and well-being of individuals with ASD and their families. In summary, the major challenges of ML-based interventions for ASD include: (a) Generalisability and robustness: ML models can struggle to generalize learned patterns to new settings or populations. This is especially relevant in the case of ASD, where variability in clinical presentation and individual characteristics can be considerable. Models must be able to adapt to this variability to be effective in different clinical contexts; (b) Interpretability and explainability: Often, ML models, especially more complex ones such as deep neural networks, can be difficult to interpret, making it difficult to understand how they arrive at their decisions. This lack of explainability can lead to mistrust and limit the adoption of these tools in clinical practice; (c) Privacy and data security: The collection and processing of sensitive ASD patient data raise ethical and legal concerns related to privacy and data security. It is essential to ensure that patient data are protected from unauthorized access and that privacy standards set by regulations and ethical norms are upheld; (d) Equity and algorithmic bias: There is a risk that ML models reproduce and amplify existing biases in training data, which could lead to unfair disparities in care and treatment for people with ASD. It is important to address these biases through algorithmic bias mitigation and evaluation techniques to ensure equity in access to ML-based interventions; (e) Integration with clinical practice: For ML-based interventions to be effective in clinical settings, they must be seamlessly integrated into existing workflows and accepted by healthcare professionals and patients. Achieving this may require specialized training, updates to technological infrastructure, and careful consideration of the human and social aspects of implementation. 4.3.Limitations (RQ5, RQ6, and RQ8) Regarding the type of data (RQ5), most studies that applied ML techniques to the problem of autism prediction used brain data, such as those extracted from electroencephalograms. Consequently, the data extraction can sometimes be very challenging, which justifies the use of public databases (58 % overall). ADIBE is one of the most common databases that uses brain data, and it was used by R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 12 38 % of the studies that used public databases. Studies that create their own databases tend to include fewer patients because it is sometimes difficult to obtain the opportunity and permission from the ethics committee to work with a larger number of patients (RQ6). Most of the studies considered between 0 and 50 people and between 51 and 100 people (21 % for both cases). To date, structured data are the most frequently used. As suggested in Ref. [82], it would be interesting to consider the possibility of using less structured data with context-sensitive information in the assessment and intervention tasks. Therefore, the lack of a larger number of databases available to researchers when reproducing their experiments must be noted. One promising trend that will likely take some time to develop is the application of digital twins in healthcare and their generation of synthetic data, which would greatly benefit artificial intelligence [83]. Moreover, in many cases, the genetic basis of ASD must be further explored. However, this task is not straightforward and is challenging as well. Despite the strong genetic basis of ASD, the complete complement of ASD-associated genes is still lacking [84]. In this context, ML and the advances in deep learning may allow researchers to shed some light on this computational complexity. Finally, regarding the performance of ML techniques in the ASD field (RQ8), various ML techniques have been applied over the years to address challenges in the medical and biomedical domains. This indicates that ML algorithms have been improved, and training corpora have also been enriched qualitatively and quantitatively. Therefore, the continuous number of research works that apply ML techniques in these fields enhances the knowledge of new researchers who use this feedback to improve their new ML techniques. Thus, ML techniques relevant to ASD continue to undergo continuous improvement and evolution. 4.4.Innovative approaches To conclude this section, we summarise the limitations of ML models applied in ASD: •Data scarcity: Lack of specific, quality data affects the ability of models to learn representative patterns. This is a constant in the area of health and in particular in the area of autism. •Data Bias: Biases arise from collecting data from homogeneous populations and this leads to biased models. For example, there are many more autism studies on males than females and this biases the data set and its models. •Generalisation of models: For these reasons it is very important to always look for the highest generalisability, otherwise models, when tested outside their training samples, will not adapt well to new populations. In light of these limitations, some possible innovative solutions would include: •Federated Learning: can present a solution for sharing knowledge without compromising the privacy of patient data, allowing for more robust and general models. •Development of benchmarking datasets: Suggests collaborative initiatives to create representative datasets, which include a di- versity of population samples and ensure that the model is applicable to diverse autism spectrum characteristics. •Fine tuning and cross-validation techniques: Highlights the use of techniques that can help models fit new data, helping to improve generalisability. 5. Conclusion In this paper, a systematic review of research conducted between 2017 and 2023 was presented, focusing on the application of ML techniques to ASD. The studies reviewed show promising developments, particularly in supervised methods; however, the potential for hybrid methods that integrate unsupervised learning, deep learning, and fuzzy logic is significant and represents an exciting avenue for future research. Hybrid methods could make it possible to discover new patterns and clusters within ASD, an important advancement given the complexity and variability of the spectrum. Rather than simply identifying the presence of autism, these approaches could help define diverse ASD subgroups and degrees of dependency, supporting a more nuanced approach to diagnosis. To ensure hybrid methods become more accessible and interpretable, future research will need to address the integration of diverse data sources. This will require attention to three key areas: (1) the use of less structured data that includes context-sensitive infor- mation, (2) digital twins and synthetic data generation, and (3) genetic data analysis combined with deep learning to uncover new features and unknown correlations. Future ML approaches should also incorporate new data collection methods, including wearable technologies (such as biometric sensors) that enable continuous and non-intrusive monitoring of behaviour and physiological re- sponses in natural environments. Additionally, integrating multimodal data —clinical records, genetic data, medical imaging, and behavioural records— will enable a holistic understanding of ASD. However, this integration introduces technical and ethical chal- lenges that must be addressed. Further exploration is necessary to understand how evolving ML techniques can meet specific challenges in ASD diagnosis and treatment. This includes advancements in deep learning, natural language processing for analysing textual data, and semi-supervised or reinforcement learning methods adapted to clinical datasets. The potential for interpretable and explainable ML models is particularly crucial in autism research, as these could improve confidence among clinicians and patients in the decision-making processes supported by these tools. Specifically, in the last two years (2024 and 2025), the application of ML and deep learning techniques to ASD detection and assessment has become a consolidated trend [85–90]. Some studies explore models with a reduced number of features to improve efficiency, while others rely on neural networks and bidirectional long short-term memory architectures for classification. Additional approaches incorporate functional magnetic resonance imaging data or eye-tracking information to R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 13 enhance diagnostic performance. Considering the most common evaluation metrics (accuracy, F1-score, sensitivity and specificity), these recent works consistently demonstrate that the use of ML and deep learning for ASD diagnosis remains a robust and promising research direction. Finally, while ML offers opportunities to improve ASD detection and diagnosis, applications extend beyond these areas. Emerging ML applications in public health, treatment support, and clinical research show strong potential, as initial results indicate. The ongoing expansion of ML tools requires multidisciplinary teams —including experts in computer science, psychology, neuroscience, and medicine— to collaboratively tackle the complexities of ASD and to develop customized, context-sensitive approaches for personalised diagnoses, treatments, and interventions. The insights presented in this study establish a foundation for future research and progress in this growing field, providing a roadmap for advancing ASD research through machine learning. CRediT authorship contribution statement Rafael Mu ̃ noz-Terol: Writing – review & editing, Writing – original draft, Investigation, Funding acquisition. Jesús Peral: Writing – review & editing, Writing – original draft, Investigation, Funding acquisition. Sandra Amador: Writing – review & editing, Writing – original draft, Investigation. David Gil: Writing – review & editing, Writing – original draft, Investigation, Funding acquisition. Ethics declaration Review and/or approval by an ethics committee as well as informed consent was not required for this study because this literature review only used existing data from published studies and did not involve any direct experimentation/studies on living beings. Data availability statement No data was used for the research described in the article. Funding This research has been funded by BALLADEER Project (PROMETEO/2021/088) from the Conselleria de Innovaci ́ on, Universidades, Ciencia y Sociedad Digital, Generalitat Valenciana (Valencia, Spain). Furthermore, it has been supported by the KOSMOS-UA project (PID2024-155363OB-C43), funded by the Spanish Ministry of Science and Innovation, the BALIDA-A project (CIPROM/2024/13), funded by Conselleria de Educaci ́ on, Cultura, Universidades y Empleo (Generalitat Valenciana), the IAEAV project (INREIA/2024/ 176), funded by the Conselleria de Innovaci ́ on, Industria, Comercio y Turismo (Generalitat Valenciana), the ENIA Chair of Artificial Intelligence (TSI-100927-2023-6), the AgroVAL (TSI-100122-2024-10), the Sophia (TSI-100130-2024-10) and European mobility for efficient planning and new business opportunities (TSI-100121-2024-10) projects, funded by the Recovery, Transformation and Resilience Plan from the European Union Next Generation through the Ministry for Digital Transformation and the Civil Service, and Grant RED2022-134656-T, funded by MCIN/AEI/10.13039/501100011033. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. References [1]K. Kei Mak, K. Lee, C. Park, Applications of machine learning in addiction studies, A systematic review (2019), https://doi.org/10.1016/j.psychres.2019.03.001. [2]K.K. Hyde, M.N. Novack, N. LaHaye, C. Parlett-Pelleriti, R. Anden, D.R. Dixon, E. Linstead, Applications of supervised machine learning in autism spectrum disorder research: a review, Rev. J. Autism Dev. Disord. 6 (2019) 128–146, https://doi.org/10.1007/s40489-019-00158-x. [3]A.B.R. Shatte, D.M. Hutchinson, S.J. Teague, Psychological Medicine Machine learning in mental health: a scoping review of methods and applications. https:// doi.org/10.1017/S0033291719000151, 2019. [4]S. Lloyd, M. Mohseni, P. Rebentrost, Quantum Algorithms for Supervised and Unsupervised Machine Learning, 2013, https://doi.org/10.48550/ arXiv.1307.0411 arXivForum2013. [5]A. Crippa, C. Salvatore, P. Perego, S. Forti, M. Nobile, M. Molteni, I. Castiglioni, Use of machine learning to identify children with autism and their motor abnormalities, J. Autism Dev. Disord. 45 (2015) 2146–2156, https://doi.org/10.1007/s10803-015-2379-8. [6]M. Duda, R. Ma, N. Haber, D.P. Wall, Use of machine learning for behavioral distinction of autism and ADHD, Transl. Psychiatry 6 (2016) 732, https://doi.org/ 10.1038/tp.2015.221. [7]D. Bone, S.L. Bishop, M.P. Black, M.S. Goodwin, C. Lord, S.S. Narayanan, Use of machine learning to improve autism screening and diagnostic instruments: effectiveness, efficiency, and multi-instrument fusion. https://doi.org/10.1111/jcpp.12559, 2016. [8]M. Baris Usta, K. Karabekiroglu, B. Sahin, M. Aydin, A. Bozkurt, T. Karaosman, A. Aral, C. Cobanoglu, A. Duman Kurt, N. Kesim, ̇ I. Sahin, E. Ürer, Psychiatry and Clinical psychopharmacology use of machine learning methods in prediction of short-term outcome in autism spectrum disorders use of machine learning methods in prediction of short-term outcome in autism spectrum disorders. https://doi.org/10.1080/24750573.2018.1545334, 2018. [9]O. Rudovic, J. Lee, M. Dai, B. Schuller, R.W. Picard, Personalized machine learning for robot perception of affect and engagement in autism therapy, Sci. Robot. 3 (2018), https://doi.org/10.1126/scirobotics.aao6760. [10]F. Thabtah, D. Peebles, A new machine learning model based on induction of rules for autism detection, Health Inf. J. 26 (2020) 264–286, https://doi.org/ 10.1177/1460458218824711. [11]Q. Tariq, J. Daniels, J.N. Schwartz, P. Washington, H. Kalantarian, D.P. Wall, Mobile detection of autism through machine learning on home video: a development and prospective validation study, PLoS Med. 15 (2018), https://doi.org/10.1371/journal.pmed.1002705. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 14 [12]W. Liu, M. Li, L. Yi, Identifying children with autism spectrum disorder based on their face processing abnormality: a machine learning framework, Autism Res. 9 (2016) 888–898, https://doi.org/10.1002/aur.1615. [13]P. Kazienko, E. Lughofer, B. Trawi ́ nski, Hybrid and ensemble methods in machine learning J.UCS special issue, J. Univers. Comput. Sci. 19 (2013) 457–461, https://doi.org/10.3217/jucs-019-04. [14]X. Peng, Y. Shuai, Y. Gan, Y. Chen, Hybrid feature selection model based on machine learning and knowledge graph, J. Phys. Conf. Ser. 2079 (2021), https:// doi.org/10.1088/1742-6596/2079/1/012028. [15]S. Rani, Hybrid model using stack-based ensemble classifier and dictionary classifier to improve classification accuracy of Twitter sentiment analysis, Int. J. Emerg. Trends Eng. Res. 8 (2020) 2893–2900, https://doi.org/10.30534/ijeter/2020/02872020. [16]S. Alharbi, M.S. Hossain, A.A. Monrat, A belief rule based expert system to assess autism under uncertainty, WCECS 1 (2015). [17]M.E. Alam, M. Shamim Kaiser, M.S. Hossain, K. Andersson, An IoT-belief rule base smart system to assess autism. https://doi.org/10.1109/CEEICT.2018. 8628131, 2019. [18]O. Mugzach, M. Peleg, S.C. Bagley, S.J. Guter, E.H. Cook, R.B. Altman, An ontology for autism spectrum disorder (ASD) to infer ASD phenotypes from autism diagnostic interview-revised data, J. Biomed. Inf. 56 (2015) 333–347, https://doi.org/10.1016/j.jbi.2015.06.026. [19]L. Gong, Y. Yan, J. Xie, H. Liu, X. Sun, Prediction of autism susceptibility genes based on association rules, J. Neurosci. Res. 90 (2012) 1119–1125, https://doi. org/10.1002/jnr.23015. [20]D.P. Wall, J. Kosmicki, T.F. Deluca, E. Harstad, V.A. Fusaro, Use of machine learning to shorten observation-based screening and diagnosis of autism, Transl. Psychiatry 2 (2012), https://doi.org/10.1038/tp.2012.10. [21]D. Moher, A. Liberati, J. Tetzlaff, D.G. Altman, Reprint—preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement, Phys. Ther. 89 (2009) 873–880, https://doi.org/10.1093/ptj/89.9.873. [22]M.J. Page, J.E. McKenzie, P.M. Bossuyt, I. Boutron, T.C. Hoffmann, C.D. Mulrow, L. Shamseer, J.M. Tetzlaff, E.A. Akl, S.E. Brennan, R. Chou, J. Glanville, J. M. Grimshaw, A. Hr ́ objartsson, M.M. Lalu, T. Li, E.W. Loder, E. Mayo-Wilson, S. McDonald, L.A. McGuinness, L.A. Stewart, J. Thomas, A.C. Tricco, V.A. Welch, P. Whiting, D. Moher, The PRISMA 2020 statement: an updated guideline for reporting systematic reviews, BMJ 372 (2021), https://doi.org/10.1136/bmj.n71. [23]A.S. Nunes, F. Mamashli, N. Kozhemiako, S. Khan, N.M. McGuiggan, A. Losh, R.M. Joseph, J. Ahveninen, S.M. Doesburg, M.S. H ̈ am ̈ al ̈ ainen, T. Kenet, Classification of evoked responses to inverted faces reveals both spatial and temporal cortical response abnormalities in Autism spectrum disorder, NeuroImage Clin 29 (2021), https://doi.org/10.1016/j.nicl.2020.102501. [24]N. Chaitra, P.A. Vijaya, G. Deshpande, Diagnostic prediction of autism spectrum disorder using complex network measures in a machine learning framework, Biomed. Signal Process Control 62 (2020) 102099, https://doi.org/10.1016/j.bspc.2020.102099. [25]A. Alivar, C. Carlson, A. Suliman, S. Warren, P. Prakash, D.E. Thompson, B. Natarajan, Smart bed based daytime behavior prediction in children with autism spectrum disorder - a pilot study, Med. Eng. Phys. 83 (2020) 15–25, https://doi.org/10.1016/j.medengphy.2020.07.004. [26]E. Ferrari, P. Bosco, S. Calderoni, P. Oliva, L. Palumbo, G. Spera, M.E. Fantacci, A. Retico, Dealing with confounders and outliers in classification medical studies: the autism spectrum disorders case study, Artif. Intell. Med. 108 (2020) 101926, https://doi.org/10.1016/j.artmed.2020.101926. [27]I. Bilgen, G. Guvercin, I. Rekik, Machine learning methods for brain network classification: application to autism diagnosis using cortical morphological networks, J. Neurosci. Methods 343 (2020) 108799, https://doi.org/10.1016/j.jneumeth.2020.108799. [28]N. Goel, B. Grover, Anuj, D. Gupta, A. Khanna, M. Sharma, Modified grasshopper optimization algorithm for detection of autism spectrum disorder, Phys. Commun. 41 (2020) 101115, https://doi.org/10.1016/j.phycom.2020.101115. [29]D. Fabiano, S. Canavan, H. Agazzi, S. Hinduja, D. Goldgof, Gaze-based classification of autism spectrum disorder, Pattern Recognit. Lett. 135 (2020) 204–212, https://doi.org/10.1016/j.patrec.2020.04.028. [30]M. Cordova, K. Shada, D.V. Demeter, O. Doyle, O. Miranda-Dominguez, A. Perrone, E. Schifsky, A. Graham, E. Fombonne, B. Langhorst, J. Nigg, D.A. Fair, E. Feczko, Heterogeneity of executive function revealed by a functional random forest approach across ADHD and ASD, NeuroImage Clin 26 (2020) 102245, https://doi.org/10.1016/j.nicl.2020.102245. [31]M. Raki ́ c, M. Cabezas, K. Kushibar, A. Oliver, X. Llad ́ o, Improving the detection of autism spectrum disorder by combining structural and functional MRI information, NeuroImage Clin 25 (2020) 102181, https://doi.org/10.1016/j.nicl.2020.102181. [32]J. Li, Y. Zhong, J. Han, G. Ouyang, X. Li, H. Liu, Classifying ASD children with LSTM based on raw videos, Neurocomputing 390 (2020) 226–238, https://doi. org/10.1016/j.neucom.2019.05.106. [33]R.O. Bahado-Singh, S. Vishweswaraiah, B. Aydas, N.K. Mishra, A. Yilmaz, C. Guda, U. Radhakrishna, Artificial intelligence analysis of newborn leucocyte epigenomic markers for the prediction of autism, Brain Res. 1724 (2019), https://doi.org/10.1016/j.brainres.2019.146457. [34]O. Graa, I. Rekik, Multi-view learning-based data proliferator for boosting classification using highly imbalanced classes, J. Neurosci. Methods 327 (2019) 108344, https://doi.org/10.1016/j.jneumeth.2019.108344. [35]A.S. Heinsfeld, A.R. Franco, R.C. Craddock, A. Buchweitz, F. Meneguzzi, Identification of autism spectrum disorder using deep learning and the ABIDE dataset, NeuroImage Clin 17 (2018) 16–23, https://doi.org/10.1016/j.nicl.2017.08.017. [36]A. Gaspar, D. Oliva, S. Hinojosa, I. Aranguren, D. Zaldivar, An optimized Kernel extreme learning machine for the classification of the autism spectrum disorder by using gaze tracking images, Appl. Soft Comput. 120 (2022) 108654, https://doi.org/10.1016/j.asoc.2022.108654. [37]P. Wei, D. Ahmedt-Aristizabal, H. Gammulle, S. Denman, M.A. Armin, Vision-based activity recognition in children with autism-related behaviors, Heliyon 9 (2023) e16763, https://doi.org/10.1016/j.heliyon.2023.e16763. [38]A. Dickinson, M. Daniel, A. Marin, B. Gaonkar, M. Dapretto, N.M. McDonald, S. Jeste, Multivariate neural connectivity patterns in early infancy predict later autism symptoms, Biol. Psychiatry Cogn. Neurosci. Neuroimaging 6 (2021) 59–69, https://doi.org/10.1016/j.bpsc.2020.06.003. [39]Y. Fu, J. Zhang, Y. Li, J. Shi, Y. Zou, H. Guo, Y. Li, Z. Yao, Y. Wang, B. Hu, A novel pipeline leveraging surface-based features of small subcortical structures to classify individuals with autism spectrum disorder, Prog. Neuropsychopharmacol. Biol. Psychiatry 104 (2021), https://doi.org/10.1016/j.pnpbp.2020.109989. [40]T. Wu, H. Wang, W. Lu, Q. Zhai, Q. Zhang, W. Yuan, Z. Gu, J. Zhao, H. Zhang, W. Chen, Potential of gut microbiome for detection of autism spectrum disorder, Microb. Pathog. 149 (2020) 104568, https://doi.org/10.1016/j.micpath.2020.104568. [41]N.N. Capriola-Hall, A.T. Wieckowski, D. Swain, V. Tech, S. Aly, A. Youssef, A.L. Abbott, S.W. White, Group differences in facial emotion expression in autism: evidence for the utility of machine classification, Behav. Ther. 50 (2019) 828–838, https://doi.org/10.1016/j.beth.2018.12.004. [42]E. Feczko, N.M. Balba, O. Miranda-Dominguez, M. Cordova, S.L. Karalunas, L. Irwin, D.V. Demeter, A.P. Hill, B.H. Langhorst, J. Grieser Painter, J. Van Santen, E. J. Fombonne, J.T. Nigg, D.A. Fair, Subtyping cognitive profiles in autism spectrum disorder using a functional random forest algorithm, Neuroimage 172 (2018) 674–688, https://doi.org/10.1016/j.neuroimage.2017.12.044. [43]F. Zhang, P. Savadjiev, W. Cai, Y. Song, Y. Rathi, B. Tunç, D. Parker, T. Kapur, R.T. Schultz, N. Makris, R. Verma, L.J. O'Donnell, Whole brain white matter connectivity analysis using machine learning: an application to autism, Neuroimage 172 (2018) 826–837, https://doi.org/10.1016/j.neuroimage.2017.10.029. [44]E. Moradi, B. Khundrakpam, J.D. Lewis, A.C. Evans, J. Tohka, Predicting symptom severity in autism spectrum disorder based on cortical thickness measures in agglomerative data, Neuroimage 144 (2017) 128–141, https://doi.org/10.1016/j.neuroimage.2016.09.049. [45]W.H.L. Pinaya, A. Mechelli, J.R. Sato, Using deep autoencoders to identify abnormal brain structural patterns in neuropsychiatric disorders: a large-scale multi- sample study, Hum. Brain Mapp. 40 (2019) 944–954, https://doi.org/10.1002/hbm.24423. [46]R.W. Emerson, C. Adams, T. Nishino, H.C. Hazlett, J.J. Wolff, L. Zwaigenbaum, J.N. Constantino, M.D. Shen, M.R. Swanson, J.T. Elison, S. Kandala, A.M. Estes, K.N. Botteron, L. Collins, S.R. Dager, A.C. Evans, G. Gerig, H. Gu, R.C. Mckinstry, S. Paterson, R.T. Schultz, M. Styner, B.L. Schlaggar, J.R. Pruett, J. Piven, Functional neuroimaging of high-risk 6-month-old infants predicts a diagnosis of autism at 24 months of age, Sci. Transl. Med. 9 (2017), https://doi.org/ 10.1126/scitranslmed.aag2882. [47]J.C.Y. Lau, S. Patel, X. Kang, K. Nayar, G.E. Martin, J. Choy, P.C.M. Wong, M. Losh, Cross-linguistic patterns of speech prosodic differences in autism: a machine learning study, PLoS One 17 (2022) 1–16, https://doi.org/10.1371/journal.pone.0269637. [48]S.M. Mahedy Hasan, M.P. Uddin, M. Al Mamun, M.I. Sharif, A. Ulhaq, G. Krishnamoorthy, A machine learning framework for early-stage detection of autism spectrum disorders, IEEE Access 11 (2022) 15038–15057, https://doi.org/10.1109/ACCESS.2022.3232490. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 15 [49]J. Talukdar, D.K. Gogoi, T.P. Singh, A comparative assessment of most widely used machine learning classifiers for analysing and classifying autism spectrum disorder in toddlers and adolescents, Healthc. Anal. 3 (2023) 100178, https://doi.org/10.1016/j.health.2023.100178. [50]M. Lai, J. Lee, S. Chiu, J. Charm, W.Y. So, F.P. Yuen, C. Kwok, J. Tsoi, Y. Lin, B. Zee, A machine learning approach for retinal images analysis as an objective screening method for children with autism spectrum disorder, eClinicalMedicine 28 (2020) 100588, https://doi.org/10.1016/j.eclinm.2020.100588. [51]Y. Li, M.A. Mache, T.A. Todd, Automated identification of postural control for children with autism spectrum disorder using a machine learning approach, J. Biomech. 113 (2020) 110073, https://doi.org/10.1016/j.jbiomech.2020.110073. [52]J. Kang, X. Han, J.F. Hu, H. Feng, X. Li, The study of the differences between low-functioning autistic children and typically developing children in the processing of the own-race and other-race faces by the machine learning approach, J. Clin. Neurosci. 81 (2020) 54–60, https://doi.org/10.1016/j. jocn.2020.09.039. [53]J. Kang, X. Han, J. Song, Z. Niu, X. Li, The identification of children with autism spectrum disorder by SVM approach on EEG and eye-tracking data, Comput. Biol. Med. 120 (2020) 103722, https://doi.org/10.1016/j.compbiomed.2020.103722. [54]F. Huang, E.L. Tan, P. Yang, S. Huang, L. Ou-Yang, J. Cao, T. Wang, B. Lei, Self-weighted adaptive structure learning for ASD diagnosis via multi-template multi- center representation, Med. Image Anal. 63 (2020), https://doi.org/10.1016/j.media.2020.101662. [55]M.M. Hassan, H.M.O. Mokhtar, Investigating autism etiology and heterogeneity by decision tree algorithm, Inform. Med. Unlocked 16 (2019) 100215, https:// doi.org/10.1016/j.imu.2019.100215. [56]A. Retico, S. Arezzini, P. Bosco, S. Calderoni, A. Ciampa, S. Coscetti, S. Cuomo, L. De Santis, D. Fabiani, M.E. Fantacci, A. Giuliano, E. Mazzoni, P. Mercatali, G. Miscali, M. Pardini, M. Prosperi, F. Romano, E. Tamburini, M. Tosetti, F. Muratori, ARIANNA: a research environment for neuroimaging studies in autism spectrum disorders, Comput. Biol. Med. 87 (2017) 1–7, https://doi.org/10.1016/j.compbiomed.2017.05.017. [57]E. Grossi, C. Olivieri, M. Buscema, Diagnosis of autism through EEG processed by advanced computational algorithms: a pilot study, Comput. Methods Progr. Biomed. 142 (2017) 73–79, https://doi.org/10.1016/j.cmpb.2017.02.002. [58]A.S. Albahri, A.A. Zaidan, H.A. AlSattar, R.A. Hamid, O.S. Albahri, S. Qahtan, A.H. Alamoodi, Towards physician's experience: development of machine learning model for the diagnosis of autism spectrum disorders based on complex T-spherical fuzzy-weighted zero-inconsistency method, Comput. Intell. 39 (2023) 225–257, https://doi.org/10.1111/coin.12562. [59]M. Schulte-Rüther, T. Kulvicius, S. Stroth, N. Wolff, V. Roessner, P.B. Marschik, I. Kamp-Becker, L. Poustka, Using machine learning to improve diagnostic assessment of ASD in the light of specific differential and co-occurring diagnoses, J. Child Psychol. Psychiatry Allied Discip. 64 (2023) 16–26, https://doi.org/ 10.1111/jcpp.13650. [60]L. Xu, Q. Hua, J. Yu, J. Li, Classification of autism spectrum disorder based on sample entropy of spontaneous functional near infra-red spectroscopy signal, Clin. Neurophysiol. 131 (2020) 1365–1374, https://doi.org/10.1016/j.clinph.2019.12.400. [61]E. Stevens, D.R. Dixon, M.N. Novack, D. Granpeesheh, T. Smith, E. Linstead, Identification and analysis of behavioral phenotypes in autism spectrum disorder via unsupervised machine learning, Int. J. Med. Inf. 129 (2019) 29–36, https://doi.org/10.1016/j.ijmedinf.2019.05.006. [62]F. Thabtah, F. Kamalov, K. Rajab, A new computational intelligence approach to detect autistic features for autism screening, Int. J. Med. Inf. 117 (2018) 112–124, https://doi.org/10.1016/j.ijmedinf.2018.06.009. [63]J. Castelhano, P. Tavares, S. Mouga, G. Oliveira, M. Castelo-Branco, Stimulus dependent neural oscillatory patterns show reliable statistical identification of autism spectrum disorder in a face perceptual decision task, Clin. Neurophysiol. 129 (2018) 981–989, https://doi.org/10.1016/j.clinph.2018.01.072. [64]L. Xu, Y. Guo, J. Li, J. Yu, H. Xu, Classification of autism spectrum disorder based on fluctuation entropy of spontaneous hemodynamic fluctuations, Biomed. Signal Process Control 60 (2020) 101958, https://doi.org/10.1016/j.bspc.2020.101958. [65]Q. Wei, X. Xu, X. Xu, Q. Cheng, Early identification of autism spectrum disorder by multi-instrument fusion: a clinically applicable machine learning approach, Psychiatry Res. 320 (2023) 115050, https://doi.org/10.1016/j.psychres.2023.115050. [66]T.H. Pham, J. Vicnesh, J.K.E. Wei, S.L. Oh, N. Arunkumar, E.W. Abdulhay, E.J. Ciaccio, U.R. Acharya, Autism spectrum disorder diagnostic system using HOS bispectrum with EEG signals, Int. J. Environ. Res. Publ. Health 17 (2020) 1–15, https://doi.org/10.3390/ijerph17030971. [67]J. Peral, D. Gil, S. Rotbei, S. Amador, M. Guerrero, H. Moradi, A machine learning and integration based architecture for cognitive disorder detection used for early autism screening, Electronics 9 (2020) 516, https://doi.org/10.3390/electronics9030516. [68]I.A. Ahmed, E.M. Senan, T.H. Rassem, M.A.H. Ali, H.S.A. Shatnawi, S.M. Alwazer, M. Alshahrani, Eye tracking-based diagnosis and early detection of autism spectrum disorder using machine learning and deep learning techniques, Electron 11 (2022), https://doi.org/10.3390/electronics11040530. [69]N. Wolff, M. Eberlein, S. Stroth, L. Poustka, S. Roepke, I. Kamp-Becker, V. Roessner, Abilities and disabilities—Applying machine learning to disentangle the role of intelligence in diagnosing autism spectrum disorders, Front. Psychiatr. 13 (2022), https://doi.org/10.3389/fpsyt.2022.826043. [70]M.R. Kanhirakadavath, M.S.M. Chandran, Investigation of eye-tracking scan path as a biomarker for autism screening using machine learning algorithms, Diagnostics 12 (2022), https://doi.org/10.3390/diagnostics12020518. [71]M.J. Uddin, M.M. Ahamad, P.K. Sarker, S. Aktar, N. Alotaibi, S.A. Alyami, M.A. Kabir, M.A. Moni, An integrated statistical and clinically applicable machine learning framework for the detection of autism spectrum disorder, Computers 12 (2023), https://doi.org/10.3390/computers12050092. [72]M.A. Khatun, M.A. Ali, M.R. Ahmed, S.R.H. Noori, A. Sahayadhas, Empirical Study of Computational Intelligence Approaches for the Early Detection of Autism Spectrum Disorder, Springer Singapore, 2021, https://doi.org/10.1007/978-981-15-5566-4_14. [73]E.K. Ruzzo, L. P ́ erez-Cano, J.Y. Jung, L. kai Wang, D. Kashef-Haghighi, C. Hartl, C. Singh, J. Xu, J.N. Hoekstra, O. Leventhal, V.M. Lepp ̈ a, M.J. Gandal, K. Paskov, N. Stockham, D. Polioudakis, J.K. Lowe, D.A. Prober, D.H. Geschwind, D.P. Wall, Inherited and De Novo Genetic Risk for Autism Impacts Shared Networks, Cell 178 (2019) 850–866.e26, https://doi.org/10.1016/j.cell.2019.07.015. [74]M. Wang, C. Doenyas, J. Wan, S. Zeng, C. Cai, J. Zhou, Y. Liu, Z. Yin, W. Zhou, Virulence factor-related gut microbiota genes and immunoglobulin A levels as novel markers for machine learning-based classification of autism spectrum disorder, Comput. Struct. Biotechnol. J. 19 (2021) 545–554, https://doi.org/ 10.1016/j.csbj.2020.12.012. [75]S. Ibrahim, R. Djemal, A. Alsuwailem, Electroencephalography (EEG) signal processing for epilepsy and autism spectrum disorder diagnosis, Biocybern. Biomed. Eng. 38 (2018) 16–26, https://doi.org/10.1016/j.bbe.2017.08.006. [76]M. Liao, H. Duan, G. Wang, Application of machine learning techniques to detect the children with autism spectrum disorder, J. Healthc. Eng. (2022) 2022, https://doi.org/10.1155/2022/9340027. [77]R. Asheghi, S.A. Hosseini, M. Saneie, A.A. Shahri, Updating the neural network sediment load models using different sensitivity analysis methods: a regional application, J. Hydroinform. 22 (2020) 562–577, https://doi.org/10.2166/hydro.2020.098. [78]A. Abbaszadeh Shahri, C. Shan, S. Larsson, A novel approach to uncertainty quantification in groundwater table modeling by automated predictive deep learning, Nat. Resour. Res. 31 (2022) 1351–1373, https://doi.org/10.1007/s11053-022-10051-w. [79]M. Kohli, A.K. Kar, A. Bangalore, P. Ap, Machine learning-based ABA treatment recommendation and personalization for autism spectrum disorder: an exploratory study, Brain Informatics 9 (2023) 1–25, https://doi.org/10.1186/S40708-022-00164-6/TABLES/7. [80]J. Albert Mayan, V.S. Reddy, B.T. Varma, T. Mohamed, P. Vishal, AutismSense: early prediction of autism in children using machine learning, in: 4th Int. Conf. Electron. Sustain. Commun. Syst., ICESC 2023 - Proc., 2023, p. 1719–1725, https://doi.org/10.1109/ICESC57686.2023.10193374. [81]J. Peng, M. Zhao, J. Havrilla, C. Liu, C. Weng, W. Guthrie, R. Schultz, K. Wang, Y. Zhou, Natural language processing (NLP) tools in extracting biomedical concepts from research articles: a case study on autism spectrum disorder, BMC Med. Inf. Decis. Making 20 (2020) 1–9, https://doi.org/10.1186/s12911-020- 01352-2. [82]A.B.R. Shatte, D.M. Hutchinson, S.J. Teague, Machine learning in mental health: a scoping review of methods and applications, Psychol. Med. 49 (2019) 1426–1448, https://doi.org/10.1017/S0033291719000151. [83]K.P. Venkatesh, M.M. Raza, J.C. Kvedar, Health digital twins as tools for precision medicine: considerations for computation, implementation, and regulation, Npj Digit. Med. 5 (2022), https://doi.org/10.1038/s41746-022-00694-7. [84]A. Krishnan, R. Zhang, V. Yao, C.L. Theesfeld, A.K. Wong, A. Tadych, N. Volfovsky, A. Packer, A. Lash, O.G. Troyanskaya, Genome-wide prediction and functional characterization of the genetic basis of autism spectrum disorder, Nat. Neurosci. 19 (2016) 1454–1462, https://doi.org/10.1038/n.4353. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 16 [85]S.S. Rajagopalan, Y. Zhang, A. Yahia, K. Tammimies, Machine learning prediction of autism spectrum disorder from a minimal set of medical and background information, JAMA Netw. Open 7 (2024), https://doi.org/10.1001/jamanetworkopen.2024.29229. [86]K. Khan, R. Katarya, MCBERT: a multi-modal framework for the diagnosis of autism spectrum disorder, Biol. Psychol. 194 (2024), https://doi.org/10.1016/j. biopsycho.2024.108976. [87]K. Khan, R. Katarya, WS-BiTM: integrating White shark optimization with Bi-LSTM for enhanced autism spectrum disorder diagnosis, J. Neurosci. Methods 413 (2024), https://doi.org/10.1016/j.jneumeth.2024.110319. [88]A. Massoodi, M. Taghavijelodar, M. Erfanipour, F. Jouybari, Z. Ghasempour, Detecting autism spectrum disorders from resting-state fMRI in young children using bidirectional long-short term memory neural networks, InfoScience Trends 2 (2025) 25–35, https://doi.org/10.61186/ist.202502.04.03. [89]W. Kasri, Y. Himeur, A. Copiaco, W. Mansoor, A. Albanna, V. Eapen, Hybrid vision transformer-mamba framework for autism diagnosis via eye-tracking analysis, in: International Conference on Communication, Computing, Networking, and Control in Cyber-Physical Systems, 2025, p. 343–348, https://doi.org/ 10.1109/CCNCPS66785.2025.11135843. [90]X. Liu, R. Hasan, T. Gedeon, Z. Hossain, MADE-for-ASD: a multi-atlas deep ensemble network for diagnosing autism spectrum disorder, Comput. Biol. Med. 182 (2024), https://doi.org/10.48550/arXiv.2407.07076. R. Mu ̃ noz-Terol et al. Heliyon 12 (2026) e44359 17