Paper deep dive
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
Dong Chen, Kenneth M. C. Cheung
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/23/2026, 1:27:07 AM
Summary
The paper introduces TokenSTFormer, a novel spatial-temporal attention model for screening Adolescent Idiopathic Scoliosis (AIS) using gait video analysis. The authors present the ScoliGait dataset, comprising 1,516 gait video clips paired with X-ray records, and propose a system called ScoliDetect. TokenSTFormer utilizes Spatial-Temporal Tokenization (STT) to enhance feature representation, achieving an accuracy of 0.79, which surpasses the vanilla Vision Transformer encoder. The study highlights the potential of this approach for scalable, cost-effective, and non-invasive AIS screening.
Entities (7)
Relation Signals (6)
ScoliGait → contains → gait video clips
confidence 98% · ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records.
ScoliGait → pairedwith → X-ray records
confidence 98% · ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records.
TokenSTFormer → achievesaccuracy → 0.79
confidence 95% · Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79.
TokenSTFormer → outperforms → Vision Transformer
confidence 94% · Compared to Vision Transformer encoder, the proposed TokenSTFormer consistently outperformed the baseline across all evaluation metrics.
ScoliDetect → utilizes → TokenSTFormer
confidence 93% · we propose ScoliDetect TM , a novel system for AIS screening... introduce TokenSTFormer
YOLOv8 → usedfor → pose_estimation
confidence 90% · we utilized pose estimation technology (YoLoV8) to derive 2D joint coordinates
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.
Tags
Links
- Source: https://arxiv.org/abs/2608.16122v1
- Canonical: https://arxiv.org/abs/2608.16122v1
Trouble viewing inline? Open PDF directly →
Full Text
16,711 characters extracted from source content.
Expand or collapse full text
TokenSTFormer:ATokenizedSpatial-temporal AttentionModelforHolisticMotionAnalysisin AdolescentIdiopathicScoliosisScreening DongChen 1 , KennethM.C.Cheung 1 1 TheUniversityofHongKong olichen@connect.hku.com;oliver.cd@outlook.com Abstract.AdolescentIdiopathicScoliosis(AIS)isaprevalentspinaldeformity inadolescentsthat,ifleftuntreated,canresultinseverehealthoutcomes. Traditionalscreeningmethodsarelimitedbysubjectiveinterpretation,reliance onprofessionalexpertiseandlowscalability.Toaddressthesechallenges,we presentScoliGaitdataset,whichcomprises1,516gaitvideoclipspairedwith correspondingX-rayrecords.WealsointroduceTokenSTFormer,anovel modelthattokenizesspatialandtemporalsemanticstoenhancefeature representationandconvergence.Ourmodelachievesstate-of-the-art performance,surpassingvanillaVisionTransformerencoderacrosskey metrics,includingaccuracyof0.79.Thisstudyhighlightsthepotentialof leveragingholisticmotionfeaturesderivedfromgaitvideoandattention-based modelsforscalable,cost-effectiveAISscreening,pavingthewayforfuture clinicalapplicationsinscoliosisdetection. Keywords:AdolescentIdiopathicScoliosis,GaitAnalysis,Kinematic KnowledgeMap,Spatial-TemporalTokenization. 1Introduction Adolescentidiopathicscoliosis(AIS)isastructural,lateralcurvatureofthespine accompaniedbyvertebralrotation,typicallydiagnosedduringadolescence.Affecting approximate5%ofchildrenglobally[1,2],AIScanleadtoseverehealth consequencesifleftuntreated,suchaschronicbackpainandpsychosocialdistress[3, 4].ThecurrentgoldstandardfordiagnosingAISrequiresradiographicimagingand measuringthecoronalCobbAngle(CA)>10°,whichindicatesscoliosis[3]. However,repeatedexposuretoX-raysraisesconcernsaboutcumulativeradiation risk,highlightingtheneedfornon-invasiveandscalablescreeningmethods. Earlyscreeningisthereforecriticaltopreventingcurveprogressionandenabling timelyintervention,especiallyduringadolescencewhenthespineisstillgrowing. TraditionalAISscreeningmethods,suchastheAdams’ForwardBendingTestand ScoliometerMeasurement,arewidelyusedinclinicalandschool-basedsettings[3]. However,theireffectivenessandefficiencyvaryacrossstudies[5,6],oftenbeing influencedbysubjectiveinterpretationandfactorslikeparticipants’obesity[7]. Moreover,thesemethodsfacesignificantchallengesinscalabilityforlarge-scale 2AnonymizedAuthoretal. screeningprograms,astheyrelyonprofessionalexpertise,involvehighequipment costs,andraiseprivacyconcerns[3]. AdvancesintechnologyhaveintroducedalternativeapproachesforAISscreening, suchasusingsingle-cameraphotographstoanalyzebackasymmetry[8].However, methodsofusingstaticinformationfailtoincorporatekinematicfeaturesproviding bio-mechanicalinsights[9].Additionally,methodslikeGaitEdge[10]and SkeletonGait[11]extractspatiotemporalfeaturesusingsyntheticsilhouettemapsand skeletonmaps.Despitetheireffectiveness,thecomplexpreprocessingpipelines requiredbythesemethodslimittheirpracticalityinreal-worldapplications. Toaddressthesechallenges,weproposeScoliDetect TM ,anovelsystemforAIS screeningbasedongaitvideoanalysis,asillustratedinFig.1.Specifically,we(1) establishScoliGaitdatasetcontaining1,516gaitvideospairedwithcorresponding spinalX-rayimagesandCAmeasurements.(2)designade-identifiedkinematic knowledgemaptorepresentthekinematicfeaturesofholisticgaitmotion,enabling scalableandprivacy-consciousscreeningonmobiledevices.(3)introduce TokenSTFormer,amodelequippedwithSpatial-TemporalTokenization(STT)to enhancefeaturesrepresentationandmodelconvergence. Fig.1.WorkflowofScoliDetect TM system.Agaitvideorecordedusingamobilephonecamera isprocessedtoconstructakinematicknowledgemap,representingholisticmotionfeatures. TokenSTFormermodelincorporatesSpatial-TemporalTokenization(STT). 2Dataset 2.1DatasetDescription TheScoliGaitdatasetwascollectedusingamobilephonecameraat***Hospitalto supportscoliosisscreeningstudy.Atotalof758participantswhosignedinformed consentformwereenrolledinthisstudy.Basicdemographicdataissummarizedin TokenSTFormer3 Table1.Thedatasetwasexpandedbysegmentingnon-overlappingvideoclipsto enhanceinferencerobustnessacrossdifferentwalkingperiods.Eachclip,recordedat 30Hzframerateand1080presolution,captured5secondsofholisticwalkingmotion. Intotal,thefinaldatasetcomprises1,516videoclips. Tothebestofourknowledge,ScoliGaitisthefirstdatasettoincludebothgait videosandcorrespondingspinalX-rays,whichserveasthegoldenstandardmedical labels.Theannotationqualitywasvalidatedbyseniormedicaldoctors.Notably,the groundtruthforscoliosisdiagnosisreliesonradiographicmeasurements,unlike traditionalscreeningmethodswhichlacksufficientevidencetoserveasreliable labels. Table1.SummaryofdemographicandclinicalattributesintheScoliGaitdataset AttributesPositive(Cobbangle>10°)Negative(Cobbangle≤10°) Numberofparticipants758subjectshavinginformedconsentform Non-overlappedvideoclips 1516non-overlappingclips (150frameswith30Hzframerate) Gender(F/M) 722/320275/198 Age(mean,std)13.86,2.4411.59,2.86 2.2Datacollectionandpreprocessing ThesetupandrecordingprocessareillustratedinFig.1.Thecamerawaspositioned ataheightof2.5meterstocaptureshoulder-pelvicangles.Participantswere instructedtowalkatanaturalpace,completingoneforward-and-backwardcycle alonga4-meter-longpath,starting2metersawayfromthecamera. 3Methodology Thissectioninvolvesthewayofconstructingkinematicknowledgemaps(Fig.1) fromposeestimationdataanddesigningSTTmodulestoeffectivelycapturegait features. Givenavideoclip V i (t,w,h,c) =f 1 ,f 2 ,...,f n (1) where 퐀 퐀 representsthen th frameofi th subject,(t,w,h,c)representperiod,width, heightandchannelofframes. M i (t,v) =∅(τ∗F(V i (t,w,h,c) )) (2) where 퐀 퐀 (퐀,퐀) representsthekinematicknowledgemapwithtperiodandvvariates; 퐀 isthe2Dposeestimationtechnology; 퐀 isanamplifyingfactorwhichis1000inour setting; ∅ isapriorknowledgefunctionoflandmarkcoordinatestransformation. 4AnonymizedAuthoretal. 3.1KinematicKnowledgeMap Toextractkinematicfeatures,weutilizedposeestimationtechnology(YoLoV8)to derive2Djointcoordinates(x,y)fromthegaitvideos[12].Priorstudieshave[13, 14]approvedthatscolioticgaitmotionexhibitsdetectabledeviationscomparedto normalgait.Thesedeviationsareprimarilyinducedbymusculoskeletal,perceptronor post-adaptiveissues.Intermsofthesepriorknowledge,kinematicknowledgemapis constructedinthreedomains:(1)featuresrepresentingtheoverallgaitpatternin motionspace,(2)featurescapturingthesubject'sskeletalstructureinself-skeleton space,(3)featuresderivedfrommotionlaggingandsignalrelationships. Thekinematicknowledgemap,asshowninFig.2,comprises238featuresthat representholisticmotion,composing140featuresinmotionspace,32featuresinself- skeletonspace,and66featuresforsignalcorrelation.Thenumericalvaluesineach sectionaredependentlynormalized.Specifically,pairedjoint-relatedfeaturesare calculatedusingEuclideandistance,whilemotionanglesbetweenvectorsare determinedusingtrigonometricfunctions.Motionlaggingsectionsarederived throughsignalcross-correlation,calculatedusingtheSciPypackage. Fig.2.Akinematicknowledgemaprepresentingholisticmotionfeaturescomprises238 featuresthatrepresentholisticmotion,composing140featuresinmotionspace,32featuresin self-skeletonspace,and66featuresforsignalcorrelation. 3.2ModelArchitecture ThegeneralarchitectureofTokenSTFormerisinspiredbytheVisionTransformer, havingresidualblockscomposedofMultiheadedSelf-Attention(MSA)andMulti- TokenSTFormer5 LayerPerceptron(MLP)[15].TheproposedSpatial-TemporalTokenizationis illustratedinFig.3. Buildingoninsightsfromapriorstudy[16],aDenselayerisappliedafterspatial tokenstoenhancefeaturerepresentationacrossvariates.Additionally,otherkey modules,suchasLayerScale[17]andStochasticDepth[18],areintegratedin standardconfigurations. Spatial-TemporalTokenization.Thekinematicknowledgemap 퐀 퐀 (퐀,퐀) ,witht periodandvvariates,istokenizedintospatialandtemporaltokensusing2D convolutionallayerswithcolumn-sizeandrow-sizekernels,respectively.Temporal tokens 퐀 퐀萀퐀䠀퐀栀퐀琀 (퐀,퐀) andspatialtokens 퐀 퐀䠀퐀琀 (퐀,퐀) havedoutputdimension.Bothtokens followedLayerNorm(LN)areconcatenatedas 퐀 퐀䠀氀퐀 (퐀+퐀,퐀) . 퐀 퐀䠀氀퐀 (퐀+퐀,퐀) =퐀(퐀ئج 퐀 퐀萀퐀䠀퐀栀퐀琀 퐀,퐀 ,퐀ئج(Dense(퐀 퐀䠀퐀琀 퐀,퐀 )))(3) Themainlossfunctionisbinarycrossentropy(BCM).CLStokensarerespectively appliedfortemporalandspatialembeddingsinourexperiments.Auxiliarylossis calculatedbyMeanSquaredErrorofthesetwoCLStokens. 퐀 퐀 =퐀 퐀萀퐀䠀 ,퐀 퐀䠀퐀 (4) 퐀=퐀 퐀 +퐀 퐀 (5) Fig.3.ThefigureisSpatial-TemporalTokenizationmodule 3.3Metrics Keymetricsincludeaccuracy,sensitivity,specificity,PositivePredictiveValue (PV+),andNegativePredictiveValue(PV−).Thesemetricsprovideacomprehensive understandingofthemodel'sabilitytodistinguishbetweenconditionsofinterest(e.g., diseasevs.nodisease)basedontestoutcomes.Here,TNrepresentstruenegative;TP representstruepositive;FNrepresentsfalsenegative;FNrepresentsfalsenegative. Accuracy:Measurestheproportionofcorrectlyclassifiedinstances(bothpositive andnegative)amongallsamples. 6AnonymizedAuthoretal. Accuracy=(TN+TP)/(TN+FN+TP+FP) (6) PositivePredictiveValue(PV+):Indicatestheproportionoftruepositive predictionsamongallpositivepredictions. PositivePredictiveValue=TP/(TP+FP) (7) NegativePredictiveValue(PV−):Representstheproportionoftruenegative predictionsamongallnegativepredictions. NegativePredictiveValue=TN/(TN+FN) (8) Sensitivity:Measurestheabilityofthemodeltocorrectlyidentifytruepositive cases(i.e.,theproportionofactualpositivescorrectlypredicted). Sensitivity=TP/(TP+FN) (9) Specificity:Reflectstheabilityofthemodeltocorrectlyclassifytruenegative cases(i.e.,theproportionofactualnegativescorrectlypredicted). Specificity=TN/(TN+FP) (10) 4Experiments 4.1Experimentalsetup Inthissection,theproposedTokenSTFormermodeliscomparedwithavanilla VisionTransformerencoder[15],configuredwitha6by6patchsizeandthesame hyperparameterstothoseoftheTokenSTFormermodel.Thetraining,validationand testingdatasetsconsistedof1216,150,150samples,respectively.Toreduceclass imbalance,westratifiedthedatasampleswitha2.2:1positive-to-negativeratio.Both categorieswereadequatelyrepresentedandminimizedbiasduringmodelevaluation. Additionally,thecontributionsofSSTwereanalyzedinablationstudy.Table2is thedetailsofhyperparameterssettingusedinthisstudy. Table2.Hyperparameterdetailsformodelconfigurationandtraining ModelparametersTrainingparameters MLPdimension=384Learningrate=2e-5 Numberofheads=6Warmupratio=0.1 Numberoflayers=5Cosinelearningrateschedule MHAdimension=6*256Optimizer:Adam Dropout=0.1Batchsize=64 TokenSTFormer7 4.2Results ComparedtoVisionTransformerencoder[15],theproposedTokenSTFormer consistentlyoutperformedthebaselineacrossallevaluationmetrics.Asshownin Table3,TokenSTFormerachievedaccuracyof0.787,demonstratingitssuperior overallclassificationcapability.Additionally,itachievedasensitivityof0.845and specificityof0.660,underscoringitseffectivenessincorrectlyidentifyingboth positiveandnegativecases. Moreover,TokenSTFormerexcelledinpredictivevalues,withPV+(0.845)andPV− (0.660),highlightingitsreliabilityinmakingaccurateandbalancedpredictions.These resultshighlightedtherobustnessandefficiencyofTokenSTFormerforAIS screening. Table3.ComparisonofevaluationmetricsamongTokenSTFormer,VisionTransformer encoder,andtraditionalmethods 4.3Ablationstudies Weconductedablationstudiestoanalyzecertainfeatureeffectsonperformance.As showninTable4,LayerNormandindependentpositionalencodingplaycriticalroles inSSTmodulestoensurethemodel’saccuracyandrobustness. Table4.MetricscomparisonamongTokenSTFormer,w/oSpatial-TemporalTokenization. AccuracySensitivitySpecificityPV+PV- SSTw/oLayerNorm0.7200.7380.6810.8350.542 Singleposencoding0.6870.7280.6000.7980.500 TokenSTFormer0.7870.8450.6600.8450.660 Spatial-TemporalTokenization.ThepurposeofSTTistoseparateandnormalize temporalandspatialtokens,therebyminimizingthedistancebetweentwotypesof tokens.Tofurtheranalyzeitsimpact,weanalyzedthecosinesimilarityofspatial- temporaltokensbetweenbaselinemodelandthatwithoutLayerNorminSTT,as AccuracySensitivitySpecificityPV+PV- PaperReport[5,6] 0.460.840.30 0.510.960.950.53 0.370.900.800.59 Average0.4470.9000.6830.560 Transformerencoder0.7400.7960.6170.8200.580 TokenSTFormer0.7870.8450.6600.8450.660 8AnonymizedAuthoretal. illustratedinFig.4.Themajorityofpointslieabovethereddashedline(slope=1), indicatingthatthecosinesimilarityinthebaselinemodelissignificantlysmaller. TheseresultsdemonstratethatSTTeffectivelylearnstemporal-spatialspecific transformationstoreprojecttokensintoanenhancedsemanticspace. Numberoflayers.Weevaluatedtheperformancemetricsofmodelswithvarying numbersofattentionblocks.AsillustratedinFig.5,themodelachievesthehighest accuracywhenthenumberofattentionblocksissetto5.Thissuggeststhatan optimalbalancebetweenmodelcomplexityandperformanceisachievedatthis configuration. Fig.4.ScatterplotcomparingtokencosinesimilaritydistancesbetweentheTokenSTFormer andthatwithoutSTT. Fig.5.Theplotshowsthevariationinperformanceasthenumberofattentionblock(layers) increases. TokenSTFormer9 5Conclusions Inthisstudy,weintroducedScoliDetect TM ,anAI-assistedholisticmotionanalysis systemutilizingasmartphonecameraforscalableandcost-effectivescoliosis screening.ByleveragingSpatial-TemporalTokenization(SST),theproposed TokenSTFormereffectivelylearnsrobustspatial-temporalfeatures.Thiswork representsapromisingsteptowardthedeploymentofaccessible,accurate,and efficientdiagnostictoolsforscoliosisscreeningandmonitoring. References 1.Hengwei,F.,Zifang,H.,Qifei,W.,Weiqing,T.,Nali,D.,Ping,Y.,Junlin,Y.:Prevalence ofIdiopathicScoliosisinChineseSchoolchildren:ALarge,Population-BasedStudy. Spine,41(3),259-64(2016) 2.Catanzariti,JF.,Rimetz,A.,Genevieve,F.,Renaud,G.,Mounet,N.:Idiopathicadolescent scoliosisandobesity:prevalencestudy.EurSpineJ,32(6),2196-2202(2023) 3.Luk,K.D.,Lee,C.F.,Cheung,K.M.,Cheng,J.C.,Ng,B.K.,Lam,T.P.,Mak,K.H., Yip,P.S.,Fong,D.Y.:Clinicaleffectivenessofschoolscreeningforadolescentidiopathic scoliosis:alargepopulation-basedretrospectivecohortstudy.Spine,35(17),1607-14 (2010) 4.Cheng,J.C.,Castelein,R.M.,Chu,W.C.,Danielsson,A.J.,Dobbs,M.B.,Grivas,T.B., Gurnett,C.A.,Luk,K.D.,Moreau,A.,Newton,P.O.,Stokes,I.A.,Weinstein,S.L.,& Burwell,R.G.:Adolescentidiopathicscoliosis.NatRevDisPrimers,1,15030(2015) 5.Amendt,L.E.,Ause-Ellias,K.L.,Eybers,J.L.,Wadsworth,C.T.,Nielsen,D.H.,Weinstein, S.L.:ValidityandreliabilitytestingoftheScoliometer®.Physicaltherapy,70(2),108-117 (1990) 6.Coelho,D.M.,Bonagamba,G.H.,Oliveira,A.S.:Scoliometermeasurementsofpatients withidiopathicscoliosis.Brazilianjournalofphysicaltherapy,17(2),179-184(2013) 7.Margalit,A.,McKean,G.,Constantine,A.,Thompson,C.B.,Lee,R.J.,Sponseller,P.D.: BodyMassHidestheCurve:ThoracicScoliometerReadingsVarybyBodyMassIndex Value.Journalofpediatricorthopedics,37(4),e255-e260(2017) 8.Zhang,T.,Zhu,C.,Zhao,Y.,Zhao,M.,Wang,Z.,Song,R.,Meng,N.,Sial,A.,Diwan, A.,Liu,J.,Cheung,J.P.Y.:DeepLearningModeltoClassifyandMonitorIdiopathic ScoliosisinAdolescentsUsingaSingleSmartphonePhotograph.JAMANetworkOpen,6 (8),e2330617(2023) 9.Pesenti,S.,Prost,S.,Pomero,V.,Authier,G.,Roscigni,L.,Viehweger,E.,Blondel,B., Jouve,J.L.:Doesstatictrunkmotionanalysisreflectitstruepositionduringdaily activitiesinadolescentwithidiopathicscoliosis?.Orthopaedics&traumatology,surgery& research:OTSR,106(7),1251–1256(2020) 10.Liang,J.,Fan,C.,Hou,S.,Shen,C.,Huang,Y.,Yu,S.:Gaitedge:Beyondplainend-to- endgaitrecognitionforbetterpracticality.In:EuropeanConferenceonComputerVision, p.375-390.SpringerNatureSwitzerland,(2022) 11.Fan,C.,Ma,J.,Jin,D.,Shen,C.,Yu,S.:SkeletonGait:GaitRecognitionUsingSkeleton Maps.In:ProceedingsoftheAAAIConferenceonArtificialIntelligence,p.1662-1669, (2024). 12.GlennJocher,AyushChaurasia,andJ.Qiu."UltralyticsYOLOv8." https://github.com/ultralytics/ultralytics(accessed2024) 10AnonymizedAuthoretal. 13.Ji,R.,Liu,X.,Liu,Y.,Yan,B.,Yang,J.,Lee,W.Y.,Wang,L.,Tao,C.,Kuai,S.,Fan,Y.: Kinematicdifferenceandasymmetriesduringlevelwalkinginadolescentpatientswith differenttypesofmildscoliosis.Biomedicalengineeringonline,23(1),22(2024) 14.Boulcourt,S.,Badel,A.,Pionnier,R.,Neder,Y.,Ilharreborde,B.,Simon,A.L.:Agait functionalclassificationofadolescentidiopathicscoliosis(AIS)basedonspatio-temporal parameters(STP).Gait&Posture,102,50–55(2023) 15.Dosovitskiy,A.,etal.:Animageisworth16x16words:Transformersforimage recognitionatscale.arXivpreprintarXiv:2010.11929(2020). 16.Liu,Y.,Hu,T.,Zhang,H.,Wu,H.,Wang,S.,Ma,L.,Long,M.:itransformer:Inverted transformersareeffectivefortimeseriesforecasting.arXivpreprint arXiv:2310.06625(2023) 17.Touvron,H.,Cord,M.,Sablayrolles,A.,Synnaeve,G.,Jégou,H.:Goingdeeperwith imagetransformers.In:ProceedingsoftheIEEE/CVFinternationalconferenceon computervision,p.32-42.IEEE(2021) 18.Huang,G.,Sun,Y.,Liu,Z.,Sedra,D.,Weinberger,K.Q.:Deepnetworkswithstochastic depth.In:ComputerVision–ECCV2016:14thEuropeanConference,PartIV14,p.646- 661.SpringerInternationalPublishing,Amsterdam,TheNetherlands(2016)