Paper deep dive
KGS-GCN: Enhancing Sparse Skeleton Sensing via Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Action Recognition
Yuhan Chen, Yicui Shi, Guofa Li, Liping Zhang, Jie Li, Jiaxin Gao, Wenbo Chu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:16:32 AM
Summary
KGS-GCN is a novel graph convolutional network for skeleton-based action recognition that addresses data sparsity and topological rigidity by integrating kinematics-driven Gaussian splatting and probabilistic topology. It transforms discrete joints into continuous generative representations using anisotropic covariance matrices derived from joint velocities and constructs an adaptive prior adjacency matrix via the Bhattacharyya distance between joint Gaussian distributions.
Entities (5)
Relation Signals (3)
Probabilistic Topology Construction → uses → Bhattacharyya distance
confidence 99% · quantifies statistical correlations via the Bhattacharyya distance between joint Gaussian distributions
KGS-GCN → implements → Probabilistic Topology Construction
confidence 95% · KGS-GCN introduces a probabilistic topology construction strategy
KGS-GCN → utilizes → Kinematics-Driven Gaussian Splatting Module
confidence 95% · KGS-GCN, a graph convolutional network that integrates kinematics-driven Gaussian splatting
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Skeleton-based action recognition is widely utilized in sensor systems including human-computer interaction and intelligent surveillance. Nevertheless, current sensor devices typically generate sparse skeleton data as discrete coordinates, which inevitably discards fine-grained spatiotemporal details during highly dynamic movements. Moreover, the rigid constraints of predefined physical sensor topologies hinder the modeling of latent long-range dependencies. To overcome these limitations, we propose KGS-GCN, a graph convolutional network that integrates kinematics-driven Gaussian splatting with probabilistic topology. Our framework explicitly addresses the challenges of sensor data sparsity and topological rigidity by transforming discrete joints into continuous generative representations. Firstly, a kinematics-driven Gaussian splatting module is designed to dynamically construct anisotropic covariance matrices using instantaneous joint velocity vectors. This module enhances visual representation by rendering sparse skeleton sequences into multi-view continuous heatmaps rich in spatiotemporal semantics. Secondly, to transcend the limitations of fixed physical connections, a probabilistic topology construction method is proposed. This approach generates an adaptive prior adjacency matrix by quantifying statistical correlations via the Bhattacharyya distance between joint Gaussian distributions. Ultimately, the GCN backbone is adaptively modulated by the rendered visual features via a visual context gating mechanism. Empirical results demonstrate that KGS-GCN significantly enhances the modeling of complex spatiotemporal dynamics. By addressing the inherent limitations of sparse inputs, our framework offers a robust solution for processing low-fidelity sensor data. This approach establishes a practical pathway for improving perceptual reliability in real-world sensing applications.
Tags
Links
- Source: https://arxiv.org/abs/2603.16943v1
- Canonical: https://arxiv.org/abs/2603.16943v1
Trouble viewing inline? Open PDF directly →
Full Text
48,132 characters extracted from source content.
Expand or collapse full text
1 X-X©XXXXIEEE.Personaluseispermitted,butrepublication/redistributionrequiresIEEEpermission. Seehttp://w.ieee.org/publications_standards/publications/rights/index.htmlformoreinformation. Abstract—Skeleton-basedactionrecognitioniswidely utilizedinsensorsystemsincludinghuman-computer interactionandintelligentsurveillance.Nevertheless, currentsensordevicestypicallygeneratesparseskeleton dataasdiscretecoordinates,whichinevitablydiscards fine-grainedspatiotemporaldetailsduringhighlydynamic movements.Moreover,therigidconstraintsofpredefined physicalsensortopologieshinderthemodelingoflatent long-rangedependencies.Toovercometheselimitations, weproposeKGS-GCN,agraphconvolutionalnetworkthat integrateskinematics-drivenGaussiansplattingwith probabilistictopology.Ourframeworkexplicitlyaddressesthechallengesofsensordatasparsityandtopological rigiditybytransformingdiscretejointsintocontinuousgenerativerepresentations.Firstly,akinematics-driven Gaussiansplattingmoduleisdesignedtodynamicallyconstructanisotropiccovariancematricesusinginstantaneous jointvelocityvectors.Thismoduleenhancesvisualrepresentationbyrenderingsparseskeletonsequencesinto multi-viewcontinuousheatmapsrichinspatiotemporalsemantics.Secondly,totranscendthelimitationsoffixed physicalconnections,aprobabilistictopologyconstructionmethodisproposed.Thisapproachgeneratesan adaptiveprioradjacencymatrixbyquantifyingstatisticalcorrelationsviatheBhattacharyyadistancebetweenjoint Gaussiandistributions.Ultimately,theGCNbackboneisadaptivelymodulatedbytherenderedvisualfeaturesviaa visualcontextgatingmechanism.EmpiricalresultsdemonstratethatKGS-GCNsignificantlyenhancesthemodeling ofcomplexspatiotemporaldynamics.Byaddressingtheinherentlimitationsofsparseinputs,ourframeworkoffersa robustsolutionforprocessinglow-fidelitysensordata.Thisapproachestablishesapracticalpathwayforimproving perceptualreliabilityinreal-worldsensingapplications. IndexTerms—Skeleton-basedactionrecognition;Gaussianspaltting;Probabilistictopologylearning I.INTRODUCTION APIDadvancementsinmicro-electro-mechanicalsystems and3Dsensingtechnologieshavepavedthewayfor motion-capture-basedperceptionincriticaldomains,suchas intelligentmedicalrehabilitation,human-robotcollaboration, andpervasivecomputing[1-3].Inthesecontexts,skeletondata servesasahigh-levelsensorsignalextractedfromdepth cameras,radars,orinertialmeasurementunits.Itcan effectivelyencapsulatehumanbiomechanicalstructureswith ThisworkissupportbyNationalKeyR&DProgramofChina 2024YFB2505500.(Correspondingauthor:GuofaLi.) YuhanChen,YicuiShi,GuofaLiandJieLiarewiththeCollegeof MechanicalandVehicleEngineering,ChongqingUniversity,Chongqing 400044,China(e-mail:20240701028@stu.cqu.edu.cn; yicuishi@cqu.edu.cn;liguofa@cqu.edu.cn;jieli@cqu.edu.cn). LipingZhangiswiththeDepartmentofMathematicalSciences, TsinghuaUniversity,Beijing100084,China(e-mail: lipingzhang@tsinghua.edu.cn). JiaxinGaoiswiththeSchoolofVehicleandMobility,Tsinghua University,Beijing100084,China(e-mail:gaojiaxin2017@163.com). WenboChuiswiththeNationalInnovationCenterofIntelligentand ConnectedVehicles,Beijing100089,China(e-mail: chuwenbo@wicv.cn). minimalbandwidthrequirementswhileensuringsuperior privacyprotectioncomparedtorawRGBvideos.Consequently, thesemeritshaveestablishedskeletondataasafundamental modalityforactionrecognitiontasks[2-4]. Traditionalsensorsignalprocessingmethodsrelyonmanual featureengineeringorRNN/CNN-basedtime-seriesanalysis. However,theseapproachesstruggletoaccommodatethe non-Euclideantopologicalstructuresinherentinhuman skeletons.Toaddressthis,ST-GCN[5]pioneeredastrategyto formulateskeletonsensordataasspatiotemporalgraphs, utilizinggraphconvolutionalnetworkstoaggregatesignals alongphysicallimbconnections.Thisparadigmeffectively establishesstructureddependenciesamongsensornodes, establishinggraphneuralnetworksasthedominantarchitecture forskeleton-basedperceptiontasks. Subsequentworkshavefocusedonovercomingthe limitationsoffixedphysicaltopologybyexploringdata-driven topologylearningandspatiotemporalinteractionmechanisms. Forinstance,2s-AGCN[6]introducedadaptivegraph convolutiontolearndata-specificadjacencymatrices end-to-end,significantlyenhancingfeaturediscriminability. Similarly,Shift-GCN[10]employedshiftoperationstoreduce computationalcomplexitywhileextendingspatiotemporal KGS-GCN:EnhancingSparseSkeletonSensing viaKinematics-DrivenGaussianSplattingand ProbabilisticTopologyforActionRecognition YuhanChen,YicuiShi,GuofaLiSeniorMember,IEEE,LipingZhang,JieLi,JiaxinGao,WenboChu R 2 receptivefields.DeGCN[20]proposeddeformablegraph convolutiontodynamicallyadjustneighborhoodaggregation ranges,accommodatingdeformationvariationsacrossdistinct actions.Meanwhile,Transformer-basedarchitectureshave leveragedself-attentionmechanismstocaptureglobal long-rangedependencies,continuouslypushingthe performanceboundariesofthistask[1,21]. Despitetheseadvancements,currentskeletonperception methodsfacetwofundamentalbottlenecksinsamplingand modelingcomplexmotionsignals.First,discretepoint samplingfailstoadequatelyrepresentcontinuousmotion signals.Conventionalpipelinestypicallytreatjointsasisolated pointsandlearnfeaturesdirectlyfromtheircoordinate sequences.Whilethisapproachremainseffectiveforslowor predictablemovements,itfailstocapturetheintricate dynamicsofcomplexactions.Forrapidorexplosive movements,instantaneousjointvelocities,directions,and momentumserveascriticaldiscriminativefactors.Being restrictedtosparsecoordinatesattenuatesspatialexpansion cuesalongthemotiondirection.Consequently,themodel becomespronetoconfusingactionsthatexhibitsimilar trajectoriesbutpossessdistinctdynamiccharacteristics. Second,currentsensortopologyconstructionlacksstatistical interpretability.Althoughadaptivegraphlearningcan transcendphysicalconnectivityconstraintstoincorporatelatent long-rangedependencies[6],theedgeweightsaretypically determinedthroughimplicitend-to-endtraining.Sucha heuristicapproachlacksexplicitstatisticalsignificanceand controllability,oftenleadingtoablack-boxoptimization processthatobscurestheunderlyingstructuralcorrelations betweensensornodes.Furthermore,thejointoptimizationof adjacencymatricesandnetworkparameterscanleadto topologicalforgettingorrelationalinstability,thereby degradingthefidelityoftopologyawareness[19]. Consequently,constructingtopologicalpriorswithrigorous statisticalsignificance,interpretability,andtransferability, whilesimultaneouslyenhancingcontinuousmotion representation,remainsapivotalchallengeinskeleton-based actionrecognition. Toaddressthesechallenges,thispaperproposesKGS-GCN, aGraphConvolutionalNetworkintegratedwithKinematic GaussianSplattingandProbabilisticTopology.Ourframework re-envisionsskeletonjointrepresentationandtopology constructionthroughthelensesofkinematicsandprobability. Bydoingso,weeffectivelyresolvethechallengesofsensor datasparsityandtopologicalrigidityinherentinconventional methods.Drawinginspirationfromthesuccessof3DGaussian Splattingincontinuousradiancefieldrepresentationand efficientrendering[24],theproposedmethodtransforms discretejointsfromdeterministicpointsintoprobability distributions,explicitlyencodingspatialuncertaintywithinthe featurespace.Incontrasttocomputergraphics,which prioritizesstaticgeometricrepresentation,ourapproach emphasizesthekinematics-drivennatureofsensordata.As jointvelocityincreases,thecorrespondingspatialdistribution elongatesanisotropicallyalongthedirectionofmotion.This mechanismensuresthatvelocityandorientationarenaturally encodedasintrinsicfeaturesoftherepresentation.Furthermore, formulatingnodesasprobabilitydistributionsenablesthe rigorousquantificationofinter-noderelationshipsviastatistical distance.Thisestablishesaninterpretablepriorforconstructing semantictopologiesthattranscendphysicalconnections. Specifically,themaincontributionsofthisworkare summarizedasfollows: 1.Akinematics-drivenGaussiansplattingmoduleis designedtodynamicallyconstructanisotropiccovariance matricesbasedoninstantaneousjointvelocities.Thismodule renderssparseskeletonsequencesintomulti-viewcontinuous heatmapsenrichedwithspatiotemporalsemantics,significantly enhancingtherepresentationofrapidmotionsanduncertainty. 2.KGS-GCNintroducesaprobabilistictopology constructionstrategythatquantifiesstatisticalcorrelationsvia theBhattacharyyadistancebetweenjointGaussian distributions.Thisgeneratesaninterpretableadaptiveprior adjacencymatrixtocomplementphysicaltopologiesand capturelatentlong-rangedependencies. 3.KGS-GCNincorporatesavisualcontextgating mechanismthatleveragesrenderedvisualfeaturestomodulate featurepropagationwithintheGCNbackbone,facilitatingthe synergisticmodelingofcontinuousvisualrepresentationsand graphstructurelearning.Extensiveexperimentsonmultiple benchmarkdatasetsvalidatetheframework'sefficacyin capturingcomplexspatiotemporaldynamics. I.RELATEDWORK A.Skeleton-BasedActionRecognition Skeleton-basedactionrecognitionaimstoparse spatiotemporaldynamicpatternsfromsequencesofhuman jointcoordinates.ComparedwithRGBvideos,skeleton representationsarecharacterizedbystructuralcompactnessand robustnessagainstappearanceperturbations,establishingthem asacriticalmodalityforactionunderstanding.Early approachesprimarilyreliedonhandcraftedfeaturesor sequencemodelsbasedonRNNsandCNNs,yettheyfaced inherentlimitationsinexplicitlyrepresentingthe non-Euclideantopologyofthehumanbody.ST-GCN[5] pioneeredthemodelingofskeletonsequencesas spatiotemporalgraphs,performingfeaturepropagationalong physicalskeletalconnections.Thisparadigmestablishedthe dominanceofGraphConvolutionalNetworks(GCNs)inthe field. FollowingtheST-GCNparadigm,subsequentworkshave primarilyevolvedalongthreedirections:adaptivetopology learning,thedesignofefficientspatiotemporaloperators,and theenhancementoftopologyawarenessstability.Regarding adaptivetopology,2s-AGCN[6]introducedend-to-end learningofdata-drivenadjacencymatrices.Thismechanism allowsgraphstructurestoadaptivelyadjustaccordingto specificsamplesornetworklayers,therebyovercomingthe limitationsoffixedphysicalconnections.AS-GCN[7]and MS-AAGCN[9]reinforcedspatiotemporalrepresentation capabilitiesthroughstructuredrelationshipreasoningand multi-streamfeaturefusion,respectively.Furthermore,DGNN [8]utilizeddirectedgraphneuralnetworkstomineasymmetric dependenciesamongskeletaljoints.Targetingefficient modeling,Shift-GCN[10]replacedcomputationallyintensive graphconvolutionswithlightweightspatialandtemporalshift mechanisms,reducingcomplexitywhileexpandingeffective receptivefields.GCN-NAS[13]leveragedneuralarchitecture 3 searchtoautomaticallydiscoveroptimalnetworkstructures, aimingtoenhancecomputationalefficiency.Furthermore, recentworkshavefocusedonconstructingrobustandefficient baselinestoreducetraininganddeploymentcosts[16]. Toenhanceoperatorexpressiveness,existingworkshave focusedondesigningrefinedfeatureaggregationmechanisms. Forinstance,DisentanglingGCN[11]proposedadisentangled graphconvolutionstrategytounifymulti-scalefeature aggregationandeliminateredundantdependencies.Similarly, Context-AwareGCN[12]incorporatedacontext-aware mechanism,enablingthenetworktoeffectivelycapture frame-wiseglobalcontextinformation.Subsequently, CTR-GCN[14]refinedtopologymodelingalongthechannel dimensiontolearnchannel-specificrelationalstructures. InfoGCN[15]introducedaninformationbottleneckobjective topromotediscriminativerepresentationlearning.Furthermore, DeGCN[20]integrateddeformablesamplingintospatialand temporalgraphconvolutions,adaptingtointra-class deformationsbylearningdynamicreceptivefields. Hierarchicaldecompositionandlong-rangeconnection modelingwereemployedtoenhancemulti-scalerelational representations[17].Meanwhile,Transformerarchitectures havefurtherpushedperformanceboundariesbyleveraging self-attentionmechanismstocapturegloballong-range dependencies[21]. Despitecontinuousadvancementsintopologylearningand spatiotemporalinteractionmodeling,mainstreamframeworks typicallytreatjointsasdiscretepointcoordinates.This limitationhinderstheexplicitcharacterizationofkinematic uncertaintyinrapidmotionsandtheassociatedspatial distributionalongthevelocitydirection.Furthermore,although adaptivegraphlearningmethods[6]capturelatentlong-range dependencies,theedgeweightsaretypicallylearnedimplicitly asnetworkparameters,resultinginalackofexplicitstatistical significanceandcontrollability.Duringjointoptimization,the emphasisontopologicalinformationmaybecompromisedor subjecttoforgetting,leadingtothedegradationoftopology awareness[18-19].Motivatedbytheseobservations,the proposedKGS-GCNformulatesjointsasprobability distributionsratherthandeterministicpoints.Byleveraging statisticaldistancetoconstructinterpretableprobabilistic topologypriors,thisapproachprovidesaunifiedand interpretableperspectiveforcontinuousmotionrepresentation andtopologylearning. B.GaussianSplatting Thefieldofneuralrenderingfocusesonrepresentingand renderingscenesinadifferentiablemanner.NeuralRadiance Fields(NeRF)[22]achievedhigh-qualitynovelviewsynthesis viaimplicitneuralfields,buttheyincurhightrainingand inferencecosts.Meanwhile,relatedimplicitneural representationsemployingperiodicactivationfunctions[23] havedemonstratedstrongfittingcapabilitiesforcontinuous signals.Incontrast,3DGaussianSplatting(3DGS)[24]adopts explicitGaussianprimitivestorepresentscenes,integrating differentiablerasterizerstoachieveefficientrenderingand high-qualityreconstruction.Thisapproachhassignificantly advancedthepoint-basedexplicitrenderingparadigm. Regardinggeometricconsistency,2DGaussianSplatting (2DGS)[25]modelsobjectsurfacesusingsurfel-based2D Gaussiandisks.Thismethodimprovesgeometricaccuracy, establishingasolidfoundationforsubsequentextensions. Benefitingfromtheefficiencyofexplicitrepresentationand differentiablerendering,Gaussiansplattingtechniqueshave beenrapidlyappliedtodiversetasks.Intherealmof2Dimage representationandcompression,approachessuchas GaussianImage[26]andLargeImagesAreGaussians[27] utilize2DGaussianprimitivestoconstructefficientimage representations.Furthermore,InstantGaussianImage[28] exploresimagerepresentationcapabilitieswithenhanced generalizationandadaptivity.Beyondgenerationand representationtasks,sparseGaussianrepresentationshavebeen appliedtodatacompressionandknowledgedistillationto enhanceefficiencyandscalability[29].Furthermore, approacheslikeSpeedy-Splat[30]acceleratethe3DGS pipelinethroughrenderingandoptimizationimprovements, whileGaussianPro[31]refinestrainingstrategiesvia progressivepropagation.Forlarge-scalescenereconstruction, CityGaussian[32]improvestrainingefficiencyandreal-time renderingcapabilitiesincomplexscenariosbyemploying divide-and-conquerandlevel-of-detailstrategies.Inthecontext ofdynamicscenemodeling,approachessuchasStreet Gaussians[33],MVSGaussian[34],andDrivingGaussian[35] incorporatetemporaldimensionsordynamicattributesto capturemovingobjectsandenvironmentalvariations. Additionally,Momentum-GS[36]reinforcesthequalityand stabilityofreconstructionthroughself-distillationand consistencyconstraints. Ingeneral,existingGaussiansplattingapproachesprimarily targetgenerativeorreconstructivetasks,suchasvisual reconstruction,novelviewsynthesis,andcompression.Evenin dynamicsettings,currentworksfocuspredominantlyon improvingthereconstructionqualityoftime-varying appearanceandgeometry[33-36].Incontrast,for discriminativetaskssuchasskeleton-basedactionrecognition, systematicexplorationislackingregardingthetransformation ofanisotropicdeformationfeaturesofGaussianprimitivesinto learnablekinematiccuesandtheirsynergisticmodelingwith graphstructurelearning.TheproposedKGS-GCNintroduces kinematics-drivenanisotropicGaussiansplattingtoexplicitly encodejointvelocityanddirectionalinformationinto spatiotemporalcontinuousjointheatmaps.Simultaneously, interpretableprobabilistictopologypriorsareconstructed basedonthestatisticaldistancebetweenjointGaussian distributions.Thisapproachachievestheunifiedmodelingof continuousmotionrepresentationsandinterpretabletopology learningforskeleton-basedactionrecognition. I.PROPOSEDMETHOD ThissectiondetailsKGS-GCN,agraphconvolutional networkframeworkintegratingkinematicGaussiansplatting andprobabilistictopology.AsillustratedinFig.1,theproposed methodaimstomitigatetheinherentlimitationsofdiscrete skeletonrepresentationsthroughkinematics-awarecontinuous fieldmodeling. A.ProblemDefinitionandOverallFrameWork Theinputsequenceisrepresentedasafive-dimensional tensorwhereineachsamplecomprises 퐀 pedestrianswith 퐀 4 jointsspanningasequencelength 퐀 with 퐀 coordinate channels: 퐀∈ ℜ 퐀×퐀×퐀×퐀×퐀 , 퐀∈[1,퐀]∩ℤ (1) whereNdenotesthebatchsize,andKrepresentsthenumberof classes.TheKGS-GCNframeworkspecificallytargetstwo criticalchallenges:1.theinadequacyofdiscrete,sparsejoint coordinatesincharacterizingfine-grainedmotionblurand velocity;and2.therestrictionsofpredefinedskeleton topologies,whichhindertheadaptivemodelingoflatent long-rangedependencies. Twocorecomponentsareintroducedtoaddressthisissue. Sparseskeletonsareinnovativelyrenderedintocontinuous multi-viewheatmapsbytheKinematics-drivenGaussian SplattingModule(KGSM)alongsidetheexplicitconstruction ofvelocity-drivenanisotropiccovariance.Aninnovative probabilistictopologyconstructionstrategyisconcurrently proposed.Specifically,KGS-GCNformulateseachjointasa Gaussiandistributionandquantifiesstatisticalcorrelationsvia theBhattacharyyadistancetogenerateasample-adaptiveprior adjacencymatrix.Finally,theframeworkincorporatesavisual contextgatingmechanismtoinjecttherenderedvisualcontext intotheGCNbackbone. Thedetailedproceduralstepsoftheentireframeworkare summarizedinAlgorithm1toelucidatetheoveralltraining procedureofKGS-GCN.Skeletonsequencesarespecifically transformedintomulti-viewheatmapsviakinematic computationsandtheGaussiansplattingmodule. Algorithm1outlinesthedetailedtrainingworkflowofthe proposedKGS-GCN.Initially,theframeworktransforms skeletonsequencesintomulti-viewheatmapsbyleveraging kinematiccomputationsandtheGaussianSplattingModule. Subsequently,theprobabilistictopologyinferredviathe Bhattacharyyadistanceinitializestheadjacencymatrixofthe GCN.Duringfeatureaggregation,visualfeaturesmodulate skeletonrepresentationslayer-by-layerthroughagating mechanism.Finally,themodelisupdatedbyjointly minimizingtheclassificationlossandthetopologyconstraint loss. Fig.1.TheoverallframeworkofKGS-GCN.Inputskeletonsequencesareprocessedtoextracthybridspatialandkinematicfeatures.Anisotropic heatmapsandprobabilisticjointdistributionsaregeneratedviatheKinematics-drivenGaussianSplattingModuledrivenbyhybridfeatures. Probabilistictopologyisconstructedutilizingstatisticaldistancemetricstocapturelatentlong-rangedependencies.Discreteskeletonfeaturesfrom thebackbonenetworkandcontinuousvisualcuesfromGaussianmapsarefinallyintegratedtopredictactionclasses. 5 B.Kinematics-DrivenGaussianSplatting Traditionalsparseanddiscreteskeletonrepresentationsoften overlookmotionblureffectsinducedbyvelocityvariations, whichencapsulaterichtemporaldynamicinformation.To addressthislimitation,theKinematics-DrivenGaussian SplattingModule(KGSM)leveragesinstantaneouskinematic statestodynamicallyconstructanisotropicGaussian distributions. KinematicStateExtractionandNormalization:Tomitigate scalediscrepanciesacrossdatasets,weimplementanadaptive skeletonnormalizationstrategy.Specifically,foragivenframe 퐀 ,wecalculatethegeometriccenterandtranslatetheskeletonto theorigin.Subsequently,wescalethecoordinatestothe interval [−퐀,퐀] basedonthemaximumboundingsphere radiustoobtainthenormalizedcoordinates 퐀 ( 퐀 issetto0.8). Furthermore,toexplicitlymodelmotiontrends,wecompute theinstantaneousvelocityvector퐀 퐀谀퐀 ∈ ℜ 퐀×퐀×퐀 foreachjoint: 퐀 퐀,퐀 =퐀 퐀+1,퐀 −퐀 퐀,퐀 (2) where 퐀 퐀,퐀 and 퐀 퐀,퐀 representthevelocityandpositionofthe 퐀 jointinthe 퐀 frame. DynamicAnisotropicCovarianceConstruction:This representsthecoreinnovationofKGS-GCN.Unliketraditional Gaussianheatmapsthatrelyonfixedvariances,we dynamicallyadapttheshapeandorientationofGaussian kernelsaccordingtojointvelocities.Specifically,foreachjoint, weconstructa2Dcovariancematrix Σ∈ℜ 2×2 definedbythe interplayofrotation 퐀 andscalingmatrices 퐀 ,formulatedas: Σ=퐀 퐀 퐀 퐀 (3) Regardingtheconstructionofthescalingmatrix 퐀= 퐀堀(퐀 퐀 ,퐀 퐀 ) ,wesimulatemotionblurbystretchingthescale 퐀 퐀 alongthedirectionofmotionasthevelocitymagnitude ‖퐀‖ increases,whilepreservingthebasescale 퐀 base inthe perpendiculardirection 퐀 퐀 .Weformulatethisprocessas: 퐀 퐀 =퐀 base ⋅(1+퐀tanh(‖퐀‖)),퐀 퐀 =퐀 base (4) Thestretchingdegreeiscontrolledbythehyperparameter 퐀 set to2inthisworkwhereastheadaptivebaseline 퐀 base islearned bythenetwork.Therotationmatrix 퐀 isdeterminedbythe directionofthevelocityvector (cos퐀,sin퐀) andis formulatedasfollowsutilizingthenormalizedvelocity direction: 퐀= cos퐀−sin퐀 sin퐀cos퐀 (5) Accordingtotheabovedefinition,theelementsofafter expansioncanberepresentedas: 퐀 = 퐀 퐀 2 cos 2 퐀+퐀 퐀 2 sin 2 퐀 퐀 = 퐀 퐀 2 sin 2 퐀+퐀 퐀 2 cos 2 퐀 퐀 = (퐀 퐀 2 −퐀 퐀 2 )sin퐀cos퐀 (6) Throughthisformulation,stationaryjointsmanifestas isotropiccirculardistributions,whilerapidlymovingjoints appearasellipseselongatedalongtheirmotiontrajectories. Consequently,thismechanismexplicitlyencodestemporal motionintensitydirectlywithinthespatialdomain. Multi-viewRendering:The3Dspaceisprojectedontothree orthogonalplanes (퐀Ⰰ,Ⰰ , 퐀) toprocess3Dskeletondata wherein2DGaussianSplattingisindependentlyexecutedon eachview.Theresponseintensity 퐀 퐀,퐀 (퐀) generatedbythe 퐀 jointatthe 퐀 framefollowsamultivariateGaussiandistribution foranarbitrarypixelpoint 퐀∈ ℜ 2 ontheplane: 퐀 퐀,퐀 (퐀)=exp(− 1 2 (퐀−퐀 퐀,퐀 ) 퐀 Σ 퐀,퐀 −1 (퐀−퐀 퐀,퐀 ))(7) where 퐀 퐀,퐀 = 퐀 퐀,퐀 denotesthemeanvectoroftheGaussian distributioncorrespondingphysicallytothecenterpositionof thejointwithintheimagecoordinatesystem.Theheatmap 퐀 퐀 is formulatedastheaggregationofresponsesfromalljointsfor thefinalrepresentation: 퐀 퐀 (퐀)=A( 퐀 퐀,퐀 (퐀)) 퐀=1 퐀 (8) Sparseskeletonsequencesaretransformedintomulti-channel continuousvisualrepresentations 퐀 퐀谀퐀谀퐀 ∈ ℜ 퐀×퐀×퐀×퐀 viathis process. C.ProbabilisticTopologyConstruction Latentdependenciesamongphysicallyunconnectedjoints arefrequentlyoverlookedbyphysicalconnectiongraphs.We proposetheconstructionofaprobabilistictopologyutilizing statisticalparametersgeneratedbyKGSMtoaddressthis limitation.Eachjointismodeledasaprobabilitydistribution N 퐀 퐀 , Σ 퐀 whereinthecorrelationbetweenjoints 퐀 and 퐀 is quantifiedviatheBhattacharyyadistancebetweentheir respectivedistributions.Theanalyticalformisexpressedas: 퐀 퐀 (퐀,퐀)= 1 8 퐀 퐀 − 퐀 퐀 퐀 Σ 퐀堀 −1 퐀 퐀 − 퐀 퐀 + 1 2 ln detΣ 퐀堀 det Σ 퐀 detΣ 퐀 (9) Thefirsttermof(7)quantifiesthespatialEuclideandistance betweenjointswhereasthesecondtermmeasurestheshape discrepancybetweentwodistributions. Σ 퐀堀 isexpressedas: Σ 퐀堀 = Σ 퐀 + Σ 퐀 2 (10) BasedontheBachdistance 퐀 퐀 ,weconstructedanadaptive prioradjacencymatrix 퐀 prior ∈ ℜ 퐀×퐀 : 퐀 prior (퐀,퐀)= 1 퐀 퐀=1 퐀 exp −퐀 퐀 ( 퐀 퐀 ,퐀 퐀 )(11) Long-rangedependenciesamongjointsbasedonmotion statisticalcharacteristicsareadaptivelycapturedbythismatrix 퐀 퐀簀퐀 .Thematrixissubsequentlyinjectedintothefollowing GraphConvolutionalNetworkaspriorknowledge. D.Visual-ContextModulatedGCN Weconstructtheskeletonrecognitionnetworkbystacking 퐀 Spatio-TemporalGraphConvolutionalModules(ST-Blocks). AsillustratedinFig.1,eachST-Blockadoptstheclassic spatial-temporaldecouplingdesign,consistingoftwo sequentialsub-stages:theVisually-EnhancedSpatialGCNand theMulti-ScaleTemporalTCN.Thespatialmodelingphase targetsthecaptureofintra-framejointdependencies.Tothis end,weemployachannel-leveltopologyrefinement mechanismandintegratevisualcontextgatingtoachievedeep multi-modalfeaturefusion. Initially,weprocesstherenderedinput 퐀 render viaa lightweightCNNvisualbranch.Thismoduleincorporates 6 multipleconvolutionallayersanddownsamplingoperationsto extracthigh-levelvisualsemanticfeatures.Subsequently,we applyglobalaveragepoolingandlinearprojectiontotheoutput featurestoyieldthevisualcontextfeature 퐀 퐀 ∈ ℜ 퐀×퐀 ' ×퐀 . VisualContextGating:Thisworkdesignavisualcontext gatingmechanismwithintheGCNlayertoachievedeepfusion ofskeletonandvisualfeatures.Let Ⰰ 堀倀퐀 ∈ ℜ 퐀×퐀 簀瀀퐀 ×퐀×퐀 denote theintermediatefeatureoftheGCNlayer. 퐀 퐀 isalignedand expandedtotheidenticaldimensiontogeneratemodulation coefficientsviaanonlineargatingnetwork: 퐀=퐀( 퐀 堀 ⋅ 퐀 퐀 +퐀 堀 )(12) where 퐀 isthesigmoidactivationfunctionand 퐀 堀 isthe learnableprojectionweights.Thefinalfusedfeaturesare calculatedasfollows: Ⰰ = Ⰰ 堀倀퐀 ⊗(1+퐀)+Ⰰ 퐀谀퐀 (13) where ⊗ denoteselement-wisemultiplication.Specific skeletonchannelfeaturesareadaptivelyenhancedor suppressedaccordingtothecurrentactioncontextviathis residualgatingmechanismtoachievethecomplementarityof multi-modalinformation. GraphConvolutionandTopologyFusion:Thetotal adjacencymatrix 퐀 withineachgraphconvolutionlayeris composedofthepre-definedphysicalgraph 퐀 퐀ℎ퐀 ,the network-learnedgraph 퐀 learn andtheprobabilistictopology 퐀 prior generatedbyourmethod.Thefeatureaggregation processisformulatedas: Ⰰ 堀谀퐀 = 퐀 ( 퐀 퐀ℎ퐀 (퐀) +퐀 퐀谀퐀 (퐀) +퐀 퐀簀퐀 )퐀 퐀 (14) where 퐀 isalearnablescalingfactorusedtodynamicallyadjust theimportanceofthepriorprobabilitytopology. Multi-ScaleTemporalConvolution:Thisworkdesigna Multi-ScaleTemporalConvolutionModule(MS-TCN)along thetemporaldimensiontocaptureactionpatternsofvarying durations.Thismodulecomprisesmultipleparallelconvolution branchesutilizingdistinctdilationrates 퐀 퐀 ∈1,2,3,4 .The temporaloutput Ⰰ 簀瀀퐀 isdefinedastheaggregationofoutputs fromrespectivebranchesforthefusedfeature Ⰰ : Ⰰ 簀瀀퐀 = 퐀 Conv 퐀 퐀 (ReLU(BN(Conv 1×1 (Ⰰ ))))+Ⰰ (15) MS-TCNsimultaneouslycapturesshort-termtransientchanges suchaskickinginstantsandlong-termactiondependencieslike walkingperiodicityviathecombinationofconvolutionkernels withdiversereceptivefields.Unifiedmodelingofcomplex spatiotemporaldynamicsistherebyachieved. E.LossFunction Amulti-tasklossfunctionisadoptedformodeltraining.The standardcross-entropyloss 퐀 倀谀 servesastheprimarylossforthe classificationtask.Weintroduceatopologyconsistency regularizationterm 퐀 퐀簀퐀簀 toconstrainthelearnedtopology 퐀 퐀谀퐀 withintheGCNfromdeviatingfromstatisticaldata regularities.Thetermisformulatedas: 퐀 퐀簀퐀簀 = 1 퐀 퐀=1 퐀 ‖퐀( 퐀 learn (퐀) )−퐀 prior ‖ 퐀 2 (16) where 퐀 denotesthenumberofGCNlayerswhereas 퐀 representsthesigmoidactivationfunction. 퐀 prior istreatedas thepseudolabel.Thetotallossfunctionisformulatedas: 퐀 퐀簀퐀 =퐀 倀谀 +퐀(퐀)⋅ 퐀 퐀簀퐀簀 (17) where 퐀(퐀) denotestheweightcoefficientdependentonepoch 퐀 . 퐀(퐀) isinitializedtoasmallvalueduringtheinitialtraining phasetoallowforfreenetworkexplorationwhereastheweight isgraduallyincreasedastrainingproceedstoenforcethe constrainingeffectofthestatisticalprior. IV.EXPERIMENTS A.ExperimentalSetup Datasets:Twowidelyusedbenchmarkdatasetsintheaction understandingdomainareselectedtocomprehensively evaluatetheeffectivenessandgeneralizationcapabilityofthe EMS-GCNmodelforactionrecognitiontasks:PennAction andNTURGB+D.ThePennActiondatasetcentersonroutine sportssequencesandtypicallyprovidesrichhumanpose annotationsalongsideactioncategoryinformation.Itissuitable forverifyingmodelperformanceregardingposevariationsand motiondetailmodeling.NTURGB+Drepresentsoneofthe mostwidelyadoptedlarge-scalebenchmarksfor3Dskeleton actionrecognition,featuringdiversemotioncategoriesand scenevariations.Weutilizethisdatasettosystematicallyassess therobustnessandcross-scenariogeneralizationofEMS-GCN undercomplexconditions.Consequently,thejointevaluation onthesebenchmarksallowsustoobjectivelyverifythemodel’ sperformanceintermsofbothfine-grainedactionmodeling andstabilitywithinrealistic,complexenvironments. EvaluationMetrics:Toalignwithstandardcomparison protocolsinactionrecognition,weexclusivelyemployTop-1 Accuracyasthequantitativemetric.Thischoicehighlights thecoreperformanceofthemodelintermsofclassification correctness.Thismetricmeasurestheproportionofsamples wherethehighest-probabilitypredictionmatchestheground truth.Itdirectlyindicatesthemodel'srecognitioncapability understandardclassificationsettings,therebyfacilitating consistentandfaircomparisonsacrossdiverseexperimental configurations. ImplementationDetails:WeimplementedtheEMS-GCN architecturebasedonthePyTorchframework.Alltraining andinferencephaseswereexecutedonasingleNVIDIARTX 4060GPU.Tooptimizecomputationalefficiency,we employedmixed-precisiontraining.Toensureexperimental reproducibility,wedetailtheparametersettingsandtraining strategiesasfollows: NetworkArchitectureConfiguration:Weconstructthe KGS-GCNbackboneusing10stackedspatial-temporalgraph convolutionmodules.Thefeaturechannelsaresetto64,128, and256forlayers1-4,5-7,and8-10,respectively.Toexpand thetemporalreceptivefieldandreducecomputational overhead,weapplyatemporalconvolutionstrideof2inthe 5thand8thlayers,whilemaintainingastrideof1elsewhere. Furthermore,wefixtherenderedheatmapresolutionat 32× 32 andinitializethelearnablelog-scaleparameter 퐀簀堀(scale) to-2.0.Tofacilitategradientbackpropagationduringthe initialtrainingphase,weconfigurethevelocitystretching coefficient 퐀 to2.0,ensuringsmallvarianceinthegenerated heatmaps.Fortheclassificationhead,weprojecttheencoded visualfeaturesinto128-dimensionalvectors.Subsequently, weperformglobalaveragepoolingandconcatenatethese 7 visualvectorswiththe256-dimensionalskeletonfeatures derivedfromthe10thlayer,forwardingthefused representationtothefullyconnectedlayer. TrainingStrategy:Wetrainthenetworkend-to-endusingthe SGDoptimizer,withtheNesterovmomentumsetto0.9and weightdecayfixedat 4×10 −4 .Weinitializethebaselearning rateat0.05andapplyamulti-stepdecayschedule,scalingthe ratebyafactorof0.1atthe40thand60thepochs.Toprevent gradientinstabilityduringtheearlyphase,weimplementa linearwarm-upstrategyforthefirst10epochs,increasingthe learningratelinearlyfrom0tothebasevalue.Regardingthe lossfunction,weadoptadynamicadjustmentstrategy,where 퐀(퐀) isformulatedas: 퐀(퐀)= 퐀 base ×min (1.0,퐀/5)(18) where 퐀 퐀谀 =0.2 .Thisimpliesthattopologicalconstraintsare graduallyimposedduringtheinitialtrainingphasetoprovidea bufferperiodforadaptivenetworkadjustment. B.PerformanceComparison Webenchmarktheproposedmethodagainststate-of-the-art approachesusingtheNTU[44],NW-UCLA[45],andPenn Action[46]datasets,withdetailedcomparisonsprovidedin TableI.Theexperimentalresultshighlightthesignificant advantagesofourapproachacrossalldatasets.This performanceisparticularlynoteworthyaswepresentthefirst frameworktointegrateGaussianSplattingintothisdomain. Specifically,KGS-GCNachievessuperiorperformance acrossvariousdatasetswhilerequiringonly1.4Mparameters and1.3GFLOPs.OntheNTU-60benchmark,themodelattains 92.8%accuracyonthex-subsplitand97.2%onthex-view split.Thisperformancerankssecond,trailingFreqMixFormer byamarginalgapofonly0.2%.Furthermore,ontheNTU-120 x-subandx-viewbenchmarks,KGS-GCNachievesaccuracies of88.9%and90.8%,respectively,comparabletoleading state-of-the-artmethods.Giventhesubstantialscaleofthe NTU-120datasetandthelightweightnatureofourmodel,these resultsvalidatetheefficacyoftheproposedGaussianSplatting moduleandtheprobabilistictopologyconstructionstrategy. KGS-GCNexhibitsremarkablegeneralizationcapabilities onsmaller-scaledatasets,suchasNW-UCLAandPennAction. Specifically,ontheNW-UCLAbenchmark,ourmodel achievesexceptionalperformance,showingamarginal differenceofonly0.1%relativetotherunner-up.Similarly,on thePennActiondataset,KGS-GCNsecuresthesecondrank, trailingFreqMixFormerbyamere0.2%. Insummary,KGS-GCNstrikesanoptimalbalancebetween modelcomplexityandinferenceperformance,securingthe secondrankinparametercountandthetoprankin computationalefficiency.Theseresultsvalidatetheefficacyof theproposedGaussianSplattingmoduleandprobabilistic topologyconstructionstrategy,establishingadistinct competitiveadvantageoverstate-of-the-artmethods. C.AblationStudy Torigorouslyevaluatethecontributionofeachcomponent withinKGS-GCN,weconductcomprehensiveablationstudies TABLEI QUANTITATIVEPERFORMANCECOMPARISONRESULTSONDATASETS.REDANDBLUEINDICATETHEFIRSTANDSECONDBESTRESULTSRESPECTIVELYFOR EACHINDIVIDUALMETRIC . Methods NTU-60(%)NTU-120(%) PennAction(%)NW-UCLA(%)Params(M)Flops(G) x-subx-viewx-subx-set MS-G3D[11]91.596.286.988.496.1-2.85.2 CTR-GCN[14]92.496.488.990.496.996.51.52.0 EfficientGCN[16]91.795.788.389.196.7-2.015.2 InfoGCN[15]92.896.789.290.796.596.61.61.8 FRHead[38]93.196.889.590.997.096.82.0- BlockGCN[19]92.497.090.391.596.896.91.31.6 DeGCN[20]93.397.491.092.197.697.25.6- ST-TR[37]90.896.385.187.196.3-12.1259.4 TranSkeleton[39]92.897.089.490.596.7-2.29.2 Hyperformer[40]92.996.589.991.397.196.72.79.6 SkeMixFormer[41]93.097.190.191.399.297.42.14.8 SkateFormer[43]93.597.489.891.498.498.32.03.6 FreqMixFormer[42]93.697.490.591.999.797.42.064.4 KGS-GCN92.897.288.990.899.597.31.41.3 TABLEII Q UANTITATIVE P ERFORMANCE C OMPARISON R ESULTSOF C ORE C OMPONENT C ONTRIBUTIONS .R EDANDBLUEINDICATETHEFIRSTANDSECONDBEST RESULTSRESPECTIVELYFOREACHINDIVIDUALMETRIC . Method+KGSM+PT+VCGPennAction(%)NW-UCLA(%) Baseline×96.996.5 Baseline+KGSM√×98.096.7 Baseline+KGSM+PT√×99.196.9 Baseline+KGSM+VCG√×√98.797.1 KGS-GCN√99.597.3 8 onthePennActionandNW-UCLAdatasets.Weutilize CTR-GCN[14]asthebaselinemodelforourbackbone. Specifically,weexaminetheeffectivenessofthreekey elements:theKinematics-DrivenGaussianSplatting,the probabilistictopologyconstructionstrategy,andthevisual contextgatingmechanism.Toensurefaircomparisons,we maintainconsistenttrainingconfigurationsacrossall experiments. ContributionofIndividualComponents:Weinitially evaluatetheimpactofthecoremoduleswithinKGS-GCN, specificallyverifyingtheeffectivenessofthe Kinematics-DrivenGaussianSplattingModule(KGSM),the ProbabilisticTopologyConstructionstrategy(PT),andthe VisualContextGatingmechanism(VCG).Asindicatedin TableII,eachcomponentyieldsconsistentperformance improvements.Mostnotably,thecompleteframework outperformsthebaselinenetworkby2.6%onthePennAction datasetand0.8%onNW-UCLA.Theseoutcomesempirically validatetheefficacyofourarchitecturaldesign. SplattingStrategyAnalysis:AcoremotivationofKGS-GCN liesinexplicitlymodelingmotionblureffectsviaanisotropic covariancematrices.Tovalidatethisdesign,wecomparethe proposedkinematics-drivenstrategyagainstastandard isotropiccounterpart.Theisotropicapproachdisregards velocityvectors,limitingthecovariancematrixtoascaled identitymatrix.AsindicatedinTableIII,ourkinematics-driven methodsurpassestheisotropicbaselineby1.6%,empirically confirmingtheefficacyoftheproposedstrategy. MechanismofVisualFeatureFusion:Toidentifythe optimalstrategyforintegratingvisualcontextintotheGCN backbone,wecomparethreedistinctfusionmechanisms: Element-wiseAddition,Concatenation,andourproposed VisualContextGating(VCG).AspresentedinTableIV, simpleadditionandconcatenationyieldonlymarginal performanceimprovements.Weattributethislimitationto potentialbackgroundnoiseandspatialredundancyinthe renderedheatmaps,whichcancorrupthigh-levelskeleton semanticsduringdirectfusion.Incontrast,theVCG mechanismeffectivelymodulatesthefeatureflow, outperformingtheadditionandconcatenationstrategiesby 0.6%and0.3%,respectively.Consequently,theseresults validateVCGasthesuperiorchoiceforourframework. MetricforTopologyConstruction:Weexaminethemetrics usedtoconstructprobabilistictopologypriors.Specifically,we benchmarktheBhattacharyyadistanceadoptedinthiswork againsttheconventionalEuclideandistancederivedfromjoint coordinates.AsshowninTableV,theBhattacharyyadistance yieldssuperiorperformance,validatingitsselectionfor KGS-GCN. V.CONCLUSION WeproposeKGS-GCN,agraphconvolutionalnetworkthat integrateskinematics-drivenGaussianSplattingwith probabilistictopology.Byconceptualizingskeletondataas generativesourcesofcontinuousvisualsignals,ourapproach significantlyenhancesthecapacitytomodelcomplex spatiotemporaldynamics.Ultimately,thisframeworkprovides anovelperspectivefortheunifiedmodelingofskeletonand visualfeatures.Ourcorecontributionliesintheproposalofthe Kinematics-drivenGaussianSplattingModule.Byleveraging instantaneousjointvelocities,thismoduledynamically constructsanisotropiccovariancematrices.Thismechanism effectivelyrecoversthemotionblurandspatiotemporal continuitylostindiscretecoordinates,therebyproviding semanticinputssignificantlyricherthanrawpositional information.Furthermore,wechallengetheconventionoffixed orattention-basedtopologiesbyintroducingaprobabilistic topologyconstructionstrategy.BymodelingjointsasGaussian distributionsandquantifyingtheiroverlapviathe Bhattacharyyadistance,wederiveagraphstructurethat capturesintrinsicstatisticalmotioncorrelations.This formulationprovesrobustagainstnoisearisingfrom non-physicallyconnectedjoints.Ultimately,byemploying renderedvisualcontexttomodulategeometricgraph convolutions,ourunifiedarchitectureestablishesanovel synergybetweenlow-levelkinematicsandhigh-levelvisual semantics.Whiletheseresultsareencouraging,thereremains potentialforfurtherinvestigation.Infuturework,weaimto developlightweightapproximationmethodsforprobabilistic topologytoenhancecomputationalefficiency.Additionally, weintendtoexploretheapplicationofKGS-GCNingenerative tasks,wherecontinuousGaussianrepresentationscanprovide superiorsmoothnessandenhancedinterpretability. TABLEIII A BLATION R ESULTSOF S PLATTING S TRATEGY A NALYSIS .R EDINDICATES THEBESTRESULTFOREACHINDIVIDUALMETRIC. StrategyCovarianceTypeMotion-AwarePennAction(%) Isotropic Splatting ScaledIdentity×97.9 Kinematics- Driven Anisotropic√99.5 TABLEIV ABLATIONRESULTSOFTHEVISUALFEATUREFUSIONMECHANISM.RED INDICATESTHEBESTRESULTFOREACHINDIVIDUALMETRIC . Fusion Mechanism FormulaPropertiesPennAction(%) Element-wise Addition Ⰰ=퐀+퐀 퐀 Equal Weight 98.9 ConcatenationⰀ=Concat(퐀,퐀 퐀 ) Channel Expansion 99.2 VCG Ⰰ=퐀⋅(1+퐀(퐀 퐀 )) Adaptive Selection 99.5 TABLEV ABLATIONRESULTSOFMETRICFORTOPOLOGYCONSTRUCTION.RED INDICATESTHEBESTRESULTFOREACHINDIVIDUALMETRIC . MetricNW-UCLA(%)PennAction(%) EuclideanDistance96.998.9 Bhattacharyya97.399.5 9 REFERENCES [1]W.Xin,R.Liu,Y.Liu,etal.,“Transformerforskeleton-basedaction recognition:Areviewofrecentadvances,”Neurocomputing,vol.537,p. 164–186,2023. [2]R.Yue,Z.Tian,andS.Du,“ActionrecognitionbasedonRGBand skeletondatasets:Asurvey,”Neurocomputing,vol.512,p.287–306, 2022. [3]Y.KongandY.Fu,“Humanactionrecognitionandprediction:A survey,”Int.J.Comput.Vis.,vol.130,no.5,p.1366–1401,2022. [4]J.Zhang,L.Lin,S.Yang,etal.,“Self-supervisedskeleton-basedaction representationlearning:Abenchmarkandbeyond,”arXivpreprint arXiv:2406.02978,2024. [5]S.Yan,Y.Xiong,andD.Lin,“Spatialtemporalgraphconvolutional networksforskeleton-basedactionrecognition,”inProc.AAAIConf. Artif.Intell.,vol.32,no.1,2018. [6]L.Shi,Y.Zhang,J.Cheng,etal.,“Two-streamadaptivegraph convolutionalnetworksforskeleton-basedactionrecognition,”inProc. IEEE/CVFConf.Comput.Vis.PatternRecognit.(CVPR),2019,p. 12026–12035. [7]M.Li,S.Chen,X.Chen,etal.,“Actional-structuralgraphconvolutional networksforskeleton-basedactionrecognition,”inProc.IEEE/CVF Conf.Comput.Vis.PatternRecognit.(CVPR),2019,p.3595–3603. [8]L.Shi,Y.Zhang,J.Cheng,etal.,“Skeleton-basedactionrecognitionwith directedgraphneuralnetworks,”inProc.IEEE/CVFConf.Comput.Vis. PatternRecognit.(CVPR),2019,p.7912–7921. [9]L.Shi,Y.Zhang,J.Cheng,etal.,“Skeleton-basedactionrecognitionwith multi-streamadaptivegraphconvolutionalnetworks,”IEEETrans. ImageProcess.,vol.29,p.9532–9545,2020. [10]K.Cheng,Y.Zhang,X.He,etal.,“Skeleton-basedactionrecognition withshiftgraphconvolutionalnetwork,”inProc.IEEE/CVFConf. Comput.Vis.PatternRecognit.(CVPR),2020,p.183–192. [11]Z.Liu,H.Zhang,Z.Chen,etal.,“Disentanglingandunifyinggraph convolutionsforskeleton-basedactionrecognition,”inProc.IEEE/CVF Conf.Comput.Vis.PatternRecognit.(CVPR),2020,p.143–152. [12]X.Zhang,C.Xu,andD.Tao,“Contextawaregraphconvolutionfor skeleton-basedactionrecognition,”inProc.IEEE/CVFConf.Comput. Vis.PatternRecognit.(CVPR),2020,p.14333–14342. [13]W.Peng,X.Hong,H.Chen,etal.,“Learninggraphconvolutional networkforskeleton-basedhumanactionrecognitionbyneural searching,”inProc.AAAIConf.Artif.Intell.,vol.34,no.3,2020,p. 2669–2676. [14]Y.Chen,Z.Zhang,C.Yuan,etal.,“Channel-wisetopologyrefinement graphconvolutionforskeleton-basedactionrecognition,”inProc. IEEE/CVFInt.Conf.Comput.Vis.(ICCV),2021,p.13359–13368. [15]H.Chi,M.H.Ha,S.Chi,etal.,“InfoGCN:Representationlearningfor humanskeleton-basedactionrecognition,”inProc.IEEE/CVFConf. Comput.Vis.PatternRecognit.(CVPR),2022,p.20186–20196. [16]Y.-F.Song,Z.Zhang,C.Shan,etal.,“Constructingstrongerandfaster baselinesforskeleton-basedactionrecognition,”IEEETrans.Pattern Anal.Mach.Intell.,vol.45,no.2,p.1474–1488,2022. [17]J.Lee,M.Lee,D.Lee,etal.,“Hierarchicallydecomposedgraph convolutionalnetworksforskeleton-basedactionrecognition,”inProc. IEEE/CVFInt.Conf.Comput.Vis.(ICCV),2023,p.10444–10453. [18]Y.Zhou,Z.-Q.Cheng,J.-Y.He,etal.,“Overcomingtopology agnosticism:Enhancingskeleton-basedactionrecognitionthrough redefinedskeletaltopologyawareness,”arXivpreprintarXiv:2305.11468, 2023. [19]Y.Zhou,X.Yan,Z.-Q.Cheng,etal.,“BlockGCN:Redefinetopology awarenessforskeleton-basedactionrecognition,”inProc.IEEE/CVF Conf.Comput.Vis.PatternRecognit.(CVPR),2024,p.2049–2058. [20]W.Myung,N.Su,J.-H.Xue,etal.,“DE-GCN:Deformablegraph convolutionalnetworksforskeleton-basedactionrecognition,”IEEE Trans.ImageProcess.,vol.33,p.2477–2490,2024. [21]J.DoandM.Kim,“Skateformer:Skeletal-temporaltransformerfor humanactionrecognition,”inProc.Eur.Conf.Comput.Vis.(ECCV), 2024,p.401–420. [22]B.Mildenhall,P.P.Srinivasan,M.Tancik,J.T.Barron,R.Ramamoorthi, andR.Ng,“NeRF:Representingscenesasneuralradiancefieldsforview synthesis,”Commun.ACM,vol.65,no.1,p.99–106,2021. [23]V.Sitzmann,J.Martel,A.Bergman,D.Lindell,andG.Wetzstein, “Implicitneuralrepresentationswithperiodicactivationfunctions,”in Adv.NeuralInf.Process.Syst.(NeurIPS),vol.33,2020,p.7462–7473. [24]B.Kerbl,G.Kopanas,T.Leimkühler,andG.Drettakis,“3DGaussian splattingforreal-timeradiancefieldrendering,”ACMTrans.Graph.,vol. 42,no.4,Art.no.139,2023. [25]B.Huang,Z.Yu,A.Chen,A.Geiger,andS.Gao,“2DGaussiansplatting forgeometricallyaccurateradiancefields,”inACMSIGGRAPHConf. Papers,2024,p.1–11. [26]X.Zhang,X.Ge,T.Xu,etal.,“GaussianImage:1000FPSimage representationandcompressionby2DGaussiansplatting,”inProc.Eur. Conf.Comput.Vis.(ECCV),2024,p.327–345. [27]L.Zhu,G.Lin,J.Chen,etal.,“LargeimagesareGaussians:High-quality largeimagerepresentationwithlevelsof2DGaussiansplatting,”inProc. AAAIConf.Artif.Intell.(AAAI),2025,p.10977–10985. [28]Z.Zeng,Y.Wang,C.Yang,T.Guan,andL.Ju,“InstantGaussianImage: Ageneralizableandself-adaptiveimagerepresentationvia2DGaussian splatting,”arXivpreprintarXiv:2506.23479,2025. [29]C.Jiang,Z.Li,H.Zhao,Q.Shan,S.Wu,andJ.Su,“Beyondpixels: EfficientdatasetdistillationviasparseGaussianrepresentation,”arXiv preprintarXiv:2509.26219,2025. [30]A.Hanson,A.Tu,G.Lin,V.Singla,M.Zwicker,andT.Goldstein, “Speedy-Splat:Fast3DGaussiansplattingwithsparsepixelsandsparse primitives,”inProc.IEEE/CVFConf.Comput.Vis.PatternRecognit. (CVPR),2025,p.21537–21546. [31]K.Cheng,X.Long,K.Yang,Y.Yao,W.Yin,Y.Ma,etal.,“GaussianPro: 3DGaussiansplattingwithprogressivepropagation,”inProc.Int.Conf. Mach.Learn.(ICML),2024. [32]Y.Liu,C.Luo,L.Fan,N.Wang,J.Peng,andZ.Zhang,“CityGaussian: Real-timehigh-qualitylarge-scalescenerenderingwithGaussians,”in Proc.Eur.Conf.Comput.Vis.(ECCV),2024,p.265–282. [33]Y.Yan,H.Lin,C.Zhou,W.Wang,H.Sun,K.Zhan,etal.,“Street Gaussians:ModelingdynamicurbansceneswithGaussiansplatting,”in Proc.Eur.Conf.Comput.Vis.(ECCV),2024,p.156–173. [34]T.Liu,G.Wang,S.Hu,L.Shen,X.Ye,Y.Zang,etal.,“MVSGaussian: FastgeneralizableGaussiansplattingreconstructionfrommulti-view stereo,”inProc.Eur.Conf.Comput.Vis.(ECCV),2024,p.37–53. [35]X.Zhou,Z.Lin,X.Shan,Y.Wang,D.Sun,andM.-H.Yang, “DrivingGaussian:CompositeGaussiansplattingforsurrounding dynamicautonomousdrivingscenes,”inProc.IEEE/CVFConf.Comput. Vis.PatternRecognit.(CVPR),2024,p.21634–21643. [36]J.Fan,W.Li,Y.Han,T.Dai,andY.Tang,“Momentum-GS:Momentum Gaussianself-distillationforhigh-qualitylargescenereconstruction,”in Proc.IEEE/CVFInt.Conf.Comput.Vis.(ICCV),2025,p. 25250–25260. [37]C.Plizzari,M.Cannici,andM.Matteucci,“Skeleton-basedaction recognitionviaspatialandtemporaltransformernetworks,”Comput.Vis. ImageUnderstand.,vol.208,Art.no.103219,2021. [38]H.Zhou,Q.Liu,andY.Wang,“Learningdiscriminativerepresentations forskeletonbasedactionrecognition,”inProc.IEEE/CVFConf.Comput. Vis.PatternRecognit.(CVPR),2023,p.10608–10617. [39]H.Liu,Y.Liu,Y.Chen,C.Yuan,B.Li,andW.Hu,“TransSkeleton: Hierarchicalspatial–temporaltransformerforskeleton-basedaction recognition,”IEEETrans.CircuitsSyst.VideoTechnol.,vol.33,no.8,p. 4137–4148,2023. [40]Y.Zhou,Z.-Q.Cheng,C.Li,Y.Fang,Y.Geng,X.Xie,andM.Keuper, “Hypergraphtransformerforskeleton-basedactionrecognition,”arXiv preprintarXiv:2211.09590,2022. [41]W.Xin,Q.Miao,Y.Liu,R.Liu,C.-M.Pun,andC.Shi,“Skeleton MixFormer:Multivariatetopologyrepresentationforskeleton-based actionrecognition,”inProc.ACMInt.Conf.Multimedia(ACMMM), 2023,p.2211–2220. [42]W.Wu,C.Zheng,Z.Yang,C.Chen,S.Das,andA.Lu,“Frequency guidancematters:Skeletalactionrecognitionbyfrequency-awaremixed transformer,”inProc.ACMInt.Conf.Multimedia(ACMMM),2024,p. 4660–4669. [43]J.DoandM.Kim,“Skateformer:Skeletal-temporaltransformerfor humanactionrecognition,”inProc.Eur.Conf.Comput.Vis.(ECCV), 2024,p.401–420. [44]A.Shahroudy,J.Liu,T.-T.Ng,etal.,“NTURGB+D:Alargescale datasetfor3Dhumanactivityanalysis,”inProc.IEEEConf.Comput.Vis. PatternRecognit.(CVPR),2016,p.1010–1019. [45]W.Zhang,M.Zhu,andK.G.Derpanis,“Fromactemestoaction:A strongly-supervisedrepresentationfordetailedactionunderstanding,”in Proc.IEEEInt.Conf.Comput.Vis.(ICCV),2013,p.2248–2255. [46]J.Wang,X.Nie,Y.Xia,etal.,“Cross-viewactionmodeling,learningand recognition,”inProc.IEEEConf.Comput.Vis.PatternRecognit. (CVPR),2014,p.2649–2656. 10 YuhanChen receivedhismaster'sdegreein 2024fromtheCollegeofMechanical EngineeringatChongqingUniversityof Technology.Heiscurrentlypursuingthe Ph.D.degreeinCollegeofMechanicaland VehicleEngineeringatChongqingUniversity, China.Hisresearchinterestsincludedeep learning,Low-levelVisionandGaussian Splatting. YicuiShireceivedtheB.Edegreemajoringin AutomotiveEngineeringatChongqing Universityin2025.Heiscurrentlypursuing theM.S.degreeinAutomotiveEngineeringat ChongqingUniversity,Chongqing,China.His researchinterestsincludecomputervision andGaussianSplatting. GuofaLireceivedthePh.D.degreein MechanicalEngineeringfromTsinghua University,China,in2016.Heiscurrentlya ProfessorwithChongqingUniversity,China. Hisresearchinterestsincludeenvironment perception,driverbehavioranalysis,and smartdecision-makingbasedonartificial intelligencetechnologiesinautonomous vehiclesandintelligenttransportation systems.HeservesastheAssociateEditor forIEEETransactionsonIntelligent TransportationSystems,IEEETransactions onAffectiveComputing,andIEEESensors Journal. LipingZhang iscurrentlyatenured ProfessorinDepartmentofMathematical Sciences,TsinghuaUniversity.Shereceived herPh.D.degreefromtheAcademyof MathematicsandSystemsSciences,Chinese AcademyofSciencesin2001.Herresearch interestsincludecontinuousoptimization, tensoranalysisandcomputation,machine learning.Shehaspublishedmorethan70 researchpapersininternationaljournalssuch asMathematicalProgramming,SIAMJournal onOptimization,Mathematicsof Computation,MathematicsofOperational Research,SIAMJournalonMatrixAnalysis andApplications,JournalofMachine LearningResearch,ExpertSysytemwith Applications,etc. JieLireceivedthePh.D.degreein mechanicalengineeringfromTsinghua University,Beijing,China,2024.Heis currentlyanAssociateProfessorwiththe CollegeofMechanicalandVehicle Engineering,ChongqingUniversity, Chongqing,China.Hisresearchinterests includemodelpredictivecontrol,adaptive dynamicprogrammingandreinforcement learning. JiaxinGao receivedhisB.S.andPh.D. degreesfromtheUniversityofScience& TechnologyBeijingin2017and2023, respectively.HeiscurrentlyanAssistant ResearcherattheSchoolofVehicleand Mobility,TsinghuaUniversity.Hiscurrent researchinterestsfocusonreinforcement learning,decisionandcontrolforautonomous vehicles,andperceptualdatagenerationfor autonomousdriving. WenboChu receivedhisB.S.degree majoredinAutomotiveEngineeringfrom TsinghuaUniversity,China,in2008,andhis M.S.degreemajoredinAutomotive EngineeringfromRWTH-Aachen,German andPh.D.degreemajoredinMechanical EngineeringfromTsinghuaUniversity,China, in2014.Heiscurrentlyaresearchfellowat WesternChinaScienceCityInnovation CenterofIntelligentandConnectedVehicles (Chongqing)Co,Ltd.,andNationalInnovation CenterofIntelligentandConnectedVehicles.